Simplified entropy coding of sub-block based motion information list

By optimizing video encoding and decoding through affine coding and decoding technology, and selecting multiple control points for affine motion transformation and conditional bypass coding and decoding, the redundancy problem in the construction of the Merge mode candidate list and motion vector prediction in the HEVC standard is solved, thereby improving the encoding and decoding efficiency and quality of high-resolution video.

CN116471403BActive Publication Date: 2026-01-27DOUYIN VISION CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211599587.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-10-23
Filing Date
2019-10-23
Publication Date
2026-01-27
Estimated Expiration
2039-10-23

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies still have room for improvement in terms of encoding and decoding efficiency and bandwidth requirements when processing high-resolution video. In particular, in the HEVC standard, there are issues of computational complexity and redundancy in the construction of the candidate list for Merge mode and motion vector prediction.

Method used

Affine encoding and decoding technology is adopted. Affine motion transformation is performed by selecting multiple control points of the current block, and the Merge index is determined by using conditional bypass encoding and decoding. Motion vector prediction is optimized by combining spatial and temporal candidate lists, which reduces redundant checks and improves parallel processing capabilities.

Benefits of technology

It improves the efficiency and quality of video encoding and decoding, reduces computational complexity, enhances the ability to process high-resolution video, and optimizes the parallelism of the encoding process and the performance of encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116471403B_ABST
    Figure CN116471403B_ABST
Patent Text Reader

Abstract

Devices, systems, and methods related to digital video coding involving adaptive control point selection including affine coding are described, in particular, to simplified entropy coding based on sub-block based motion information list. An example method for video processing includes, for a conversion between a current block of a video and a bitstream representation of the video, selecting a plurality of control points for the current block, the plurality of control points including at least one non-corner point of the current block, and each of the plurality of control points representing an affine motion for the current block, and performing the conversion between the current block and the bitstream representation based on the plurality of control points.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese Patent Application No. 201980004125.2, filed on October 23, 2019. Technical Field

[0002] This patent document relates to video encoding and decoding technologies, equipment, and systems. Background Technology

[0003] Despite advancements in video compression, digital video still accounts for the largest share of bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention

[0004] Apparatus, systems, and methods for digital video coding and decoding involving adaptive control point selection, including affine coding and decoding, are described. The methods can be applied to both existing video coding and decoding standards (e.g., High Efficiency Video Coding (HEVC)) and future video coding and decoding standards, or to video codecs.

[0005] In one representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes: selecting a plurality of control points for a current block of video and a bitstream representation of the video, the plurality of control points including at least one non-corner point of the current block, and each of the plurality of control points representing an affine motion of the current block; and performing the conversion between the current block and the bitstream representation based on the plurality of control points.

[0006] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes: determining, for a conversion between a current block of video and a bitstream representation of the video, at least one of a plurality of binary bits used for conditionally using bypass encoding / decoding of a Merge index of a Merge candidate list for encoding / decoding a Merge-based sub-block; and performing the conversion based on the determination.

[0007] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes: for a conversion between a current block of video and a bitstream representation of the video, selecting a plurality of control points for the current block, the plurality of control points including at least one non-corner point of the current block, and each of the plurality of control points representing an affine motion of the current block; deriving motion vectors of one or more of the plurality of control points based on control point motion vectors (CPMVs) of one or more adjacent blocks of the current block; and performing the conversion between the current block and the bitstream representation based on the plurality of control points and the motion vectors.

[0008] In another representative aspect, the above method is implemented in the form of processor-executable code and stored in a computer-readable program medium.

[0009] In another representative aspect, a device configured or operable to perform the methods described above is disclosed. This device may include a processor programmed to implement the methods.

[0010] In another representative aspect, video decoder devices can implement the methods described in this paper.

[0011] The above-described aspects and features, as well as other aspects and features, of the disclosed technology are described in more detail in the accompanying drawings, description and claims. Attached Figure Description

[0012] Figure 1 An example of constructing a Merge candidate list is shown.

[0013] Figure 2 An example of a candidate location for the airspace is shown.

[0014] Figure 3 An example of a candidate pair that has undergone a redundancy check of the spatial merge candidate is shown.

[0015] Figure 4A and Figure 4B An example of the position of the second prediction unit PU based on the size and shape of the current block is shown.

[0016] Figure 5 An example of motion vector scaling for temporal Merge candidates is shown.

[0017] Figure 6 An example of candidate locations for temporal Merge candidates is shown.

[0018] Figure 7 An example of generating bidirectional prediction Merge candidates using a combination is shown.

[0019] Figure 8 An example of constructing motion vector prediction candidates is shown.

[0020] Figure 9 An example of motion vector scaling for spatial motion vector candidates is shown.

[0021] Figure 10 An example of motion prediction using an alternative temporal motion vector prediction (ATMVP) algorithm for codec units (CUs) is shown.

[0022] Figure 11An example is shown with a codec unit (CU) for sub-blocks and adjacent blocks used by the spatial-temporal motion vector prediction (STMVP) algorithm.

[0023] Figure 12 An example flowchart is shown for encoding with different MV precisions.

[0024] Figure 13A and Figure 13B An example of dividing a codec unit (CU) into two triangular prediction units (PUs) is shown.

[0025] Figure 14 An example of the location of adjacent blocks is shown.

[0026] Figure 15 An example of applying the first weighting factor group to CU is shown.

[0027] Figure 16 An example of motion vector storage is shown.

[0028] Figure 17A and Figure 17B An example snapshot of a sub-block is shown when using the Overlapping Block Motion Compensation (OBMC) algorithm.

[0029] Figure 18 An example of neighboring samples used to derive the parameters of the Local Luminance Compensation (LIC) algorithm is shown.

[0030] Figure 19A and Figure 19B Simplified examples of affine motion models with 4 and 6 parameters are shown respectively.

[0031] Figure 20 An example of the affine motion vector field (MVF) for each sub-block is shown.

[0032] Figure 21 An example of motion vector prediction (MVP) for the AF_INTER affine motion pattern is shown.

[0033] Figure 22A and Figure 22B Examples of affine models with 4 and 6 parameters are shown respectively.

[0034] Figure 23A and Figure 23B Example candidates for the AF_Merge affine motion mode are shown.

[0035] Figure 24 An example of candidate positions for the affine Merge pattern is shown.

[0036] Figure 25An example of bilateral matching in the Pattern Matching Motion Vector Derivation (PMMVD) mode is shown, which is a special Merge mode based on the Frame Rate Upconversion (FRUC) algorithm.

[0037] Figure 26 An example of template matching in the FRUC algorithm is shown.

[0038] Figure 27 An example of one-sided motion estimation in the FRUC algorithm is shown.

[0039] Figure 28 An example of the optical flow trajectory used by the bidirectional optical flow (BIO) algorithm is shown.

[0040] Figure 29A and Figure 29B An example snapshot is shown using the bidirectional optical flow (BIO) algorithm without block expansion.

[0041] Figure 30 An example of a decoder-side motion vector refinement (DMVR) algorithm based on bilateral template matching is shown.

[0042] Figure 31 Examples of different control points of the encoding / decoding unit are shown.

[0043] Figure 32 An example of a control point using time-domain motion information is shown.

[0044] Figure 33 An example is shown of deriving motion vectors used in OBMCs using the control point motion vectors (CPMV) of adjacent CUs.

[0045] Figure 34 An example of using the CPMV of adjacent blocks to derive a motion vector predictor is shown.

[0046] Figures 35A to 35C A flowchart of an example method for video encoding and decoding is shown.

[0047] Figure 36 This is a block diagram of an example hardware platform used to implement the visual media decoding or visual media encoding techniques described in this document.

[0048] Figure 37 This is a block diagram of an example video processing system that can implement the disclosed technology. Detailed Implementation

[0049] Due to the increasing demand for higher resolution video, video encoding and decoding methods and technologies are ubiquitous in modern technology. Video codecs typically consist of electronic circuitry or software that compresses or decompresses digital video and are constantly being improved to provide higher encoding and decoding efficiency. A video codec converts uncompressed video into a compressed format and vice versa. There is a complex relationship between video quality, the amount of data used to represent the video (determined by the bit rate), the complexity of the encoding and decoding algorithms, sensitivity to data loss and errors, ease of editing, random access, and end-to-end latency (delay). Compression formats typically conform to standard video compression specifications, such as the High Efficiency Video Codec (HEVC) standard (also known as H.265 or MPEG-H Part 2), a universally applicable video codec standard yet to be finalized, or other current and / or future video codec standards.

[0050] Embodiments of the disclosed techniques can be applied to existing video codec standards (e.g., HEVC, H.265) and future standards to improve compression performance. Section headings are used in this document to improve the readability of the description and do not in any way limit the discussion or embodiments (and / or implementations) to the relevant sections only.

[0051] 1. Example of inter-frame prediction in HEVC / H.265

[0052] Over the years, video codec standards have improved significantly and now offer, in part, high encoding and decoding efficiency and support for higher resolutions. Latest standards such as HEVC and H.265 are based on a hybrid video codec architecture, which uses temporal prediction plus transform encoding and decoding.

[0053] 1.1 Example of inter-frame prediction mode

[0054] Each inter-frame prediction PU (prediction unit) has motion parameters for one or two lists of reference images. In some embodiments, the motion parameters include motion vectors and reference image indices. In other embodiments, inter_pred_idc can also be used to signal the use of one of the two reference image lists. In still other embodiments, the motion vectors can be explicitly encoded as deltas relative to the predictor.

[0055] When using skip mode to encode and decode a CU, a PU is associated with that CU, and there are no significant residual coefficients, no encoding / decoding motion vector increments, or reference picture indices. A merge mode is specified, thereby obtaining motion parameters for the current PU from neighboring PUs—including spatial and temporal candidates. The merge mode can be applied to any PU for inter-frame prediction, not just skip mode. An alternative to the merge mode is explicit transmission of motion parameters, where the motion vectors, the corresponding reference picture index for each reference picture list, and the reference picture list usage are explicitly signaled to each PU.

[0056] When signaling indicates that one of two lists of reference images should be used, the PU is generated from a sample block. This is called "one-way prediction". One-way prediction can be used for both P and B strips.

[0057] When signaling indicates that two lists of reference images should be used, the PU is generated from two sample blocks. This is called "bidirectional prediction". Bidirectional prediction can only be used for B-bands.

[0058] 1.1.1 Implementation Examples for Constructing Candidate Merge Patterns

[0059] When predicting a PU using the Merge pattern, indices pointing to entries in the Merge Candidates list are parsed from the bitstream, and these indices are used to retrieve motion information. The construction of this list can be summarized in the following steps:

[0060] Step 1: Initial Candidate Derivation

[0061] Step 1.1: Spatial Candidate Derivation

[0062] Step 1.2: Redundancy check of airspace candidates

[0063] Step 1.3: Time-domain candidate derivation

[0064] Step 2: Add candidate insertions

[0065] Step 2.1: Create bidirectional prediction candidates

[0066] Step 2.2: Insert zero-motion candidates

[0067] Figure 1An example of constructing the Merge candidate list based on the steps summarized above is shown. For spatial Merge candidate derivation, up to four Merge candidates are selected from candidates located at five different positions. For temporal Merge candidate derivation, up to one Merge candidate is selected from two candidates. Since the number of candidates per PU is assumed to be constant at the decoder, additional candidates are generated when the number of candidates does not reach the maximum number of Merge candidates (MaxNumMergeCand) signaled in the stripe header. Because the number of candidates is constant, a binary unary truncation (TU) is used to encode the index of the best Merge candidate. If the CU size is equal to 8, all PUs of the current CU share a single Merge candidate list, which is identical to the Merge candidate list of a 2N×2N prediction unit.

[0068] 1.1.2 Constructing Spatial Merge Candidates

[0069] In the derivation of the spatial Merge candidate, located in Figure 2 Up to four merge candidates are selected from the candidates at the positions depicted. The derivation order is A1, B1, B0, A0, and B2. Position B2 is considered only if any PU at positions A1, B1, B0, or A0 is unavailable (e.g., because the PU belongs to another slice or tile) or during intra-frame encoding / decoding. After adding the candidate at position A1, a redundancy check is performed on the remaining candidates to ensure that candidates with the same motion information are excluded from the list, thereby improving encoding / decoding efficiency.

[0070] To reduce computational complexity, not all possible candidate pairs were considered in the aforementioned redundancy check. Instead, only pairs with... Figure 3 The arrows in the list link pairs, and a candidate is added to the list only if the corresponding candidate used for redundancy checking has different motion information. Another source of duplicate motion information is a "second PU" associated with a segmentation other than 2Nx2N. As an example, Figure 4A and 4B The second prediction units (PUs) are depicted for the N×2N and 2N×N cases, respectively. When the current PU is segmented into N×2N, the candidate at position A1 is not considered for list construction. In some embodiments, adding this candidate can result in two prediction units with the same motion information, which is redundant for an encoding / decoding unit with only one PU. Similarly, when the current PU is segmented into 2N×N, position B1 is not considered.

[0071] 1.1.3 Constructing Temporal Merge Candidates

[0072] In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal merge candidate, the scaled motion vector is derived based on the co-located PU, which belongs to the image with a POC difference relative to the current image within a given list of reference images. The list of reference images used for the derivation of the co-located PU is explicitly signaled in the strip header.

[0073] Figure 5 An example of the derivation of the motion vector for scaling temporal merge candidates (e.g., dashed lines) is shown. This scaling motion vector is scaled from the motion vector of the co-located PU using POC distances tb and td, where tb is defined as the POC difference between the current image and the reference image, and td is defined as the POC difference between the co-located image and the reference image. The reference image index of the temporal merge candidate is set to zero. For the B-strip, two motion vectors are obtained and combined to produce bidirectional predicted merge candidates, one for reference image list 0 and the other for reference image list 1.

[0074] like Figure 6 As shown, among the corresponding PUs (Y) belonging to the reference frame, a position for the temporal candidate is selected between candidates C0 and C1. If the PU at position C0 is unavailable, intra-frame encoded, or outside the current CTU, position C1 is used. Otherwise, position C0 is used in the derivation of the temporal merge candidate.

[0075] 1.1.4 Constructing Merge Candidates for Additional Types

[0076] In addition to the spatial-temporal merge candidates, two additional types of merge candidates exist: combined bidirectional prediction merge candidates and zero merge candidates. Combined bidirectional prediction merge candidates are generated by utilizing spatial-temporal merge candidates. These combined bidirectional prediction merge candidates are only used for B-strips. Combined bidirectional prediction candidates are generated by combining the motion parameters of the first reference image list of the initial candidate with the motion parameters of the second reference image list of another candidate. If these two tuples provide different motion hypotheses, they form a new bidirectional prediction candidate.

[0077] Figure 7An example of this process is shown, in which two candidates in the original list (710, on the left) with mvL0 and refIdxL0 or mvL1 and refIdxL1 are used to create a combined bidirectional prediction Merge candidate, which is then added to the final list (720, on the right).

[0078] Zero-motion candidates are inserted to populate the remaining entries in the Merge candidate list, thus reaching the MaxNumMergeCand capacity. These candidates have zero spatial displacements and reference image indices that start at zero and increase whenever a new zero-motion candidate is added to the list. The number of reference frames used by these candidates is 1 and 2, respectively, for unidirectional and bidirectional prediction. In some embodiments, redundancy checks are not performed on these candidates.

[0079] 1.1.5 Example of motion estimation region for parallel processing

[0080] To accelerate the encoding process, motion estimation can be performed in parallel, thereby simultaneously deriving the motion vectors of all prediction units within a given region. The derivation of spatially adjacent merge candidates can interfere with parallel processing, as a prediction unit cannot derive its motion parameters from neighboring PUs until its associated motion estimation is complete. To mitigate the trade-off between encoding / decoding efficiency and processing latency, a Motion Estimation Region (MER) can be defined. The size of the MER can be signaled in the Picture Parameter Set (PPS) using the "log2_parallel_merge_level_minus2" syntax element. When a MER is defined, merge candidates falling into the same region are marked as unavailable and therefore not considered in list construction.

[0081] 1.2 Examples of Advanced Motion Vector Prediction (AMVP)

[0082] AMVP utilizes the spatial-temporal correlation between motion vectors and adjacent PUs, which is used for the explicit transmission of motion parameters. The motion vector candidate list is constructed by first checking the availability of temporally adjacent PU positions to the left and above, removing redundant candidates, and adding zero vectors to make the candidate list a constant length. The encoder can then select the best predictor from the candidate list and transmit the corresponding index indicating the selected candidate. Similar to the Merge index signaling, a unary truncation is used to encode the index of the best motion vector candidate. The maximum value to be encoded in this case is 2 (see [link to documentation]). Figure 8 The following sections provide details of the derivation process for the motion vector prediction candidates.

[0083] 1.2.1 Example of constructing motion vector prediction candidates

[0084] Figure 8 The derivation process for motion vector prediction candidates is summarized and can be implemented for each list of reference images with refidx as input.

[0085] In motion vector prediction, two types of motion vector candidates are considered: spatial domain motion vector candidates and temporal domain motion vector candidates. For example... Figure 2 As previously shown, the derivation of spatial motion vector candidates ultimately derives two motion vector candidates based on the motion vectors of each PU located at five different positions.

[0086] For temporal motion vector candidate derivation, one motion vector candidate is selected from two candidates derived based on two different co-locations. After creating a first list of spatial-temporal candidates, duplicate motion vector candidates are removed from the list. If the number of potential candidates is greater than 2, motion vector candidates whose reference image index is greater than 1 in the associated reference image list are removed from the list. If the number of spatial-temporal motion vector candidates is less than 2, additional zero motion vector candidates are added to the list.

[0087] 1.2.2 Constructing Candidate Spatial Motion Vectors

[0088] In the derivation of the spatial motion vector candidates, at most two candidates are considered from five potential candidates, and these five potential candidates come from locations such as... Figure 2 The previously shown positions of the PU, these positions are related to movement

[0089] The positions of the merges are the same. The derivation order to the left of the current PU is defined as A0, A1, and scaled A0, scaled A1. The derivation order to the top of the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. Therefore, for each side, there are four possible motion vector candidates, two of which do not require spatial scaling, and two of which do. The four different cases are summarized below.

[0090] • No spatial scaling

[0091] (1) Same list of reference images and same reference image indices (same POC) (2) Different list of reference images, but the same reference images (same POC)

[0092] • Spatial scaling

[0093] (3) Same list of reference images, but different reference images (different POCs) (4) Different list of reference images, and different reference images (different POCs)

[0094] First, we check for cases without spatial scaling, then we check for cases where spatial scaling is allowed. Regardless of the reference image list, spatial scaling is considered when the POC differs between the reference image of an adjacent PU and the reference image of the current PU. If all PUs in the left-hand candidate list are unavailable or intra-frame encoded / decoded, scaling of the upper motion vector is allowed to aid in the parallel derivation of the left-hand and upper MV candidates. Otherwise, spatial scaling is not allowed for the upper motion vector.

[0095] like Figure 9 As shown in the example, for spatial scaling, the motion vectors of adjacent PUs are scaled in a manner similar to temporal scaling. One difference is that the list of reference images and indices for the current PU are given as input; the actual scaling process is the same as the temporal scaling process.

[0096] 1.2.3 Constructing Temporal Motion Vector Candidates

[0097] Except for the derivation of the reference image index, all the procedures for the derivation of the temporal merge candidate are the same as those for the derivation of the spatial motion vector candidate (e.g., ...). Figure 6 (As shown in the example). In some embodiments, reference image index signaling is notified to the decoder.

[0098] 2. Examples of inter-frame prediction methods in the Joint Exploration Model (JEM)

[0099] In some embodiments, reference software called the Joint Exploration Model (JEM) is used to explore future video coding and decoding technologies. In the JEM, sub-block-based predictions are employed across several coding and decoding tools, such as affine prediction, optional temporal motion vector prediction (ATMVP), spatial-temporal motion vector prediction (STMVP), bidirectional optical flow (BIO), frame rate upconversion (FRUC), locally adaptive motion vector resolution (LAMVR), overlapping block motion compensation (OBMC), local luma compensation (LIC), and decoder-side motion vector refinement (DMVR).

[0100] 2.1 Example of motion vector prediction based on sub-CU

[0101] In a JEM with a quadtree plus binary tree (QTBT), each CU can have a set of at most one set of motion parameters for each prediction direction. In some embodiments, two sub-CU level motion vector prediction methods are considered in the encoder by dividing the large CU into sub-CUs and deriving the motion information of all sub-CUs of the large CU. The optional temporal motion vector prediction (ATMVP) method allows each CU to obtain multiple sets of motion information from multiple blocks smaller than the current CU in the co-located reference image. In the spatial-temporal motion vector prediction (STMVP) method, the motion vectors of the sub-CUs are recursively derived using temporal motion vector prediction and spatially adjacent motion vectors. In some embodiments, and to preserve a more accurate motion field for sub-CU motion prediction, motion compression of the reference frame can be disabled.

[0102] 2.1.1 Example of Optional Temporal Motion Vector Prediction (ATMVP)

[0103] In the ATMVP method (also known as Sub-Block Temporal Motion Vector Prediction (sbTMVP)), the Temporal Motion Vector Prediction (TMVP) method is modified by retrieving multiple sets of motion information (including motion vectors and reference indices) from blocks smaller than the current CU.

[0104] Figure 10 An example of the ATMVP motion prediction process for CU 1000 is shown. The ATMVP method predicts the motion vectors of sub-CUs 1001 within CU 1000 in two steps. The first step is to identify the corresponding block 1051 in reference image 1050 using temporal vectors. Reference image 1050 is also called the motion source image. The second step is to divide the current CU 1000 into sub-CUs 1001 and obtain the motion vector and reference index of each sub-CU from the block corresponding to each sub-CU.

[0105] In the first step, reference image 1050 and the corresponding block are determined by the motion information of spatially adjacent blocks of the current CU 1000. To avoid repeated scanning of adjacent blocks, the first Merge candidate in the Merge candidate list of the current CU 1000 is used. The first available motion vector and its associated reference index are set to the index of the temporal vector and the motion source image. In this way, the corresponding block can be identified more accurately than TMVP, where the corresponding block (sometimes called the co-occurring block) is always located in the lower right or center position relative to the current CU.

[0106] In the second step, the corresponding block of sub-CU 1051 is identified from the temporal vector in the motion source image 1050 by adding a temporal vector to the coordinates of the current CU. For each sub-CU, the motion information of its corresponding block (e.g., the smallest motion grid covering the center sample) is used to derive the motion information of the sub-CU. After identifying the motion information of the corresponding N×N block, it is converted into the motion vector and reference index of the current sub-CU in the same manner as the TMVP of HEVC, in which motion scaling and other processes are applied. For example, the decoder checks whether a low-latency condition is met (e.g., the POC of all reference images of the current image is less than the POC of the current image) and may use the motion vector MVx (e.g., the motion vector corresponding to the reference image list X) to predict the motion vector MVy of each sub-CU (e.g., where X equals 0 or 1 and Y equals 1-X).

[0107] 2.1.2 Example of Spatial-Time Motion Vector Prediction (STMVP)

[0108] In the STMVP method, the motion vectors of the sub-CUs are recursively derived according to the raster scan order. Figure 11 An example of a CU with four sub-blocks and adjacent blocks is shown. Consider an 8×8 CU 1100 comprising four 4×4 sub-CUs A(1101), B(1102), C(1103), and D(1104). The adjacent 4×4 blocks in the current frame are labeled as b(1111), b(1112), c(1113), and d(1114).

[0109] Motion derivation of sub-CU A begins by identifying its two spatial neighbors. The first neighbor is the N×N block (block c 1113) above sub-CU A 1101. If block c (1113) is unavailable or intra-frame encoded, the other N×N blocks above sub-CU A (1101) are checked (from left to right, starting with block c 1113). The second neighbor is the block to the left of sub-CU A 1101 (block b 1112). If block b (1112) is unavailable or intra-frame encoded, the other blocks to the left of sub-CU A 1101 are checked (from top to bottom, starting with block b 1112). Motion information obtained from neighboring blocks for each list is scaled to the first reference frame for the given list. Next, the temporal motion vector prediction (TMVP) of sub-block A 1101 is derived by following the same procedure as the TMVP derivation specified in HEVC. Motion information for the co-located block at block D 1104 is retrieved and scaled accordingly. Finally, after extracting and scaling the motion information, all available motion vectors are averaged separately for each reference list. The averaged motion vector is specified as the motion vector for the current sub-CU.

[0110] 2.1.3 Example of Sub-CU Motion Prediction Mode Signaling Notification

[0111] In some embodiments, the sub-CU mode is enabled as an additional Merge candidate, and no additional syntax elements are required to signal the mode. Two additional Merge candidates are added to merge the candidate lists for each CU to represent the ATMVP and STMVP modes. In other embodiments, up to seven Merge candidates can be used if the sequence parameter set indicates that ATMVP and STMVP are enabled. The encoding logic for the additional Merge candidates is the same as the encoding / decoding logic for the Merge candidates in the HM, meaning that for each CU in a P or B stripe, the two additional Merge candidates may require two more RD checks. In some embodiments, such as JEM, all bits of the Merge index are context-coded using CABAC (Context-Based Adaptive Binary Arithmetic Codec). In other embodiments, such as HEVC, only the first bit (bin) is context-coded, and the remaining bits are context-bypass encoded / decoded.

[0112] 2.2 Example of Adaptive Motion Vector Difference Resolution

[0113] In some embodiments, when the use_integer_mv_flag in the stripe header is equal to 0, the motion vector difference (MVD) is signaled in units of quarter-luminance samples (between the motion vector of the PU and the predicted motion vector). Local Adaptive Motion Vector Resolution (LAMVR) is introduced in JEM. In JEM, MVD can be encoded and decoded in units of quarter-luminance samples, integer luminance samples, or four luminance samples. MVD resolution is controlled at the codec unit (CU) level, and the MVD resolution flag is conditionally signaled to each CU having at least one non-zero MVD component.

[0114] For a CU with at least one non-zero MVD component, signaling informs a first flag to indicate whether quarter-sample MV precision is used in the CU. When the first flag (equal to 1) indicates that quarter-sample MV precision is not used, signaling informs another flag to indicate whether integer MV precision or four-sample MV precision is used.

[0115] When the first MVD resolution flag of the CU is zero, or when no encoding / decoding is performed for the CU (meaning all MVDs in the CU are zero), a quarter-luminance sample MV resolution is used for that CU. When the CU uses integer luminance sample MV precision or four luminance sample MV precision, the MVPs in the CU's AMVP candidate list are rounded to the corresponding precision.

[0116] In the encoder, CU-level RD checking is used to determine which MVD resolution will be used for a CU. In other words, for each MVD resolution, the CU-level RD checking is performed three times. To speed up the encoder, the following coding scheme is applied in JEM:

[0117] · During the RD checking of a CU with a normal quarter-luminance sample MVD resolution, the motion information (integer luminance sample precision) of the current CU is stored. For the same CU with integer luminance samples and a 4-luminance sample MVD resolution, the stored motion information (after rounding) is used as the starting point for further small-range motion vector refinement during the RD checking, so that the time-consuming motion estimation process is not repeated three times.

[0118] · Conditionally call the RD checking of a CU with a 4-luminance sample MVD resolution. For a CU, when the RD cost of the integer luminance sample MVD resolution is much greater than the RD cost of the quarter-luminance sample MVD resolution, skip the RD checking of the 4-luminance sample MVD resolution of this CU.

[0119] The encoding process is shown in Figure 12 First, test the 1 / 4-pixel MV and calculate the RD cost and represent the RD cost as RDCost0, then test the integer MV and represent the RD cost as RDCost1. If RDCost1 < th * RDCost0 (where th is a positive threshold), then test the 4-pixel MV; otherwise, skip the 4-pixel MV. Basically, when checking the integer or 4-pixel MV, the motion information and RD cost of the 1 / 4-pixel MV are known, etc., which can be reused to accelerate the encoding process of the integer or 4-pixel MV.

[0120] 2.3 Examples of Higher Motion Vector Storage Precision

[0121] In HEVC, the motion vector precision is a quarter pixel (pel) (quarter-luminance samples and eighth-chrominance samples for 4:2:0 video). In JEM, the internal motion vector storage and

[0122] the precision of Merge candidates are increased to 1 / 16 pixel. The higher motion vector precision (1 / 16 pixel) is used for motion-compensated inter prediction of CUs encoded / decoded in skip mode / Merge mode. For CUs encoded / decoded using the normal AMVP mode, integer pixels or quarter pixels of motion are used.

[0123] An SHVC upsampling interpolation filter with the same filter length and normalization factor as the HEVC motion-compensated interpolation filter is used as the motion-compensated interpolation filter for the additional fractional pixel positions. In JEM, the chroma component motion vector precision is 1 / 32 sample, and the additional interpolation filter for the 1 / 32 pixel fractional position is derived by averaging the filters at two adjacent 1 / 16 pixel fractional positions.

[0124] 2.4 Examples of Triangular Prediction Unit Patterns

[0125] The concept of the triangular prediction unit pattern is to introduce a new triangular segmentation for motion compensation prediction. For example... Figure 13A and 13B As shown, the triangular prediction unit mode divides the CU into two triangular prediction units along either the diagonal or anti-diagonal direction. Each triangular prediction unit in the CU performs inter-frame prediction using its own unidirectional prediction motion vector derived from the unidirectional prediction candidate list and a reference frame index. After prediction for the triangular prediction unit, adaptive weighting is applied to the diagonal edges. Then, transform and quantization are applied to the entire CU. It should be noted that this mode is only applicable to skip and merge modes.

[0126] 2.4.1 One-way prediction candidate list

[0127] The unidirectional prediction candidate list includes five unidirectional prediction motion vector candidates. For example... Figure 14 As shown, it is derived from seven neighboring blocks, which include five spatially adjacent blocks (1 to 5) and two temporally co-located blocks (6 to 7). Motion vectors from the seven neighboring blocks are collected and placed into a unidirectional prediction candidate list in the order of unidirectional predicted motion vector, L0 motion vector of bidirectional predicted motion vector, L1 motion vector of bidirectional predicted motion vector, and the average motion vector of the L0 and L1 motion vectors of bidirectional predicted motion vector. If the number of candidates is less than five, a zero motion vector is added to the list.

[0128] 2.4.2 Adaptive Weighted Processing

[0129] After predicting for each triangular prediction unit, an adaptive weighting process is applied to the diagonal edges between two triangular prediction units to derive the final prediction for the entire CU. The two weighting factor groups are defined as follows:

[0130] First weighting factor group: {7 / 8, 6 / 8, 4 / 8, 2 / 8, 1 / 8} and {7 / 8, 4 / 8, 1 / 8} are used for luminance and chrominance samples, respectively;

[0131] The second weighting factor group: {7 / 8, 6 / 8, 5 / 8, 4 / 8, 3 / 8, 2 / 8, 1 / 8} and {6 / 8, 4 / 8, 2 / 8} are used for luminance and chromaticity samples, respectively.

[0132] A weighting factor group is selected based on a comparison of the motion vectors of two triangulated prediction units. The second weighting factor group is used when the reference images of the two triangulated prediction units are different from each other or when the difference in their motion vectors is greater than 16 pixels. Otherwise, the first weighting factor group is used. Figure 15 An example of this adaptive weighting process is shown.

[0133] 2.4.3 Motion Vector Storage

[0134] The motion vector of the triangular prediction unit ( Figure 16 Mv1 and Mv2 in the CU are stored in a 4×4 grid. For each 4×4 grid, whether to store a unidirectional or bidirectional predicted motion vector depends on the position of the 4×4 grid within the CU. Figure 16 As shown, the unidirectional predicted motion vector (Mv1 or Mv2) is stored in a 4×4 grid located in the unweighted region. On the other hand, the bidirectional predicted motion vector is stored in a 4×4 grid located in the weighted region. The bidirectional predicted motion vector is derived from Mv1 and Mv2 according to the following rules:

[0135] 1) When Mv1 and Mv2 have motion vectors from different directions (L0 or L1), Mv1 and Mv2 can be simply combined to form a bidirectional predicted motion vector.

[0136] 2) When Mv1 and Mv2 both originate from the same L0 (or L1) direction,

[0137] 2a) If the reference image for Mv2 is the same as an image in the L1 (or L0) reference image list, then scale Mv2 to that image. Combine Mv1 and the scaled Mv2 to form a bidirectional predicted motion vector.

[0138] 2b) If the reference image for Mv1 is the same as an image in the L1 (or L0) reference image list, then scale Mv1 to that image. Combine the scaled Mv1 and Mv2 to form a bidirectional predicted motion vector.

[0139] 2c) Otherwise, only store Mv1 for the weighted region.

[0140] 2.5 Example of Overlapping Block Motion Compensation (OBMC)

[0141] In JEM, OBMC can be toggled at the CU level using syntax elements. When OBMC is used in JEM, it is performed on all motion compensation (MC) block boundaries except for the right and bottom boundaries of the CU. Furthermore, it is applied to the luma and chroma components. In JEM, MC blocks correspond to codec blocks. When encoding and decoding a CU using subCU modes (including subCU Merge, affine, and FRUC modes), each sub-block of the CU is an MC block. To uniformly handle CU boundaries, when the sub-block size is set to 4x4, OBMC is performed at the sub-block level on all MC block boundaries, such as... Figure 17A and Figure 17B As shown.

[0142] Figure 17A The sub-block at the CU / PU boundary is shown, and the shaded sub-block is where the OBMC is applied. Similarly, Figure 17B The sub-PU in the ATMVP pattern is shown.

[0143] When OBMC is applied to the current sub-block, in addition to the current MV, the vectors of the four adjacent sub-blocks (if available and not exactly the same as the current motion vector) are also used to derive the prediction block for the current sub-block. These multiple prediction blocks based on multiple motion vectors are combined to generate the final prediction signal for the current sub-block.

[0144] The predicted block based on the motion vectors of neighboring sub-blocks is represented as PN, where N indicates the indices of the neighboring above, below, left, and right sub-blocks, and the predicted block based on the motion vector of the current sub-block is represented as PC. OBMC is not performed from PN when PN is based on the motion information of neighboring sub-blocks and that motion information is the same as that of the current sub-block. Otherwise, each sample from PN is added to the same sample in PC, i.e., four rows / columns of PN are added to PC. Weighting factors {1 / 4, 1 / 8, 1 / 16, 1 / 32} are used for PN, and weighting factors {3 / 4, 7 / 8, 15 / 16, 31 / 32} are used for PC. An exception is for small MC blocks (i.e., when the height or width of the codec block is equal to 4 or the CU uses sub-CU mode encoding / decoding), where only two rows / columns of PN are added to PC. In this case, weighting factors {1 / 4, 1 / 8} are used for PN, and weighting factors {3 / 4, 7 / 8} are used for PC. For a PN generated based on the motion vectors of vertically (horizontally) adjacent sub-blocks, samples from the same row (column) of the PN are added to a PC with the same weighting factor.

[0145] In JEM, for CUs with a size of 256 lumen samples or less, signaling informs the CU level flag to indicate whether OBMC should be applied to the current CU. For CUs with a size greater than 256 lumen samples or not using AMVP mode encoding / decoding, OBMC is applied by default. At the encoder, when OBMC is applied to the CU, its effects are considered during the motion estimation phase. The predicted signal formed by OBMC using motion information from the top and left adjacent blocks is used to compensate for the top and left boundaries of the original signal of the current CU, and then the normal motion estimation process is applied.

[0146] 2.6 Example of Local Luminance Compensation (LIC)

[0147] LIC is based on a linear model for brightness variations, using a scaling factor a and an offset b. Furthermore, the codec unit (CU) adaptively enables or disables LIC for each inter-frame mode encoding / decoding.

[0148] When LIC is applied to CU, the least squares error method is used to derive parameters a and b by using the neighboring samples of the current CU and their corresponding reference samples. Figure 18 An example of neighboring samples used to derive the parameters of the IC algorithm is shown. Specifically, and as... Figure 18 As shown, neighboring samples (2:1 subsampling) of the CU and corresponding samples in the reference image (identified by motion information of the current CU or subCU) were used. IC parameters were derived and applied to each prediction direction respectively.

[0149] When encoding and decoding a CU using Merge mode, the LIC flag is copied from the adjacent block in a manner similar to motion information copying in Merge mode; otherwise, the CU signaling is notified of the LIC flag to indicate whether LIC is applicable.

[0150] When LIC is enabled for an image, an additional CU-level RD check is required to determine whether LIC should be applied to the CU. When LIC is enabled for the CU, mean-removed sum of absolute difference (MR-SAD) and mean-removed sum of absolute Hadamard-transformed difference (MR-SATD) are used instead of SAD and SATD for integer pixel motion search and fractional pixel motion search, respectively.

[0151] To reduce coding complexity, the following coding scheme is applied in JEM:

[0152] --Disable LIC for the entire image when there is no significant brightness change between the current image and its reference images. To identify this situation, at the encoder, a histogram of the current image and each of its reference images is calculated. If the histogram difference between the current image and each of its reference images is less than a given threshold, LIC is disabled for the current image; otherwise, LIC is enabled for the current image.

[0153] 2.7 Examples of Affine Motion Compensation Prediction

[0154] In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). However, cameras and objects can exhibit a variety of motions, such as zooming in / out, rotation, perspective motion, and other irregular motions. On the other hand, JEM applies a simplified affine transformation motion compensation prediction. Figure 19A An example of the affine motion field of block 1900, described by two control point motion vectors V0 and V1, is shown. The motion vector field (MVF) of block 1900 can be described by the following equation:

[0155]

[0156] like Figure 19A As shown, (v 0x ,v 0y ) is the motion vector of the top left control point, and (v 1x ,v 1y () is the motion vector of the upper right control point. Similarly, for a 6-parameter affine model, the MVF of a block can be described as:

[0157]

[0158] Here, as Figure 19B As shown, (v 0x ,v 0y ) is the motion vector of the top left control point, and (v 1x ,v 1y ) is the motion vector of the upper right control point. And (v 2x ,v 2y (x, y) represents the motion vector of the lower left control point, and (x, y) represents the coordinates of the representative point relative to the upper left sample within the current block. In VTM, the representative point is defined as the center position of the sub-block. For example, when the coordinates of the upper left corner of the sub-block relative to the upper left sample within the current block are (xs, ys), the coordinates of the representative point are defined as (xs+2, ys+2).

[0159] To simplify motion compensation prediction, a sub-block-based affine transformation prediction can be applied. The sub-block size M×N is derived as follows:

[0160] Here, MvPre is the fractional precision of the motion vector (e.g., 1 / 16 in JEM). (v 2x ,v 2y ) is the motion vector of the lower left control point, calculated according to equation (1). If necessary, M and N can be adjusted downwards to make them the divisors of w and h, respectively.

[0161] Figure 20 An example of the affine MVF for each sub-block of block 2000 is shown. To derive the motion vector for each M×N sub-block, the motion vector of the center sample of each sub-block can be calculated according to equation (1a) and rounded to the fractional precision of the motion vector (e.g., 1 / 16 in JEM). A motion-compensated interpolation filter can then be applied to generate a prediction for each sub-block using the derived motion vector. After MCP, the high-precision motion vector for each sub-block is rounded and saved with the same precision as the normal motion vector.

[0162] 2.7.1 Example of AF_INTER mode

[0163] In JEM, there are two affine motion modes: AF_INTER mode and AF_MERGE mode. AF_INTER mode can be applied to CUs with both width and height greater than 8. Signaling in the bitstream informs the affine flag at the CU level to indicate whether AF_INTER mode is used. In AF_INTER mode, adjacent blocks are used to construct motion vector pairs {(v0,v1)|v0={v...}. A ,v B ,v c},v1={v D ,v E The candidate list of}}.

[0164] Figure 21 An example of motion vector prediction (MVP) for block 1700 in AF_INTER mode is shown. Figure 21As shown, v0 is selected from the motion vectors of sub-blocks A, B, or C. Motion vectors from neighboring blocks can be scaled according to a reference list. Motion vectors can also be scaled based on the relationship between the Picture Order Count (POC) of the references of neighboring blocks, the POC of the current CU's reference, and the POC of the current CU. The method for selecting v1 from neighboring sub-blocks D and E is similar. If the number of candidates in the candidate list is less than two, the list is populated by motion vector pairs constructed by repeating each AMVP candidate. When the candidate list is greater than two, candidates are first categorized based on adjacent motion vectors (e.g., based on the similarity of two motion vectors in a candidate pair). In some implementations, the first two candidates are retained. In some implementations, a rate distortion (RD) cost check is used to determine which motion vector pair is selected as the Control Point Motion Vector Prediction (CPMVP) for the current CU. The index indicating the position of the CPMVP in the candidate list can be signaled in the bitstream. After determining the CPMVP of the current affine CU, affine motion estimation is applied, and the Control Point Motion Vector (CPMV) is found. The difference between the CPMV and the CPMVP is then signaled in the bitstream.

[0165] In AF_INTER mode, when using the 4 / 6 parameter affine mode, 2 / 3 control points are required, and therefore 2 / 3 MVD encoding / decoding is needed for these control points, such as... Figure 22A and 22B As shown. In existing implementations, MV can be derived as follows, for example, it predicts mvd1 and mvd2 from mvd0.

[0166]

[0167] In this article, mvd i mv1 and mv1 are the predicted motion vector, motion vector difference, and motion vector of the left top pixel (i=0), right top pixel (i=1), or left bottom pixel (i=2), respectively. Figure 22B As shown. In some embodiments, the sum of two motion vectors (e.g., mvA(xA, yA) and mvB(xB, yB)) is equal to the sum of the two individual components. For example, newMV = mvA + mvB means that the two components of newMV are set to (xA + xB) and (yA + yB), respectively.

[0168] 2.7.2 Example of the Fast ME Algorithm in AF_INTER Mode

[0169] In some embodiments of affine patterns, it is necessary to jointly determine the MVs of 2 or 3 control points. Directly and jointly searching multiple MVs is computationally complex. In one example, a fast affine ME algorithm is proposed and applied to VTM / BMS.

[0170] For example, the fast affine ME algorithm is described for 4-parameter affine models, and the idea can be extended to 6-parameter affine models:

[0171]

[0172] Replacing (a-1) with a' allows the motion vector to be rewritten as:

[0173]

[0174] If we assume that the motion vectors of the two control points (0, 0) and (0, w) are known, then the affine parameters can be derived from equation (5) as follows:

[0175]

[0176] The motion vector can be rewritten in vector form as:

[0177]

[0178] Here, P = (x, y) is the pixel position.

[0179] And equation (8)

[0180]

[0181] In some embodiments, and at the encoder, the MVD of AF_INTER can be iteratively derived. Let MVi(P) represent the MV derived in the i-th iteration at position P, and let dMV... C i This represents the increment for the MVC update in the i-th iteration. Then, in the (i+1)-th iteration,

[0182]

[0183] Pic ref Show as reference image and Pic cur This is represented as the current image and Q = P + MV. i (P). If MSE is used as the matching criterion, the function to be minimized can be written as:

[0184]

[0185] If we assume If it is small enough, it can be based on a first-order Taylor expansion. Rewrite as an approximation, such as:

[0186]

[0187] here, If E is used i+1 (P)=Pic cur (P)-Pic ref (Q), then:

[0188]

[0189] the term This can be achieved by setting the derivative of the error function to zero, and then according to... The incremental MV of control points (0, 0) and (0, w) is derived as follows:

[0190]

[0191] In some embodiments, the MVD derivation process can be iterated n times, and the final MVD can be calculated as follows:

[0192]

[0193] In the aforementioned implementation, the incremental MV from the control point (0, 0) represented by mvd0 predicts the incremental MV of the prediction control point (0, w) represented by mvd1, resulting in... It is encoded only for mvd1.

[0194] 2.7.3 Example of AF_Merge pattern

[0195] When the CU is applied in AF_Merge mode, it retrieves the first block encoded and decoded in affine mode from the valid adjacent reconstructed blocks. Figure 23A This shows an example of the current selection order of candidate blocks for the CU 2300. (Example:) Figure 23A As shown, the selection order can be from the left (2301), top (2302), top right (2303), bottom left (2304) to top left (2305) of the current CU 2300. Figure 23B Another example of a current CU 2300 candidate block in AF_Merge mode is shown. Figure 23B As shown, if the adjacent lower left block 2301 is encoded and decoded in affine mode, the motion vectors v2, v3, and v4 of the CU containing the sub-block 2301 at its upper left, upper right, and lower left corners are derived. The motion vector v0 of the upper left corner of the current CU 2300 is calculated based on v2, v3, and v4. The motion vector v1 of the upper right corner of the current CU can be calculated accordingly.

[0196] After calculating the CPMV v0 and v1 of the current CU based on the affine motion model in equation (1a), the MVF of the current CU can be generated. In order to identify whether the current CU uses AF_Merge mode encoding and decoding, when there is at least one adjacent block encoded and decoded in affine mode, the affine flag can be signaled in the bit stream.

[0197] In some embodiments, an affine Merge candidate list can be constructed using the following steps:

[0198] 1) Insertion of affine candidates for inheritance

[0199] Inherited affine candidates refer to candidates derived from the affine motion models of their effective neighboring affine codec blocks. On a common basis, such as... Figure 24 As shown, the scanning order of the candidate positions is: A1, B1, B0, A0, and B2.

[0200] After exporting the candidates, a full pruning is performed to check if the same candidate has already been inserted into the list. If the same candidate exists, the exported candidate is discarded.

[0201] 2) Insertion to construct affine candidates

[0202] If the number of candidates in the affine Merge candidate list is less than MaxNumAffineCand (set to 5 in this paper), then an affine candidate is constructed and inserted into the candidate list. Constructing an affine candidate means that the candidate is built by combining the adjacent motion information of each control point.

[0203] Firstly, from Figure 24 The motion information of the control points is derived from the specified spatial and temporal domains. CPk (k = 1, 2, 3, 4) represents the k-th control point. A0, A1, A2, B0, B1, B2, and B3 are used to predict the spatial position of CPk (k = 1, 2, 3); T is used to predict the temporal position of CP4.

[0204] The coordinates of CP1, CP2, CP3 and CP4 are (0, 0), (W, 0), (H, 0) and (W, H) respectively, where W and H are the width and height of the current block.

[0205] The motion information for each control point is obtained in the following priority order:

[0206] For CP1, the check priority is B2->B3->A2. If B2 is available, then B2 is used. Otherwise, if B2 is unavailable, then B3 is used. If neither B2 nor B3 is available, then A2 is used. If none of the three candidates are available, motion information for CP1 cannot be obtained.

[0207] For CP2, the check priority is B1->B0;

[0208] For CP3, the inspection priority is A1->A0;

[0209] For CP4, use T.

[0210] Secondly, combinations of control points are used to construct affine Merge candidates.

[0211] Motion information from three control points is needed to construct a 6-parameter affine candidate. These three control points can be selected from one of the following four combinations ({CP1, CP2, CP4}, {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4}). The combinations {CP1, CP2, CP3}, {CP2, CP3, CP4}, and {CP1, CP3, CP4} will be transformed into a 6-parameter motion model represented by the top-left, top-right, and bottom-left control points.

[0212] Motion information from two control points is needed to construct a 4-parameter affine candidate. These two control points can be selected from one of the following six combinations ({CP1,CP4},{CP2,CP3},{CP1,CP2},{CP2,CP4},{CP1,CP3},{CP3,CP4}). The combination {CP1,CP4},{CP2,CP3},{CP2,CP4},{CP1,CP3},{CP3,CP4} will be converted into a 4-parameter motion model represented by the top-left and top-right control points.

[0213] The combinations that construct affine candidates are inserted into the candidate list in the following order:

[0214] {CP1,CP2,CP3},{CP1,CP2,CP4},{CP1,CP3,CP4},{CP2,CP3,CP4},{CP1,CP2},{CP1,CP3},{CP2,CP3},{CP1,CP4},{CP2,CP4},{CP3,CP4}.

[0215] For a combined reference list X (X is 0 or 1), the reference index with the highest usage rate among the control points is selected as the reference index of list X, and the motion vector pointing to different reference images is scaled.

[0216] After exporting the candidates, a full pruning process is performed to check if the same candidate has already been inserted into the list. If the same candidate exists, the exported candidate is discarded.

[0217] 3) Fill in the zero motion vector

[0218] If the number of candidates in the affine Merge candidate list is less than 5, then insert a zero motion vector with a zero reference index into the candidate list until the list is full.

[0219] 2.8 Example of Pattern Matching Motion Vector Derivation (PMMVD)

[0220] PMMVD mode is a special merge mode based on the Frame Rate Upconversion (FRUC) method. Using this mode, block motion information is not communicated via signaling but is derived on the decoder side.

[0221] When the Merge flag of the CU is true, the FRUC flag can be signaled to that CU. When the FRUC flag is false, the Merge index can be signaled, and the normal Merge mode can be used. When the FRUC flag is true, an additional FRUC mode flag can be signaled to indicate which method (e.g., bilateral matching or template matching) will be used to derive the block's motion information.

[0222] On the encoder side, the decision on whether to use the FRUC Merge mode for the CU is based on the RD cost selection, as done for normal Merge candidates. For example, multiple matching modes for the CU (e.g., bilateral matching and template matching) are examined using RD cost selection. The matching mode resulting in the minimum cost is further compared with other CU modes. If the FRUC matching mode is the most efficient, the FRUC flag is set to true for the CU, and the relevant matching mode is used.

[0223] Typically, the motion derivation process in the FRUC Merge pattern involves two steps. First, a CU-level motion search is performed, followed by sub-CU-level motion refinement. At the CU level, initial motion vectors are derived for the entire CU based on bilateral matching or template matching. First, a candidate MV list is generated, and the candidate resulting in the minimum matching cost is selected as the starting point for further CU-level refinement. Then, a local search based on bilateral matching or template matching is performed around the starting point. The MV resulting in the minimum matching cost is taken as the MV for the entire CU. Subsequently, the motion information is further refined at the sub-CU level, with the derived CU motion vectors serving as the starting point.

[0224] For example, the following derivation process is performed for the motion information derivation of W×HCU. In the first stage, the MV of the entire W×HCU is derived. In the second stage, the CU is further divided into M×M sub-CUs. The value of M is calculated as in equation (3), where D is a predefined partitioning depth, which is set to 3 by default in JEM. Then the MV of each sub-CU is derived.

[0225]

[0226] Figure 25 An example of bilateral matching used in the Frame Rate Upconversion (FRUC) method is shown. Bilateral matching is used to derive motion information of the current CU (2000) by finding the closest match between two blocks along the motion trajectory of the current CU (2000) in two different reference images (2510, 2511). Under the assumption of a continuous motion trajectory, the motion vectors MV0 (2501) and MV1 (2502) pointing to the two reference blocks are proportional to the temporal distances between the current image and the two reference images—e.g., TD0 (2503) and TD1 (2504). In some embodiments, when the current image (2500) is in the temporal domain between the two reference images (2510, 2511) and the temporal distances from the current image to the two reference images are the same, bilateral matching becomes a mirror-based bidirectional MV.

[0227] Figure 26 An example of template matching used in the Frame Rate Upconversion (FRUC) method is shown. Template matching can be used to derive motion information of the current CU 2600 by finding the closest match between a template in the current image (e.g., the top adjacent block and / or left adjacent block of the current CU) and a block in the reference image 2610 (e.g., having the same size as the template). In addition to the FRUC Merge mode described above, template matching can also be applied to the AMVP mode. In both JEM and HEVC, AMVP has two candidates. Using the template matching method, new candidates can be derived. If the newly derived candidate by template matching is different from the first existing AMVP candidate, it is inserted at the beginning of the AMVP candidate list, and then the list size is set to 2 (e.g., by removing the second existing AMVP candidate). When applied to AMVP mode, only the CU-level search is applied.

[0228] The set of MV candidates at the CU level may include: (1) the original AMVP candidate if the current CU is in AMVP mode, (2) all Merge candidates, (3) several MVs in the interpolated MV field (described later), and (4) the top and left adjacent motion vectors.

[0229] When using bilateral matching, each valid MV of the Merge candidate can be used as input to generate MV pairs under the assumption of bilateral matching. For example, a valid MV of the Merge candidate is (MVa, ref) in reference list A. a Then, find the reference image ref for its paired bilateral MV in other reference list B. b , making ref a and ref b Located on a different side of the current image in the time domain. If referring to a ref like this in list B... bIf unavailable, then ref b Determined to be related to ref a Different references, and ref b The temporal distance to the current image is the smallest in list B. When determining the ref... b Then, based on the current image and ref a ref b The temporal distance between them is derived by scaling MVA to produce MVb.

[0230] In some implementations, four MVs from the interpolated MV field can also be added to the CU-level candidate list. More specifically, the interpolated MVs at positions (0,0), (W / 2,0), (0,H / 2), and (W / 2,H / 2) of the current CU are added. When FRUC is applied to the AMVP pattern, the original AMVP candidates are also added to the CU-level MV candidate set. In some implementations, at the CU level, 15 MVs for the AMVP CU and 13 MVs for the Merge CU can be added to the candidate list.

[0231] The candidate set of MVs at the sub-CU level includes:

[0232] (1) Search for the determined MV at the CU level.

[0233] (2) The adjacent MVs of the top, left, left top and right top,

[0234] (3) A scaled version of the juxtaposed music video from the reference image.

[0235] (4) One or more ATMVP candidates (e.g., up to four), and

[0236] (5) One or more STMVP candidates (e.g., up to four).

[0237] The scaled MV from the reference images is exported as follows. Reference images are iterated through in both lists. MVs at adjacent positions in the sub-CUs within the reference images are scaled to the reference of the starting CU-level MV. ATMVP and STMVP candidates can be the top four candidates. At the sub-CU level, one or more MVs (e.g., up to seventeen) are added to the candidate list.

[0238] Generation of interpolated MV fields

[0239] Before encoding and decoding the frames, an interpolated motion field is generated for the entire image based on a one-sided ME. The motion field can then be used later as a CU-level or sub-CU-level MV candidate.

[0240] In some embodiments, the motion field of each reference image in the two reference lists is traversed at a 4×4 block level. Figure 27An example of unidirectional motion estimation (ME)2700 in the FRUC method is shown. For each 4×4 block, if the motion associated with the block passes through 4×4 blocks in the current image and that block has not yet been assigned any interpolated motion, the motion of the reference block is scaled to the current image according to the temporal distances TD0 and TD1 (in the same way that the MV is scaled in TMVP in HEVC), and the scaled motion is assigned to the block in the current frame. If an unscaled MV is assigned to a 4×4 block, the block's motion is marked as unavailable in the interpolated motion field.

[0241] Interpolation and matching costs

[0242] When the motion vector points to the location of a fractional sample, motion-compensated interpolation is required. To reduce complexity, both bilateral matching and template matching can use bilinear interpolation instead of ordinary 8-tap HEVC interpolation.

[0243] The calculation of the matching cost varies slightly at different steps. When selecting candidates from the candidate set at the CU level, the matching cost can be the sum of absolute differences (SAD) of bilateral matching or template matching. After determining the starting MV, the matching cost C of bilateral matching for sub-CU level search is calculated as follows:

[0244]

[0245] Where w is a weighting factor. In some embodiments, w is empirically set to 4, and MV and MV s These indicate the current MV and the starting MV, respectively. SAD can still be used as the matching cost for template matching in sub-CU level searches.

[0246] In FRUC mode, the motion signature (MV) is derived solely using luma samples. The derived motion is then used for both luma and chroma predictions in the inter-frame prediction (MC). After determining the MV, the final MC is performed using an 8-tap interpolation filter for luma and a 4-tap interpolation filter for chroma.

[0247] MV refinement is a pattern-based MV search based on bilateral matching cost or template matching cost. JEM supports two search modes: unrestricted center-biased diamond search (UCBDS) and adaptive cross search, for MV refinement at the CU level and sub-CU level, respectively. For CU-level and sub-CU-level MV refinement, the MV is directly searched with MV precision at one-quarter of the luminance samples, followed by refinement with MV precision at one-eighth of the luminance samples. The search range for MV refinement at the CU and sub-CU steps is set to equal 8 luminance samples.

[0248] In the bilateral matching Merge mode, bidirectional prediction is applied because the motion information of the CU is derived based on the closest match between two blocks along the motion trajectory of the current CU in two different reference images. In the template matching Merge mode, the encoder can choose between unidirectional prediction from list 0, unidirectional prediction from list 1, or bidirectional prediction for the CU. The choice can be based on the template matching cost, as follows:

[0249] If costBi <= factor * min(cost0, cost1)

[0250] Use two-way forecasting;

[0251] Otherwise, if cost0 <= cost1

[0252] Use one-way prediction from list 0;

[0253] otherwise,

[0254] Use one-way prediction from List 1;

[0255] Here, cost0 is the SAD of template matching in list 0, cost1 is the SAD of template matching in list 1, and costBi is the SAD of bidirectional prediction template matching. For example, when the value of factor equals 1.25, this means that the selection process is biased towards bidirectional prediction. Inter-frame prediction direction selection can be applied to the CU-level template matching process.

[0256] 2.9 Examples of Bidirectional Optical Flow (BIO)

[0257] The Bidirectional Optical Flow (BIO) method is a per-sample motion refinement performed on top of block-by-block motion compensation used for bidirectional prediction. In some implementations, sample-level motion refinement does not use signaling notification.

[0258] Let I (k) It is the brightness value of the reference k (k=0,1) after block motion compensation, and will Represented as I (k) The horizontal and vertical components of the gradient. Assuming optical flow is effective, the motion vector field (v) x ,v y ) is given by the following equation.

[0259]

[0260] Combining this optical flow equation with the Hermite interpolation used for the motion trajectory of each sample, the result is a matching function value I at the end. (k) and derivative The only third-order polynomial. The value of this polynomial at t=0 is predicted by BIO:

[0261]

[0262] Figure 28 An example of optical flow trajectories in the Two-Way Optical Flow (BIO) method is shown. Here, τ0 and τ1 represent the distances to the reference frames. The distances τ0 and τ1 are calculated based on the Proof-of-Concept (POC) of Ref0 and Ref1: τ0 = POC(current) - POC(Ref0), τ1 = POC(Ref1) - POC(current). If both predictions originate from the same temporal direction (either from the past or the future), the signs are different (e.g., τ0·τ1<0). In this case, if the predictions do not originate from the same time (e.g., τ0≠τ1), BIO is applied. Both reference regions have non-zero motion (e.g., MVx0, MVy0, MVx1, MVy1≠0) and the block motion vector is proportional to the temporal distance (e.g., MVx0 / MVx1 = MVy0 / MVy1 = -τ0 / τ1).

[0263] The motion vector field (v) is determined by minimizing the difference Δ between the values ​​at points A and B. x ,v y ). Figure 9 An example of the intersection of the motion trajectory and the reference frame plane is shown. The model uses only the first linear term of the local Taylor expansion for Δ:

[0264]

[0265] All values ​​in the above equations depend on the sample location, denoted as (i′,j′). Assuming the motion is consistent in the local surrounding region, Δ can be minimized within a (2M+1)×(2M+1) square window Ω centered at the current prediction point (i,j), where M equals 2:

[0266]

[0267] For this optimization problem, JEM uses a simplification method, first minimizing in the vertical direction and then minimizing in the horizontal direction. This leads to...

[0268]

[0269] in,

[0270]

[0271] To avoid division by zero or very small values, regularization parameters r and m are introduced in equations (28) and (29).

[0272] r = 500·4 d-8 Equation (31)

[0273] m = 700·4 d-8 Equation (32) where d is the bit depth of the video sample.

[0274] To keep BIO memory access the same as normal bidirectional predictive motion compensation, all predicted and gradient values ​​I are calculated for the position within the current block. (k) , Figure 29A An example of an access location outside block 2900 is shown. Figure 29A As shown, in equation (30), a (2M+1)×(2M+1) square window Ω centered on the current prediction point on the boundary of the prediction block needs to access locations outside the block. In JEM, the location outside the block is I. (k) , The value is set to be equal to the nearest available value within the block. For example, this could be implemented to fill region 2901, such as... Figure 29B As shown.

[0275] Using BIO, it is possible to refine the motion field for each sample. To reduce computational complexity, a block-based BIO design can be used in JEM. Motion refinement can be calculated based on 4×4 blocks. In block-based BIO, all samples in the 4×4 block can be aggregated in equation (30) s n The value, and then the aggregated s n The value is used for the derived BIO motion vector offset of a 4×4 block. More specifically, the following formula can be used for block-based BIO derivation:

[0276]

[0277] Where b k Let represent the set of samples belonging to the k-th 4×4 block of the prediction block. (s) n,bk )>>4) Replace s in equations (28) and (29) n To derive the associated motion vector offset.

[0278] In some cases, the MV clique (regiment) in BIO may be unreliable due to noise or irregular motion. Therefore, in BIO, the magnitude of the MV clique is pruned to a threshold. The threshold is determined based on whether all reference images of the current image come from the same direction. For example, if all reference images of the current image come from the same direction, the threshold is set to 12×2. 14 -d Otherwise, it is set to 12×2 13-d .

[0279] It can simultaneously calculate the gradient and motion-compensated interpolation of BIO, and this motion-compensated interpolation uses...

[0280] The operation is consistent with the HEVC motion compensation process (e.g., 2D separable finite impulse response (FIR)). In some embodiments, the input to the 2D separable FIR is a reference frame sample that is identical to the fractional position (fracX, fracY) of the fractional portion of the block motion vector and the motion compensation process. For horizontal gradients... First, corresponding to the fractional position fracY with a descaling offset of d-8, vertical interpolation of the signal is performed using BIOfilters. Then, corresponding to the fractional position fracX with a descaling offset of 18-d, a gradient filter BIOfilterG is applied in the horizontal direction. For the vertical gradient... First, for the fractional position fracY with a descaling offset of d-8, a gradient filter is applied vertically using BIOfilterG. Then, for the fractional position fracX with a descaling offset of 18-d, signal displacement is performed horizontally using BIOfilterS. The lengths of the interpolation filters used for gradient calculation (BIOfilterG) and signal displacement (BIOfilterF) can be relatively short (e.g., 6 taps) to maintain reasonable complexity. Table 1 shows example filters that can be used for gradient calculation at different fractional positions of block motion vectors in BIO. Table 2 shows example interpolation filters that can be used for predictive signal generation in BIO.

[0281] Example filters for gradient computation in Table 1 BIO

[0282] Fractional pixel position Gradient interpolation filter (BIOfilterG) 0 {8,-39,-3,46,-17,5} 1 / 16 {8,-32,-13,50,-18,5} 1 / 8 {7,-27,-20,54,-19,5} 3 / 16 {6,-21,-29,57,-18,5} 1 / 4 {4,-17,-36,60,-15,4} 5 / 16 {3,-9,-44,61,-15,4} 3 / 8 {1,-4,-48,61,-13,3} 7 / 16 {0,1,-54,60,-9,2} 1 / 2 {-1,4,-57,57,-4,1}

[0283] Exemplary interpolation filters for predicting signal generation in Table 2BIO

[0284] Fractional pixel position Interpolation filters for predicted signals (BIOfilters) 0 {0,0,64,0,0,0} 1 / 16 {1,-3,64,4,-2,0} 1 / 8 {1,-6,62,9,-3,1} 3 / 16 {2,-8,60,14,-5,1} 1 / 4 {2,-9,57,19,-7,2} 5 / 16 {3,-10,53,24,-8,2} 3 / 8 {3,-11,50,29,-9,2} 7 / 16 {3,-11,44,35,-10,3} 1 / 2 {3,-10,35,44,-11,3}

[0285] In JEM, BIO can be applied to all bidirectional prediction blocks when two predictions come from different reference images. BIO can be disabled when Local Luminance Compensation (LIC) is enabled for the CU.

[0286] In some embodiments, OBMC is applied to the block after the normal MC process. To reduce computational complexity, BIO cannot be applied during the OBMC process. This means that BIO is applied to the MC process of a block when its own MV is used, but not when the MV of an adjacent block is used during the OBMC process. 2.10 Example of Decoder-Side Motion Vector Refinement (DMVR)

[0287] In bidirectional prediction, to predict a block region, two prediction blocks formed using motion vectors (MV) from list 0 and MV from list 1 are combined to form a single prediction signal. In the decoder-side motion vector refinement (DMVR) method, the two bidirectional prediction motion vectors are further refined through a bilateral template matching process. Bilateral template matching is applied in the decoder to perform a distortion-based search between the bilateral templates and reconstructed samples in the reference image to obtain the refined MV without transmitting additional motion information.

[0288] like Figure 30 As shown, in DMVR, bilateral templates are generated as a weighted combination (i.e., average) of two prediction blocks from the initial MV0 of list 0 and MV1 of list 1, respectively. The template matching operation involves calculating a cost metric between the generated template and the sample regions in the reference images (around the initial prediction block). For each of the two reference images, the MV that produces the minimum template cost is considered the updated MV for that list to replace the original template. In JEM, for each list, nine MV candidates are searched. These nine MV candidates include the original MV and eight surrounding MVs, each surrounding MV having an offset relative to the original MV by a brightness sample in the horizontal or vertical direction, or both. Finally, as... Figure 30 As shown, two new MVs, MV0' and MV1', are used to generate the final bidirectional prediction result. The sum of absolute differences (SAD) is used as the cost metric.

[0289] DMVR is applied to the Merge pattern of bidirectional prediction, using one MV from a past reference picture and another MV from a future reference picture without transmitting additional syntax elements. In JEM, DMVR is not applied when LIC, affine motion, FRUC, or subCU Merge candidate is enabled for a CU.

[0290] 2.11 Example of a Merge candidate list based on sub-blocks

[0291] In some embodiments (e.g., JVET-L0369), a sub-block-based Merge candidate list is generated by moving the ATMVP from the regular Merge list to the first position in the affine Merge list. The maximum number of sub-block-based Merge candidates remains 5, and the list length for the regular Merge list is reduced from 6 to 5.

[0292] The steps for constructing the candidate list for sub-block merge are as follows:

[0293] 1) Insert ATMVP candidate

[0294] 2) Inserting affine candidates for inheritance

[0295] 3) Insertion of constructed affine candidates

[0296] 4) Fill in the zero motion vector

[0297] Merge indices in a sub-block-based Merge list are binarized using truncated unary codes.

[0298] 3. Deficiencies of existing implementation methods

[0299] In existing implementations of affine inter-frame mode, corner points are always chosen as control points for the 4 / 6-parameter affine model. However, since the affine parameters are represented by CPMV, this can lead to large errors when deriving motion vectors using equation (1) for larger x and y values.

[0300] In other existing implementations of sub-block-based Merge candidate lists, all binary bits are context-encoded when encoding the Merge index, increasing the overhead of parsing. Furthermore, how to reconcile affine mapping with UMVE / triangulation is unknown. Current affine designs lack an aspect that utilizes temporal affine information.

[0301] 4. Example method for control point selection in affine mode encoding and decoding

[0302] The currently disclosed embodiments of the technology overcome the shortcomings of existing implementations, thereby providing video codecs with higher encoding and decoding efficiency. The following examples, described for various implementations, illustrate the selection of control points for affine mode codecs that can enhance existing and future video codec standards based on the disclosed technology. The examples of the disclosed technology provided below explain general concepts and are not intended to be construed as limiting. In the examples, the various features described in these examples can be combined unless explicitly stated otherwise. For the examples described below, the block size is represented as W×H, and the sub-block size in the affine mode is represented as sW×sH.

[0303] Example 1. A 4-parameter affine model is proposed that can use different sets of control points instead of always using the positions (0, 0) and (W, 0).

[0304] (a) For example, Figure 31 Any two points highlighted (shaded) in the middle can form a control point pair.

[0305] (b) In one example, the locations (0, 0) and (W / 2, 0) are selected as control points.

[0306] (c) In one example, the locations (0, 0) and (0, H / 2) are selected as control points.

[0307] (d) In one example, the locations (W, 0) and (0, H) are selected as control points.

[0308] (e) In one example, the locations (W / 2, H / 2) and (W, 0) are selected as control points.

[0309] (f) In one example, the locations (W / 2, H / 2) and (0, 0) are selected as control points.

[0310] (g) In one example, the locations (0, H / 2) and (W, H / 2) are selected as control points.

[0311] (h) In one example, the locations (W / 4, H / 2) and (3*W / 4, H / 2) are selected as control points.

[0312] (i) In one example, the locations (W / 2, H / 4) and (W / 2, 3*H / 4) are selected as control points.

[0313] (j) In one example, the locations (W, 0) and (W, H) are selected as control points.

[0314] (k) In one example, the locations (W / 2, 0) and (W / 2, H) are selected as control points.

[0315] (l) In one example, the positions (W / 4, H / 4) and (3*W / 4, 3*H / 4) are selected as control points.

[0316] (m) In one example, the locations (3*W / 4, H / 4) and (W / 4, 3*H / 4) are selected as control points.

[0317] (n) In one example, the positions (sW / 2, 0) and (W / 2+sW / 2, 0) are selected as control points.

[0318] (o) In one example, the positions (0, sH / 2) and (0, H / 2+sH / 2) are selected as control points.

[0319] (p) In one example, the locations (0, H / 2) and (W / 2, 0) are selected as control points.

[0320] (q) In one example, the locations (0, H) and (W, 0) are selected as control points.

[0321] (r) In one example, the locations (0, 0) and (W, H) are selected as control points.

[0322] Example 2. A 6-parameter affine model is proposed that can use different sets of control points instead of always using (0,0), (W,0), and (0,H).

[0323] (a) For example, Figure 31 Any three of the highlighted red dots that are not on a straight line can form a set of control points.

[0324] (b) In one example, the locations (W / 2, 0), (0, H), and (W, H) are selected as control points.

[0325] (c) In one example, the locations (0, 0), (W, 0), and (W / 2, H) are selected as control points.

[0326] (d) In one example, the locations (0, 0), (W, H / 2), and (0, H) are selected as control points.

[0327] (e) In one example, the locations (0, H / 2), (W, 0), and (W, H) are selected as control points.

[0328] (f) In one example, the locations (W / 4, H / 4), (3*W / 4, H / 4), and (3*W / 4, 3*H / 4) are selected as control points.

[0329] (g) In one example, select the locations (3*W / 4, H / 4), (W / 4, 3*H / 4), and (3*W / 4, 3*H / 4) as control points.

[0330] (h) In one example, the locations (0, 0), (W, 0), and (0, H / 2) are selected as control points.

[0331] (i) In one example, the locations (0, 0), (W, 0), and (W, H / 2) are selected as control points.

[0332] (j) In one example, the locations (3*W / 4–sW / 2, H / 4), (W / 4–sW / 2, 3*H / 4), and (3*W / 4–sW / 2, 3*H / 4) are selected as control points.

[0333] (k) In one example, the locations (3*W / 4–sW / 2, H / 4–sH / 2), (W / 4–sW / 2, 3*H / 4–sH / 2), and (3*W / 4–sW / 2, 3*H / 4–sH / 2) are selected as control points.

[0334] Example 3. The selection of control points can depend on motion information, strip / piece / image type, block size, and block shape.

[0335] (a) In one example, for blocks with width ≥ height, (W / 2, 0) and (W / 2, H) are selected as control points for the 4-parameter model.

[0336] (b) In one example, for blocks with width ≥ height, (W / 4, H / 2) and (W*3 / 2, H / 2) are selected as control points for the 4-parameter model.

[0337] (c) In one example, for blocks with width ≥ height, (0, H / 2) and (W, H / 2) are selected as control points for the 4-parameter model.

[0338] (d) In one example, for blocks with width ≤ height, (W / 2, 0) and (W / 2, H) are selected as control points for the 4-parameter model.

[0339] (e) In one example, for blocks with width ≤ height, (W / 2, H / 4) and (W / 2, H*3 / 4) are selected as control points for the 4-parameter model.

[0340] (f) In one example, for blocks with width ≤ height, (0, H / 2) and (W, H / 2) are selected as control points for the 4-parameter model.

[0341] (g) In one example, for blocks with width ≥ height, (0, 0), (W / 2, H) and (W, H) are selected as control points for the 4-parameter model.

[0342] (h) In one example, for blocks with width ≥ height, (0, 0), (0, H) and (W, H / 2) are selected as control points for the 4-parameter model.

[0343] (i) The motion information of a large set of control points can be examined first, for example, in terms of similarity, and those control points whose motion information is more closely related to each other can be selected as control points.

[0344] Example 4. The motion vector at the proposed control point can be stored separately from the normal motion vector buffer. Alternatively, the motion vector at the proposed control point can be stored in the normal motion vector buffer.

[0345] Example 5. In affine Merge or affine inter-frame modes, if some selected control points are not in the top row or left column, temporal motion information can be used to derive the constructed affine Merge / AMVP candidates, such as... Figure 32 As shown.

[0346] (a) In one example, for such control points, their associated motion information can be derived from one or more temporal blocks that are not located in the same slice / strip / picture.

[0347] (i) In one example, the co-location reference image and co-location block of the selected control point are first identified, and then the motion vector of the co-location block is scaled (if necessary) to generate a motion vector predictor. This process is performed in the same way as TMVP.

[0348] (ii) Alternatively, a co-reference picture may be included in the signaling notification in the SPS / PPS / strip header / piece group header.

[0349] (iii) Alternatively, a co-location reference picture can be determined by examining the adjacent motion information of the block. For example, the first reference picture with available adjacent motion information in list X or 1–X can be selected as the co-location picture for the predicted direction X, where X is 0 or 1.

[0350] (iv) Alternatively, the picture most frequently referenced by the neighboring blocks of CU in list X or 1–X is selected as the co-position picture for predicting direction X.

[0351] (v) Alternatively, co-location blocks can be identified using spatially adjacent motion information. For example, neighboring blocks of the CU can be scanned sequentially, and first available motion information from a reference co-location image can be used to identify co-location blocks.

[0352] (vi) Alternatively, if there is no adjacent motion information for reference co-images, the adjacent blocks of the CU are scanned sequentially, and the first available motion information is scaled to the co-image to identify the co-blocks.

[0353] (vii) can identify and check the movement of multiple co-located blocks.

[0354] (b) Alternatively, for such control points, the motion information of the nearest few spatially adjacent blocks can be checked in turn, and the first available motion information of the reference target image can be used as a motion vector predictor.

[0355] (i) If there is no available adjacent motion information, the control point is considered unavailable.

[0356] (ii) If there is no adjacent motion information for the reference target image, the control point is considered unusable.

[0357] (iii) Alternatively, if there is no adjacent motion information for the target reference image, the first available adjacent motion information is scaled to the target reference image and used as a motion vector predictor.

[0358] (iv) In one example, when generating a motion vector predictor in list X for the control point, for each available bidirectional adjacent motion information, the motion vectors in list X are checked first, and then the motion vectors in list 1–X are checked.

[0359] (v) In one example, if the control point is closer to a corner point than other corner points (using N) cp If (indicated), then proceed with the process described in Section 2.7.3 to derive N. cp Motion vector predictor,

[0360] Then it is used for that control point.

[0361] Example 6. In affine inter-frame mode, the MVD of a control point can be predicted from another control point. In one example, positions (W / 4, H / 4) and (3*W / 4, 3*H / 4) are chosen as control points, and the MVD of control point (3*W / 4, 3*H / 4) is predicted from the MVD of control point (W / 4, H / 4).

[0362] Example 7. The proposed CPMV of neighboring blocks can be used to derive motion vectors of one or more locations within and / or adjacent to the current block (e.g., representative points outside neighboring blocks but within and / or adjacent to the current block), and the derived MV can be used to generate the final prediction block of the current block (e.g., the derived MV can be used for OBMC processing).

[0363] (a) such as Figure 33 As shown, if an adjacent block (e.g., block A) is encoded and decoded in affine mode, its associated CPMV is used to derive the MV of the adjacent sub-blocks (e.g., sub-blocks B0-B3) of block A in the current block, for example, the sub-blocks highlighted in yellow.

[0364] (b) The CPMV of adjacent blocks can be used to derive motion information of the top-left corner position of the current block. For example, if using equations (1a) and / or (1b), assuming that the size of the adjacent affine codec block is represented by W'×H', the representative point (x, y) can be set to (W'+sW / 2, H'+sH / 2).

[0365] (c) The CPMV of adjacent blocks can be used to derive motion information of the center position of the current block or the sub-block located in the top row and / or left column of the current block. For example, if using equations (1a) and / or (1b), assuming that the size of the adjacent affine codec block is represented by W'×H', the representative point (x, y) can be set as (W'+W / 2+sW / 2, H'+H / 2+sH / 2) or (W'+W / 2-sW / 2, H'+H / 2-sH / 2).

[0366] (d) The CPMV of adjacent blocks can be used to derive motion information of the lower right position of the current block. For example, if using equations (1a) and / or (1b), assuming that the size of the adjacent affine codec block is represented by W'×H', the representative point (x, y) can be set to (W'+W-sW / 2, H'+H-sH / 2) or (W'+W+sW / 2, H'+H+sH / 2) or (W'+W, H'+H).

[0367] Example 8. This paper proposes that, in AMVP / Merge / UMVE / triangle mode, if adjacent blocks are encoded and decoded using affine mode, the motion vector predictor for the current block can be derived using the CPMV of adjacent blocks by setting one or more representative points to positions within the current block. Examples are provided below. Figure 34 As shown.

[0368] (a) In one example, the center position of the current block is used as the representative point for deriving the MV predictor.

[0369] (b) In one example, the top-left corner of the current block is used as the representative point for deriving the MV predictor.

[0370] (c) The exported MV predictor can be additionally inserted into the AMVP / Merge list or the UMVE-based Merge list or the triangular Merge list.

[0371] (d) In one example, the derived MV predictor can be used to replace the MV of the corresponding neighboring block. Therefore, the motion information of the neighboring blocks is not used as a predictor for encoding and decoding the current block. For example, in Figure 34 In this process, the derived MV predictor is used to replace the motion information associated with block B2.

[0372] (e) In one example, if several adjacent blocks come from the same affine codec CU / PU, the exported MV is used only to replace the MV of one adjacent block.

[0373] (f) Alternatively, if several adjacent blocks come from the same affine codec CU / PU, the derived MV is used and the MVs of all these adjacent blocks are not inserted into the list.

[0374] (g) Alternatively, if several adjacent blocks come from the same affine codec CU / PU, multiple MV predictors can be derived using different positions within the current block, and these MV predictors are further used to replace the MVs of all these adjacent blocks.

[0375] (h) In one example, if M neighboring blocks come from N affine codecs' CUs / PUs, then only K derived MV predictors (which can be used at different locations of the blocks) are generated and used to replace the MVs of the K neighboring blocks, where K = 1, 2, ..., or M. Alternatively, K = 1, 2, ..., or N.

[0376] (i) In one example, the derived MV predictors are given a higher priority than other normal MV predictors. That is, they can be added to the motion vector candidate list earlier than other normal MV predictors.

[0377] (j) Alternatively, the derived MV predictors are given a lower priority than other MV predictors. That is, they can be added to the motion vector candidate list after other normal MV predictors.

[0378] (k) In one example, the derived MV predictors have the same priority as other normal MV predictors. That is, they can be added to the motion vector candidate list in an interleaved manner with other normal MV predictors.

[0379] Example 9. A method is proposed to bypass encoding / decoding certain binary bits when encoding the Merge index of a sub-block-based Merge candidate list. The maximum length of the sub-block-based Merge candidate list is denoted as maxSubMrgListLen.

[0380] (a) In one example, only the first bit is context-coded, and all other bits are bypass-coded.

[0381] (i) In one example, the first binary bit is encoded using a context.

[0382] (ii) In one example, the first binary bit is encoded using more than one context. For example, using three contexts, as shown below:

[0383] (1)ctxIdx=aboveBlockIsAffineMode+leftBlockIsAffineMode;

[0384] (2) If the upper adjacent block is encoded and decoded in affine mode, then aboveBlockIsAffineMode equals 1; otherwise, aboveBlockIsAffineMode equals 0; and

[0385] (3) If the left adjacent block is encoded and decoded in affine mode, then leftBlockIsAffineMode equals 1; otherwise, leftBlockIsAffineMode equals 0.

[0386] (b) In one example, only the first K bits are encoded and decoded using context, while all other bits are encoded and decoded using bypass, where K = 0, 1, ... or maxSubMrgListLen–1.

[0387] (i) In one example, all binary bits encoded and decoded by context share a single context, except for the first binary bit.

[0388] (ii) In one example, each binary bit encoded and decoded by a context uses a context, except for the first binary bit.

[0389] Example 10. Whether to enable or disable the adaptive control point selection method can be indicated in signaling notifications such as SPS / PPS / VPS / sequence header / image header / strip header / piece group header / CTU group.

[0390] (a) In one example, if the block is encoded and decoded using affine Merge mode, adaptive control point selection is not applied.

[0391] (b) In one example, if there is no affine MVD for block encoding and decoding, adaptive control point selection is not applied.

[0392] (c) In one example, if there is no non-zero affine MVD for block encoding and decoding, adaptive control point selection is not applied.

[0393] Example 11. Whether and how adaptive control point selection is applied on an affine codec block depends on the reference picture of the current block. In one example, if the reference picture is the current picture, adaptive control point selection is not applied; for example, intra-block copying is applied in the current block.

[0394] Example 12. The proposed method can be applied under certain conditions, such as block size, encoding mode, motion information, and strip / image / piece type.

[0395] (a) In one example, the above method is not allowed when the block size contains samples smaller than M×H (e.g., 16, 32 or 64 luminance samples).

[0396] (b) In one example, the above method is not allowed when the block size contains more than M×H samples (e.g., 16, 32 or 64 luminance samples).

[0397] (c) Alternatively, the above method is not allowed when the minimum size of the block width or height is less than or no greater than X. In one example, X is set to 8.

[0398] (d) In one example, when the width or height of a block, or both the width and height, are greater than (or equal to) a threshold L, the block can be divided into multiple sub-blocks. Each sub-block is treated in the same way as a normal codec block with a size equal to the sub-block size.

[0399] (i) In one example, L is 64. A 64x128 / 128x64 block is divided into two 64x64 sub-blocks, and a 128x128 block is divided into four 64x64 sub-blocks. However, a Nx128 / 128xN block (where N<64) is not divided into sub-blocks.

[0400] (ii) In one example, L is 64. A 64x128 / 128x64 block is divided into two 64x64 sub-blocks, and a 128x128 block is divided into four 64x64 sub-blocks. At the same time, an Nx128 / 128xN block (where N<64) is divided into two Nx64 / 64xN sub-blocks.

[0401] Example 13. When the CPMV is stored in a separate buffer, different from the regular MV used for motion compensation, the CPMV can be stored with the same precision as the MV used for motion compensation.

[0402] (a) In one example, for regular MV and CPMV used for motion compensation, they are stored with a precision of 1 / M pixels, such as M=16 or 4.

[0403] (b) Alternatively, the precision of the CPMV storage can differ from the precision used for motion compensation. In one example, the CPMV can be stored with 1 / N pixel precision, and the MV used for motion compensation can be stored with 1 / M pixel precision, where M is not equal to N. In one example, M is set to 4, and N is set to 16.

[0404] Example 14. It is proposed that the CPMV used for generating affine fields can be modified before it is used to derive affine fields.

[0405] (a) In one example, the CPMV can be shifted to a higher precision first. After deriving the affine motion information for a position / block, it may be necessary to further scale the exported motion information to align with the precision of the regular MV used for motion compensation.

[0406] (b) Alternatively, the CPMV can be first rounded (scaled) to align with the precision of the normal MV used for motion compensation, and then the rounded CPMV can be used to derive the affine motion field. For example, this method can be applied when the CPMV is stored with 1 / 16 pixel precision and the normal MV is stored with 1 / 4 pixel precision.

[0407] (c) The rounding process can be defined as f(x) = (x + offset) >> L, for example, offset = 1 << (L - 1), and L can depend on the precision of the stored CPMV and the normal MV.

[0408] (d) The shift process can be defined as f(x) = x << L, for example, L = 2.

[0409] The examples described above can be incorporated into the context of the methods described below, such as methods 3500, 3530, and 3560, which can be implemented at a video decoder or a video encoder.

[0410] Figure 35A A flowchart of an example method for video processing is shown. Method 3500 includes: at step 3502, for the conversion between the current block of a video and the bitstream representation of the video, selecting multiple control points for the current block, the multiple control points including at least one non-corner point of the current block, and each of the multiple control points representing the affine motion of the current block.

[0411] Method 3500 includes: at step 3504, performing the conversion between the current block and the bitstream representation based on the multiple control points.

[0412] Figure 35B A flowchart of an example method for video processing is shown. Method 3530 includes: at step 3532, for the conversion between the current block of a video and the bitstream representation of the video, determining that at least one of the multiple binary bits for encoding / decoding the Merge index of the Merge candidate list for a Merge-based sub-block uses bypass encoding / decoding based on a condition.

[0413] Method 3530 includes: at step 3534, performing the conversion based on the determination.

[0414] Figure 35C A flowchart of an example method for video processing is shown. Method 3560 includes: at step 3562, for the conversion between the current block of a video and the bitstream representation of the video, selecting multiple control points for the current block, the multiple control points including at least one non-corner point of the current block, and each of the multiple control points representing the affine motion of the current block.

[0415] Method 3560 includes: in step 3564, deriving motion vectors of one or more control points of multiple control points based on the control point motion vectors (CPMV) of one or more adjacent blocks of the current block.

[0416] Method 3560 includes: in step 3566, performing a conversion between the current block and the bitstream representation based on multiple control points and motion vectors.

[0417] In some embodiments, the following technical solutions can be implemented:

[0418] A1. A method for video processing (e.g., Figure 35A Method 3500 in the video includes: selecting (3502) a plurality of control points of the current block for a conversion between a current block of the video and a bitstream representation of the video, wherein the plurality of control points includes at least one non-corner point of the current block, and wherein each of the plurality of control points represents an affine motion of the current block; and performing (3504) a conversion between the current block and the bitstream representation based on the plurality of control points.

[0419] A2. The method according to solution A1, wherein the size of the current block is W×H, where W and H are positive integers, wherein the current block is encoded and decoded using a four-parameter affine model, and wherein the plurality of control points do not include (0, 0) and (W, 0).

[0420] A3. The method according to solution A2, wherein the plurality of control points include (0, 0) and (W / 2, 0).

[0421] A4. The method according to solution A2, wherein the plurality of control points include (0, 0) and (0, H / 2).

[0422] A5. The method according to solution A2, wherein the plurality of control points include (W, 0) and (0, H).

[0423] A6. The method according to solution A2, wherein the plurality of control points include (W, 0) and (W, H).

[0424] A7. The method according to solution A2, wherein the plurality of control points include (W / 2, 0) and (W / 2, H).

[0425] A8. The method according to solution A2, wherein the plurality of control points include (W / 4, H / 4) and (3×W / 4, 3×H / 4).

[0426] A9. The method according to solution A2, wherein the plurality of control points include (W / 2, H / 2) and (W, 0).

[0427] A10. The method according to solution A2, wherein the plurality of control points include (W / 2, H / 2) and (0, 0).

[0428] A11. The method according to solution A2, wherein the plurality of control points include (0, H / 2) and (W, H / 2).

[0429] A12. The method according to solution A2, wherein the plurality of control points include (W / 4, H / 2) and (3×W / 4, H / 2).

[0430] A13. The method according to solution A2, wherein the plurality of control points include (W / 2, H / 4) and (W / 2, 3×H / 4).

[0431] A14. The method according to solution A2, wherein the plurality of control points include (3×W / 4, H / 4) and (W / 4, 3×H / 4).

[0432] A15. The method according to solution A2, wherein the current block includes sub-blocks, wherein the size of the sub-blocks is sW×sH, and wherein sW and sH are positive integers.

[0433] A16. The method according to solution A15, wherein the plurality of control points include (sW / 2, 0) and (W / 2+sW / 2, 0).

[0434] A17. The method according to solution A15, wherein the plurality of control points include (0, sH / 2) and (0, H / 2+sH / 2).

[0435] A18. The method according to solution A2 or A15, wherein the plurality of control points include (0, H / 2) and (W / 2, 0).

[0436] A19. The method according to solution A2 or A15, wherein the plurality of control points include (0, H) and (W, 0).

[0437] A20. The method according to solution A2 or A15, wherein the plurality of control points include (0, 0) and (W, H).

[0438] A21. The method according to solution A1, wherein the size of the current block is W×H, where W and H are positive integers, wherein the current block is encoded and decoded using a six-parameter affine model, and wherein the plurality of control points do not include (0, 0), (W, 0) and (0, H).

[0439] A22. The method according to solution A21, wherein the plurality of control points include (W / 2, 0), (0, H) and (W, H).

[0440] A23. The method according to solution A21, wherein the plurality of control points include (0, 0), (W, 0) and (W / 2, H).

[0441] A24. The method according to solution A21, wherein the plurality of control points include (0, 0), (W, H / 2) and (0, H).

[0442] A25. The method according to solution A21, wherein the plurality of control points include (0, H / 2), (W, 0) and (W, H).

[0443] A26. The method according to solution A21, wherein the plurality of control points include (W / 4, H / 4), (3×W / 4, H / 4) and (3×W / 4, 3×H / 4).

[0444] A27. The method according to solution A21, wherein the plurality of control points include (3×W / 4, H / 4), (W / 4, 3×H / 4) and (3×W / 4, 3×H / 4).

[0445] A28. The method according to solution A21, wherein the plurality of control points include (0, 0), (W, 0) and (0, H / 2).

[0446] A29. The method according to solution A21, wherein the plurality of control points include (0, 0), (W, 0) and (W, H / 2).

[0447] A30. The method according to solution A21, wherein the current block includes sub-blocks, wherein the size of the sub-blocks is sW×sH, and wherein sW and sH are positive integers.

[0448] A31. The method according to solution A30, wherein the plurality of control points include (3×W / 4-sW / 2, H / 4), (W / 4-sW / 2, 3×H / 4) and (3×W / 4-sW / 2, 3×H / 4).

[0449] A32. The method according to solution A30, wherein the plurality of control points include (3×W / 4-sW / 2, H / 4-sH / 2), (W / 4-sW / 2, 3×H / 4-sH / 2) and (3×W / 4-sW / 2, 3×H / 4-sH / 2).

[0450] A33. The method according to solution A1, wherein the selection of the plurality of control points is based on at least one of motion information, the size or shape of the current block, strip type, picture type, or piece type.

[0451] A34. The method according to solution A33, wherein the size of the current block is W×H, where W and H are positive integers, and wherein the current block is encoded and decoded using a four-parameter affine model.

[0452] A35. The method described in solution A34, wherein W≥H.

[0453] A36. The method according to solution A35, wherein the plurality of control points include (W / 2, 0) and (W / 2, H).

[0454] A37. The method according to solution A35, wherein the plurality of control points include (W / 4, H / 2) and (3 × W / 2, H / 2).

[0455] A38. The method according to solution A35, wherein the plurality of control points include (0, H / 2) and (W, H / 2).

[0456] A39. The method described in solution A34, where W≤H.

[0457] A40. The method according to solution A39, wherein the plurality of control points include (W / 2, 0) and (W / 2, H).

[0458] A41. The method according to solution A39, wherein the plurality of control points include (W / 2, H / 4) and (W / 2, 3×H / 4).

[0459] A42. The method according to solution A39, wherein the plurality of control points include (0, H / 2) and (W, H / 2).

[0460] A43. The method according to solution A33, wherein the size of the current block is W×H, where W and H are positive integers, and wherein the current block is encoded and decoded using a six-parameter affine model.

[0461] A44. The method according to solution A43, wherein W≥H, and wherein the plurality of control points include (0,0), (W / 2,H) and (W,H).

[0462] A45. The method according to solution A43, wherein W≥H, and wherein the plurality of control points includes (0,0), (0,H) and (W,H / 2).

[0463] A46. The method according to any one of solutions A1 to A45, wherein one or more motion vectors corresponding to the plurality of control points are stored in a buffer different from the normal motion vector buffer.

[0464] A47. The method according to any one of solutions A1 to A45, wherein one or more motion vectors corresponding to the plurality of control points are stored in a normal motion vector buffer.

[0465] A48. The method according to solution A1, wherein the current block is encoded and decoded using an affine Merge mode or an affine inter-frame mode, and wherein the method further comprises: using temporal motion information to derive one or more constructed affine Merge or advanced motion vector prediction (AMVP) candidates based on the positions of the plurality of control points in the current block.

[0466] A49. The method according to solution A48, wherein the positions of the plurality of control points do not include the top row and left column of the current block.

[0467] A50. The method according to solution A48, wherein the temporal motion information is derived from a block located within a first slice, picture, or strip that is different from the second slice, picture, or strip that includes the current block.

[0468] A51. The method according to solution A48, wherein the temporal motion information is derived from a co-location reference image or co-location reference block.

[0469] A52. The method according to solution A51, wherein the co-located reference image is signaled in the Picture Parameter Set (PPS), Sequence Parameter Set (SPS), strip header or slice header.

[0470] A53. The method according to solution A51, wherein the co-position reference image is determined by examining the adjacent motion information of the current block.

[0471] A54. The method according to solution A51, wherein the co-position reference image corresponds to the reference image most frequently used by the neighboring blocks of the current block.

[0472] A55. The method according to solution A1, wherein the current block is encoded and decoded using an affine inter-frame mode, and wherein the method further comprises: predicting the MVD of a second control point of the plurality of control points based on the motion vector difference (MVD) of a first control point of the plurality of control points.

[0473] A56. The method according to solution A55, wherein the first control point is (W / 4, H / 4), and wherein the second control point is (3×W / 4, 3×H / 4).

[0474] A57. The method according to solution A1, wherein enabling the selection of the plurality of control points is based on signaling notifications in the Sequence Parameter Set (SPS), Video Parameter Set (VPS), Picture Parameter Set (PPS), Sequence Header, Picture Header, Strip Header, or Slice Header.

[0475] A58. The method according to solution A57, wherein the encoding / decoding mode of the current block does not include the affine Merge mode.

[0476] A59. The method according to solution A57, wherein at least one affine motion vector difference (MVD) is encoded and decoded for the current block.

[0477] A60. The method according to solution A57, wherein at least one non-zero affine motion vector difference (MVD) is encoded and decoded for the current block.

[0478] A61. The method according to solution A1, wherein the selection of the plurality of control points is enabled based on a reference image of the current block.

[0479] A62. The method according to solution A61, wherein the reference image is different from the current image including the current block.

[0480] A63. The method according to solution A1, wherein the number of samples in the current block is greater than or equal to K, and wherein K is a positive integer.

[0481] A64. The method according to solution A1, wherein the number of samples in the current block is less than or equal to K, and wherein K is a positive integer.

[0482] A65. The method described according to solution A63 or A64, wherein K = 16, 32 or 64.

[0483] A66. The method according to solution A1, wherein the height or width of the current block is greater than K, and wherein K is a positive integer.

[0484] A67. The method according to solution A66, wherein K = 8.

[0485] A68. The method according to solution A1, wherein when it is determined that the height or width of the current block is greater than or equal to a threshold (L), the selection of the plurality of control points is performed in at least one of the plurality of sub - blocks of the current block.

[0486] A69. The method according to solution A68, wherein L = 64, wherein the size of the current block is 64×128, 128×64 or 128×128, and wherein the size of each of the plurality of sub - blocks is 64×64.

[0487] A70. The method according to solution A68, wherein L = 64, wherein the size of the current block is N×128 or 128×N, and wherein the sizes of two of the plurality of sub - blocks are N×64 or 64×N respectively.

[0488] A71. The method according to solution A1, wherein one or more motion vectors corresponding to the plurality of control points are stored in a first buffer different from a second buffer, and the second buffer includes conventional motion vectors for motion compensation.

[0489] A72. The method according to solution A71, wherein the one or more motion vectors are stored with 1 / N pixel accuracy, the conventional motion vectors are stored with 1 / M pixel accuracy, and wherein M and N are integers.

[0490] A73. The method according to solution A72, wherein M = N, and wherein M = 4 or 16.

[0491] A74. The method according to solution A72, wherein M = 4, and N = 16.

[0492] A75. The method according to solution A1, further comprising: modifying the plurality of control points before performing the transformation, wherein the transformation includes deriving an affine motion field; and transforming the affine motion field to align the accuracy of the affine motion field with the accuracy of the motion vectors for motion compensation.

[0493] A76. The method according to solution A75, wherein the modification includes applying a shift operation to at least one of the plurality of control points.

[0494] A77. The method according to solution A76, wherein the shift operation is defined as: f(x)=x << L, where L is a positive integer.

[0495] A78. The method described in solution A77, where L = 2.

[0496] A79. The method according to solution A75, wherein the modification includes applying a rounding or scaling operation to align the precision of at least one of the plurality of control points with the precision of the motion vector used for motion compensation.

[0497] A80. The method described in solution A79, wherein the rounding operation is defined as: f(x) = (x + offset) >> L, where offset = (1 << (L - 1)), and where L is a positive integer.

[0498] A81. The method according to solution A80, wherein L is based on the accuracy of a common motion vector or a motion vector of one of the plurality of control points.

[0499] A82. The method according to any one of solutions A1 to A81, wherein the transformation generates the current block from the bitstream representation.

[0500] A83. The method according to any one of solutions A1 to A81, wherein the transformation generates the bitstream representation from the current block.

[0501] A84. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein when the processor executes the instructions, the instructions cause the processor to implement the method of any one of solutions A1 to A83.

[0502] A85. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for implementing the method of any one of solutions A1 to A83.

[0503] In some embodiments, the following technical solutions can be implemented:

[0504] B1. A method for video processing (e.g., Figure 35B Method 3530 in the video includes: for the conversion between the current block of the video and the bitstream representation of the video, determining (3502) at least one of a plurality of binary bits for encoding and decoding the Merge index of the Merge candidate list of the Merge sub-blocks based on conditional bypass encoding and decoding; and performing (3534) the conversion based on the determination.

[0505] B2. The method according to solution B1, wherein the first binary bit of the plurality of binary bits is encoded / decoded using at least one context, and wherein all other binary bits of the plurality of binary bits are bypassed and encoded / decoded.

[0506] B3. The method according to solution B2, wherein the at least one context comprises a context.

[0507] B4. The method according to solution B2, wherein the at least one context comprises three contexts.

[0508] B5. According to the method described in solution B4, wherein the three contexts are defined as: ctxIdx = aboveBlockIsAffineMode + leftBlockIsAffineMode; wherein, if the upper adjacent block of the current block uses the first affine mode encoding and decoding, then aboveBlockIsAffineMode = 1, otherwise aboveBlockIsAffineMode = 0; and wherein, if the left adjacent block of the current block uses the second affine mode encoding and decoding, then leftBlockIsAffineMode = 1, otherwise leftBlockIsAffineMode = 0.

[0509] B6. The method according to solution B1, wherein each of the first K bits of the plurality of binary bits is encoded using at least one context, wherein all other bits of the plurality of binary bits are encoded using a bypass, wherein K is a non-negative integer, wherein 0 ≤ K ≤ maxSubMrgListLen–1, and wherein maxSubMrgListLen is the maximum length of the plurality of binary bits.

[0510] B7. The method described in solution B6, wherein, except for the first bit of the first K bits, the first K bits share a context.

[0511] B8. The method described in solution B6, wherein each of the first K bits, except for the first bit of the aforementioned K bits, uses a context.

[0512] B9. The method according to any one of solutions B1 to B8, wherein the Merge-based sub-block Merge candidate list includes sub-block temporal motion vector prediction (SbTMVP) candidates.

[0513] B10. The method according to any one of solutions B1 to B8, wherein the Merge-based sub-block Merge candidate list includes one or more inherited affine candidates, one or more constructed affine candidates, or one or more zero candidates.

[0514] B11. The method according to any one of solutions B1 to B10, wherein the transformation generates the current block from the bitstream representation.

[0515] B12. The method according to any one of solutions B1 to B10, wherein the transformation generates the bitstream representation from the current block.

[0516] B13. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein when the processor executes the instructions, the instructions cause the processor to implement the method of any one of solutions B1 to B12.

[0517] B14. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for implementing the method of any one of solutions B1 to B12.

[0518] In some embodiments, the following technical solutions can be implemented:

[0519] C1. A method for video processing (e.g., Figure 35C Method 3560 in the video includes: for a conversion between a current block of the video and a bitstream representation of the video, selecting (3562) a plurality of control points of the current block, wherein the plurality of control points includes at least one non-corner point of the current block, and wherein each of the plurality of control points represents an affine motion of the current block; deriving (3564) motion vectors of one or more of the plurality of control points based on motion vectors of control points of one or more adjacent blocks of the current block; and performing (3566) a conversion between the current block and the bitstream representation based on the plurality of control points and the motion vectors.

[0520] C2. The method according to solution C1, wherein the one or more adjacent blocks include a first block encoded and decoded using an affine pattern, and wherein the motion vector is derived from one or more sub-blocks in the current block that are directly adjacent to one or more sub-blocks in the first block.

[0521] C3. The method according to solution C1, wherein the one or more control points include a control point at the top left corner of the current block.

[0522] C4. The method according to solution C1, wherein the one or more control points include a control point located at the center of the current block, or a control point located at the center of a sub-block of the current block in the top row or left column of the current block.

[0523] C5. The method according to solution C1, wherein the one or more control points include a control point at the lower right corner of the current block.

[0524] C6. The method according to solution C1, wherein the derived motion vectors of the one or more control points are used for Overlapping Block Motion Compensation (OMBC) processing of the current block.

[0525] C7. The method according to solution C1, wherein the current block is encoded and decoded using at least one of Merge mode, Advanced Motion Vector Prediction (AMVP) mode, Triangulation mode, or Ultimate Motion Vector Expression (UMVE) mode, wherein the UMVE mode includes a motion vector expression, the motion vector expression including the starting point, motion amplitude, and motion direction of the current block.

[0526] C8. The method described in solution C7, wherein the deriving of the motion vector is based on representative points in the current block.

[0527] C9. The method described in solution C8, wherein the representative point is the center position or the upper left corner position of the current block.

[0528] C10. The method according to solution C7 further includes: inserting the motion vector into an AMVP list, or a Merge list, or a UMVE-based Merge list, or a triangular Merge list.

[0529] C11. The method according to solution C7 further includes: replacing at least one of the motion vectors of the one or more adjacent blocks with the derived motion vector.

[0530] C12. The method according to solution C11, wherein the one or more adjacent blocks come from the same affine codec unit (CU) or prediction unit (PU).

[0531] C13. The method according to solution C7 further includes: inserting each of the derived motion vectors into the motion vector candidate list before inserting the normal motion vector predictor into the motion vector candidate list.

[0532] C14. The method according to solution C7 further includes: after inserting the normal motion vector predictor into the motion vector candidate list, inserting each of the derived motion vectors into the motion vector candidate list.

[0533] C15. The method according to solution C7 further includes: inserting the motion vector into the motion vector candidate list in an interleaved manner with the normal motion vector predictor.

[0534] C16. The method according to any one of solutions C1 to C15, wherein the transformation generates the current block from the bitstream representation.

[0535] C17. The method according to any one of solutions C1 to C15, wherein the transformation generates the bitstream representation from the current block.

[0536] C18. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein when the processor executes the instructions, the instructions cause the processor to implement the method of any one of solutions C1 to C17.

[0537] C19. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for implementing the method of any one of solutions C1 to C17.

[0538] 5. Example implementation of the publicly available technology

[0539] Figure 36 This is a block diagram of a video processing apparatus 3600. Apparatus 3600 can be used to implement one or more methods described herein. Apparatus 3600 can be implemented in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. Processor 3602 can be configured to implement one or more methods described herein (including, but not limited to, methods 3500, 3530, and 3560). Memory (multiple memories) 3604 can be used to store data and code for implementing the methods and techniques described herein. Video processing hardware 3606 can be used to implement some of the techniques described herein in hardware circuitry.

[0540] In some embodiments, implementations such as those regarding Figure 36 The hardware platform described describes a device for implementing video encoding and decoding methods.

[0541] Some embodiments of the disclosed technology include making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in the processing of video blocks, but will not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when a video processing tool or mode is enabled based on a decision or determination, the conversion from a video block to a bitstream representation of the video will use that video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that it has been modified based on the video processing tool or mode. That is, the conversion from a bitstream representation of the video to a video block will be performed using a video processing tool or mode enabled based on a decision or determination.

[0542] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, the encoder will not use the tool or mode in the conversion of video blocks into a bitstream representation of video. In another example, when a video processing tool or mode is disabled, the decoder will process the bitstream knowing that the bitstream has not been modified using a video processing tool or mode enabled based on the decision or determination.

[0543] Figure 37 This is a block diagram illustrating an example video processing system 3700, in which various techniques disclosed herein can be implemented. Various implementations may include some or all of the components of system 3700. System 3700 may include an input 3702 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or in a compressed or encoded format. Input 3702 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.

[0544] System 3700 may include codec component 3704, which can implement the various encoding or character encoding methods described in this document. Codec component 3704 can reduce the average bit rate of the video from input 3702 to the output of codec component 3704 to generate a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of codec component 3704 can be stored or transmitted via connected communication, as shown in component 3706. The stored or transmitted bitstream (or codec) representation of the video received at input 3702 can be used by component 3708 to generate pixel values ​​or displayable video sent to display interface 3710. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “encoding” operations or tools, it should be understood that encoding tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the encoding / decoding results will be performed by the decoder.

[0545] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be implemented in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0546] As can be understood from the foregoing, specific embodiments of the disclosed technology have been described for illustrative purposes, but various modifications can be made without departing from the scope of the invention. Therefore, the technology disclosed in this invention is not limited except for the appended claims.

[0547] The subject matter and functional operations described in this patent document can be implemented in various systems, in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations thereof. The subject matter described in this specification can be implemented as one or more computer program products, i.e., one or more computer program instruction modules encoded on a tangible and non-transitory computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances affecting machine-readable propagation signals, or a combination thereof. The terms "data processing unit" or "data processing apparatus" encompass all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof.

[0548] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as a portion of a file containing other programs or data (e.g., one or more scripts stored in a markup language file), as a single file dedicated to the program in question, or as multiple coordinated files (e.g., a file storing one or more modules, subroutines, or code portions). Computer programs can be deployed to execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communication network.

[0549] The processes and logic flows described in this specification can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processes and logic flows can also be executed by dedicated logic circuitry, and the device can be implemented as dedicated logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).

[0550] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors of any kind of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, to receive data from, transfer data to, or both receive and transfer data from such mass storage devices. However, a computer does not need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.

[0551] The purpose is to combine the instruction manual with the accompanying Figure 1 The above is to be considered illustrative only, where illustrative means example. As used herein, unless the context clearly indicates otherwise, the use of "or" is intended to include "and / or".

[0552] While this patent document contains numerous details, these details should not be construed as limiting any invention or the scope of the claims, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features described in the context of separate embodiments in this patent document may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Furthermore, although the features described above may be described as functioning in certain combinations and even initially claimed in this way, in some cases, one or more features from a claimed combination may be removed from that combination, and the claimed combination may refer to a sub-combination or a variation of a sub-combination.

[0553] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order shown or sequentially, or to perform all the shown operations to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0554] Only a few implementations and examples are described, and other implementations, enhancements and variations can be made based on what is described and shown in this patent document.

Claims

1. A video data processing method, comprising: For the conversion between the current video block and the bitstream of the video, at least one of a plurality of binary bits used for encoding and decoding the Merge index of the Merge-based sub-block Merge candidate list is determined to use bypass encoding and decoding, wherein the Merge-based sub-block Merge candidate list is constructed based on sub-block temporal motion vector prediction SbTMVP candidates; and The conversion is performed based on the determination. The SbTMVP candidate is derived from the spatial neighboring blocks of the current video block based on temporal motion displacement, and the reference image associated with the temporal motion displacement is the same as the co-located image of the current video block. The Merge-based sub-block Merge candidate list is also constructed based on inherited affine candidates, constructed affine candidates, or zero candidates. The inherited affine candidate control point motion vectors are derived from the control point motion vectors of spatially adjacent blocks encoded in affine mode, wherein all control point motion vectors and all motion vectors in the inter-frame prediction mode are stored with 1 / 16 pixel precision. Wherein, the temporal motion displacement is used to locate at least one region in an image different from the current image including the current video block, and wherein the first binary bit of the plurality of binary bits is encoded and decoded with at least one context, and wherein all other binary bits of the plurality of binary bits are encoded and decoded with bypass.

2. The method according to claim 1, wherein, The at least one context consists of a single context.

3. The method according to claim 1, wherein, Each of the first K bits of the plurality of binary bits is encoded and decoded with at least one context, and all other bits of the plurality of binary bits are encoded and decoded with a bypass, where K is a non-negative integer, where 0 ≤ K ≤ maxSubMrgListLen – 1, and where maxSubMrgListLen is the maximum length of the plurality of binary bits.

4. The method according to any one of claims 1 to 3, wherein, The conversion includes generating the current video block from the bitstream.

5. The method according to any one of claims 1 to 3, wherein, The conversion includes generating the bitstream from the current video block.

6. A video data processing apparatus, comprising a processor and a non-transitory memory having instructions thereon, wherein, When the instruction is executed by the processor, the processor: For the conversion between the current video block and the bitstream of the video, at least one of a plurality of binary bits used for encoding and decoding the Merge index of the Merge-based sub-block Merge candidate list is determined to use bypass encoding and decoding, wherein the Merge-based sub-block Merge candidate list is constructed based on sub-block temporal motion vector prediction SbTMVP candidates; and The conversion is performed based on the determination. The SbTMVP candidate is derived from the spatial neighboring blocks of the current video block based on temporal motion displacement, and the reference image associated with the temporal motion displacement is the same as the co-located image of the current video block. The Merge-based sub-block Merge candidate list is also constructed based on inherited affine candidates, constructed affine candidates, or zero candidates. The inherited affine candidate control point motion vectors are derived from the control point motion vectors of spatially adjacent blocks encoded in affine mode, wherein all control point motion vectors and all motion vectors in the inter-frame prediction mode are stored with 1 / 16 pixel precision. Wherein, the temporal motion displacement is used to locate at least one region in an image different from the current image including the current video block, and wherein the first binary bit of the plurality of binary bits is encoded and decoded with at least one context, and wherein all other binary bits of the plurality of binary bits are encoded and decoded with bypass.

7. The apparatus according to claim 6, wherein, The at least one context consists of a single context.

8. A non-transitory computer-readable storage medium having instructions stored thereon, the instructions causing a processor to: For the conversion between the current video block and the bitstream of the video, at least one of the multiple binary bits used for encoding and decoding the Merge index of the Merge candidate list based on the Merge sub-block is determined to use bypass encoding and decoding, wherein, The Merge-based sub-block Merge candidate list is constructed based on the sub-block temporal motion vector prediction SbTMVP candidates; and The conversion is performed based on the determination. The SbTMVP candidate is derived from the spatial neighboring blocks of the current video block based on temporal motion displacement, and the reference image associated with the temporal motion displacement is the same as the co-located image of the current video block. The Merge-based sub-block Merge candidate list is also constructed based on inherited affine candidates, constructed affine candidates, or zero candidates. The inherited affine candidate control point motion vectors are derived from the control point motion vectors of spatially adjacent blocks encoded in affine mode, wherein all control point motion vectors and all motion vectors in the inter-frame prediction mode are stored with 1 / 16 pixel precision. Wherein, the temporal motion displacement is used to locate at least one region in an image different from the current image including the current video block, and wherein the first binary bit of the plurality of binary bits is encoded and decoded with at least one context, and wherein all other binary bits of the plurality of binary bits are encoded and decoded with bypass.

9. A method for storing a video bitstream, comprising: At least one of the multiple binary bits used to determine the Merge index for encoding and decoding a Merge-based sub-block Merge candidate list is used for bypass encoding and decoding, wherein the Merge-based sub-block Merge candidate list is constructed based on sub-block temporal motion vector prediction SbTMVP candidates. The bit stream is generated based on the determination; and The bitstream is stored in a non-transitory computer-readable recording medium. The SbTMVP candidate is derived from the spatial neighboring blocks of the current video block of the video based on temporal motion displacement, and the reference image associated with the temporal motion displacement is the same as the co-located image of the current video block. The Merge-based sub-block Merge candidate list is also constructed based on inherited affine candidates, constructed affine candidates, or zero candidates. The inherited affine candidate control point motion vectors are derived from the control point motion vectors of spatially adjacent blocks encoded in affine mode, wherein all control point motion vectors and all motion vectors in the inter-frame prediction mode are stored with 1 / 16 pixel precision. Wherein, the temporal motion displacement is used to locate at least one region in an image different from the current image including the current video block, and wherein the first binary bit of the plurality of binary bits is encoded and decoded with at least one context, and wherein all other binary bits of the plurality of binary bits are encoded and decoded with bypass.

Citation Information

Patent Citations

  • Method and apparatus for context adaptive binary arithmetic coding of syntax elements

    US20140328396A1

  • Sub-prediction unit temporal motion vector prediction (sub-PU TMVP) for video coding

    US20180288430A1