Sub - block Motion Vector Derivation Based on a Fall - back Motion Vector Field
Through the fallback-based motion vector field (RMVF) method, the motion information of non-adjacent airspace and time domain neighboring blocks is used to deduce the affine model and motion candidate list of video blocks, solving the problem of inefficiency in existing video encoding and decoding technology, and achieving more efficient video encoding and decoding.
Patent Information
- Application Number
- CN202080017067.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-02-27
- Filing Date
- 2020-02-27
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2040-02-27
AI Technical Summary
The existing video encoding and decoding technology still has inefficient problems in dealing with video bandwidth requirements, especially in the processing of high-resolution and high-frame rate video data. The existing motion vector prediction methods are difficult to effectively reduce the bit rate and improve the encoding efficiency.
The fallback-based motion vector field (RMVF) method is used to derive the affine model and motion candidate list of the current video block by using the motion information of the non-adjacent airspace and time domain neighboring blocks, and improve coding efficiency by updating and reordering the motion information.
Through the RMVF method, the motion information of video blocks can be predicted more accurately, the bit rate can be reduced, and the encoding efficiency can be improved. It is suitable for efficient video encoding and codec standards such as HEVC and future video codecs, reducing network bandwidth requirements.
Smart Images

Figure CN113508593B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] In accordance with applicable patent laws and / or in accordance with the rules of the Paris Convention, this application is intended to claim the priority and benefits of International Patent Application PCT / CN2019 / 076300, filed on February 27, 2019, in a timely manner. Its entire disclosure is incorporated by reference as part of the disclosure of this application. Technical field
[0003] This patent document relates to video coding and decoding technologies, devices, and systems. Background art
[0004] Despite the progress in video compression, digital video still accounts for the largest bandwidth usage on the Internet and other digital communication networks. As the number of networked user devices capable of receiving and displaying video increases, the bandwidth demand for digital video usage is expected to continue to grow. Summary of the invention
[0005] Devices, systems, and methods related to digital video coding and decoding, in particular, related to deriving motion vectors are described. The described methods can be applied to existing video coding and decoding standards (e.g., High Efficiency Video Coding (HEVC) or Versatile Video Coding) and future video coding and decoding standards or video codecs.
[0006] In a representative aspect, a method for video processing is disclosed, including: deriving at least one motion model of a current video block based on motion information of at least one non - adjacent spatial neighboring block or at least one temporal neighboring block of the current video block; deriving motion information of the current video block or at least one sub - block of the current video block based on at least one motion model; and performing a transformation of the current video block according to the derived motion information.
[0007] In another representative aspect, a method for video processing is disclosed, including: deriving one or more control - point motion - vector predictors (CPMVPs) of an affine model of a current video block from at least one set of neighboring blocks in a fallback - based motion - vector - field (RMVF) scheme; updating a motion candidate list of the current video block based on the one or more CPMVPs, wherein the one or more CPMVPs are associated with a specific affine motion pattern; and performing a transformation of the current video block according to the motion candidate list.
[0008] In yet another representative aspect, a method for video processing is disclosed, including: deriving at least one set of affine parameters of an affine model for a current video block by utilizing a reverse motion vector field (RMVF) scheme based at least on at least one set of motion information stored in a lookup table; and performing a transformation of the current video block based on the at least one set of affine parameters.
[0009] In yet another representative aspect, a method for video processing is disclosed, including: maintaining an affine motion candidate list for a current video block, the affine motion candidate list including a plurality of affine motion candidates, wherein the current video block is encoded and decoded in an affine mode; reordering the plurality of affine motion candidates in the affine motion candidate list based on at least one set of motion information associated with one or more previously transformed blocks to update the affine motion candidate list; and performing a transformation of the current video block according to the updated affine motion candidate list.
[0010] Furthermore, in a representative aspect, any one or more of the disclosed methods are encoder-side implementations.
[0011] And, in a representative aspect, any one or more of the disclosed methods are decoder-side implementations.
[0012] The above-described one or more methods are embodied in processor-executable code and stored in a computer-readable program medium.
[0013] In yet another representative aspect, a device in a video system is disclosed, which includes a processor and a non-transitory memory having instructions thereon. When the instructions are executed by the processor, the processor is caused to implement any one or more of the disclosed methods.
[0014] The above and other aspects and features of the disclosed technology are described in more detail in the drawings, the specification, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 An example of constructing a Merge candidate list is shown.
[0016] Figure 2 An example of the position of a spatial candidate is shown.
[0017] Figure 3 An example of a candidate pair for performing a redundancy check on a spatial Merge candidate is shown.
[0018] Figure 4A And Figure 4B An example of the position of a second prediction unit (PU) based on the size and shape of a current block is shown.
[0019] Figure 5Shows an example of motion vector scaling for temporal Merge candidates.
[0020] Figure 6 Shows an example of candidate positions for temporal Merge candidates.
[0021] Figure 7 Shows an example of generating combined bi - directional prediction Merge candidates.
[0022] Figure 8 Shows an example of constructing motion vector prediction candidates.
[0023] Figure 9 Shows an example of motion vector scaling for spatial motion vector candidates.
[0024] Figure 10 Shows an example of the encoding / decoding process for history - based motion vector prediction (HMVP) candidates.
[0025] Figure 11 Shows an example of updating a table by the HMVP method.
[0026] Figure 12 Shows an example of motion information prediction.
[0027] Figure 13 Shows an example of an affine motion model.
[0028] Figure 14 Shows an example of an affine motion vector field by sub - blocks.
[0029] Figure 15 Shows example candidate positions for the affine Merge mode.
[0030] Figure 16A and Figure 16B Show a 4 - parameter affine model and a 6 - parameter affine model, respectively.
[0031] Figure 17 Shows an example of neighboring motion vectors used in the derivation of RMVF motion information.
[0032] Figure 18 Shows an example of a reduced set of neighboring motion vectors used in the derivation of RMVF motion information.
[0033] Figure 19 Is a block diagram of an example of a hardware platform for implementing the visual media decoding or visual media encoding techniques described in this document.
[0034] Figure 20 Shows a flowchart of an example method for video processing.
[0035] Figure 21A flowchart showing another example method for video processing.
[0036] Figure 22 A flowchart showing another example method for video processing.
[0037] Figure 23 A flowchart showing another example method for video processing. Detailed implementation
[0038] 1. Video coding and decoding in HEVC / H.265
[0039] Video coding and decoding standards have evolved mainly through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed H.262 / MPEG-2 video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards. Since H.262, video coding and decoding standards have been based on a hybrid video coding and decoding structure, where temporal prediction plus transform coding is used. To explore future video coding and decoding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Exploration Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the VVC standard, with the goal of reducing the bit rate by 50% compared to HEVC.
[0040] 2.1 Inter-frame prediction in HEVC / H.265
[0041] Each inter-frame predicted PU has motion parameters for one or two reference picture lists. The motion parameters include motion vectors and reference picture indices. The use of one of the two reference picture lists can also be signaled using inter_pred_idc. The motion vectors can be explicitly coded and decoded as increments relative to a predictor.
[0042] When coding or decoding a CU in skip mode, a PU is associated with the CU and there are no significant residual coefficients, coded motion vector deltas, or reference picture indices. The Merge mode is defined, whereby the motion parameters of the current PU are obtained from neighboring PUs including both spatial and temporal candidates. The Merge mode can be applied to any inter-predicted PU, not just for skip mode. An alternative to the Merge mode is the explicit transmission of motion parameters, where each PU explicitly signals the motion vector (more precisely, the motion vector difference (MVD) compared to the motion vector predictor), the corresponding reference picture index for each reference picture list, and the use of the reference picture list. In this patent document, this mode is referred to as Advanced Motion Vector Prediction (AMVP).
[0043] When signaling indicates to use one of the two reference picture lists, a PU is generated from a single sample block. This is referred to as "unidirectional prediction". Unidirectional prediction can be used for P slices and B slices.
[0044] When signaling indicates to use both of the two reference picture lists, a PU is generated from two sample blocks. This is referred to as "bidirectional prediction". Bidirectional prediction can only be used for B slices.
[0045] The following text provides details on the inter-prediction modes defined in HEVC. The description will start with the Merge mode.
[0046] 2.1.1 Reference Picture Lists
[0047] In HEVC, the term inter-prediction is used to denote prediction derived from data elements (e.g., sample values or motion vectors) of reference pictures other than the current decoded picture. As in H.264 / AVC, a picture can be predicted from multiple reference pictures. The reference pictures used for inter-prediction are organized in one or more reference picture lists. A reference index identifies which reference picture in the list should be used to create the prediction signal.
[0048] A single reference picture list (list 0) is used for P slices, while two reference picture lists (list 0 and list 1) are used for B slices. In terms of capture / display order, the reference pictures included in list 0 / 1 can be pictures from the past and the future.
[0049] 2.1.2 Merge Mode
[0050] 2.1.2.1 Derivation of Merge Mode Candidates
[0051] When predicting a PU using the Merge mode, an index pointing to an entry in the Merge candidate list is parsed from the bitstream and used to retrieve motion information. The construction of this list is specified in the HEVC standard and can be outlined as a sequence of the following steps:
[0052] · Step 1: Initial candidate derivation
[0053] o Step 1.1: Spatial candidate derivation
[0054] o Step 1.2: Redundancy check of spatial candidates
[0055] o Step 1.3: Temporal candidate derivation
[0056] · Step 2: Additional candidate insertion
[0057] o Step 2.1: Create bi-prediction candidates
[0058] o Step 2.2: Insert zero-motion candidates
[0059] These steps are also schematically depicted in Figure 1 . For spatial Merge candidate derivation, up to four Merge candidates are selected among candidates located at five different positions. For temporal Merge candidate derivation, up to one Merge candidate is selected among two candidates. Since a constant number of candidates per PU is assumed at the decoder, additional candidates are generated when the number of candidates obtained from Step 1 does not reach the maximum number of Merge candidates (MaxNumMergeCand) signaled in the slice header. Since the number of candidates is constant, the index of the best Merge candidate is encoded using Truncated Unarybinarization (TU). If the size of the CU is equal to 8, all PUs of the current CU share a single Merge candidate list, which is the same as the Merge candidate list of a 2N×2N prediction unit.
[0060] In the following, the operations associated with the above steps will be described in detail.
[0061] 2.1.2.2 Spatial candidate derivation
[0062] In the derivation of spatial Merge candidates, among those located at Figure 2Among the candidates for the positions depicted, up to four Merge candidates are selected. The derivation order is A1, B1, B0, A0, and B2. Position B2 is considered only if any of the PUs at positions A1, B1, B0, A0 are unavailable (e.g., because it belongs to another strip or slice) or are intra-coded. After adding the candidate at position A1, a redundancy check is performed on the addition of the remaining candidates, which ensures that candidates with the same motion information are excluded from the list, thereby improving the coding and decoding efficiency. To reduce the computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the pairs linked by the arrow in Figure 3 are considered, and the candidate is added to the list only if the corresponding candidate used for the redundancy check does not have the same motion information. Another source of duplicate motion information is the "second PU" associated with a partition different from 2N×2N. As an example, Figure 4A and Figure 4B depict the cases where the second PU is N×2N and 2N×N. When the current PU is partitioned into N×2N, the candidate at position A1 is not considered for list construction. In fact, adding this candidate would result in two prediction units having the same motion information, which is redundant for a coding and decoding unit with only one PU. Similarly, when the current PU is partitioned into 2N×N, position B1 is not considered.
[0063] 2.1.2.3 Temporal Candidate Derivation
[0064] In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal Merge candidate, a scaled motion vector is derived based on the co-located PU belonging to the picture within the given reference picture list that has the smallest POC (Picture Order Count) difference from the current picture. The reference picture list to be used for deriving the co-located PU is signaled explicitly in the slice header. As shown by the dashed line in Figure 5 , the scaled motion vector for the temporal Merge candidate is obtained by scaling the motion vector of the co-located PU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal Merge candidate is set to be equal to zero. The actual implementation of the scaling process is described in the HEVC specification. For B slices, two motion vectors (one for reference picture list 0 and the other for reference picture list 1) are obtained and combined to generate a bi-predictive Merge candidate.
[0065] In the co-located PU (Y) belonging to the reference frame, the position of the temporal candidate is selected between candidates C0 and C1, as shown in Figure 6If the PU at position C0 is not available, is intra-coded, or is outside the current codec tree unit (CTU, also known as LCU, the largest codec unit) row, position C1 is used. Otherwise, position C0 is used in the derivation of the time domain merge candidate.
[0066] 2.1.2.4 Additional Candidate Insertion
[0067] In addition to spatial and temporal Merge candidates, there are two additional types of Merge candidates: combined bi-predictive Merge candidates and zero Merge candidates. Combined bi-predictive Merge candidates are generated by utilizing spatial and temporal Merge candidates. Combined bi-predictive Merge candidates are only used for B slices. Combined bi-predictive candidates are generated by combining the first reference picture list motion parameters of an initial candidate with the second reference picture list motion parameters of another. If the two tuples provide different motion hypotheses, they will form a new bi-predictive candidate. As an example, Figure 7 Shown is when two candidates in the original list (on the left) (which have mvL0 and refIdxL0 or mvL1 and refIdxL1) are used to create a combined bi-predictive Merge candidate that is added to the final list (on the right). There are many rules for combining that are considered to generate these additional Merge candidates.
[0068] Zero motion candidates are inserted to fill the remaining entries in the Merge candidate list and thus reach the MaxNumMergeCand capacity. These candidates have zero spatial displacement and a reference picture index that starts at zero and increases each time a new zero motion candidate is added to the list. Finally, no redundancy check is performed on these candidates.
[0069] 2.1.3AMVP
[0070] AMVP exploits the spatiotemporal correlation of motion vectors with neighboring PUs, which is used for explicit transmission of motion parameters. For each reference picture list, a motion vector candidate list is constructed by first checking the availability of the temporally neighboring PU positions to the left and above, removing redundant candidates and adding zero vectors to make the candidate list a constant length. The encoder can then select the best predictor from the candidate list and send the corresponding index indicating the selected candidate. Similar to Merge index signaling, the index of the best motion vector candidate is encoded using truncated unary. In this case, the maximum value to be encoded is 2 (see Figure 8 ). In the following sections, details about the derivation process of motion vector prediction candidates will be provided.
[0071] 2.1.3.1 Derivation of AMVP Candidates
[0072] Figure 8 Outlines the derivation process of motion vector prediction candidates.
[0073] In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. For spatial motion vector candidate derivation, ultimately two motion vector candidates are derived based on the motion vectors of each PU located at five different positions as depicted in Figure 2 Two motion vector candidates are derived based on the motion vectors of each PU located at five different positions as depicted in
[0074] For temporal motion vector candidate derivation, one motion vector candidate is selected from two candidates derived based on two different co-located positions. After generating the first spatio-temporal candidate list, duplicate motion vector candidates in the list are removed. If the number of potential candidates is greater than two, motion vector candidates with a reference picture index greater than 1 in the associated reference picture list are removed from the list. If the number of spatio-temporal motion vector candidates is less than two, additional zero motion vector candidates are added to the list.
[0075] 2.1.3.2 Spatial Motion Vector Candidates
[0076] In the derivation of spatial motion vector candidates, at most two candidates are considered among the five potential candidates derived from the PUs located at the positions as depicted in Figure 2 Those positions are the same as the positions for motion Merge. The derivation order for the left side of the current PU is defined as A0, A1, and scaled A0, scaled A1. The derivation order for the upper side of the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. Thus for each side, there are four cases that can be used as motion vector candidates, where two cases do not require the use of spatial scaling and two cases use spatial scaling. The four different cases are outlined as follows:
[0077] · Without spatial scaling
[0078] -(1) The same reference picture list and the same reference picture index (same POC)
[0079] -(2) Different reference picture lists, but the same reference picture (same POC)
[0080] · With spatial scaling
[0081] -(3) The same reference picture list, but different reference pictures (different POCs)
[0082] -(4) Different reference picture lists and different reference pictures (different POCs)
[0083] First, check the non-spatial domain scaling situation, and then the situation allowing spatial domain scaling. Spatial domain scaling is considered when the POC is different between the reference pictures of the neighboring PUs and the reference picture of the current PU regardless of the reference picture list. If all the left candidate PUs are unavailable or are intra-coded, allowing scaling for the upper motion vector helps in the parallel derivation of the left and upper MV candidates. Otherwise, spatial domain scaling is not allowed for the upper motion vector.
[0084] During the spatial domain scaling process, the motion vectors of neighboring PUs are scaled in a similar way to the temporal domain scaling, as Figure 9 depicted. The main difference is that the reference picture list and the index of the current PU are given as inputs; the actual scaling process is the same as that of the temporal domain scaling.
[0085] 2.1.3.3 Temporal Motion Vector Candidates
[0086] Except for the reference picture index derivation, all the processes for deriving temporal Merge candidates are the same as those for deriving spatial domain motion vector candidates (see Figure 6 ). The reference picture index is signaled to the decoder to indicate which reference pictures to use.
[0087] 2.2 Inter - frame Prediction Methods in VVC
[0088] There are several new coding tools for inter - frame prediction improvement, such as Adaptive motion vector difference resolution (AMVR) for signaling MVD, affine prediction mode, Triangle prediction mode (TPM), Advanced TMVP (ATMVP, also known as SbTMVP), Generalized Bi - direction prediction (GBI), and Bi - direction optical flow (BIO).
[0089] In VVC, the quadtree / binary tree / octree (QT / BT / TT) structure is adopted to divide the picture into square or rectangular blocks.
[0090] In addition to QT / BT / TT, a separate tree (also known as a dual - coding tree) is also adopted for I - frames in VVC. With the separate tree, the coding block structure is signaled separately for the luminance and chrominance components.
[0091] 2.2.1 History - based Motion Vector Prediction
[0092] A history-based MVP (HMVP) method is proposed, where HMVP candidates are defined as the motion information of previous coded / decoded blocks. During the encoding / decoding process, a table with multiple HMVP candidates is maintained. When a new slice is encountered, the table is cleared. Whenever there is an inter-coded / decoded block, the associated motion information is added to the last entry of the table as a new HMVP candidate. The entire coding / decoding process is as Figure 10 depicted.
[0093] In one example, the table size is set to 1 (e.g., L = 16 or 6, or 44), which indicates that up to L HMVP candidates can be added to the table.
[0094] (1) In one embodiment, if there are more than L HMVP candidates from previous coded blocks, the First-In-First-Out (FIFO) rule is applied so that the table always contains the latest L motion candidates from the previous coding. Figure 11 depicts an example of applying the FIFO rule to remove HMVP candidates and add new candidates to the table used in the proposed method.
[0095] (2) In another embodiment, whenever a new motion candidate is added (such as when the current block is inter-coded and in a non-affine mode), a redundancy check process is first applied to identify whether there are the same or similar motion candidates in the LUT.
[0096] 2.2.2 Sub-CU Based Motion Vector Prediction
[0097] In the encoder, two sub-CU level motion vector prediction methods are considered by dividing a large CU into sub-CUs and deriving the motion information of all sub-CUs of the large CU. The Alternate Time Motion Vector Prediction (ATMVP) method allows each CU to extract multiple sets of motion information from multiple blocks smaller than the current CU in the co-located reference picture. In the Spatial-Temporal Motion Vector Prediction (STMVP) method, the motion vectors of sub-CUs are recursively derived by using the temporal motion vector predictor and the spatial neighboring motion vectors.
[0098] To maintain a more accurate motion field for sub-CU motion prediction, motion compression of the reference frames is currently disabled.
[0099] 2.2.2.1 Alternate Time Motion Vector Prediction
[0100] In the optional temporal motion vector prediction (ATMVP) method, the motion vector of the temporal motion vector prediction (TMVP) is modified by extracting multiple sets of motion information (including motion vectors and reference indices) from blocks smaller than the current CU. As Figure 12 shown, the sub-CU is a square N×N block (N is default set to 4).
[0101] ATMVP predicts the motion vectors of sub-CUs within a CU in two steps. The first step is to identify the corresponding block in the reference picture using a so-called temporal vector. The reference picture is also referred to as the motion source picture. The second step is to divide the current CU into sub-CUs and obtain the motion vectors and the reference index for each sub-CU from the corresponding blocks, as Figure 12 shown.
[0102] In the first step, the reference picture and the corresponding block are determined by the motion information of the spatial neighboring blocks of the current CU. To avoid the repeated scanning process of neighboring blocks, the first Merge candidate in the Merge candidate list of the current CU is used. The first available motion vector and its associated reference index are set as the temporal vector and the index of the motion source picture. Thus, in ATMVP, the corresponding block can be identified more accurately compared to TMVP, where the corresponding block (sometimes called the co-located block) is always in the lower-right or center position relative to the current CU.
[0103] In the second step, the corresponding block of the sub-CU is identified by the temporal vector in the motion source picture by adding the temporal vector to the coordinates of the current CU. For each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the central sample) is used to derive the motion information of the sub-CU. After identifying the motion information of the corresponding N×N block, it is converted into the motion vector and reference index of the current sub-CU in the same way as in the TMVP of HEVC, where motion scaling and other processes apply. For example, the decoder checks whether the low-latency condition is satisfied (i.e., the POC of all reference pictures of the current picture is less than the POC of the current picture), and may use the motion vector MV x (e.g., the motion vector corresponding to reference picture list X) to predict the motion vector MV y (e.g., where X is equal to 0 or 1, and Y is equal to 1 - X).
[0104] 2.2.3 Affine motion compensation prediction
[0105] In HEVC, only the translational motion model is applied to Motion Compensation Prediction (MCP). However, in the real world, there are many types of motions, such as zooming in / out, rotation, perspective motion, and other irregular motions. The Simplified Affine Transform Motion Compensation Prediction is applied. As Figure 13 shown, the affine motion field of a block is described by two control point motion vectors.
[0106] The Motion Vector Field (MVF) of a block is described by the following equation:
[0107]
[0108] where (v 0x , v 0y ) is the motion vector of the upper left control point, and (v 1x , v 1y ) is the motion vector of the upper right control point.
[0109] To derive the motion vector of each M×N sub-block, as Figure 14 shown, the motion vector of the center sample of each sub-block is calculated according to Equation 1 and rounded to a fractional precision of 1 / 16.
[0110] There are two affine motion modes: AF_INTER mode and sub-block Merge mode.
[0111] 2.2.3.1. AF_INTER mode
[0112] For a CU with both width and height greater than 8, the AF_INTER mode can be applied. The affine flag at the CU level is signaled in the bitstream to indicate whether the AF_INTER mode is used. In the AF_INTER mode, the control point motion vector predictor (CPMVP) is derived in the affine AMVP list. The coordinates of the three control points CP1, CP2, and CP3 are (0,0), (W,0), and (H,0) respectively, where W and H are the width and height of the current block.
[0113] The affine AMVP list is constructed as follows:
[0114] 1) Insert the inherited affine AMVP candidates
[0115] The inherited affine candidates refer to candidates that are derived from the affine motion models of their valid neighboring affine coded blocks. Up to 2 candidates can be derived in this step. As Figure 15As shown, check A0 and A1 to generate the first candidate, and check B0, B1, and B2 to generate the second candidate.
[0116] For each available candidate, its affine parameters are used to derive the CPMVP, and the CPMVP is inserted into the affine AMVP list.
[0117] 2) Insert the constructed affine AMVP candidates
[0118] If the number of candidates in the affine AMVP list is less than the maximum affine AMVP list size (represented by MaxAffineAmvpListSize), then the constructed affine candidates are inserted into the candidate list. The constructed affine candidates refer to constructing the candidate CPMVP by combining the neighboring motion information of each control point.
[0119] The motion information of the control points is first derived from Figure 15 the specified spatial neighbors as shown. CPk (k = 1, 2, 3) represents the k-th control point. A0, A1, A2, B0, B1, B2, and B3 are the spatial positions for predicting CPk (k = 1, 2, 3).
[0120] The MVP for each control point is obtained according to the following priority order:
[0121] · For CP1, the check priority is B2 -> B3 -> A2. If B2 is available, use B2. Otherwise,
[0122] if B2 is not available, use B3. If both B2 and B3 are not available, use A2.
[0123] If all three candidates are not available, the motion information of CP1 cannot be obtained.
[0124] · For CP2, the check priority is B1 -> B0.
[0125] · For CP3, the check priority is A1 -> A0.
[0126] If the MVPs of CP1 and CP2 are available for the 4-parameter affine model, or the MVPs of CP1, CP2, and CP3 are available for the 6-parameter affine model, then the CPMVP is inserted into the affine AMVP list.
[0127] If the number of candidates in the affine AMVP list is still less than MaxAffineAmvpListSize, the same MVP can be assigned to all control points as follows until the size of the affine AMVP list is equal to MaxAffineAmvpListSize:
[0128] a) If the MVP of CP2 is available, set it as the MVP of CP1 and CP0, and insert this CPMVP into the affine AMVP list.
[0129] b) If the MVP of CP1 is available, set it as the MVP of CP0 and CP2, and insert this CPMVP into the affine AMVP list.
[0130] c) If the MVP of CP0 is available, set it as the MVP of CP1 and CP2, and insert this CPMVP into the affine AMVP list.
[0131] 3) Fill with zero motion vectors
[0132] If the number of candidates in the affine Merge candidate list is less than MaxAffineAmvpListSize, zero motion vectors with zero reference indices are used as the MVP for all control points until the list is full.
[0133] After determining the CPMVP of the current affine CU, apply affine motion estimation and find the control point motion vector (CPMV). Then, signal the difference between the CPMV and the CPMVP in the bitstream.
[0134] In AF_INTER mode, when using the 4 / 6 parameter affine mode, 2 / 3 control points are required, and thus 2 / 3 MVDs need to be coded and decoded for these control points, as Figure 16A shown. In the example, it is proposed to derive the MVs as follows: predict mvd1 and mvd2 from mvd0.
[0135]
[0136]
[0137]
[0138] where mvd i and mv1 are the predicted motion vector, motion vector difference, and motion vector of the upper left pixel (i = 0), upper right pixel (i = 1), or lower left pixel (i = 2), respectively, as Figure 16B shown. The addition of two motion vectors (e.g., mvA(xA,yA) and mvB(xB,yB)) is equal to the separate summation of the two components, i.e., newMV = mvA + mvB, and the two components of newMV are set to (xA + xB) and (yA + yB), respectively.
[0139] 2.2.3.2 Fast Affine ME Algorithm in AF_INTER Mode
[0140] In the affine mode, it is necessary to jointly determine the MVs of 2 or 3 control points. Directly jointly searching for multiple MVs is computationally complex. A fast affine ME algorithm is proposed and applied to VTM / BMS.
[0141] The fast affine ME algorithm is described for the 4-parameter affine model, and this idea can be extended to the 6-parameter affine model.
[0142]
[0143]
[0144] Replacing (a - 1) with a', the motion vector can be rewritten as:
[0145]
[0146] Assuming that the motion vectors of two control points (0, 0) and (0, w) are known, from equation (5), the affine parameters can be derived.
[0147]
[0148] The motion vector can be rewritten in vector form as:
[0149]
[0150] where
[0151]
[0152]
[0153] P = (x, y) is the pixel position.
[0154] At the encoder, the MVD of AF_INTER is iteratively derived. Denote the MV i (P) as the MV derived in the i-th iteration at position P, and denote dMV C i as the increment updated for the MV C in the i-th iteration. Then, in the (i + 1)-th iteration,
[0155]
[0156] Denote Pic ref as the reference picture, and denote Pic cur as the current picture, and denote Q = P + MV i(P). Assuming the use of MSE as the matching criterion, the following equation can be written:
[0157]
[0158] Assume is small enough, and it can be approximately rewritten using the first-order Taylor expansion as as follows.
[0159]
[0160] where represents E i+1 (P) = Pic cur (P) - Pic ref (Q),
[0161]
[0162] It can be derived by setting the derivative of the error function to zero Then, the incremental MVs of the control points (0, 0) and (0, w) can be calculated according to :
[0163]
[0164]
[0165]
[0166]
[0167] Assuming such an MVD derivation process is iterated n times, the final MVD is calculated as follows:
[0168]
[0169]
[0170]
[0171]
[0172] In the example, that is, the incremental MV of the control point (0, w) represented by mvd1 is predicted from the incremental MV of the control point (0, 0) represented by mvd0. Now, only mvd1 is actually encoded
[0173] 2.2.3.3 Sub-block Merge Mode
[0174] Construct sub - block merge candidates according to the following steps:
[0175] 1) Insert ATMVP candidates
[0176] 2) Insert inherited affine Merge candidates
[0177] Inherited affine Merge candidates refer to candidates that are derived from the affine motion models of their valid neighboring affine encoding - decoding blocks. Up to 2 candidates can be derived in this step. As Figure 15 shown, check A0, A1 to generate the first candidate, and check B0, B1, B2 to generate the second candidate.
[0178] For each available candidate, its affine parameters are used to derive CPMVP, and CPMVP is inserted into the sub - block Merge list.
[0179] 3) Insert constructed affine Merge candidates
[0180] If the number of candidates in the sub - block Merge candidate list is less than MaxNumAffineCand (set to 5 here), insert the constructed affine Merge candidates into the candidate list. Constructed affine Merge candidates refer to candidates constructed by combining the neighboring motion information of each control point.
[0181] The motion information of the control points is first derived from Figure 15 the specified spatial and temporal neighbors as shown. CPk (k = 1, 2, 3, 4) represents the k - th control point. A0, A1, A2, B0, B1, B2, and B3 are the spatial positions for predicting CPk (k = 1, 2, 3); T is the temporal position for predicting CP4.
[0182] The coordinates of CP1, CP2, CP3, and CP4 are (0, 0), (W, 0), (H, 0), and (W, H) respectively, where W and H are the width and height of the current block.
[0183] Obtain the motion information of each control point according to the following priority order:
[0184] · For CP1, the check priority is B2 -> B3 -> A2. If B2 is available, use B2. Otherwise, if B2 is not available, use B3. If both B2 and B3 are not available, use A2. If all three candidates are not available, the motion information of CP1 cannot be obtained.
[0185] · For CP2, the check priority is B1 -> B0.
[0186] · For CP3, the check priority is A1 -> A0.
[0187] ·For CP4, use T.
[0188] Secondly, use a combination of control points to construct an affine Merge candidate.
[0189] Constructing a 6-parameter affine candidate requires the motion information of three control points. The three control points can be selected from one of the following four combinations ({CP1, CP2, CP4}, {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4}).
[0190] Constructing a 4-parameter affine candidate requires the motion information of two control points. These two control points can be selected from one of the following two combinations ({CP1, CP2}, {CP1, CP3}).
[0191] The combinations of the constructed affine candidates are inserted into the candidate list in the following order:
[0192] {CP1, CP2, CP3}, {CP1, CP2, CP4}, {CP1, CP3, CP4}, {CP2, CP3, CP4}, {CP1, CP2}, {CP1, CP3}.
[0193] For the combined reference list X (X is 0 or 1), if different control points use the same reference picture, this combination is regarded as "usable", otherwise, it is regarded as "unusable".
[0194] 4) Fill with zero motion vectors
[0195] If the number of candidates in the sub-block Merge candidate list is less than 5, zero motion vectors with zero reference indices are inserted into the candidate list until the list is full.
[0196] 2.3 Regression-based Motion Vector Field
[0197] In the example, a sub-block motion vector derivation based on a regression-based motion vector field (RMVF) is proposed. This tool attempts to model the motion vector of each block at the sub-block level based on neighboring motion vectors in the spatial domain.
[0198] Figure 17 The neighboring 4×4 motion blocks for the motion parameter derivation of the proposed method are shown. As shown, during the regression process, the neighboring motion vectors of one line and one row (and their central positions) based on 4×4 sub-blocks from each side of the block are used.
[0199] To reduce the amount of neighboring motion information used for RMVF parameter derivation, use Figure 18A method in which almost half of the neighboring 4×4 motion blocks are used for motion parameter derivation.
[0200] When collecting motion information for motion parameter derivation, five conventional regions (lower left, left, upper left, upper, upper right) as shown in the figure are used. The upper right (with a length of W / 2) and lower left (with a length of H / 2) reference motion regions are limited to half of the corresponding width or height of the current block.
[0201] In the RMVF mode, the motion of a block is defined by a 6-parameter motion model. These parameters a xx , a xy , a yx , a yy , b x and b y are calculated by solving a linear regression model in the sense of mean square error (MSE). The inputs to the regression model include the center positions (x, y) of the available neighboring 4×4 sub-blocks described above and the motion vectors (mv x and mv y ).
[0202] Then, the motion vectors (MV subPU , MV subPU ) of the 8×8 sub-block with the center position at (X X_subPU , Y Y_subPU ) are calculated as:
[0203]
[0204] The motion vectors are calculated for the 8×8 sub-blocks with respect to the center position of each sub-block. Therefore, in the RMVF mode, motion compensation is also applied with 8×8 sub-block precision.
[0205] To effectively model motion information (e.g., motion vector field), RMVF can be applied only when at least one motion vector from at least three candidate regions is available.
[0206] Further details:
[0207] · The minimum decoding unit size of the RMVF mode is 8×8.
[0208] · The determination of uni-directional or bi-directional prediction is made based on the picture type (P / B picture) and the availability of neighboring motion information in that prediction direction.
[0209] · The reference indices in both prediction directions are set to 0.
[0210] · When calculating motion parameters, apply the conventional motion vector scaling used in VTM-3.0 to motion vectors for which the reference picture is different from the current block.
[0211] · Complete sub-block motion compensation according to the regulations in VTM-3.0 (i.e., 8×8 sub-block size).
[0212] 3. Disadvantages of existing RMVF implementations
[0213] The current RMVF method may have the following problems:
[0214] 1. It cannot be applied to the affine inter prediction mode.
[0215] 2. It uses neighboring motion vectors to derive affine parameters. For small-resolution sequences such as WVGA, WQVGA, or 1080P sequences or smaller block sizes, there are few available MVs for the block, and RMVF may not be efficient.
[0216] 3. Only derive 6-parameter affine parameters
[0217] 4. Example techniques and embodiments
[0218] The detailed embodiments described below should be regarded as examples to explain the general concept. These embodiments should not be interpreted narrowly. In addition, these embodiments can be combined in any way.
[0219] The RMVF embodiments described below can represent the techniques described in Section 2.3, or any variation of RMVF or any codec tool that depends on the motion information of neighboring blocks to derive a motion model for encoding and decoding the current block.
[0220] I. RMVF can utilize motion information from non-adjacent blocks.
[0221] a. In one example, RMVF can utilize the codec information of a temporal block located in the reference picture.
[0222] b. Additionally, alternatively, the non-adjacent blocks should be within the same CTU.
[0223] c. Additionally, alternatively, the non-adjacent blocks should be in the same CTU row.
[0224] d. In one example, the neighboring / non-adjacent / temporal blocks to be utilized can be changed between blocks.
[0225] e. In one example, whether to utilize non-adjacent / temporal blocks in RMVF can depend on the availability of motion information from neighboring blocks.
[0226] 2. RMVF can be used for the inter - frame affine mode, where some control point predictors can be derived from the RMVF.
[0227] a. In one example, the RMVF method can be used to derive the CPMVP, and the derived CPMVP (i.e., the RMVF - affine AMVP candidate) can be inserted into the existing affine AMVP list.
[0228] i. Alternatively, a new affine AMVP list can be constructed based on the CPMVP derived from the RMVF. Additionally, alternatively, it can be signaled explicitly or implicitly at the block / CU / CTU / slice level whether to use the existing affine AMVP list together with the candidates derived from the RMVF or to use a separate affine AMVP list.
[0229] b. In one example, the 4 - parameter affine parameters can be derived by RMVF and can be used to generate the CPMVP for all control points of the block.
[0230] c. In one example, the 6 - parameter affine parameters can be derived by RMVF and can be used to generate the CPMVP for all control points of the block.
[0231] d. In one example, RMVF can derive more than one set of 4 - parameter affine parameters or / and 6 - parameter affine parameters by using different groups of neighboring adjacent or / and non - adjacent MVs.
[0232] e. In one example, the CPMVP derived by RMVF can be inserted before the inherited affine AMVP candidates.
[0233] f. In one example, the CPMVP derived by RMVF can be inserted after some of the inherited affine AMVP candidates or all of the inherited affine candidates.
[0234] i. Additionally, alternatively, before the constructed affine AMVP candidates.
[0235] ii. The CPMVP derived by RMVF can be interleaved with the inherited affine AMVP candidates.
[0236] g. In one example, the CPMVP derived by RMVF can be inserted after some of the constructed affine AMVP candidates or all of the constructed affine AMVP candidates.
[0237] i. Additionally, alternatively, before the zero MV.
[0238] ii. The CPMVP derived by RMVF can be interleaved with the constructed affine AMVP candidates.
[0239] h. The CPMVP derived through RMVF can be interleaved with zero MVs.
[0240] 3. The RMVF method can be used to derive the CPMVP, and the derived CPMVP (i.e., the RMVF affine Merge candidate) can be inserted into the sub-block Merge list / affine Merge list.
[0241] a. In one example, 4-parameter affine parameters can be derived through RMVF. Additionally, alternatively, those parameters can be used to generate the CPMVP for all control points of a block.
[0242] b. In one example, 6-parameter affine parameters can be derived through RMVF. Additionally, alternatively, those parameters can be used to generate the CPMVP for all control points of a block.
[0243] c. In one example, RMVF can derive more than one set of 4-parameter affine parameters or / and 6-parameter affine parameters by using different groups of neighboring adjacent or / and non-adjacent MVs.
[0244] d. In one example, the prediction direction of the RMVF affine Merge candidate can depend on the slice type.
[0245] i. For example, for a P slice, the prediction direction can be list 0.
[0246] ii. For example, for a B slice, the prediction direction can be list 0 and list 1.
[0247] e. In one example, the reference picture of the RMVF affine Merge candidate can always be the first reference picture in each reference list.
[0248] i. Alternatively, the reference picture of the RMVF affine Merge candidate can be signaled in the VPS / SPS / PPS / picture group header / picture header / slice header, etc.
[0249] f. In one example, the CPMVP derived through RMVF can be inserted before the inherited affine Merge candidate.
[0250] g. In one example, the CPMVP derived through RMVF can be inserted after some or all of the inherited affine Merge candidates.
[0251] i. Additionally, alternatively, before the constructed affine Merge candidate.
[0252] ii. The CPMVP derived through RMVF can be interleaved with the inherited affine Merge candidate.
[0253] h. In one example, the CPMVP derived by RMVF can be inserted after some constructed affine Merge candidates or all constructed affine Merge candidates.
[0254] i. Additionally, alternatively, before zero MV.
[0255] ii. The CPMVP derived by RMVF can be interleaved with the constructed affine Merge candidates.
[0256] i. The CPMVP derived by RMVF can be interleaved with zero MV.
[0257] 4. Motion information (such as CPMV from a previous decoded block) and / or the affine model parameters stored in the HMVP lookup table can be used to derive affine parameters using the RMVF method.
[0258] a. Alternatively, an independent MV lookup table (referred to as the RMVF lookup table) can be maintained for RMVF, similar to the HMVP lookup table.
[0259] b. In one example, the neighboring adjacent or / and non - adjacent MVs or / and the MVs stored in the HMVP lookup table or / and the MVs stored in the RMVF lookup table can be used to derive affine parameters using the RMVF method.
[0260] c. In one example, only the MVs stored in the RMVF lookup table can be used to derive affine parameters using the RMVF method.
[0261] d. In one example, only the MVs stored in the HMVP lookup table can be used to derive affine parameters using the RMVF method.
[0262] 5. The motion information of neighboring and / or non - adjacent blocks, or / and the motion information stored in the HMVP lookup table or / and the motion information stored in the RMVF lookup table can be used to re - order the affine motion candidates in a candidate list (such as the affine Merge candidate list).
[0263] a. In one example, for each affine Merge candidate, its affine parameters are used to derive the MVs at K (K >= L) neighboring adjacent or / and non - adjacent positions. For these positions, the distance between the derived MV (DMV) and the signaled MV (SMV) is calculated, denoted as dist at position i i (i <= K). The cost value is based on dist [[ID=;3]] i and is calculated, for example, as cost = f(dist1,..., dist K ). The affine Merge candidates can be sorted in ascending order of the cost value.
[0264] b. In one example, the distance between DMV and SMV can be defined as the sum of the squared differences between DMV and SMV, as follows:
[0265] dist(DMV, SMV) = (DMV x - DMV x ) 2 + (DMV y - SMV y ) 2 .
[0266] Where Mv x and Mv y are the horizontal and vertical components of MV.
[0267] c. In one example, the distance between DMV and SMV can be defined as the sum of the absolute differences between DMV and SMV, as follows:
[0268] dist(DMV, SMV) = |DMV x - SMV x | + |DMV y - SMV y |.
[0269] d. In one example, the cost value is defined as the sum of dist i .
[0270] e. In one example, fixed neighboring adjacent or / and non - adjacent positions are predefined and used.
[0271] i. If the block in a position is not encoded or decoded using the inter - frame mode, it may be considered "unavailable" and thus not used.
[0272] ii. In one example, non - adjacent positions can be the positions of MVs stored in the HMVP lookup table or / and the RMVF lookup table.
[0273] f. If the prediction direction or / and the reference picture of the affine Merge candidate is different from that of the neighboring block, MV scaling can be performed to generate a derived MV. The MV scaling process can depend on the POC (Picture Order Count) of the reference picture of the affine Merge candidate and the POC of the reference picture of the neighboring block.
[0274] 6. If the neighboring adjacent or / and non - adjacent MV reference has a different reference picture from the target reference picture, it may be considered "unavailable" and may not be used for RMVF derivation.
[0275] 7. RMVF can be enabled / disabled according to rules regarding block dimensions.
[0276] a. In one example, when the block size contains fewer than M*H samples (e.g., 16 or 32 or 64 luma samples), RMVF is not allowed.
[0277] b. In one example, when the block size contains more than M*H samples (e.g., 16 or 32 or 64 luma samples), RMVF is not allowed.
[0278] c. Alternatively, when the minimum dimension of the width or height of the block is less than or not greater than X, RMVF is not allowed. In one example, X is set to 8.
[0279] d. Alternatively, when the width of the block > th1 or >= th1 and / or the height of the block > th2 or >= th2, then RMVF is not allowed. In one example, X is set to 64.
[0280] i. For example, RMVF is disabled for MxM (e.g., 128x128) blocks.
[0281] ii. For example, RMVF is disabled for NxM / MxN blocks, e.g., where N >= 64 and M = 128.
[0282] iii. For example, RMVF is disabled for NxM / MxN blocks, e.g., where N >= 4 and M = 128.
[0283] e. Alternatively, when the width of the block < th1 or <= th1 and / or the height of the block < th2 or <= th2, then RMVF is not allowed. In one example, th1 and / or th2 are set to 8.
[0284] 8. An indication for informing whether the proposed method is enabled can be signaled in the PPS / VPS / picture header / slice header / slice group header / CTU group.
[0285] a. In one example, if a separate tree partitioning structure (e.g., dual tree) is applied, then for a picture / slice / slice group / CTU, this indication can be signaled multiple times.
[0286] b. In one example, if a separate tree partitioning structure (e.g., dual tree) is applied, then this indication can be signaled separately for different color components.
[0287] i. Alternatively, this indication is signaled only for one color component and applied to all other color components.
[0288] 5 Example Embodiments of the Disclosed Technology
[0289] Figure 19It is a block diagram of a video processing apparatus 1900. The apparatus 1900 can be used to implement one or more methods described herein. The apparatus 1900 can be embodied in a smart phone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 1900 can include one or more processors 1902, one or more memories 1904, and video processing hardware 1906. The (multiple) processors 1902 can be configured to implement one or more methods described in this document. The memory (multiple memories) 2704 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 1906 can be used to implement some of the techniques described in this document in hardware circuits and can be partially or fully part of the processor 1902 (e.g., a graphics processing unit core GPU or other signal processing circuit).
[0290] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. As defined by grammar, the bitstream representation of the current video block can, for example, correspond to bits co-located or scattered at different positions within the bitstream. For example, a macroblock can be encoded based on the transform and coding / decoding error residual values and also using bits in the header and other fields in the bitstream.
[0291] It should be understood that the disclosed methods and techniques will be beneficial to video encoder and / or decoder embodiments incorporated within video processing devices such as smart phones, laptop computers, desktop computers, and similar devices by allowing the use of the techniques disclosed in this document.
[0292] Figure 20 It is a flowchart of an example method 2000 for video processing. The method 2000 includes: at 2010, deriving at least one motion model for a current video block based on motion information of at least one non-adjacent spatial neighboring block or at least one temporal neighboring block of the current video block. The method further includes: at 2020, deriving motion information for the current video block or at least one sub-block of the current video block based on the at least one motion model; and at 2030, performing a transform of the current video block based on the derived motion information.
[0293] Figure 21It is a flowchart of an example method 2100 for video processing. Method 2100 includes: at 2110, deriving one or more control point motion vector predictors (CPMVPs) of an affine model for a current video block from at least one set of neighboring blocks according to a back-off based motion vector field (RMVF) scheme; at 2120, updating a motion candidate list for the current video block based on the one or more CPMVPs, wherein the one or more CPMVPs are associated with a specific affine motion pattern; and at 2130, performing a transformation on the current video block based on the motion candidate list.
[0294] Figure 22 It is a flowchart of an example method 2200 for video processing. Method 2200 includes: at 2210, deriving at least one set of affine parameters of an affine model for a current video block by using a back-off based motion vector field (RMVF) scheme based at least on at least one set of motion information stored in a lookup table; at 2220, performing a transformation on the current video block based on the at least one set of affine parameters.
[0295] Figure 23 It is a flowchart of an example method 2300 for video processing. Method 2300 includes: at 2310, maintaining an affine motion candidate list for a current video block, the affine motion candidate list including a plurality of affine motion candidates, wherein the current video block is encoded and decoded in an affine mode; at 2320, reordering the plurality of affine motion candidates in the affine motion candidate list based on at least one set of motion information associated with one or more previously transformed blocks to update the affine motion candidate list; at 2330, performing a transformation on the current video block based on the updated affine motion candidate list.
[0296] Some embodiments can be described using the following examples.
[0297] In one aspect, a method for video processing is disclosed, including: deriving at least one motion model for a current video block based on motion information of at least one non-adjacent spatial neighboring block or at least one temporal neighboring block of the current video block; deriving motion information of the current video block or at least one sub-block of the current video block based on the at least one motion model; performing a transformation on the current video block based on the derived motion information.
[0298] In an example, the motion model includes a 4-parameter affine model.
[0299] In an example, the motion model includes a 6-parameter affine model.
[0300] In an example, the motion model is derived according to a back-off based vector field (RMVF) scheme.
[0301] In an example, at least one temporal neighboring block is located in at least one reference picture.
[0302] In an example, at least one non-adjacent spatial neighboring block is located within the same coding tree unit (CTU) as the current video block.
[0303] In an example, at least one non-adjacent spatial neighboring block is located within the same coding tree unit (CTU) row as the current video block.
[0304] In an example, at least one non-adjacent spatial neighboring block and at least one temporal neighboring block are determined based on the characteristics of the current video block.
[0305] In an example, whether to use at least one non-adjacent spatial neighboring block or at least one temporal neighboring block to derive at least one motion model depends on the availability of the motion information of one or more adjacent spatial neighboring blocks of the current video block.
[0306] In an example, whether to enable the fallback-based vector field (RMVF) scheme to derive at least one motion model depends on the block dimension of the current video block.
[0307] In an example, if the block dimension of the current video block satisfies at least one of the following, the RMVF scheme is disabled:
[0308] The number of samples included in the current video block is less than the first threshold;
[0309] The minimum dimension of the height and width of the current video block is not greater than the second threshold;
[0310] The height of the current video block is not greater than the third threshold; and
[0311] The width of the current video block is not greater than the fourth threshold.
[0312] In an example, the first threshold is equal to one of 16, 32, and 64.
[0313] In an example, the second threshold is equal to 8.
[0314] In an example, at least one of the third threshold and the fourth threshold is equal to 8.
[0315] In an example, if the block dimension of the current video block satisfies at least one of the following, the RMVF scheme is disabled:
[0316] The number of samples included in the current video block is greater than the fifth threshold;
[0317] The height of the current video block is not less than the sixth threshold; and
[0318] The width of the current video block is not less than the seventh threshold.
[0319] In an example, the fifth threshold is equal to one of 16, 32, and 64.
[0320] In an example, at least one of the sixth threshold and the seventh threshold is equal to 64.
[0321] In an example, if the size of the current video block is M×M, M×N, or N×M, the RMVF scheme is disabled, where M = 128, and N = 64 or 4.
[0322] In an example, whether to enable the fallback-based motion vector field (RMVF) scheme to derive at least one motion model depends on an indication signaled in at least one of a video parameter set (VPS), a picture parameter set (PPS), a picture header, a slice header, a slice group header, and a CTU group.
[0323] In an example, if a separate tree segmentation structure is applied to the current video block, the indication is signaled more than once for at least one of a picture, a slice, a slice group, and a CTU covering the current video block.
[0324] In an example, if a separate tree segmentation structure is applied to different color components, the indication is signaled separately for different color components of the current video block.
[0325] In an example, at least one non-adjacent spatial neighboring block or at least one temporal neighboring block does not include any neighboring block whose motion vector reference is different from the target reference picture.
[0326] On the other hand, a method for video processing is disclosed, including: deriving one or more control point motion vector predictors (CPMVPs) of an affine model of a current video block from at least one set of neighboring blocks in a fallback-based motion vector field (RMVF) scheme; updating a motion candidate list of the current video block based on the one or more CPMVPs, where the one or more CPMVPs are associated with a specific affine motion pattern; and performing a transformation of the current video block based on the motion candidate list.
[0327] In an example, at least one set of neighboring blocks includes at least one of an adjacent block and a non-adjacent block of the current video block.
[0328] In an example, the specific affine motion pattern is an affine advanced motion vector prediction (AMVP) pattern, and the motion candidate list is an affine AMVP list, and the one or more CPMVPs are inserted as RMVF-based affine candidates.
[0329] In an example, based on an indication signaled at a level of at least one of a block, a coding unit (CU), a coding tree unit (CTU), and a slice, an affine AMVP list is selected from one of an existing affine AMVP list and a separate affine AMVP list, where the existing affine AMVP list includes at least one affine AMVP candidate not derived by the RMVF scheme, and the separate affine AMVP list includes only affine candidates based on RMVF.
[0330] In an example, a specific affine motion mode is an affine Merge mode, and a motion candidate list is an affine Merge list, and one or more CPMVPs are inserted as affine candidates based on RMVF.
[0331] In an example, the affine Merge list is based on a sub-block-based affine Merge list.
[0332] In an example, the motion candidate list includes one or more inherited affine candidates, and the affine candidates based on RMVF are arranged in the motion candidate list in one of the following orders:
[0333] Before all the inherited affine candidates;
[0334] After at least one of the one or more inherited affine candidates; or
[0335] Interleaved with the one or more inherited affine candidates.
[0336] In an example, the motion candidate list further includes one or more constructed affine candidates, and the affine candidates based on RMVF are arranged in the motion candidate list in one of the following orders:
[0337] Before all the constructed affine candidates;
[0338] After at least one of the one or more constructed affine candidates; or
[0339] Interleaved with the one or more constructed affine candidates.
[0340] In an example, the motion candidate list further includes one or more zero motion vectors (MVs), and the affine candidates based on RMVF are arranged in the motion candidate list before all the zero MVs or interleaved with the zero MVs.
[0341] In an example, in the RMVF scheme, at least one set of affine parameters of an affine model of a current video block is determined from at least one set of neighboring blocks to derive one or more CPMVPs of the current video block.
[0342] In an example, in the RMVF scheme, more than one set of affine parameters of an affine model for a current video block are determined from neighboring blocks in different groups to derive one or more CPMVPs for the current video block.
[0343] In an example, the affine model is a 6-parameter affine model or a 4-parameter affine model.
[0344] In an example, each RMVF-based affine candidate is associated with a prediction direction depending on the slice type.
[0345] In an example, the prediction direction corresponds to list 0 of P-type slices.
[0346] In an example, the prediction direction corresponds to list 0 and list 1 of B-type slices.
[0347] In an example, each RMVF-based affine candidate is associated with a reference picture included in a reference list.
[0348] In an example, the reference picture corresponds to the first reference picture in the reference list.
[0349] In an example, the reference picture is determined from indications signaled in a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a slice group header, a slice header, or a strip header.
[0350] In another aspect, a method for video processing is disclosed, including: deriving at least one set of affine parameters of an affine model for a current video block by utilizing a fallback-based motion vector field (RMVF) scheme based at least on at least one set of motion information stored in a lookup table; performing a transformation of the current video block based on the at least one set of affine parameters.
[0351] In an example, the at least one set of affine parameters of the affine model is derived based on at least one set of motion information stored in a lookup table and at least one set of motion information associated with at least one adjacent or non-adjacent neighboring block.
[0352] In an example, the lookup table includes a history-based motion vector prediction (HMVP) lookup table or an RMVF lookup table that stores motion information of at least one previously transformed block.
[0353] In an example, the HMVP lookup table is different from the RMVF lookup table.
[0354] In an example, information about the position and dimension of at least one block is stored in the lookup table together with the motion information of at least one block.
[0355] In an example, the set of motion information includes at least one of a control point motion vector predictor and an affine model parameter.
[0356] On the other hand, a method for video processing is disclosed, including: maintaining an affine motion candidate list for a current video block, the affine motion candidate list including a plurality of affine motion candidates, wherein the current video block is encoded and decoded in an affine mode; reordering the plurality of affine motion candidates in the affine motion candidate list based on at least one set of motion information associated with one or more previously transformed blocks to update the affine motion candidate list; and performing transformation of the current video block based on the updated affine motion candidate list.
[0357] In an example, the plurality of affine motion candidates are associated with a 4-parameter affine motion model or a 6-parameter affine motion model.
[0358] In an example, the one or more previously transformed blocks are one or more adjacent or non-adjacent neighboring blocks.
[0359] In an example, at least one set of motion information associated with one or more previously transformed blocks is stored in a lookup table.
[0360] In an example, the lookup table includes a history-based motion vector prediction (HMVP) lookup table or an RMVF lookup table that stores at least one set of motion information of one or more previously transformed blocks.
[0361] In an example, the HMVP lookup table is different from the RMVF lookup table.
[0362] In an example, information about the positions and dimensions of one or more previously transformed blocks is stored in the lookup table together with at least one set of motion information.
[0363] In an example, at least one set of motion information includes at least one set of motion vectors (MVs), and the plurality of affine motion candidates are reordered based on the ascending order of function values associated with the at least one set of MVs.
[0364] In an example, the function value is obtained as follows:
[0365] Derive motion vectors associated with K neighboring blocks from the affine parameters of each of the plurality of affine motion candidates. For i = 1 to K, the derived motion vectors are denoted as DMV i .
[0366] Determine each derived motion vector DMV i and the difference dist i between the derived motion vector and the signaled motion vector SMV of the associated neighboring block i ; and
[0367] Obtain the function value based on the function f(dist1, dist2,..., dist K ).
[0368] In the example, the function f(dist1, dist2, …, dist K ) = dist1 + dist2 + … + dist K .
[0369] In the example, dist i = (DMV i_x - SMV i_x )^2 + (DMV i_y - SMV i_y ), where DMV i_x and DMV i_y respectively represent the horizontal and vertical components of DMV i , and SMV i_x and SMV i_y respectively represent the horizontal and vertical components of SMV i .
[0370] In the example, dist i (DMV i , SMV i ) = |DMV i_x - SMV i_x | + |DMV i_y - SMV i_y |,
[0371] where DMV i_x and DMV i_y respectively represent the horizontal and vertical components of DMV i , and SMV i_x and SMV i_y respectively represent the horizontal and vertical components of SMV i .
[0372] In the example, the K neighboring blocks are predefined and include at least one of adjacent neighboring blocks and non - adjacent neighboring blocks.
[0373] In the example, the set of MVs is associated with the blocks stored in the lookup table.
[0374] In the example, the method further includes: if the affine Merge candidate has a different prediction direction and / or different reference picture from the neighboring block, perform scaling on the derived motion vector.
[0375] In the example, perform scaling based on the picture order count (POC) of the reference picture of the affine Merge candidate and the POC of the reference picture of the neighboring block.
[0376] In an example, the conversion includes encoding a current video block into a bitstream representation of the video and decoding the current video block from the bitstream representation of the video.
[0377] On the other hand, a device in a video system is disclosed. The device includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to implement the above method.
[0378] On the other hand, a computer program product stored on a non-transitory computer-readable medium is disclosed. The computer program product includes program code for performing the above method.
[0379] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in a combination of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a material composition implementing a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatuses, devices, and machines for processing data, e.g., including programmable processors, computers, or multiple processors or computers. In addition to hardware, the apparatus can also include code for creating a runtime environment for the computer program being discussed, e.g., code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device.
[0380] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a part of a file that holds other programs or data (e.g., one or more scripts in a markup language document), stored in a single file dedicated to the program being discussed, or stored in multiple coordinated files (e.g., files that hold one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[0381] The processes and logical flows described herein can be performed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[0382] By way of example, processors suitable for the execution of a computer program include both general and special purpose microprocessors, and any one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, e.g., magnetic disks, magneto-optical disks, or optical disks, or be operatively coupled to receive data from one or more mass storage devices or transfer data to one or more mass storage devices or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, e.g., including semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0383] Although this patent document contains many details, these should not be construed as limitations on any subject or the scope of what is claimed, but rather as descriptions of features that are specific to particular embodiments of a particular technology. Certain features that are described in this patent document in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments. In addition, although the above features may be described as acting in certain combinations and even initially claimed as such, in some cases, one or more features from a claimed combination can be deleted from the combination, and the claimed combination can be directed to a sub-combination or a variant of a sub-combination.
[0384] Similarly, although operations are depicted in the drawings in a particular order, this should not be understood as requiring that the operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In addition, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0385] Only a few embodiments and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A method for video processing, comprising: Maintaining an affine motion candidate list for a current video block, the affine motion candidate list including a plurality of affine motion candidates, wherein the current video block is encoded and decoded in an affine mode; Reordering the plurality of affine motion candidates in the affine motion candidate list based on at least one set of motion information associated with one or more previously transformed blocks to update the affine motion candidate list; and Performing transformation of the current video block based on the updated affine motion candidate list, wherein the at least one set of motion information includes at least one set of motion vectors MV, and the plurality of affine motion candidates are reordered based on an ascending order of function values associated with the at least one set of MVs, wherein the method further comprises: Deriving one or more control point motion vector predictors CPMVPs of an affine model of the current video block from at least one set of neighboring blocks according to a fallback-based motion vector field RMVF scheme; Updating the affine motion candidate list based on the one or more CPMVPs, wherein the one or more CPMVPs are associated with a specific affine motion pattern.
2. The method according to claim 1, wherein The plurality of affine motion candidates are associated with a 4-parameter affine motion model or a 6-parameter affine motion model.
3. The method according to claim 1 or 2, wherein The one or more previously transformed blocks are one or more adjacent or non-adjacent neighboring blocks.
4. The method according to claim 1 or 2, wherein At least one set of motion information associated with the one or more previously transformed blocks is stored in a lookup table.
5. The method according to claim 4, wherein The lookup table includes a history-based motion vector prediction HMVP lookup table or an RMVF lookup table storing at least one set of motion information of one or more previously transformed blocks.
6. The method according to claim 5, wherein The HMVP lookup table is different from the RMVF lookup table.
7. The method according to claim 4, wherein Information about the positions and dimensions of the one or more previously transformed blocks is stored in the lookup table together with the at least one set of motion information.
8. The method according to claim 1, wherein The function value is obtained as follows: Derive motion vectors associated with K neighboring blocks from the affine parameters of each of the plurality of affine motion candidates, and for i = 1 to K, the derived motion vectors are denoted as DMV i ; Determine each derived motion vector DMV i and the signaled motion vector SMV of an associated neighboring block i the difference dist between them i ; and Based on the function f(dist1, dist2, …, dist K ) to obtain the function value.
9. The method according to claim 8, wherein The function f(dist1, dist2, …, dist K ) = dist1 + dist2 + … + dist K .
10. The method according to claim 8 or 9, wherein dist i =(DMV i_x -SMV i_x )^2+(DMV i_y -SMV i_y )^2, where DMV i_x and DMV i_y respectively represent the horizontal and vertical components of DMV i and SMV i_x and SMV i_y respectively represent the horizontal and vertical components of SMV i and 11. The method according to claim 8 or 9, wherein dist i (DMV i ,SMV i )=|DMV i_x -SMV i_x |+|DMV i_y -SMV i_y |, Among them, DMV i_x and DMV i_y Represents DMV i The horizontal and vertical components, and SMV i_x and SMV i_y Represents SMV respectively i The horizontal and vertical components of .
12. The method according to claim 8 or 9, wherein The K neighboring blocks are predefined and include at least one of adjacent neighboring blocks and non-adjacent neighboring blocks.
13. The method according to claim 12, wherein, The set of MVs is associated with blocks stored in a lookup table.
14. The method according to claim 8 or 9, further comprising: If an affine Merge candidate has a different prediction direction and / or a different reference picture from a neighboring block, performing scaling on the derived motion vector.
15. The method according to claim 14, wherein Performing scaling based on the picture order count POC of the reference picture of the affine Merge candidate and the POC of the reference picture of the neighboring block.
16. The method according to claim 1, wherein The transformation includes encoding the current video block into a bitstream of a video and decoding the current video block from the bitstream of the video.
17. An apparatus in a video system, the apparatus including a processor and a non-transitory memory having instructions thereon, wherein, The instructions, when executed by the processor, cause the processor to implement the method according to any one of claims 1 to 16.
18. A non-transitory computer-readable medium storing program instructions, the program instructions implementing program code of the method according to any one of claims 1 to 16 when executed by a computer.
Citation Information
Patent Citations
Method and apparatus for affine inter prediction for video coding system
CN108432250A
Motion vector prediction
US20180359483A1
Motion vector generation for affine motion model for video coding
WO2018126163A1