Affine candidate refinement
Through the method of regressing the affine candidates, linear models are derived to generate predictive candidates, which solves the problem of difficult to efficiently handle complex motion and texture features in the prior art, and achieves more efficient and high-quality video encoding and codec performance.
Patent Information
- Application Number
- CN202380055868.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-03
- Filing Date
- 2023-07-18
- Publication Date
- 2025-05-13
AI Technical Summary
Existing video encoding and codec technology is difficult to achieve efficient encoding and codec performance when processing complex motion features and texture features.
Through the method of regressing the affine candidates, a linear model is derived to generate predictive candidates, thereby achieving efficient encoding and decoding of video blocks. The method includes using an affine motion model, refine affine candidates, generate refined prediction candidates, and applying these prediction candidates during encoding or decoding.
Improves the efficiency and quality of video encoding and decoding, especially in scenarios where complex motion and texture features are handled, reducing codec artifacts and improving codec performance.
Smart Images

Figure CN119999186A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a video processing method and device in a video encoding and decoding system. In particular, the present invention relates to encoding and decoding pixel blocks through inter-picture prediction using affine candidates. Background Art
[0002] High-Efficiency Video Coding (HEVC) is an international video codec standard developed by the Joint Collaborative Team on Video Coding (JCT-VC). HEVC is based on a motion-compensated DCT-like transform codec architecture based on mixed pixel blocks. The basic unit of compression is called a coding unit (CU), which is a 2Nx2N square pixel block. Each CU can be recursively divided into four smaller CUs until a predefined minimum size is reached. Each CU contains one or more prediction units (PUs).
[0003] Versatile video coding (VVC) is the latest international video coding standard developed by ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11 Joint Video Expert Team (JVET). The input video signal is predicted from the reconstructed signal, which is derived from the coded and decoded image region. The prediction residual signal is processed by pixel block transform. The transform coefficients are quantized and entropy coded along with other side information in the bitstream. The reconstructed signal is generated from the prediction signal and the reconstructed residual signal after inverse transforming the dequantized transform coefficients. The reconstructed signal is also processed by loop filtering to remove coding artifacts. The decoded image is stored in a buffer and used to predict future images in the input video signal.
[0004] In VVC, the coded image is divided into non-overlapping square pixel block areas represented by related coding tree units (CTUs). The leaf nodes of the coding tree correspond to coding units (CUs). The coded image can be represented by a set of slices, each slice including an integer number of CTUs. Each CTU in a slice is processed in raster scan order. Bi-predictive (B) slices can be decoded to predict the sample values of each pixel block using intra-picture prediction or inter-picture prediction using up to two motion vectors and reference indices. Predictive (P) slices are decoded to predict the sample values of each pixel block using intra-picture prediction or inter-picture prediction using up to one motion vector and reference index. Intra-picture (I) slices are decoded using only intra-picture prediction.
[0005] Using a quadtree (QT) with a nested multi-type-tree (MTT) structure, the CTU can be split into one or more non-overlapping codec units (CUs) to accommodate various local motion and texture features. The CU is further divided into smaller CUs using one of five partition types: quadtree partition, vertical binary tree partition, horizontal binary tree partition, vertical center-side ternary tree partition, and horizontal center-side ternary tree partition.
[0006] Each CU contains one or more prediction units (PU). The prediction unit, together with the related CU syntax, serves as a basic unit for indicating prediction sub-information. The specified prediction process is used to predict the values of related pixel samples within the PU. Each CU may contain one or more transform units (TU) representing prediction residual pixel blocks. The transform unit (TU) includes a transform pixel block (TB) of luma samples and two corresponding transform pixel blocks of chroma samples, each TB corresponding to a residual pixel block sample of a color component. An integer transform is applied to the transform pixel block. The layer values of the quantized coefficients are entropy encoded and decoded in the bitstream together with other side information. The terms Coding Tree Block (CTB), Coding Block (CB), Prediction Block (PB), and Transform Block (TB) are defined to specify a 2D sample array of a color component associated with a CTU, CU, PU, or TU, respectively. Therefore, a CTU includes a luma CTB, two chroma CTBs, and related syntax elements. Similar relationships also apply to CU, PU, and TU.
[0007] For each inter-predicted CU, motion parameters including motion vector, reference picture index and reference picture list usage index and additional information are used to generate inter-predicted samples. Motion parameters can be indicated in an explicit or implicit way. When the CU is encoded and decoded using skip mode, the CU is associated with one PU and has no valid residual coefficients, no coded motion vector delta or reference picture index. A merge mode is specified, in which the motion parameters of the current CU are obtained from neighboring CUs, including spatial candidates and temporal candidates, as well as the additional schedule introduced in VVC. Merge mode can be applied to any inter-predicted CU. An alternative to merge mode is the explicit transmission of motion parameters, in which the motion vector, the corresponding reference picture index and reference picture list usage flag for each reference picture list and other required information are explicitly indicated for each CU. Summary of the invention
[0008] Some embodiments of the present invention provide a method for refining affine candidates by regression. A video codec receives data of a pixel block of a current block of a current picture to be encoded or decoded as a video. The video codec derives a linear model to refine or generate a prediction candidate by regression to minimize the difference between a plurality of samples of a current template adjacent to the current block and a plurality of samples of a reference template identified by the prediction candidate. The video codec generates a prediction of the current block based on the derived linear model. The video codec encodes or decodes the current block by using the generated prediction.
[0009] In some embodiments, the derived prediction candidate being refined is an affine motion candidate, and the derived model is an affine motion model with four parameters based on affine motion of two checkpoints, or an affine motion model with six parameters for three checkpoints. The affine motion candidate being refined can be an MMVD candidate or an Adaptive Motion Vector Prediction (AMVP) candidate. The refined prediction candidate can also be a conventional merge candidate.
[0010] In some embodiments, the refined prediction candidate is added to the prediction candidate list of the current block. In some embodiments, the refined prediction candidate replaces the prediction candidate in the prediction candidate list of the current block.
[0011] In some embodiments, only a subset of prediction candidates in the prediction candidate list of the current block is refined through regression of the linear model. The prediction candidate subset can be MMVD candidates identified based on proximity to a specific direction. In some embodiments, for each MMVD candidate, only a base motion vector predictor (MVP) is refined through regression.
[0012] In some embodiments, the derived model is an affine motion model derived by minimizing the difference between a plurality of regression sub-block motion vectors and (i) a plurality of sub-block MVs of a plurality of sub-blocks adjacent to the current block and (ii) a plurality of sub-block MVs derived from an affine candidate of the current block to refine the affine candidate. The prediction is generated based on the refined affine candidate. In some embodiments, the affine motion model is derived by minimizing the difference between a plurality of regression sub-block motion vectors and (i) a plurality of stored sub-block MVs associated with a plurality of sub-blocks adjacent to the current block and (ii) a plurality of inherited MVs associated with a plurality of sub-blocks belonging to a reference affine codec unit (CU). The reference affine CU may not be adjacent to the current block.
[0013] Other aspects and features of the present invention will become apparent to those of ordinary skill in the art by reading the following description of specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Various embodiments of the present disclosure are described in detail with reference to the following drawings, which are presented as examples.
[0015] in:
[0016] Figure 1 The spatial and temporal candidates for the merge mode are marked.
[0017] Figure 2 The control point motion vector (CPMV) of the current block decoded by affine motion field coding is shown.
[0018] Figure 3 Neighboring blocks that provide constructed affine candidates are shown.
[0019] Figure 4 The spatial neighboring sub-blocks of the current CU used to derive the Regression-based Motion Vector Field (RMVF) motion parameters are shown.
[0020] Figure 5A -C conceptually illustrates the use of template regression to refine affine candidates.
[0021] Figure 6 The generation of an affine motion model through MV regression using neighboring sub-block MVs and inherited sub-block MVs of the current CU is conceptually illustrated.
[0022] Figure 7 Two-sided matching at the block level and sub-block level is shown.
[0023] Figure 8 A refined MV that is further refined through regression is conceptually shown.
[0024] Fig. 9 An example of a video encoder that can refine affine candidates via regression is shown.
[0025] Fig.10 A partial video encoder implementing refinement of affine candidates via regression is shown.
[0026] Fig.11 The process of refining prediction candidates by deriving a regression model is conceptually shown.
[0027] Fig.12 An example of a video decoder that can refine affine candidates is shown.
[0028] Fig.13 A partial video decoder implementing refinement of affine candidates via regression is shown.
[0029] Fig.14 The process of refining prediction candidates by deriving a regression model is conceptually shown.
[0030] Fig.15 An electronic system is conceptually illustrated that implements some embodiments of the present invention. DETAILED DESCRIPTION
[0031] It is readily understood that the components of the present invention, as generally described herein and illustrated in the accompanying drawings, may be arranged and designed in a variety of different configurations. Therefore, the following more detailed description of embodiments of the systems and methods of the present invention, as shown in the accompanying drawings, is not intended to limit the scope of the claimed invention, but is merely representative of selected embodiments of the present invention.
[0032] I. Merge Mode
[0033] Skip mode and merge mode obtain motion information from spatial neighboring blocks (spatial candidates) or temporal co-located blocks (temporal candidates). When a PU is in skip mode or merge mode, no motion information is encoded or decoded, but only the index of the selected candidate is encoded or decoded. For skip mode, the residual signal is forced to zero and is not encoded or decoded. If a particular block is encoded as skip or merge, the candidate index will be marked to indicate which candidate in the candidate set is used for merging. Each merged PU reuses the MV, prediction direction and reference picture index of the selected candidate.
[0034] Figure 1The spatial candidates and temporal candidates for the merge mode are marked. As shown in the figure, up to four spatial MV candidates are derived from the spatial neighbors A0, A1, B0 and B1, and one temporal MV candidate is derived from TBR or TCTR (TBR is used first, if TBR is not available, TCTR is used). If any of the four spatial MV candidates is not available, position B2 is used to derive the MV candidate instead. After the derivation process of the four spatial MV candidates and one temporal MV candidate, redundancy removal (i.e. pruning) is applied to remove redundant MV candidates. If the number of available MV candidates is less than 5 after redundancy removal (i.e. pruning), three types of additional candidates are derived and added to the candidate set (candidate list). Based on the rate-distortion optimization (RDO) decision, the encoder selects a final candidate in the candidate set for skip mode or merge mode and sends the index to the decoder.
[0035] II. Affine motion field
[0036] Objects in a video may have different types of motion, including translation, zooming in / out, rotation, perspective, and other irregular motions. In some embodiments, block-based affine transformation motion compensation prediction is used to account for various types of motion. VVC provides block-based affine transformation motion compensation prediction. Specifically, the affine motion field mvx, mvy of the current block at position (x, y) is in the form of the following linear model:
[0037] mv x =a*x+b*y+c
[0038] mv y =d*x+e*y+f (0)
[0040] The coefficients {a, b, c, d, e, f} are the parameters of the linear model. In some embodiments, the affine motion field at the position (x, y) can be described by the motion information of two control points (CP) (e.g., the upper right corner and the upper left corner of the block) (4-parameter model) or the motion information of three control points (e.g., the lower right corner, the upper left corner and the lower right corner of the block) (6-parameter model).
[0041] For a 4-parameter affine motion model, Equation 0 (the motion vector at the sample position (x, y) in the block) can be written as:
[0042]
[0043] For the 6-parameter affine motion model, Equation 0 can be written as:
[0044]
[0045] Among them, (mv0x, mv0y) is the motion vector of the upper left control point (upper left CPMV or mv0), (mv1x, mv1y) is the motion vector of the upper right control point (upper right CPMV or mv1), and (mv2x, mv2y) is the motion vector of the lower left control point (lower left CPMV or mv2).
[0046] In some embodiments, a 2-parameter model may be used to refine the translational inter-picture prediction candidates (eg, conventional merge candidates):
[0047] MV x regress =MV x orig +biasX
[0048] MV y regress =MV y orig +biasY (3)
[0050] Figure 2 The control point motion vector (CPMV) of the current block decoded through the affine motion field is shown. The current block has CPMVs at the upper left corner (mv0), the upper right corner (mv1), and the lower left corner (mv2). Using an affine motion model such as Equation 1 (4-parameter affine model) or Equation 2 (6-parameter affine model), the affine motion field mv' at the position (x, y) in the current block 200 can be derived.
[0051] III. Affine Merge Mode
[0052] Affine merge mode or AF_MERGE mode can be applied to CUs whose width and height are both greater than or equal to 8. In this mode, the motion vector of the current CU at the control point (CPMV) is generated based on the motion information of the spatially neighboring CUs. There can be up to five CPMVP candidates, and the index will be marked to indicate a CPMV for the current CU. The following three types of CPMV candidates are used to form the affine merge candidate list: (1) inherited affine merge candidates inferred from the CPMV of neighboring CUs; (2) constructed affine merge candidate CPMVP derived using the translation MV of neighboring CUs; (3) zero MV.
[0053] The inherited affine candidate inherits the affine model from the neighboring block to derive the affine model by directly obtaining the CPMV from the neighboring block or using the sub-block MV of the spatial candidate. The affine merge list includes at most two inherited affine candidates, and the derivation process of the inherited affine candidate follows the order of searching for available spatial merge candidates from B0 to B2 and from A0 to A1.
[0054] On the other hand, constructed affine candidates are derived from sub-block MVs from different neighboring blocks. Figure 3 Neighboring blocks that provide constructed affine candidates are shown. The constructed affine candidates are inserted into the affine merge list by selecting two or three sub-block MVs from {A, B, C, H} as input MVs for the constructed affine candidate derivation process in the following order:
[0055] A, B, C
[0056] A, B, H
[0057] A,C,H
[0058] B,C,H
[0059] A,B
[0060] A,C
[0061] The affine merge list contains at most 6 constructed affine candidates. The search process for {A, B, C, H} is to take the first available MV in {A0, A1, A2} as A, the first available MV in {B0, B1} as B, the first available MV in {C0, C1} as C, and the sub-block MV in the same block as H.
[0062] IV. Regression-based Motion Vector Field (RMVF)
[0063] The motion behavior can vary within a block. In particular, for larger CUs, it is not effective to represent the motion behavior with only one motion vector. The RMVF method models this motion behavior based on the motion vectors of spatially neighboring sub-blocks.
[0064] Figure 4The spatial neighboring sub-blocks of the current CU for RMVF motion parameter derivation are shown. The figure shows a current block 400 with a neighboring sub-block 410. The motion vectors and center positions of the neighboring sub-blocks 410 from the current CU 400 are used as input to a linear regression process to derive a set of linear model parameters {axx, axy, ayx, ayy, bx, by} for the affine motion model. The regression process minimizes the mean square error between the neighboring sub-block MV and the MV derived by the regression model. Then, using the center position (XsubPU, YsubPU), the motion vector (MVX_subPU, MVY_subPU) of the sub-block in the current CU 400 can be calculated as:
[0065]
[0066] Then, the motion field of the current CU 400 may be refined through the derived affine motion model with parameters {axx, axy, ayx, ayy, bx, by}.
[0067] IV. Affine MMVD
[0068] Unlike the regular merge mode, where the implicitly derived motion information is directly used to generate the prediction samples of the current CU, in the Merge with Motion Vector Difference (MMVD) mode, the derived motion information is further refined through the motion vector difference MVD. MMVD also extends the candidate list of the merge mode by adding additional MMVD candidates based on a predefined offset (also called MMVD offset).
[0069] In some embodiments, MMVD is extended to affine as an affine MMVD mode for further refining affine merge candidates with MV offsets. In the affine MMVD mode, the MV offsets are predefined and obtained from 8 directions (e.g., k×π / 2 horizontal and vertical angles and k×π / 4 diagonal angles) with 5 steps (i.e., {1, 2, 4, 8, 16}). Affine MMVD candidates are generated by adding the predefined translation MV offsets to the three CPMVs, and the affine MMVD list includes a total of 40 candidates. After reordering the affine MMVD list by template matching cost, only the 20 affine MMVD candidates with the smallest template matching cost are retained.
[0070] V. Affine with DMVR
[0071] To increase the accuracy of the MV in merge mode, Decoder Side Motion Vector Refinement (DMVR) can be applied to refine the MV by using methods such as bilateral-matching (BM). In the bi-prediction operation, the refined MV is searched around the initial MV in reference picture list L0 and reference picture list L1. The BM method calculates the distortion between two candidate blocks in reference picture list L0 and reference picture list L1. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bi-prediction signal.
[0072] In some embodiments, if the selected merge candidate satisfies the DMVR condition, a multi-pass decoder-side motion vector refinement (MP-DMVR) method is applied in a regular merge mode. In the first round, bilateral matching (BM) is applied to the codec block. In the second round, BM is applied to each 16x16 sub-block within the codec block. In the third round, the MV in each 8x8 sub-block is refined by applying bi-directional optical flow (BDOF). BM refines a pair of motion vectors MV0 and MV1 under the constraint that the motion vector difference MVD0 (i.e., MV0'-MV0) is exactly the opposite sign of the motion vector difference MVP1 (i.e., MV1'-MV1).
[0073] In some embodiments, MP-DMVR is used to refine the affine CPMV. The 4-parameter affine model can be described using Equation 1 above, where only two non-translation parameters are used.
[0074] Using equation 2 above, a 6-parameter affine model can be described, where (mvx,mvy) is the motion vector at position (x,y). (mv0x,mv0y) is the basis MV representing the translational motion of the affine model, and are the four non-translational parameters that define the rotation, scaling, and other non-translational motions of the affine model.
[0075] In some embodiments, the basis MV of the affine model of the codec block of the codec is refined by applying only the first step of multiple rounds of DMVR, using the affine merge mode. That is, the translation MV offset is added to all CPMV candidates in the affine merge list if the candidate meets the DMVR condition. And the MV offset is derived by minimizing the cost of bilateral matching, which is the same as traditional DMVR. And the DMVR condition is not changed.
[0076] The MV offset search process is the same as the first round of multi-round DMVR, where a 3x3 square search pattern is used to loop around the search range [-8, +8] in the horizontal direction and [-8, +8] in the vertical direction to find the best integer MV offset. Then, a half-pixel search is performed around the best integer position, and finally, an error surface estimation is performed to find the MV offset with 1 / 16 accuracy. The refined CPMV is stored for spatial and temporal motion vector prediction as a multi-round DMVR result.
[0077] VI. Regression-based Affine Candidates
[0078] A. Affine Candidate Refinement via Template Regression
[0079] Some embodiments of the present invention provide a method for refining affine candidates through a regression model, wherein the regression model is generated by minimizing the mean square error (MSE) between adjacent reconstructed samples (e.g., the first N rows of reconstructed samples in the left direction and / or the top direction, also referred to as the current sample) and the reference sample of the affine candidate to be refined. Specifically, as follows:
[0080]
[0081] or
[0082]
[0083] Where (x, y) represents the position relative to the upper left position of the reconstructed block, and (x', y') represents the position relative to the lower left position of the reference template block. B represents the current template area, and B' represents the reference template area. The reference template block is located by the regression MV generated by the regression model. The regression model can be a 4-parameter affine model (Equation 1) or a 6-parameter affine model (Equation 2), or even a 2-parameter translation model (Equation 3).
[0084] Figure 5A -C conceptually illustrates the use of template regression to refine affine candidates. Figure 5A As shown, the current CU 505 in the current picture 500 has an affine candidate 510. The affine candidate 510 is used to locate the corresponding reference template region 530 in the reference picture 501. The neighboring region 520 of the current CU 510 is used as the current template.
[0085] Figure 5BSamples of the current template 520 and samples of the reference template 530 are marked for generating a linear model 550 by regression (thus, the linear model 550 is also a regression model). The regression is used to solve the parameters of the linear model. The linear model 550 being solved can be a 4-parameter affine motion model or a 6-parameter affine motion model in the form of Equation 1 or Equation 2. According to Equation 5, the regression is driven by minimizing the MSE between the samples of the current template 520 and the samples of the reference template 530.
[0086] Figure 5A and Figure 5B The regression process is also shown. The linear motion model 550 under regression is used to generate the regression MV 511, which locates the updated reference template region 531, whose samples replace the samples of the initial reference template region 530 and are used to update the MSE calculation. The regression MV 511 can be continuously updated until the MSE calculated according to Equation 5 is minimized.
[0087] Figure 5C The completed model 550 is shown being applied to the affine candidate 510 to generate a refined affine candidate 512 which may be the final regression MV 511 .
[0088] In some embodiments, one or more affine merge candidates in the affine merge candidate list are refined through a regression model derived by minimizing the mean square error between adjacent reconstructed samples (e.g., the current template 520, which includes the first N rows of reconstructed samples in the left direction and / or top direction) and a reference template of the affine merge candidate to be refined (e.g., the reference template 531).
[0089] In some embodiments, one or more merge candidates in the merge candidate list are refined by a 2-parameter translation regression model that minimizes the mean square error between adjacent reconstructed samples and a reference sample of the merge candidate to be refined. The merge candidate can also be refined in a manner similar to bidirectional optical flow (BDOF).
[0090] In some embodiments, one or more MMVD candidates in the MMVD candidate list are refined by a 2-parameter translation regression model that minimizes the mean square error between neighboring reconstructed samples and a reference template of the MMVD candidate to be refined. MMVD candidates can also be refined in a manner similar to bidirectional optical flow (BDOF).
[0091] In some embodiments, one or more affine MMVD candidates in the affine MMVD candidate list are refined by a regression model that minimizes the mean square error between neighboring reconstructed samples and a reference template of the affine MMVD candidate to be refined. After refinement, the affine MMVD candidates are reordered by template matching cost, and only the N candidates with the smallest template matching cost are retained.
[0092] In some embodiments, the affine MMVD candidates are first reordered by template matching cost, and only the N candidates with the smallest template matching cost are retained while the other candidates are discarded. Then, the remaining N affine MMVD candidates are refined by a regression model that minimizes the mean square error between adjacent reconstructed samples and reference templates of the affine MMVD candidates to be refined.
[0093] In some embodiments, the basis affine MMVD candidates are refined through a regression model. The model is used to determine which MMVD candidate is preferred. The MMVD candidates are reordered based on the affine candidates refined by the regression model.
[0094] In some embodiments, one or more affine AMVP candidates in the affine AMVP candidate list are refined by a regression model that minimizes the mean square error between adjacent reconstructed samples and a reference template of the affine AMVP candidate to be refined. After refinement, the affine AMVP candidates in the affine AMVP candidate list can be reordered by template matching cost, and only the N candidates with the smallest template matching cost are retained.
[0095] In some embodiments, the affine AMVP candidates are first reordered by template matching cost, and only the N candidates with the smallest template matching cost are retained. Then, the retained N affine AMVP candidates are refined by a regression model that minimizes the mean square error between adjacent reconstructed samples and reference templates of the affine AMVP candidates to be refined.
[0096] In some embodiments, a regression model (e.g., using sample regression or MV regression) can be used as a decoder-side MV Refinement Process (DMVR). For example, when an affine candidate satisfies some constraints, the affine candidate is refined through the regression model. Examples of these constraints include CU size constraints (e.g., CU size / width / height>,>=,<, or<=a certain threshold), adjacent MV / sub-block MV number constraints, adjacent reconstructed sample number constraints, inter-picture prediction direction constraints, true bi-prediction constraints, etc.
[0097] B. Generating affine candidates through template regression
[0098] Some embodiments of the present invention provide regression-based affine candidates, which are derived by minimizing the mean square error between the current template samples (i.e., the neighboring reconstructed samples of the current block) and the reference template samples of the affine candidates to be refined. Figure 5C In the example of , the refined affine candidate 511 can be considered as a regression-based affine candidate, which can be added to the affine candidate list.
[0099] In some embodiments, if K affine candidates and adjacent or non-adjacent affine CUs are available, a total of K regression-based affine candidates may be included in the affine candidate list. One or more of the K regression-based affine candidates may replace any candidate in the affine candidate list, or be additionally added to the affine candidate list. The same method can be applied to derive other types of regression-based affine candidates, for example, regression-based affine merge candidates in an affine merge candidate list, or regression-based affine AMVP candidates in an affine AMVP candidate list, or regression-based MMVD AMVP candidates in an affine MMVD candidate list.
[0100] In some embodiments, the affine candidates refined by the regression model can be used to replace the affine candidates in the candidate list (e.g., the affine merge candidate list, the affine MMVD candidate list, or the affine AMVP candidate list), or can be inserted into the candidate list. For example, the affine candidates refined by the regression model can be added after the original candidates, or after the first N candidates, or after a specific type of candidate in the candidate list. The number of affine candidates refined by the added regression model can also be constrained.
[0101] C. Generate merge candidates through template regression
[0102] In some embodiments, a regression-based merge candidate is derived by using a reference template of the merge candidate or a reference template of a non-adjacent merge CU as input to derive a parametric translation regression model that minimizes the mean square error between adjacent reconstructed samples (current template) and the reference template of the merge candidate to be refined. If K merge candidates and non-adjacent merge CUs are available, a total of K regression-based merge candidates may be included in the merge candidate list. One or more of the K regression-based merge candidates may replace any candidate in the merge candidate list or be additionally added to the merge candidate list. The merge candidate may also be refined in a manner similar to BDOF.
[0103] D. Refine affine candidates through MV regression
[0104] In some embodiments, the RMVF method described in Section IV above may be used to refine one or more affine candidates in the affine candidate list through corresponding regression models. Each regression model takes as input (i) neighboring sub-block MVs and (ii) sub-block MVs of the current block (e.g., current CU) derived from the affine candidates. The regression model minimizes the mean square error between the neighboring sub-block MVs and the sub-block MVs derived by the regression model based on regression. The regression model may be a 4-parameter or 6-parameter affine model or a 2-parameter translation model. In contrast to the above-mentioned template (TM) regression, this is called motion vector (MV) regression.
[0105] In some embodiments, the affine refinement model is a linear regression model derived by minimizing the MSE between (i) the inherited sub-block MV and / or the neighboring sub-block MV and (ii) the MV derived by the regression model. The inherited sub-block MV is taken from the sub-block motion field of the previously encoded and decoded affine block as the reference affine CU, which can be a non-neighboring affine CU or a history-based affine CU. The neighboring sub-block MV is taken from the 4x4 sub-blocks neighboring the current CU.
[0106] Figure 6 The diagram conceptually illustrates the generation of an affine motion model through MV regression using the neighboring sub-block MVs and inherited sub-block MVs of the current CU. For the encoded current block 600, the diagram shows (i) the stored MVs (MVstored) associated with the neighboring sub-blocks 610 adjacent to the current block 600, and (ii) the inherited MVs (MVinherited) associated with the sub-blocks belonging to the previously encoded affine block 620 as the reference affine CU. The reference affine CU 620 can be an affine CU that is not adjacent to the current block, or an affine CU based on history. The diagram also shows (iii) the regression MV (MVregress) generated by the regression model during regression.
[0107] The linear regression model to be solved is as follows:
[0108]
[0110] Where B1 indicates the region of the reference affine coded CU (e.g., non-adjacent or CU 620), and B2 indicates the neighboring region of the current CU (e.g., neighboring sub-block 610). The linear model is constructed by finding the MVx,yregress for all sub-block positions x,y of B1 and B2 that minimizes the MSE. The regression model can be a 4-parameter (Equation 1) or a 6-parameter affine model (Equation 2), or even a 2-parameter translation model.
[0111] E. Refine the affine MMVD candidates through MV regression
[0112] In some embodiments, the basis affine MMVD candidates are refined through the regression model. The refined basis affine MMVD candidates are used to determine which MMVD candidate is preferred. The MMVD candidates are reordered according to the affine candidates refined by the regression model.
[0113] In some embodiments, the affine candidates of regression model refinement can be used to replace the affine candidates in the candidate list (e.g., the merged candidate list, the MMVD candidate list, or the AMVP candidate list), or can be inserted into the candidate list. In some embodiments, the affine candidates of regression model refinement can be added to the candidate list, after the original candidates, or after the first N candidates, or after certain types of candidates. The number of affine candidates of regression model refinement added can also be constrained.
[0114] In some embodiments, based on the RMVF method, the neighboring sub-block MVs of the current CU are taken as input to derive a regression model (denoted as Mn). In addition, the motion field in the current CU derived from each affine MMVD candidate is used to derive a regression model set (denoted as {Ma1, Ma2, ..., MaN}, where N = the number of affine MMVD candidates). By mixing Mn with {Ma1, Ma2, ..., MaN} using a specific derivation method (for example, a combination of two linear and / or nonlinear affine models based on step size and / or direction), the final regression model set ({Mf1, Mf2, ..., MfN}, where N = the number of affine MMVD candidates) can be obtained. The final regression model set is used to refine the motion field of each affine MMVD candidate.
[0115] The regression model set {Mf1, Mf2, ..., MfN} can also be used to derive N CPMV candidates for the affine MMVD candidate. In this method, the regression of the adjacent sub-block MVs only needs to be performed once.
[0116] The blending process can be applied at the sub-block MV level, the affine parameter level, or the CPMV level. The blending process can be limited by the number of spatially adjacent sub-block MVs, the number of sub-block MVs in the reference block of each affine MMVD candidate, the number of sub-blocks of the current block, the MV amplitude of each affine MMVD candidate, the step size and direction of the affine MMVD, the CU size / width / height, and any combination of the above. The blending at the affine parameter level can use the number of spatially adjacent sub-block MVs, the number of sub-block MVs in the reference block of each affine MMVD candidate, the number of sub-blocks of the current block, the MV amplitude of each affine MMVD candidate, the step size and direction of the affine MMVD, and any combination of the above.
[0117] F. Hybrid Template Regression Model and MV Regression Model
[0118] In some embodiments, a first regression model through template regression (as described in Section A above) and a second regression model through MV regression (as described in Section D above) are mixed to derive a fused regression model. Specifically, reference samples of an affine merge candidate or an affine AMVP candidate, or reference samples of a non-adjacent affine CU may be used to derive a first regression model (referred to as Ma) by minimizing the mean square error between adjacent reconstructed samples and reference templates of an affine candidate to be refined using template regression. Adjacent sub-block MVs and inherited sub-block MVs from a corresponding affine merge or affine AMVP candidate or a corresponding non-adjacent affine CU are used as input to derive a second regression model (referred to as Mb) by minimizing the mean square error between adjacent sub-block MVs and inherited sub-block MVs using MV regression.
[0119] A hybrid or fused regression model (referred to as Mf) is derived by mixing Ma and Mb in a certain derivation method (e.g., a linear and / or nonlinear combination of two affine models according to predefined weights). If K affine merge candidates or affine AMVP candidates and non-adjacent affine CUs are available, a total of K regression-based affine candidates derived from Mf are included in the affine merge candidate list or the affine AMVP candidate list. One or more of the K regression-based affine candidates may replace any candidate in the affine merge candidate list or the affine AMVP candidate list, or be additionally added to the affine merge candidate list or the affine AMVP candidate list.
[0120] In some embodiments, the affine MMVD candidates are refined by a hybrid regression model (Mf), which is a combination of a regression model that minimizes the MV error and a regression model that minimizes the template sample error. The hybrid regression model is used to refine one or more affine MMVD candidates.
[0121] In some embodiments, based on the RMVF method, the neighboring sub-block MV of the current CU is used as input to derive a regression model (denoted as Mn), and the motion field in the current CU derived from the affine MMVD candidate set is used to derive a regression model set (denoted as {Ma1, Ma2, ..., MaN}, where N = the number of affine MMVD candidates). In addition, the reference template of each affine MMVD candidate is used to derive another regression model set (denoted as {Mb1, Mb2, ..., MbN}, where N = the number of affine MMVD candidates), which minimizes the mean square error between the neighboring reconstructed samples and the reference template of each affine MMVD candidate (template regression). By mixing Mn with {Ma1, Ma2, ..., MaN} and {Mb1, Mb2, .., MbN} in a specific derivation method (e.g., a linear and / or nonlinear combination of three affine models according to step size and / or direction), a final regression model set ({Mf1, Mf2, ..., MfN}, where N = the number of affine MMVD candidates) can be obtained. The final regression model set is used to refine the motion field of each affine MMVD candidate. The regression model set {Mf1, Mf2, ..., MfN} can be used to derive N CPMV candidates for the affine MMVD candidates. In some embodiments, the regression of adjacent sub-block MVs is performed only once.
[0122] The blending process can be applied to the sub-block MV level, the affine parameter level or the CPMV level. The blending process can be limited by the number of current template samples, the number of reference template samples, the number of spatially adjacent sub-block MVs, the number of sub-block MVs in the reference block of each affine MMVD candidate, the number of sub-blocks of the current block, the MV amplitude of each affine MMAD candidate, the step size and direction of the affine MMVD, the CU size / width / height, the QP value, and any combination of the above. The blending at the affine parameter level can use the information of the number of spatially adjacent sub-block MVs, the number of sub-block MVs in the reference block of each affine MMVD candidate, the number of sub-blocks of the current block, the MV amplitude of each affine MMVD candidate, the step size and direction of the affine MMVD, and any combination of the above.
[0123] G. Reducing affine MMVD merging candidates via regression
[0124] In some embodiments, the base MVP of the affine MMVD pattern is refined through regression. The method also determines the step size and / or direction of the affine MMVD based on the CPMV difference between the base MVP of the original affine MMVD and the refined affine MMVD base MVP. The closest affine MMVD direction with the angle of the CPMV difference is represented as Dc, and the closest affine MMVD step size with the magnitude of the CPMV difference is represented as Sc. In some embodiments, through the template matching cost, only K (indicating a predefined number) affine MMVD candidates near the direction Dc and the step size Sc are reordered to reduce the template matching cost calculation or reduce the labeling overhead of the affine MMVD direction and / or step size. The regression model can minimize the mean square error between the adjacent sub-block MVs and the inherited (from the base MVP of the affine MMVD) sub-block MVs and the MVs derived by the regression model, and / or minimize the mean square error between the adjacent reconstructed samples and the reference template samples of the affine MMVD base MVP.
[0125] In some embodiments, the reduced remaining candidates are further refined through regression. Specifically, through the regression model, only K (representing a predefined number) affine MMVD candidates near the direction Dc and the step size Sc are refined to reduce the calculation of the regression matrix solution. Then, through the template matching cost, the K affine MMVD candidates are reordered, and the N candidates with the smallest template matching cost are retained. The regression model is derived by minimizing the mean square error between the adjacent sub-block MVs and the inherited sub-block MVs (i.e., the sub-block MVs derived from the base MVP of the affine MMVD or the affine MMVD candidate) and the MVs derived through the regression model, and / or minimizing the mean square error between the adjacent reconstructed samples and the reference templates of the affine MMVD base MVP or the affine MMVD candidate.
[0126] H. Reducing MMVD Merge Candidates through Regression Results
[0127] In some embodiments, the base MVP of the MMVD pattern is (only) refined through regression. The method also determines the step size and / or direction of the MMVD based on the MVP difference between the original MMVD base MVP and the refined MMVD base MVP. The closest MMVD direction with the angle of the MVP difference is represented as Dc, and the closest MMVD step size with the magnitude of the MV difference is represented as Sc. Only K (predetermined number) MMVD candidates near the direction Dc and the step size Sc are added to the MMVD candidate list to save the labeling overhead of the MMVD direction and / or step size. The regression model can minimize the mean square error between the adjacent sub-block MVs and the inherited (from the MMVD base MVP) sub-block MVs and the MVs derived through the regression model, and / or minimize the mean square error between the adjacent reconstructed samples and the reference template of the MMVD base MVP.
[0128] In some embodiments, only K (predefined number) inter-frame MMVD candidates near the direction Dt and step size St are reordered through template matching cost to reduce template matching cost calculation or reduce the labeling overhead of inter-frame MMVD direction and / or step size.
[0129] J. Refining Affine Motion via Bilateral Matching
[0130] In some embodiments, a motion offset to be tested is added to all CPMVs to become the CPMV to be tested. Next, the sub-block MVs are derived accordingly, and then the difference between the two predictors is calculated in round 1 of the bilateral matching algorithm. The CPMV to be tested with the smallest difference is selected as the refined CPMV. In some embodiments, all CPMVs of the affine codec block can be refined using the same motion offset derived through the bilateral matching algorithm with round 1, round 2, and / or round 3.
[0131] In some embodiments, first, all sub-block motions of the affine codec block are derived based on CPMV. Afterwards, the same motion offset is added to all sub-block motions, which will be tested in the bilateral matching algorithm. The difference between the two predictors is calculated and used to select the best motion offset among the motion offsets to be tested. In this method, the derivation process of the affine sub-block motions can be applied only once when testing different motion offsets to be tested.
[0132] K. Refining affine motion via bilateral matching at block and sub-block levels
[0133] In some embodiments, after all sub-block motions are derived based on CPMV, the sub-block motions in the predefined region are refined through bilateral matching. Then, based on bilateral matching, motion offsets are derived, and the derived motion offsets are added to all sub-block motions in the predefined region. For example, if the predefined region is 16x16, the 32x32 affine codec block will be split into four regions, and four motion offsets of the sub-blocks in the corresponding regions are derived.
[0134] Figure 7Bilateral matching at the block level and sub-block level is shown. In the figure, A, B, C, D...P are 8x8 sub-blocks in an affine codec CU. After all sub-block motions are derived based on CPMV, the motion in each 16x16 region (A, B, E, F)(C, D, G, H)(I, J, M, N)(K, L, O, P) is used to derive a motion offset. Sub-blocks A, B, E and F are used to derive a motion offset, and the derived motion offset is used to refine the sub-block motion in A, B, E and F. Sub-blocks C, D, G and H are used to derive a motion offset, and the derived motion offset is used to refine the sub-block motion in C, D, G and H. Sub-blocks I, J, M and N are used to derive a motion offset, and the derived motion offset is used to refine the sub-block motion in I, J, M and N. Sub-blocks K, L, O and P are used to derive a motion offset, and the derived motion offset is used to refine the sub-block motion in K, L, O and P.
[0135] In some embodiments, the above derived motion offset is added back to the corresponding CPMV to derive the refined CPMV. The final sub-block motion is derived based on the refined CPMV. In some embodiments, the above refined sub-block motion is used to derive the refined CPMV by using some regression method. The final sub-block motion is derived based on the refined CPMV.
[0136] In some embodiments, hierarchical bilateral matching motion refinement is used. For example, the two predefined regions are 32x32 and 8x8. For affine coded CUs, after all sub-block motions are derived based on CPMV, all sub-block motions in a 32x32 region are refined by the same motion offset derived by bilateral matching. Afterwards, each sub-block in a 16x16 region is refined again by the same motion offset derived by bilateral matching.
[0137] For example, a 64x64 affine codec block includes four 32x32 regions (A,B,E,F)(C,D,G,H)(I,J,M,N)(K,L,O,P). Based on all sub-block motions in regions (A,B,E,F), a motion offset is derived and added to all sub-block motions in regions (A,B,E,F). Based on all sub-block motions in regions (C,D,G,H), a motion offset is derived and added to all sub-block motions in regions (C,D,G,H). Based on all sub-block motions in regions (I,J,M,N), a motion offset is derived and added to all sub-block motions in regions (I,J,M,N). Based on all sub-block motions in regions (K,L,O,P), a motion offset is derived and added to all sub-block motions in regions (K,L,O,P). Afterwards, the four 8x8 sub-blocks in region A (16x16) are used to derive a motion offset, which is added to the four sub-block motions in region A. Similar to each 16x16 region (A, B, C, D, ... P) in the 64x64 affine codec block. The predefined region size in the above embodiment can be designed based on the CU size or the picture size.
[0138] In some embodiments, hierarchical bilateral matching motion refinement is used. Bilateral matching of deeper depths is applied only when the bilateral matching cost of shallow depth motion refinement is greater than a threshold. The threshold can be designed based on the CU size. For example, the bilateral matching cost of each 32x32 region is calculated after adding the corresponding motion offset. In a 32x32 region, bilateral matching with sub-blocks in a 16x16 region is performed only when the calculated bilateral matching cost is greater than half the CU size.
[0139] In some embodiments, the derived motion offset is added back to the corresponding CPMV to derive the refined CPMV. Based on the refined CPMV, the final sub-block motion is derived. In another embodiment, the refined sub-block motion is used to derive the refined CPMV by using a regression method. Based on the refined CPMV, the final sub-block motion is derived.
[0140] In some embodiments, bilateral matching can be replaced by template matching as described above. When template matching is used instead of bilateral matching, only the sub-blocks located at the top and left boundaries of the CU are used to determine the derived motion offset. For other sub-blocks not located at the top and left boundaries of the CU, the derived motion offset can be the same as the derived one, or set to zero.
[0141] L. Refining sbTMVP motion via bilateral matching
[0142] In some embodiments, after deriving all sub-block motions of a subblock-based temporal motion vector prediction (sbTMVP) codec block, motion refinement is derived through bilateral matching. Afterwards, the derived motion refinement is added to all sub-block motions. In the above manner, all reference sub-blocks of the sbTMVP codec block can be shifted together based on the result of bilateral matching. The bilateral matching method mentioned previously can be an MP-DMVR with round 1, round 2 and / or round 3.
[0143] In some embodiments, after deriving the motion of all sub-blocks of the sbTMVP codec block, all sub-blocks in the predefined area will be grouped to derive a motion offset. The derived motion offset will be used to refine all sub-block motions in the predefined area. For example, the area can be 16x16 or 32x32. For another example, the area should include half or a quarter of the sub-blocks within the sbTMVP codec block. For another example, the area can be designed based on the CU size or the picture size.
[0144] In some embodiments, hierarchical motion refinement is used after all sub-block motions of the sbTMVP codec block are derived. For example, a sub-block in each 32x32 of the sbTMVP codec block is used to derive a motion offset. If the current sbTMVP codec block is 64x64, 4 motion offsets will be derived based on the sub-blocks in 4 32x32 areas, respectively. Afterwards, in a 32x32 area, 4 motion offsets will be derived based on the sub-blocks in each 16x16 area, respectively, and the derived motion offsets will be added to the corresponding sub-block motions.
[0145] In some embodiments, hierarchical motion refinement is used after deriving all sub-block motions of the sbTMVP codec block. Bilateral matching of deeper depths is applied only when the bilateral matching cost of shallow depth motion refinement is greater than a threshold. The threshold can be designed based on the CU size. For example, the bilateral matching cost of each 32x32 region is calculated after adding the corresponding motion offset. In a 32x32 region, bilateral matching with sub-blocks in a 16x16 region is performed only when the calculated bilateral matching cost is greater than half the CU size.
[0146] In some embodiments, bilateral matching can be replaced by the above-mentioned template matching. When template matching is used instead of bilateral matching, only the sub-blocks located at the top and left boundaries of the CU are used to determine the motion offset. For other sub-blocks not located at the top and left boundaries of the CU, the motion offset can be the same as the derived one, or set to zero.
[0147] Any proposed invention and concept may be combined. The regression model associated with the affine inter-picture prediction mode mentioned in the present invention and the embodiments may be a 4-parameter or 6-parameter affine model, or even a 2-parameter translation model. The regression model associated with the inter-picture prediction mode mentioned in the present invention may be a 2-parameter translation model, or may be replaced by any refinement method similar to BDOF.
[0148] M. Combining Affine MP-DMVR and Regression
[0149] In some embodiments, the MV refined by the MP-DMVR related motion refinement method may be further refined by regression-based motion refinement. Figure 8 Conceptually shows the MP-DMVR refined MV further refined through regression. The figure illustrates the three CPMVs derived through the regression process taking the sub-block MV refined by MP-DMVR in the current CU as input.
[0150] First, the three CPMVs are refined separately through template matching or bilateral matching. The regression process takes the sub-block MV in the current CU as input to derive the regression model. The sub-block MV is derived from the refined CPMV and the neighboring sub-blocks. Through bilateral or template matching cost, the regression-based CPMV can be compared with the MP-DMVR refined CPMV to obtain the optimal CPMV of the current CU with the minimum cost.
[0151] Any of the aforementioned proposed methods may be implemented in an encoder and / or a decoder. For example, any of the proposed methods may be implemented in an affine inter-picture prediction module and / or a translational inter-picture prediction module of an encoder and / or a decoder. Alternatively, any of the proposed methods may be implemented as a circuit coupled to an affine inter-picture prediction module and / or a translational inter-picture prediction module of an encoder and / or a decoder.
[0152] VII. Example Video Encoder
[0153] Fig. 9An example of a video encoder 900 that can refine affine candidates through regression is shown. As shown, the video encoder 900 receives an input video signal from a video source 905 and encodes the signal into a bitstream 995. The video encoder 900 has several components or modules for encoding the signal from the video source 905, including at least a transform module 910, a quantization module 911, an inverse quantization module 914, an inverse transform module 915, an intra-frame estimation module 920, an intra-frame prediction module 925, a motion compensation module 930, a motion estimation module 935, an in-loop filter 945, a reconstructed image buffer 950, an MV buffer 965, an MV prediction module 975, and an entropy encoder 990. The motion compensation module 930 and the motion estimation module 935 are part of the inter-frame prediction module 940.
[0154] In some embodiments, modules 910-990 are modules of software instructions being executed by one or more processing units (e.g., processors) of a computing device or electronic device. In some embodiments, modules 910-990 are modules of hardware circuits implemented by one or more integrated circuits (ICs) of an electronic device. Although modules 910-990 are shown as separate modules, some of them can be combined into a single module.
[0155] The video source 905 provides a raw video signal representing pixel data for each frame of video information without compression. A subtractor 908 calculates the difference between the raw video pixel data of the video source 905 and the predicted pixel data 913 from the motion compensation module 930 or the intra-frame image prediction module 925 as a prediction residual 909. The transform module 910 transforms the difference (or residual pixel data or residual signal 908) into transform coefficients (e.g., by performing a discrete cosine transform or DCT). The quantization module 911 quantizes the transform coefficients into quantized data (or quantized coefficients) 912, which are encoded into a bitstream 995 by an entropy encoder 990.
[0156] The inverse quantization module 914 dequantizes the quantized data (or quantized coefficients) 912 to obtain transform coefficients, and the inverse transform module 915 inversely transforms the transform coefficients to generate a reconstructed residual 919. The reconstructed residual 919 is added to the predicted pixel data 913 to generate reconstructed primitive data 917. In some embodiments, the reconstructed primitive data 917 is temporarily stored in a line buffer (not shown) for intra-frame image prediction and spatial MV prediction. The reconstructed primitives are filtered by the in-loop filter 945 and stored in a reconstructed image buffer 950. In some embodiments, the reconstructed image buffer 950 is a storage external to the video codec 900. In some embodiments, the reconstructed image buffer 950 is a storage internal to the video encoder 900.
[0157] The image intra-picture estimation module 920 performs intra-picture prediction based on the reconstructed image data 917 to generate intra-picture prediction data. The intra-picture prediction data is provided to the entropy encoder 990 to be encoded into a bitstream 995. The intra-picture prediction data is also used by the intra-picture prediction module 925 to generate the predicted pixel data 913.
[0158] The motion estimation module 935 performs inter-picture prediction by generating motion vectors to reference pixel data of previously decoded frames of information stored in the reconstructed picture buffer 950. These motion vectors are provided to the motion compensation module 930 to generate predicted pixel data.
[0159] Instead of encoding the complete actual MV in the bitstream, the video codec 900 uses MV prediction to generate a predicted MV, and the difference between the MV used for motion compensation and the predicted MV is encoded as residual motion data and stored in the bitstream 995.
[0160] The motion vector prediction module 975 generates a predicted motion vector, i.e., a motion compensated motion vector used to perform motion compensation, based on the reference motion vector generated for encoding the previous video information frame. The motion vector prediction module 975 retrieves the reference motion vector from the previous video information frame from the motion vector buffer 965. The video encoder 900 stores these motion vectors generated for the current video information frame in the motion vector buffer 965 as reference motion vectors for generating the predicted motion vector.
[0161] The motion vector prediction module 975 uses the reference motion vector to create a predicted motion vector. The predicted motion vector can be calculated by spatial motion vector prediction or temporal motion vector prediction. The difference (residual motion data) between the predicted motion vector and the motion compensation motion vector (motion compensation MV, MC MV) of the current information frame is encoded into a bit stream 995 by the entropy encoder 990.
[0162] Using entropy coding techniques, such as context-adaptive binary arithmetic coding (CABAC) or Huffman coding, the entropy encoder 990 encodes various parameters and data into a bitstream 995. The entropy encoder 990 encodes various header elements, flags, and quantized transform coefficients 912 and residual motion data as syntax elements into the bitstream 995. In turn, the bitstream 995 is stored in a storage device or transmitted to a decoder via a communication medium such as a network.
[0163] The in-loop filter 945 performs filtering or smoothing operations on the reconstructed image data 917 to reduce encoding and decoding artifacts, especially artifacts located at the boundaries of pixel blocks. In some embodiments, the filtering operation or smoothing operation performed by the in-loop filter 945 includes a deblocking filter (DBF), a sample adaptive offset (SAO) and / or an adaptive loop filter (ALF).
[0164] Fig.10 A portion of a video encoder 900 implementing affine candidate refinement through regression is shown. As shown, the motion estimation module 935 searches the contents of the reconstructed picture buffer 950 to determine the MV for motion compensation. Specifically, the motion estimation module 935 can select a prediction candidate from the prediction candidate list 1020 based on the search results, and provide the selected prediction candidate to the motion compensation module 930 to generate predicted pixel data 913 for inter-picture prediction. The selection can also be provided to the entropy encoder 990 to indicate the selection in the bitstream.
[0165] The prediction candidate list 1020 may include merge candidates and affine candidates. The affine candidate may be an affine MMVD candidate and / or an affine AMVP candidate and / or an affine merge candidate. The prediction candidate is stored in the MV buffer 965. Some of the affine candidates in the list may be refined or generated by the regression model 1010.
[0166] The regression engine 1005 performs regression based on the contents of the reconstructed picture buffer 950 and the MV buffer 965 to generate a regression model 1010. The regression model 1010 may be generated by template regression, which minimizes the difference between samples of the current template adjacent to the current block and samples of the reference template identified by the refined prediction candidates. The regression model 1010 may also be generated by MV regression, which minimizes the difference between (i) sub-block MVs of sub-blocks adjacent to the current block and (ii) sub-block MVs derived from affine candidates of the current block. The regression model may also be a hybrid model of a first model created by template regression and a second model created by MV regression. In some embodiments, only a subset of the prediction candidates in the list are refined by the regression model.
[0167] Fig.11 The process 1100 of refining prediction candidates by deriving a regression model is conceptually shown. In some embodiments, one or more processing units (e.g., processors) of a computing device implementing the encoder 900 perform the process 1100 by executing instructions stored in a computer-readable medium. In some embodiments, the electronic device implementing the encoder 900 performs the process 1100.
[0168] The encoder receives (at frame 1110) data of a pixel block to be encoded as a current block of a current picture of a video. The encoder derives (at frame 1120) a linear model to refine or generate prediction candidates through regression to minimize the difference between a plurality of samples of a current template neighboring the current block and a plurality of samples of a reference template identified by the prediction candidate.
[0169] In some embodiments, the derived prediction candidate being refined is an affine motion candidate, and the derived model is an affine motion model with four parameters based on affine motion of two checkpoints, or an affine motion model with six parameters for three checkpoints. The affine motion candidate being refined is an MMVD candidate or an adaptive motion vector prediction (AMVP) candidate. The prediction candidate being refined can also be a conventional merge candidate.
[0170] In some embodiments, the refined prediction candidate is added to the prediction candidate list of the current block. In some embodiments, the refined prediction candidate replaces the prediction candidate in the prediction candidate list of the current block.
[0171] In some embodiments, only a subset of prediction candidates in the prediction candidate list for the current block is refined through regression of a linear model. The prediction candidate subset may be MMVD candidates identified based on proximity to a particular direction. In some embodiments, for each MMVD candidate, only a base motion vector predictor (MVP) is refined through regression.
[0172] The encoder generates (at frame 1130) a prediction for the current block based on the derived linear model. The prediction may be generated based on the refined prediction candidates. The encoder encodes or decodes (at frame 1140) the current block using the generated prediction to produce a prediction residual.
[0173] In some embodiments, the derived model is an affine motion model, and the affine motion model is derived by minimizing the difference between a plurality of regression sub-block motion vectors and (i) a plurality of sub-block MVs of a plurality of sub-blocks adjacent to the current block and (ii) a plurality of sub-block MVs derived from the affine candidate of the current block to refine the affine candidate. The prediction is generated based on the refined affine candidate. In some embodiments, the affine motion model is derived by minimizing the difference between the plurality of regression sub-block motion vectors and (i) a plurality of stored sub-block MVs associated with a plurality of sub-blocks adjacent to the current block and (ii) a plurality of inherited MVs associated with a plurality of sub-blocks belonging to a reference affine codec unit (CU). The reference affine CU may not be adjacent to the current block.
[0174] VIII. Example Video Decoder
[0175] In some embodiments, the encoder may indicate (or generate) one or more syntax elements in the bitstream, so that the decoder may parse the one or more syntax elements from the bitstream.
[0176] Fig.12 An example of a video decoder 1200 that can refine affine candidates is shown. As shown, the video decoder 1200 is an image decoding or video decoding circuit that receives a bitstream 1295 and decodes the content of the bitstream into pixel data of a video information frame for display. The video decoder 1200 has several components or modules for decoding the bitstream 1295, including some components selected from an inverse quantization module 1211, an inverse transform module 1210, an intra-picture prediction module 1225, a motion compensation module 1230, an in-loop filter 1245, a decoded image buffer 1250, a motion vector buffer 1265, a motion vector prediction module 1275, and a parser 1290. The motion compensation module 1230 is part of the inter-picture prediction module 1240.
[0177] In some embodiments, modules 1210-1290 are modules of software instructions being executed by one or more processing units (e.g., processors) of a computing device. In some embodiments, modules 1210-1290 are hardware circuit modules implemented by one or more ICs of an electronic device. Although modules 1210-1290 are illustrated as separate modules, some of these modules may be combined into a single module.
[0178] The parser 1290 (or entropy decoder) receives the bitstream 1295 and performs initial parsing according to the syntax defined by the video codec or image codec standard. The parsed syntax elements include various mark elements, flags, and quantized data (or quantized coefficients) 1212. The parser 1290 parses out the various syntax elements by using entropy coding techniques (such as context-adaptive binary arithmetic coding (CABAC) or Huffman coding).
[0179] The inverse quantization module 1211 dequantizes the quantized data (or quantized coefficients) 1212 to obtain transform coefficients, and the inverse transform module 1210 inversely transforms the transform coefficients 1216 to generate reconstructed residuals 1219. The reconstructed residuals 1219 are added to the predicted pixel data 1213 from the intra-picture prediction module 1225 or the motion compensation module 1230 to generate decoded pixel data 1217. The decoded pixel data is filtered by the in-loop filter 1245 and stored in the decoded image buffer 1250. In some embodiments, the decoded image buffer 1250 is a storage outside the video decoder 1200. In some embodiments, the decoded image buffer 1250 is a storage inside the video decoder 1200.
[0180] The intra-picture prediction module 1225 receives the intra-picture prediction data from the bitstream 1295 and generates predicted pixel data 1213 based on the intra-picture prediction data from the decoded pixel data 1217 stored in the decoded image buffer 1250. In some embodiments, the decoded pixel data 1217 is also stored in a line buffer (not shown) for image intra-picture prediction and spatial MV prediction.
[0181] In some embodiments, the contents of the decoded image buffer 1250 are used for display. The display device 1255 retrieves the contents of the decoded image buffer 1250 for direct display, or retrieves the contents of the decoded image buffer to a display buffer. In some embodiments, the display device receives pixel values from the decoded image buffer 1250 via pixel transfer.
[0182] Based on the motion compensated MV (MC MV), the motion compensation module 1230 generates predicted pixel data 1213 from the decoded pixel data 1217 stored in the decoded picture buffer 1250. These motion compensated MVs are decoded by adding the residual motion data received from the bitstream 1295 to the predicted MV received from the motion vector prediction module 1275.
[0183] The motion vector prediction module 1275 generates a predicted MV, for example, a motion compensated MV for performing motion compensation, based on a reference MV generated for decoding a previous video information frame. The motion vector prediction module 1275 retrieves the reference motion vector of the previous video information frame from the motion vector buffer 1265. The video decoder 1200 also stores the motion compensated motion vector generated for decoding the current video information frame in the motion vector buffer 1265 as a reference motion vector for generating a predicted motion vector.
[0184] The in-loop filter 1245 performs filtering or smoothing operations on the decoded pixel data to reduce encoding and decoding artifacts, especially artifacts located at the boundaries of pixel blocks. In some embodiments, the filtering or smoothing operations performed by the in-loop filter 1245 include a deblocking filter (DBF), a sample adaptive offset (SAO), and / or an adaptive loop filter (ALF).
[0185] Fig.13A partial video decoder 1200 is shown implementing refinement of affine candidates through regression. As shown, an entropy decoder 1290 may select a prediction candidate from a prediction candidate list 1320 based on syntax elements indicated in a bitstream 1295. The selected prediction candidate is provided to a motion compensation module 1230 to generate predicted pixel data 1213 for inter-picture prediction.
[0186] The prediction candidate list 1320 may include merge candidates and affine candidates. The affine candidate may be an affine MMVD candidate and / or an affine AMVP candidate and / or an affine merge candidate. The prediction candidates of the list 1320 are stored in the MV buffer 1265. Some of the affine candidates in the list may be refined or generated by the regression model 1310.
[0187] The regression engine 1305 performs regression based on the contents of the decoded picture buffer 1250 and the MV buffer 1265 to generate a regression model 1310. The regression model 1310 may be generated by template regression, which minimizes the difference between samples of the current template adjacent to the current block and samples of the reference template identified by the refined prediction candidates. The regression model 1310 may also be generated by MV regression, which minimizes the difference between (i) the sub-block MVs of the sub-blocks adjacent to the current block and (ii) the sub-block MVs derived from the affine candidates of the current block. The regression model may also be a hybrid model of a first model created by template regression and a second model created by MV regression. In some embodiments, only a subset of the prediction candidates in the list are refined by the regression model 1310.
[0188] Fig.14 The process 1400 of refining prediction candidates by deriving a regression model is conceptually illustrated. In some embodiments, one or more processing units (e.g., processors) of a computing device implementing the decoder 1200 perform the process 1400 by executing instructions stored in a computer-readable medium. In some embodiments, an electronic device implementing the decoder 1200 performs the process 1400.
[0189] The decoder receives (at frame 1410) data of a pixel block of a current block of a current picture of a video to be decoded. The decoder derives (at frame 1420) a linear model to refine or generate prediction candidates through regression to minimize the difference between a plurality of samples of a current template neighboring the current block and a plurality of samples of a reference template identified by the prediction candidate.
[0190] In some embodiments, the derived prediction candidate being refined is an affine motion candidate, and the derived model is an affine motion model with four parameters based on affine motion of two checkpoints, or an affine motion model with six parameters for three checkpoints. The affine motion candidate being refined is an MMVD candidate or an adaptive motion vector prediction (AMVP) candidate. The prediction candidate being refined can also be a conventional merge candidate.
[0191] In some embodiments, the refined prediction candidate is added to the prediction candidate list of the current block. In some embodiments, the refined prediction candidate replaces the prediction candidate in the prediction candidate list of the current block.
[0192] In some embodiments, only a subset of prediction candidates in the prediction candidate list for the current block is refined through regression of a linear model. The prediction candidate subset may be MMVD candidates identified based on proximity to a particular direction. In some embodiments, for each MMVD candidate, only a base motion vector predictor (MVP) is refined through regression.
[0193] The decoder generates (at frame 1430) a prediction for the current block based on the derived linear model. The prediction may be generated based on the refined prediction candidates. The decoder reconstructs (at frame 1440) the current block using the generated prediction. The decoder may then provide the reconstructed current block for display as part of the reconstructed current image.
[0194] In some embodiments, the derived model is an affine motion model, and the affine motion model is derived by minimizing the difference between a plurality of regression sub-block motion vectors and (i) a plurality of sub-block MVs of a plurality of sub-blocks adjacent to the current block and (ii) a plurality of sub-block MVs derived from the affine candidate of the current block to refine the affine candidate. The prediction is generated based on the refined affine candidate. In some embodiments, the affine motion model is derived by minimizing the difference between the plurality of regression sub-block motion vectors and (i) a plurality of stored sub-block MVs associated with a plurality of sub-blocks adjacent to the current block and (ii) a plurality of inherited MVs associated with a plurality of sub-blocks belonging to a reference affine codec unit (CU). The reference affine CU may not be adjacent to the current block.
[0195] IX. Electronic System Examples
[0196] Many of the above features and applications can be implemented as software processes, which are specified as sets of instructions recorded on a computer readable storage medium (Computer Readable Storage Medium) (also referred to as a computer readable medium). When these instructions are executed by one or more computing units or processing units (e.g., one or more processors, processor cores or other processing units), these instructions cause the processing units to perform the actions represented by these instructions. Examples of computer readable media include, but are not limited to, CD-ROMs, flash drives, random access memory (Random Access Memory, RAM) chips, hard disks, erasable programmable read only memory (Erasable Programmable Read Only Memory, EPROM), electrically erasable programmable read only memory (Electrically Erasable Programmable Read-only memory, EEPROM), etc. Computer readable media do not include carriers and telecommunications signals connected wirelessly or wired.
[0197] In this specification, the term "software" is meant to include firmware in a read-only memory or an application stored in a magnetic storage device, which can be read into the memory for processing by the processor. At the same time, in some embodiments, multiple software inventions can be implemented as sub-parts of a larger program, while retaining different software inventions. In some embodiments, multiple software inventions can be implemented as independent programs. Finally, any combination of independent programs that implement the software inventions described herein together is within the scope of the present invention. In some embodiments, when installed to operate on one or more electronic systems, the software program defines one or more specific machine implementations that execute and implement the operations of the software program.
[0198] Fig.15An electronic system 1500 is conceptually shown for implementing some embodiments of the present invention. The electronic system 1500 may be a computer (e.g., a desktop computer, a personal computer, a tablet computer, etc.), a phone, a PDA, or any other kind of electronic device. Such an electronic system includes various types of computer-readable media and interfaces for various other types of computer-readable media. The electronic system 1500 includes a bus 1505, a processing unit 1510, a graphics processing unit (GPU) 1515, a system memory 1520, a network 1525, a read-only memory (ROM) 1530, a permanent storage device 1535, an input device 1540, and an output device 1545.
[0199] Bus 1505 collectively represents all system buses, peripheral buses, and chipset buses that communicatively couple a multitude of internal devices of electronic system 1500. For example, bus 1505 communicatively couples processing unit 1510 with GPU 1515, read-only memory 1530, system memory 1520, and permanent storage device 1535.
[0200] From various memory units, the processing unit 1510 retrieves instructions to be executed and data to be processed in order to perform the method process of the present invention. In different embodiments, the processing unit can be a single processor or a multi-core processor. Certain instructions are transmitted to and executed by the GPU 1515. The GPU 1515 can offload various calculations or supplement the image processing provided by the processing unit 1510.
[0201] Read-only memory (ROM) 1530 stores static data and instructions required by processing unit 1510 or other modules of the electronic system. On the other hand, permanent storage device 1535 is a read-and-write memory device. This device is a non-volatile memory unit that stores instructions and data even when electronic system 1500 is turned off. In some embodiments of the present invention, a large-capacity storage device (such as a magnetic disk or optical disk and its corresponding magnetic disk drive) is used as permanent storage device 1535.
[0202] In other embodiments, removable storage devices (such as floppy disks, flash memory devices, etc., and their corresponding disk drives) are used as permanent storage devices. Like the permanent storage device 1535, the system memory 1520 is a read-write memory device. However, unlike the permanent storage device 1535, the system memory 1520 is a volatile read-write memory, such as a random access memory. The system memory 1520 stores some instructions and data used by the processor at runtime. In some embodiments, the method process according to the present invention is stored in the system memory 1520, the permanent storage device 1535 and / or the read-only memory 1530. For example, various memory units include instructions for processing multimedia clips according to some embodiments. For these various memory units, the processing unit 1510 retrieves the executed instructions and processed data in order to perform the processing of certain embodiments.
[0203] The bus 1505 is also connected to an input device 1540 and an output device 1545. The input device 1540 enables a user to communicate information and select instructions to the electronic system. The input device 1540 includes an alphanumeric keyboard and a pointing device (also known as a "cursor control device"), a camera (such as a webcam), a microphone or similar device for receiving voice commands, etc. The output device 1545 displays images generated by the electronic system or outputs data in other ways. The output device 1545 includes a printer and a display device, such as a cathode ray tube (CRT) or a liquid crystal display (LCD), and a speaker or similar audio output device. Some embodiments include devices such as a touch screen that serves as both an input device and an output device.
[0204] Finally, if Fig.15 As shown, bus 1505 also couples electronic system 1500 to network 1525 via a network adapter (not shown). In this manner, the computer may be part of a computer network (e.g., a local area network (LAN), a wide area network (WAN), or an intranet) or a network of networks (e.g., the Internet). Any or all components of electronic system 1500 may be used in conjunction with the present invention.
[0205] Some embodiments include electronic components, such as a microprocessor, a storage device, and a memory, which stores computer program instructions to a machine-readable medium or a computer-readable medium (optionally referred to as a computer-readable storage medium, a machine-readable medium, or a machine-readable storage medium). Some examples of computer-readable media include RAM, ROM, read-only compact discs (CD-ROM), recordable compact discs (CD-R), rewritable compact discs (CD-RW), read-only digital versatile discs (e.g., DVD-ROM, dual-layer DVD-ROM), various recordable / rewritable DVDs (e.g., DVD RAM, DVD-RW, DVD+RW, etc.), flash memory (e.g., SD card, mini SD card, micro SD card, etc.), magnetic and / or solid-state hard drives, read-only and recordable Computer readable media may store computer programs executed by at least one processing unit and include sets of instructions for performing various operations. Examples of computer programs or computer code include machine code, such as that produced by a compiler, and files containing high-level code executed by a computer, electronic component, or microprocessor using an interpreter.
[0206] While the above discussion refers primarily to microprocessors or multi-core processors that execute software, many of the above functions and applications are performed by one or more integrated circuits, such as application-specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). In some embodiments, such integrated circuits execute instructions stored on the circuits themselves. In addition, some embodiments execute software stored in programmable logic devices (PLDs), ROMs, or RAM devices.
[0207] Embodiments of the video processing method for encoding or decoding a bidirectional prediction block can be implemented in a circuit integrated into a video compression chip or a program code integrated into video compression software for performing the above-mentioned processing. For example, the selection of a weight set from a plurality of weight sets to encode a current block can be implemented in a program code to be executed on a computer processor, a digital signal processor (DSP), a microprocessor, or a field programmable gate array (FPGA). These processors can be configured to perform specific tasks according to the present invention by executing machine-readable software code or firmware code that defines a specific method specifically implemented according to the present invention.
[0208] References throughout this specification to "an embodiment," "some embodiments," or similar language mean that a particular feature, structure, or characteristic described in conjunction with the embodiment may be included in at least one embodiment of the present invention. Therefore, the phrases "in an embodiment" or "in some embodiments" appearing in various places throughout this specification do not necessarily all refer to the same embodiment, which may be implemented alone or in combination with one or more other embodiments. Moreover, the particular features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments. However, those skilled in the relevant art will recognize that the present invention may be specifically practiced in the absence of one or more of these specific details or using other methods, components, etc. In other cases, well-known structures or operations are not shown or described in detail to avoid confusing aspects of the present invention.
[0209] The present invention may be specifically implemented in other specific forms without departing from the spirit and essential features of the present invention. The examples are considered in all respects as examples and not limitations. Therefore, the scope of the present invention is indicated by the appended claims rather than the foregoing description. All changes within the meaning and scope of the equivalents of the claims will be included within the scope of these claims.
Claims
1. A video encoding and decoding method for affine candidate refinement, the method comprising the following steps: Receiving data of a pixel block to be encoded or decoded as a current block of a current picture of a video; Inducing a linear model to refine or generate prediction candidates through regression to minimize the difference between a plurality of samples of a current template adjacent to the current block and a plurality of samples of a reference template identified by the prediction candidate; generating a prediction for the current block based on the derived linear model; as well as The current block is encoded or decoded using the generated prediction.
2. The method according to claim 1, characterized in that The derived prediction candidate being refined is an affine motion candidate, and the derived model is an affine motion model with four parameters based on affine motion of two checkpoints, or an affine motion model with six parameters for three checkpoints.
3. The method according to claim 2, characterized in that The affine motion candidate being refined is a Merge candidate with Motion Vector Difference (MMVD) or an Adaptive Motion Vector Prediction (AMVP) candidate.
4. The method according to claim 1, characterized in that: The refined prediction candidate is added to the prediction candidate list of the current block.
5. The method according to claim 1, characterized in that The refined prediction candidate replaces the prediction candidate in the prediction candidate list of the current block.
6. The method according to claim 1, characterized in that Through the regression of the linear model, only a subset of prediction candidates in the prediction candidate list of the current block is refined.
7. The method according to claim 6, characterized in that The prediction candidate subset is a plurality of MMVD candidates identified based on proximity to a specific direction.
8. The method according to claim 6, characterized in that For each MMVD candidate, only the base Motion Vector Predictor (MVP) is refined through regression.
9. The method according to claim 1, characterized in that: The prediction is generated based on the refined prediction candidates.
10. The method according to claim 1, characterized in that The prediction candidate being refined is a regular merge candidate.
11. The method according to claim 1, characterized in that: The derived model is an affine motion model derived by minimizing differences between a plurality of regressed sub-block motion vectors and (i) a plurality of sub-block MVs of a plurality of sub-blocks adjacent to the current block and (ii) a plurality of sub-block MVs derived from the affine candidates of the current block to refine the affine candidates; The prediction is generated based on the refined affine candidates.
12. The method according to claim 11, characterized in that The affine motion model is derived by minimizing the difference between the plurality of regressed sub-block motion vectors and (i) a plurality of stored sub-block MVs associated with a plurality of sub-blocks adjacent to the current block and (ii) a plurality of inherited MVs associated with a plurality of sub-blocks belonging to a reference affine codec unit (CU).
13. The method according to claim 12, characterized in that The reference affine CU is not adjacent to the current block.
14. A video encoding and decoding method for affine candidate refinement, the method comprising the following steps: Receiving data of a pixel block to be encoded or decoded as a current block of a current picture of a video; Inducing a linear model to refine the affine candidate by regression to minimize the difference between a plurality of regressed sub-block motion vectors and (i) a plurality of sub-block MVs of a plurality of sub-blocks adjacent to the current block and (ii) a plurality of sub-block MVs derived from the refined affine candidate of the current block; generating a prediction for the current block based on the derived linear model; as well as The current block is encoded or decoded using the generated prediction.
15. An electronic device for video encoding and decoding for affine candidate refinement, the electronic device comprising: A video decoder circuit for performing operations including: Receiving data of a pixel block to be encoded or decoded as a current block of a current picture of a video; Inducing a linear model to refine or generate prediction candidates through regression to minimize the difference between a plurality of samples of a current template adjacent to the current block and a plurality of samples of a reference template identified by the prediction candidate, generating a prediction for the current block based on the derived linear model; and The current block is encoded or decoded using the generated prediction.
16. A video decoding method for affine candidate refinement, the method comprising the following steps: Receiving data of a pixel block to be decoded as a current block of a current picture of a video; Inducing a linear model to refine or generate prediction candidates through regression to minimize the difference between a plurality of samples of a current template adjacent to the current block and a plurality of samples of a reference template identified by the prediction candidate; generating a prediction for the current block based on the derived linear model; as well as Reconstruct the current block by using the generated prediction.
17. A video coding method for affine candidate refinement, the method comprising the following steps: Receiving data of a pixel block to be encoded as a current block of a current picture of a video; Inducing a linear model to refine or generate prediction candidates through regression to minimize the difference between a plurality of samples of a current template adjacent to the current block and a plurality of samples of a reference template identified by the prediction candidate; generating a prediction for the current block based on the derived linear model; as well as The current block is encoded using the generated prediction.