Improvements in illumination compensation in video coding
By applying the multi-parameter local lighting compensation (MPLIC) method in video data processing, the problem of video data bandwidth management is solved, and more efficient video encoding and decoding is achieved.
Patent Information
- Application Number
- CN202380079923.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-16
- Filing Date
- 2023-11-16
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, it is difficult to effectively manage bandwidth requirements when processing video data, especially in scenarios where digital video occupies network bandwidth.
The multi-parameter local lighting compensation (MPLIC) method is used to process the visual media data, and the conversion between the visual media data and the bit stream is performed through the MPLIC.
Through the MPLIC method, the bandwidth requirements of video data can be managed more efficiently, and the efficiency and quality of video encoding and decoding can be improved.
Smart Images

Figure CN120226356A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This patent application claims the benefit of International Patent Application No. PCT / CN2022 / 132234, filed on November 16, 2022, the teachings and disclosures of which are incorporated herein by reference in their entirety. Technical field
[0003] This patent document relates to the generation, storage, and consumption of digital audio - visual media information in a file format. Background art
[0004] Digital video occupies the largest bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video may continue to grow. Summary of the invention
[0005] The first aspect relates to a method for processing video data, including: determining to apply multi - parameter local illumination compensation (MPLIC) to visual media data; and performing a conversion between the visual media data and a bitstream based on MPLIC.
[0006] The second aspect relates to a device for processing video data, including: a processor; and a non - transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform any one of the foregoing aspects.
[0007] The third aspect relates to a non - transitory computer - readable medium, including a computer program product for use in a video codec device, the computer program product including computer - executable instructions stored on the non - transitory computer - readable medium such that, when executed by a processor, cause the video codec device to perform the method according to any one of the foregoing aspects.
[0008] The fourth aspect relates to a non - transitory computer - readable recording medium that stores a bitstream of a video generated by a method executed by a video processing device, wherein the method includes: determining to apply multi - parameter local illumination compensation (MPLIC) to visual media data; and generating a bitstream based on the determination.
[0009] The fifth aspect relates to a method for storing a bitstream of a video, including: determining to apply multi - parameter local illumination compensation (MPLIC) to visual media data; generating a bitstream based on the determination; and storing the bitstream in a non - transitory computer - readable recording medium.
[0010] For clarity, any one of the foregoing embodiments may be combined with any one or more of the foregoing other embodiments to create new embodiments within the scope of the present disclosure.
[0011] These and other features will become more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] For a more complete understanding of the present disclosure, reference is now made to the following brief description taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.
[0013] Figure 1 Example locations of spatial merge candidates are shown.
[0014] Figure 2 Example candidate pairs considered for redundancy check of spatial merge candidates are shown.
[0015] Figure 3 Example candidate positions of time-domain Merge candidates are shown.
[0016] Figure 4 An example Merge mode with a motion vector difference (MMVD) search point is shown.
[0017] Figure 5 An example of a symmetric motion vector difference (MVD) pattern is shown.
[0018] Figure 6 An example of an affine motion model based on control points is shown.
[0019] Figure 7 An example affine motion vector field (MVF) for each sub-block is shown.
[0020] Figure 8 Example locations of inherited affine motion predictors are shown.
[0021] Figure 9 An example of control point motion vector inheritance is shown.
[0022] Figure 10 Example positioning of candidate locations for constructing an affine Merge pattern is shown.
[0023] Figure 11 An example of motion vector (MV) usage of the combined method is shown.
[0024] Figure 12 The sub-block MV V is shown SB and pixels example.
[0025] Figure 13 An example of a sub-block based temporal motion vector prediction (sbTMVP) process in Versatile Video Codec (VVC) is shown.
[0026] Figure 14 Shows an example extended codec unit (CU) region used in bidirectional optical flow (BDOF).
[0027] Figure 15 Shows an example of decoder-side motion vector refinement.
[0028] Figure 16 Shows an example of partitioning by geometric partition mode (GPM) grouped by the same angle.
[0029] Figure 17 Shows an example of unidirectional prediction MV selection for geometric partition mode.
[0030] Figure 18 Shows an example of generating bending weights using geometric partition mode.
[0031] Figure 19 Shows examples of top and left neighboring blocks used in combined inter and intra prediction (CIIP) weight derivation.
[0032] Figure 20 Shows an example of spatial neighboring blocks for deriving spatial Merge candidates.
[0033] Figure 21 Shows an example of template matching performed on a search area around the initial MV.
[0034] Figure 22 Shows an example of a diamond region in the search area.
[0035] Figure 23 Shows examples of the frequency responses of the interpolation filter and the VVC interpolation filter at the half-pixel phase.
[0036] Figure 24 Shows examples of templates in the reference picture and reference sample points of the templates.
[0037] Figure 25 Shows examples of templates of blocks with sub-block motion having motion information of sub-blocks using the current block and reference sample points of the templates.
[0038] Figure 26 Shows an example of virtual block generation for improved DC prediction.
[0039] Figure 27 Is a block diagram showing an example video processing system.
[0040] Figure 28 Is a block diagram of an example video processing device.
[0041] Figure 29It is a flowchart of an example method for video processing.
[0042] Figure 30 It is a block diagram showing an example video codec system.
[0043] Figure 31 It is a block diagram showing an example encoder.
[0044] Figure 32 It is a block diagram showing an example decoder.
[0045] Figure 33 It is a schematic diagram of an example encoder. Detailed implementation
[0046] First, it should be understood that although the following provides exemplary implementations of one or more embodiments, any number of techniques can be used to implement the disclosed systems and / or methods, whether currently known techniques or those to be developed. The present disclosure should not be limited in any way to the exemplary implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but can be modified within the full scope of the appended claims and their equivalents.
[0047] The use of section headings in this document is for ease of understanding and does not limit the applicability of the techniques and embodiments disclosed in each section to that section. Additionally, the techniques described herein are also applicable to other video codec protocols and designs.
[0048] 1. Preliminary discussion
[0049] This document relates to video codec technology. Specifically, it relates to inter-frame prediction in video coding and decoding, with a focus on sequences with illumination changes. These ideas can be applied alone or in various combinations to image / video coding standards and / or other image / video codecs, such as next-generation image / video coding standards.
[0050] 2.1 Inter-frame prediction in VVC
[0051] For each inter-predicted CU, motion parameters including motion vectors, reference picture indices, and reference picture lists used for indexing, as well as additional information required for the new coding features of VVC are used for inter-predicted sample generation. The motion parameters can be signaled in an explicit or implicit manner. When a CU is coded / decoded in skip mode, the CU is associated with a prediction unit (PU) and does not have valid residual coefficients, coding / decoding motion vector deltas, or reference picture indices. The Merge mode is specified to obtain the motion parameters of the current CU from neighboring CUs, including spatial and temporal candidates, and additional scheduling introduced in VVC. The Merge mode can be applied to any inter-predicted CU, not just for skip mode. An alternative to the Merge mode is the explicit signaling of motion parameters, where the motion vectors, corresponding reference picture indices for each reference picture list, reference picture list usage flags, and other required information are signaled explicitly for each CU.
[0052] In addition to the inter-coding features in High Efficiency Video Coding (HEVC), VVC includes multiple inter-prediction coding tools listed below:
[0053] – Extended Merge prediction
[0054] – High-precision (1 / 16 pixel) motion compensation and motion vector storage
[0055] – Merge mode with MVD (MMVD)
[0056] – Symmetric MVD (SMVD) signaling
[0057] – Affine motion compensation prediction
[0058] – Sub-block based temporal motion vector prediction (SbTMVP)
[0059] – Adaptive motion vector resolution (AMVR)
[0060] – Bi-directional prediction with CU-level weights (BCW)
[0061] – Bi-directional optical flow (BDOF)
[0062] – Decoder-side motion vector refinement (DMVR)
[0063] – Geometric partitioning mode (GPM)
[0064] – Combined inter and intra prediction (CIIP)
[0065] – Reference picture resampling
[0066] The following text provides details on the inter-prediction methods specified in VVC.
[0067] 2.1.1 Extended Merge Prediction
[0068] In VVC, the Merge candidate list is constructed by including the following five types of candidates, in order:
[0069] (1) Spatial motion vector prediction (MVP) from spatial neighboring CUs
[0070] (2) Temporal MVP from co-located CUs
[0071] (3) History-based MVP from the first-in-first-out (FIFO) table
[0072] (4) Paired-average MVP
[0073] (5) Zero MV.
[0074] The size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 6. For each CU coding in the Merge mode, the index of the best Merge candidate is encoded using truncated unary binary (TU). The first binary bit (bin) of the Merge index is context-coded, and the other binary bits are bypass-coded.
[0075] This section will introduce the derivation process of each Merge candidate category. Similar to HEVC, VVC also supports parallel derivation of the Merge candidate list for all CUs within a certain size area.
[0076] 2.1.1.1 Spatial Candidate Derivation
[0077] Figure 1 Shows an example position of a spatial Merge candidate. Figure 2 Shows an example candidate pair considered for redundancy checking of spatial Merge candidates. The derivation of spatial Merge candidates in VVC is the same as that in HEVC, except that the positions of the first two Merge candidates are swapped. Up to four Merge candidates are selected from the candidates at the Figure 1 shown positions. The derivation order is B0, A0, B1, A1, and B2. Position B2 is considered only if one or more of the CUs at positions B0, A0, B1, A1 are unavailable (e.g., because they belong to other strips or slices) or are intra-coded. After adding the candidate at position A1, a redundancy check is performed on the addition of the remaining candidates to ensure that candidates with the same motion information are excluded from the list, thereby improving the coding efficiency. To reduce the computational complexity, not all possible candidate pairs are considered in the above redundancy check. Instead, only pairs related to Figure 2pairs linked by arrows in it, and the candidate is added to the list only when the corresponding candidates for redundancy check do not have the same motion information.
[0078] 2.1.1.2 Temporal candidate derivation
[0079] Figure 3 An example candidate position of the temporal Merge candidate including C0 and C1 is shown. In this step, only one candidate is added to the list. Specifically, when deriving this temporal Merge candidate, the scaled motion vectors are derived based on the co-located CUs belonging to the co-located reference pictures. The reference picture list and reference index for deriving the co-located CUs are explicitly signaled in the slice header. As Figure 3 shown by the dashed line in, the scaled motion vector of the temporal Merge candidate is obtained, which is scaled from the motion vector of the co-located CU using the picture order count (POC) distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal Merge candidate is set to be equal to zero.
[0080] As Figure 2 shown, the position of the temporal candidate is selected between candidates C0 and C1. If the CU at position C0 is unavailable, intra-coded / decoded, or outside the current row of the CTU, position C1 is used. Otherwise, position C0 is used when deriving the temporal Merge candidate.
[0081] 2.1.1.3 History-based Merge candidate derivation
[0082] After the spatial MVP and the temporal motion vector prediction (TMVP), the history-based MVP (HMVP) Merge candidates are added to the Merge list. In this method, the motion information of the previously coded blocks is stored in a table and used as the MVP for the current CU. A table with multiple HMVP candidates is maintained during the encoding / decoding process. When a new CTU row is encountered, the table is reset (emptied). As long as there is a non-sub-block inter-coded CU, the associated motion information is added as a new HMVP candidate to the last entry of the table.
[0083] The size S of the HMVP table is set to 6, which means that at most 5 history-based MVP (HMVP) candidates can be added to the table. When inserting a new motion candidate into the table, a constrained first-in-first-out (FIFO) rule is used, where first a redundancy check is applied to find if there is the same HMVP in the table. If found, the same HMVP is deleted from the table, then all the subsequent HMVP candidates are moved forward, and the same HMVP is inserted into the last entry of the table.
[0084] HMVP candidates can be used during the Merge candidate list construction process. The last several HMVP candidates in the table are checked in order and inserted into the candidate list after the TMVP candidates. Redundancy checks are applied to the HMVP candidates for spatial or temporal Merge candidates.
[0085] To reduce the number of redundancy check operations, the following simplifications are introduced:
[0086] 1. Redundancy checks are performed on the A1 and B1 spatial candidates for the last two entries in the table respectively.
[0087] 2. Once the total number of available Merge candidates reaches the maximum allowed Merge candidates minus 1, the process of constructing the Merge candidate list from HMVP is terminated.
[0088] 2.1.1.4 Pairwise Average Merge Candidate Derivation
[0089] Pairwise average candidates are generated by averaging predefined candidate pairs in the existing Merge candidate list using the first two Merge candidates. The first Merge candidate can be defined as p0Cand and the second Merge candidate can be defined as p1Cand respectively. For each reference list individually, the average motion vector is calculated based on the availability of the motion vectors of p0Cand and p1Cand. If both motion vectors are available in a list, they are averaged even if the two motion vectors point to different reference pictures, and their reference picture is set to one of those in p0Cand; if only one motion vector is available, that vector is directly used; if no motion vector is available, this list is kept invalid. In addition, if the half-pixel interpolation filter indices of p0Cand and p1Cand are different, they are set to 0.
[0090] When the Merge list is not full after adding pairwise average Merge candidates, zero MVPs are inserted at the end until the maximum number of Merge candidates is reached.
[0091] 2.1.1.5 Merge Estimation Region
[0092] The Merge Estimation Region (MER) allows for the independent derivation of the Merge candidate list for CUs within the same MER. When generating the Merge candidate list for the current CU, candidate blocks within the same MER as the current CU are not included. Additionally, the update process of the history-based motion vector predictor candidate list is only updated when (xCb + cbWidth) >> Log2ParMrgLevel is greater than xCb >> Log2ParMrgLevel and (yCb + cbHeight) >> log2ParMrgHevel is greater than (yCb >> log2parMrgLevel), where (xCb, yCb) is the top-left luma sample position of the current CU in the picture and (cbWidth, cbHeight) is the CU size. The MER size is selected on the encoder side and signaled in the sequence parameter set as log2_parallel_merge_level_minus2.
[0093] 2.1.2 High-precision (1 / 16 pixel) motion compensation and motion vector storage
[0094] VVC increases the MV precision to 1 / 16 luma samples to improve the prediction efficiency of slow-motion videos. For example, in the case of the affine mode, this higher motion precision is particularly helpful for video content with local variations and non-translational motion. To generate fractional position samples with higher MV precision, the 8-tap luma interpolation filter and 4-tap chroma interpolation filter of HEVC are extended to 16 phases for luma and 32 phases for chroma. This extended filter set is applied to the MC process of inter-frame coded CUs, except for CUs in the affine mode. For the affine mode, a set of 6-tap luma interpolation filters with 16 phases is used to reduce the computational complexity and save memory bandwidth.
[0095] In VVC, the highest precision of the explicitly signaled motion vectors for non-affine CUs is one-quarter luma samples. In some inter-frame prediction modes such as the affine mode, the motion vectors can be signaled with 1 / 16 luma sample precision. In all inter-frame coded CUs with implicitly inferred MVs, the MVs are derived with 1 / 16 luma sample precision and the motion compensation prediction is performed with 1 / 16 sample precision. In terms of internal motion field storage, all motion vectors are stored with 1 / 16 luma sample precision.
[0096] For the temporal motion field storage used by TMVP and SbTMVP, the motion field compression is performed at an 8×8 size granularity as compared to the 16×16 size granularity in HEVC.
[0097] 2.1.3 Merge Mode with MVD (MMVD)
[0098] In addition to the Merge mode that directly uses the implicitly derived motion information for the prediction sample generation of the current CU, the VVC also introduces a Merge mode with motion vector difference (MMVD). After sending the regular Merge flag, the MMVD flag is immediately signaled to specify whether the MMVD mode is used for the CU.
[0099] In MMVD, after selecting a Merge candidate, it is further refined by the signaled MVD information. The further information includes the Merge candidate flag, an index for specifying the motion amplitude, and an index for indicating the motion direction. In the MMVD mode, one of the first two candidates in the Merge list is selected as the MV basis. The mmvd candidate flag is signaled to specify which one of the first and second Merge candidates is used.
[0100] Figure 4 An example MMVD search point is shown. The distance index specifies the motion amplitude information and indicates a predefined offset from the starting point. As Figure 4 shown, the offset is added to either the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 3-6.
[0101] Table 3-6 - Relationship between distance index and predefined offset
[0102]
[0103] The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent four directions, as shown in Table 3-7. It should be noted that the meaning of the MVD sign can vary according to the information of the starting MV. When the starting MV is a unidirectional prediction MV or a bidirectional prediction MV where both lists point to the same side of the current picture (i.e., both reference POCs are greater than the current picture POC, or both are less than the current picture POC), the signs in Table 3-7 specify the signs of the MV offsets added to the starting MV. When the starting MV is a bidirectional prediction MV where the two MVs point to different sides of the current picture (i.e., one reference POC is greater than the current picture POC and the other is less than the current picture POC), and the difference in POCs in list 0 is greater than the difference in POCs in list 1, the signs in Table 3-7 specify the signs of the MV offsets added to the list 0 MV component of the starting MV, while the sign of the list 1 Mv has the opposite value. Otherwise, if the POC difference in list 1 is greater than list 0, the signs in Table 3-7 specify the signs of the MV offsets added to the list 1 MV component of the starting MV, while the sign of the list 0 MV has the opposite value.
[0104] The MVD is scaled according to the difference of the POCs in each direction. If the difference of the POCs in the two lists is the same, no scaling is required. Otherwise, if the difference of the POCs in list 0 is greater than the difference of the POCs in list 1, the MVD of list 1 is scaled by defining the POC difference of L0 as td and the POC difference of L1 as tb, as Figure 4 shown. If the POC difference of L1 is greater than that of L0, the MVD of list 0 is scaled in the same way. If the starting MV is unidirectionally predicted, the MVD is added to the available MVs.
[0105] Table 3-7 - Signs of MV Offsets Specified by Direction Index
[0106] Direction IDX 00 01 10 11 x-axis + - N / A N / A y-axis N / A N / A + -
[0107] 2.1.4 Symmetric MVD Coding and Decoding
[0108] In VVC, in addition to the normal unidirectional prediction and bidirectional prediction mode MVD signaling, a symmetric MVD mode for bidirectional prediction MVD signaling is also applied. In the symmetric MVD mode, the motion information including the reference picture indices of both list 0 and list 1 and the MVD of list 1 is not signaled but derived. Figure 5 An example of the symmetric MVD mode is shown.
[0109] The decoding process of the symmetric MVD mode is as follows:
[0110] 1. At the slice level, the derivation of the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 is as follows:
[0111] – If mvd_l1_zero_flag is 1, BiDirPredFlag is set to 0.
[0112] – Otherwise, if the nearest reference picture in list 0 and the nearest reference picture in list 1 form a pair of forward-backward reference pictures or a pair of backward-forward reference pictures, BiDirPredFlag is set to 1, and the reference pictures of both list 0 and list 1 are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0.
[0113] 2. At the CU level, if the CU is bidirectionally predicted coded and BiDirPredFlag is equal to 1, the symmetric mode flag indicating whether to use the symmetric mode is signaled explicitly.
[0114] When the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly signaled. The reference indexes for list 0 and list 1 are respectively set to be equal to the pair of reference pictures. MVD1 is set to be equal to (-MVD0). The final motion vectors are as shown in the following formula.
[0115] {(mvx0,mvy0)=(mvpx0+mvdx0,mvpy0+mvdy0)(mvx1,mvy1)=(mvpx1-mvdx0,mvpy1-mvdy0) (3-14)
[0116] In the encoder, symmetric MVD motion estimation starts from the initial MV evaluation. A set of initial MV candidates includes MVs obtained from unidirectional prediction search, MVs obtained from bidirectional prediction search, and MVs from the AMVP list. The one with the lowest rate-distortion cost is selected as the initial MV for the symmetric MVD motion search.
[0117] 2.1.5 Affine Motion Compensation Prediction
[0118] Figure 6 Examples of affine motion models based on control points are shown, including 4-parameter affine models and 6-parameter affine models. In HEVC, only translational motion models are applied to motion compensation prediction (MCP). In the real world, there are many types of motions, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transform motion compensation prediction is applied. As Figure 6 shown, the affine motion field of a block is described by the motion information of two control points (4-parameter) or three control point motion vectors (6-parameter).
[0119] For the 4-parameter affine motion model, the motion vector at the sample position (x,y) in the block can be derived as:
[0120]
[0121] For the 6-parameter affine motion model, the motion vector at the sample position (x,y) in the block can be derived as:
[0122]
[0123] where, (mv0x,mv0y) is the motion vector of the upper left control point, (mv1x,mv1y) is the motion vector of the upper right control point, and (mv2x,mv2y) is the motion vector of the lower left control point.
[0124] Figure 7An example affine MVF for each sub-block is shown. To simplify motion compensation prediction, block-based affine transform prediction is applied. To derive the motion vector of the center sample of each 4×4 luma sub-block, the motion vector of the center sample of each sub-block is calculated according to the above equation, as Figure 27 shown, and rounded to 1 / 16 fractional precision. Then, a motion compensation interpolation filter is applied to generate the prediction for each sub-block using the derived motion vector. The sub-block size of the chrominance component is also set to 4×4. The MV of a 4×4 chrominance sub-block is calculated as the average of the MVs of the top-left and bottom-right luma sub-blocks in the co-located 8×8 luma region.
[0125] Similar to translational motion inter prediction, there are also two affine motion inter prediction modes: affine Merge mode and affine AMVP mode.
[0126] 2.1.5.1 Affine Merge Prediction
[0127] The AF_MERGE mode is applicable to CUs with width and height both greater than or equal to 8. In this mode, the control point motion vector (CPMV) of the current CU is generated based on the motion information of spatially neighboring CUs. There can be up to five CPMV candidates, and an index is signaled to indicate the one to be used for the current CU. The following three types of CPMV candidates are used to form the affine Merge candidate list:
[0128] – Inherited affine Merge candidates inferred from the CPMVs of neighboring CUs
[0129] – Constructed affine Merge candidate CPMVs, which are derived using the translational MVs of neighboring CUs
[0130] – Zero MV
[0131] In VVC, there are at most two inherited affine candidates, which are derived from the affine motion models of neighboring blocks, one from the left neighboring CU and one from the upper neighboring CU. The candidate blocks are as Figure 8 shown. Figure 8 Shows an example position of the inherited affine predictor. For the predictor on the left, the scan order is A0->A1, and for the predictor on the upper side, the scan order is B0->B1->B2. Only the first inherited candidate from each side is selected. No deduplication check is performed between the two inherited candidates. When a neighboring affine CU is identified, its control point motion vector is used to derive CPMV candidates in the affine Merge list of the current CU. Figure 9 Shows an example of control point motion vector inheritance. As Figure 9As shown, if the adjacent lower - left block A is coded in affine mode, the motion vectors v2, v3, and v4 of the upper - left, upper - right, and lower - left corners of the CU containing block A are obtained. When block A is coded using a 4 - parameter affine model, two CPMVs of the current CU are calculated based on v2 and v3. If block A is coded using a 6 - parameter affine model, three CPMVs of the current CU are calculated based on v2, v3, and v4.
[0132] Constructing an affine candidate means that the candidate is constructed by combining the neighboring translational motion information of each control point. The motion information of the control point is derived from Figure 10 the specified spatial neighbors and temporal neighbors shown in Figure 10 An example location of constructing candidate positions for the affine Merge mode is shown. CPMVk (k = 1, 2, 3, 4) represents the k - th control point. For CPMV1, the B2->B3->A2 blocks are checked, and the MV of the first available block is used. For CPMV2, the B1->B0 blocks are checked, and for CPMV3, the A1->A0 blocks are checked. If available, TMVP is used as CPMV4.
[0133] After obtaining the MVs of the four control points, affine Merge candidates are constructed based on this motion information. The following combinations of control - point MVs are used for construction, in sequence:
[0134] {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}
[0135] Combinations of 3 CPMVs form 6 - parameter affine Merge candidates, and combinations of 2 CPMVs form 4 - parameter affine Merge candidates. To avoid the motion scaling process, if the reference indices of the control points are different, the relevant combinations of control - point MVs are discarded.
[0136] After checking the inherited affine Merge candidates and constructing the affine Merge candidates, if the list is still not full, zero MVs are inserted at the end of the list.
[0137] 2.1.5.2 Affine AMVP Prediction
[0138] The affine advanced motion vector prediction (AMVP) mode can be applied to CUs with a width and height greater than or equal to 16. In the bitstream, an affine flag at the CU level is signaled to indicate whether the affine AMVP mode is used, and then another flag is signaled to indicate whether it is 4-parameter affine or 6-parameter affine. In this mode, the difference between the CPMV of the current CU and its predictor CPMVP is signaled in the bitstream. The affine AVMP candidate list size is 2, and it is generated by using the following four types of CPVM candidates, in order:
[0139] – Inherited affine AMVP candidates inferred from the CPMV of neighboring CUs
[0140] – Constructed affine AMVP candidate CPMV, which is derived using the translational MVs of neighboring CUs
[0141] – Translational MVs from neighboring CUs
[0142] – Zero MVs
[0143] The checking order of inherited affine AMVP candidates is the same as that of inherited affine Merge candidates. The only difference is that for AVMP candidates, only affine CUs with the same reference picture in the current block are considered. When inserting the inherited affine motion predictor into the candidate list, the deduplication process is not applied.
[0144] The constructed AMVP candidates are derived from the specified spatial neighbors shown in Figure 10 The same checking order as in the construction of affine Merge candidates is used. In addition, the reference picture indices of neighboring blocks are checked. The first block in the checking order that is inter-coded and has the same reference picture as the current CU is used. There is only one. When the current CU is coded in 4-parameter affine mode and both mv0 and mv1 are available, they are added as a candidate to the affine AMVP list. When the current CU is coded in 6-parameter affine mode and all three CPMVs are available, they are added as a candidate to the affine AMVP list. Otherwise, the constructed AMVP candidate is set to unavailable.
[0145] If, after inserting the valid inherited affine AMVP candidates and the constructed AMVP candidates, the affine AMVP list candidates are still less than 2, then mv0, mv1, and mv2 (if available) are added in order as translational MVs to predict all control point MVs of the current CU. Finally, if the affine AMVP list is still not full, zero MVs are used to fill the list.
[0146] 2.1.5.3 Affine Motion Information Storage
[0147] In VVC, the CPMV of an affine CU is stored in a separate cache. The stored CPMV is only used to generate the inherited CPMVP for the most recently decoded CU in affine Merge mode and affine AMVP mode. The sub-block MVs derived from the CPMV are used for motion compensation, MV derivation for the Merge / AMVP list of translational MVs, and deblocking.
[0148] Figure 11 An example of the MV usage of the combined method is shown. To avoid the picture line cache for additional CPMVs, the affine motion data inheritance of the CU from the upper coding tree unit (CTU) is treated differently from the inheritance from normal neighboring CUs. If the candidate CU for affine motion data inheritance is in the upper CTU row, the left-bottom and right-bottom sub-block MVs in the line cache instead of the CPMV are used for affine MVP derivation. In this way, the CPMV is only stored in the local cache. If the candidate CU is 6-parameter affine coded, the affine model is degraded to a 4-parameter model. As Figure 11 shown, along the top CTU boundary, the left-bottom and right-bottom sub-block motion vectors of the CU are used for the affine inheritance of the CU in the bottom CTU.
[0149] 2.1.5.4 Prediction Refinement with Optical Flow for Affine Mode
[0150] Compared with pixel-based motion compensation, sub-block-based affine motion compensation can save memory access bandwidth and reduce computational complexity at the cost of loss of prediction accuracy. To achieve a finer motion compensation granularity, prediction refinement with optical flow (PROF) is used to refine the sub-block-based affine motion compensation prediction without increasing the memory access bandwidth of motion compensation. In VVC, after performing sub-block-based affine motion compensation, the luminance prediction samples are refined by adding the differences derived from the optical flow equation. PROF is described in the following four steps:
[0151] Step 1) Perform sub-block-based affine motion compensation to generate the sub-block prediction I(i,j).
[0152] Step 2) Calculate the spatial gradients g x (i,j) and g y (i,j) of the sub-block prediction at each sample position using a 3-tap filter [-1,0,1]. This gradient calculation is exactly the same as the gradient calculation in BDOF.
[0153] g x (i,j) = (I(i + 1,j) >> shift1) - (I(i - 1,j) >> shift1) (3-17)
[0154] g y(i,j) = (I(i,j + 1) >> shift1) - (I(i,j - 1) >> shift1) (3 - 18)
[0155] shift1 is used to control the accuracy of the gradient. The sub - block (i.e., 4×4) prediction extends one sample point on each side for gradient calculation. To avoid additional memory bandwidth and additional interpolation calculations, the extended sample points on the extended boundaries are copied from the nearest integer - pixel positions in the reference picture.
[0156] Step 3) Calculate the luminance prediction refinement through the following optical flow equation.
[0157] ΔI(i,j) = g x (i,j) * Δv x (i,j) + g y (i,j) * Δv y (i,j) (3 - 19)
[0158] Where Δv(i,j) is the difference between the sample MV (represented by v(i,j)) calculated for the sample position (i,j) and the sub - block MV of the sub - block to which the sample (i,j) belongs, as Figure 12 shown. Figure 12 Shows an example of the sub - block MV V SB and pixels. Δv(i,j) is quantized in units of 1 / 32 luminance sample accuracy.
[0159] Since the affine model parameters and the sample positions relative to the sub - block center do not change between sub - blocks, Δv(i,j) can be calculated for the first sub - block and reused for other sub - blocks in the same CU. Let dx(i,j) and dy(i,j) be the horizontal and vertical offsets from the sample position (i,j) to the center (x SB ,y SB ) of the sub - block. Δv(x,y) can be derived from the following equations:
[0160] {dx(i,j) = i - x SB dy(i,j) = j - y SB (3 - 20)
[0161] {Δv x (i,j) = C * dx(i,j) + D * dy(i,j)Δv y (i,j) = E * dx(i,j) + F * dy(i,j) (3 - 21)
[0162] To maintain accuracy, the center (x SB ,y SB)It is calculated as ((WSB - 1) / 2, (HSB - 1) / 2), where WSB and HSB are the width and height of the sub - block respectively.
[0163] For the 4 - parameter affine model,
[0164]
[0165] For the 6 - parameter affine model,
[0166]
[0167] where, (v 0x , v 0y ), (v 1x , v 1y ), (v 2x , v 2y ) are the top - left, top - right and bottom - left control - point motion vectors, and w and h are the width and height of the CU.
[0168] Step 4) Finally, add the luminance prediction refinement ΔI(i, j) to the sub - block prediction I(i, j). The final prediction I’ is generated by the following equation.
[0169] I′(i, j) = I(i, j)+ΔI(i, j)
[0170] For the CU of affine encoding and decoding, PROF is not applied in two cases: 1) all control - point MVs are the same, which indicates that the CU has only translational motion; 2) the affine motion parameters are greater than the specified limit, because the sub - block - based affine MC is degraded to CU - based MC to avoid large memory access bandwidth requirements.
[0171] Apply a fast encoding method to reduce the encoding complexity of affine motion estimation with PROF. In the following two cases, PROF is not applied in the affine motion estimation stage: a) if the CU is not a root block and its parent block has not selected the affine mode as its best mode, then PROF is not applied because the probability that the current CU selects the affine mode as the best mode is very low; b) if the magnitudes of all four affine parameters (C, D, E, F) are less than a predefined threshold and the current picture is not a low - latency picture, then PROF is not applied because the improvement introduced by PROF is very small in this case. In this way, the affine motion estimation with PROF can be accelerated.
[0172] 2.1.6 Sub - block - based Temporal Motion Vector Prediction (SbTMVP)
[0173] VVC supports the sub - block - based temporal motion vector prediction (SbTMVP) method. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the collocated picture to improve the motion vector prediction and the Merge mode of the CUs in the current picture. The same collocated picture used by TMVP is also used for SbTMVP. The main differences between SbTMVP and TMVP are in the following two aspects:
[0174] – TMVP predicts the motion at the CU level, while SbTMVP predicts the motion at the sub - CU level;
[0175] – TMVP obtains the temporal motion vector from the collocated block in the collocated picture (the collocated block is the bottom - right or central block relative to the current CU), while SbTMVP applies a motion offset before obtaining the temporal motion information from the collocated picture, where the motion offset is obtained from the motion vector of one of the spatial neighbors of the current CU.
[0176] Figure 13 An example of the SbTMVP process in VVC is shown. Figure 13 It includes a figure showing the spatial neighbors used by the optional temporal vector motion prediction (ATVMP). Figure 13 It also includes a figure showing the derivation of the sub - CU motion field by applying the motion offset from the spatial neighbors and scaling the motion information from the corresponding collocated sub - CU.
[0177] SbTMVP predicts the motion vectors of the sub - CUs within the current CU in two steps. In the first step, the spatial neighbor A1 in Figure 13 (a) is checked. If A1 has a motion vector using the collocated picture as its reference picture, then that motion vector is selected as the motion offset to be applied. If no such motion is identified, the motion offset is set to (0, 0).
[0178] In the second step, the motion offset identified in step 1 (i.e., added to the coordinates of the current block) is applied to obtain the sub - CU level motion information (motion vector and reference index) from the collocated picture, as shown in Figure 13 (b). Figure 13 (b) The example assumes that the motion offset is set to the motion of block A1. Then, for each sub - CU, the motion information of its corresponding block (the smallest motion grid covering the central samples) in the collocated picture is used to derive the motion information of the sub - CU. After identifying the motion information of the collocated sub - CU, it is converted to the action vector and reference index of the current sub - CU in a way similar to the TMVP process in HEVC, where temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current CU.
[0179] In VVC, a sub-block based Merge list containing a combination of both SbTMVP candidates and affine Merge candidates is used to signal the sub-block based Merge mode. The SbTMVP mode is enabled / disabled via a Sequence Parameter Set (SPS) flag. If the SbTMVP mode is enabled, the SbTMV predictor is added as the first entry of the sub-block based Merge candidate list, followed by the affine Merge candidates. The size of the sub-block based Merge list is signaled in the SPS, and the maximum allowed size of the sub-block Merge list is 5 in VVC.
[0180] The sub-CU size used in SbTMVP is fixed at 8×8. Similar to the affine Merge mode, the SbTMVP mode is only applicable to CUs with width and height both greater than or equal to 8.
[0181] The encoding logic for additional SbTMVP Merge candidates is the same as that for other Merge candidates. That is, for each CU in a P or B slice, an additional rate-distortion (RD) check is performed to decide whether to use the SbTMVP candidate.
[0182] 2.1.7 Adaptive Motion Vector Resolution (AMVR)
[0183] In HEVC, when the use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the motion vector of the CU and the predicted motion vector) is signaled in quarter-luma samples. In VVC, a CU-level Adaptive Motion Vector Resolution (AMVR) scheme is introduced. AMVR allows the MVD of a CU to be encoded and decoded with different precisions. Depending on the mode of the current CU (normal AMVP mode or affine AMVP mode), the MVD of the current CU can be adaptively selected as follows:
[0184] – Normal AMVP mode: quarter-luma samples, half-luma samples, integer-luma samples, or four-luma samples.
[0185] – Affine AMVP mode: quarter-luma samples, integer-luma samples, or 1 / 16-luma samples.
[0186] If the current CU has at least one non-zero MVD component, the CU-level MVD resolution indication is signaled conditionally. If all MVD components (i.e., the horizontal and vertical MVDs of reference list L0 and reference list L1) are zero, a quarter-luma sample MVD resolution is assumed.
[0187] For a CU with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter-luma sample MVD precision is used for the CU. If the first flag is 0, no further signaling is required and quarter-luma sample MVD precision is used for the current CU. Otherwise, a second flag is signaled to indicate whether half-luma sample or other MVD precision (integer or quarter-luma sample) is used for normal AMVP CUs. In the case of half-luma samples, a 6-tap interpolation filter is used for half-luma sample positions instead of the default 8-tap interpolation filter. Otherwise, a third flag is signaled to indicate whether integer-luma sample or quarter-luma sample MVD precision is used for normal AMVP CUs. In the case of affine AMVP CUs, the second flag is used to indicate whether integer-luma samples or 1 / 16-luma samples MVD precision is used. To ensure that the reconstructed MV has the expected precision (quarter-luma sample, half-luma sample, integer-luma sample, or quarter-luma sample), the motion vector predictor for the CU is rounded to the same precision as the MVD before being added to the MVD. The motion vector predictor is rounded to zero (i.e., a negative motion vector predictor is rounded towards positive infinity and a positive motion vector predictor is rounded towards negative infinity).
[0188] The encoder uses RD checking to determine the motion vector resolution of the current CU. To avoid always performing four CU-level RD checks for each MVD resolution, in the VVC Test Model (VTM) 14, the RD checks for MVD precision are only called conditionally, except for quarter-luma samples. For normal AVMP mode, first the RD costs for quarter-luma sample MVD precision and integer-luma sample MV precision are calculated. Then, the RD cost for integer-luma sample MVD precision is compared with the RD cost for quarter-luma sample MVD precision to determine whether it is necessary to further check the RD cost for quarter-luma sample MVD precision. When the RD cost for quarter-luma sample MVD precision is much smaller than the RD cost for integer-luma sample MVD precision, the RD check for quarter-luma sample MVD precision is skipped. Then, if the RD cost for integer-luma sample MVD precision is significantly larger than the best RD cost of the previously tested MVD precision, the check for half-luma sample MVD precision is skipped. For affine AMVP mode, if the affine inter-frame mode is not selected after checking the rate-distortion costs of the affine Merge / skip mode, Merge / skip mode, quarter-luma sample MVD precision normal AMVP mode, and quarter-luma sample MVD precision affine AMVP mode, the 1 / 16-luma sample MV precision and 1-pixel MV precision affine inter-frame modes are not checked. In addition, the affine parameters obtained in the quarter-luma sample MV precision affine inter-frame mode are used as the starting search points for the 1 / 16-luma sample and quarter-luma sample MV precision affine inter-frame modes.
[0189] 2.1.8 Bi - directional Prediction with CU - level Weights (BCW)
[0190] In HEVC, a bi - directional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bi - directional prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals.
[0191] P bi-pred = ((8 - w)*P0+w*P1 + 4)>>3 (3 - 24)
[0192] Five weights are allowed in weighted - average bi - directional prediction, w ∈ {-2, 3, 4, 5, 10}. For each bi - directional prediction CU, the weight w is determined in one of two ways: (1) For non - Merge CUs, the weight index is signaled after the motion - vector difference; (2) For Merge CUs, the weight index is deduced from neighboring blocks based on the Merge candidate index. BCW is only applicable to CUs with 256 or more luma samples (i.e., CU width times CU height is greater than or equal to 256). For low - latency pictures, all 5 weights are used. For non - low - latency pictures, only 3 weights (w ∈ {3, 4, 5}) are used.
[0193] – At the encoder, fast search algorithms are applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized as follows. For more details, the reader is referred to the VTM software and the document JVET - L0646. When combined with AMVR, if the current picture is a low - latency picture, unequal weights are only conditionally checked for 1 - pixel and 4 - pixel motion - vector precisions.
[0194] – When combined with affine, affine ME for unequal weights is performed if and only if the affine mode is selected as the current best mode.
[0195] – When the two reference pictures in bi - directional prediction are the same, unequal weights are only conditionally checked.
[0196] – When certain conditions are met, unequal weights are not searched, depending on the POC distance between the current picture and its reference picture, the coding - decoding QP, and the temporal level.
[0197] The BCW weight index is coded using one context - coded bit, followed by bypass - coded bits. The first context - coded bit indicates whether equal weights are used; if unequal weights are used, the bypass - coded bits are used to signal which unequal weight is used.
[0198] Weighted Prediction (WP) is a coding tool supported by the H.264 / Advanced Video Coding (AVC) and HEVC standards, which is used for efficient coding of video content with fading. Support for WP has also been added in the VVC standard. WP allows the transmission of weighting parameters (weights and offsets) for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the weights and offsets of the corresponding reference pictures are applied. WP and BCW are designed for different types of video content. To avoid interactions between WP and BCW that complicate the VVC decoder design, if a CU uses WP, the BCW weight index is not signaled, and w is presumed to be 4 (i.e., equal weights are applied). For Merge CUs, the weight index is presumed from neighboring blocks based on the Merge candidate index. This can be applied to both the normal Merge mode and the inherited affine Merge mode. For the constructed affine Merge mode, the affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index of a CU using the constructed affine Merge mode is simply set to be equal to the BCW index of the first control point MV.
[0199] In VVC, CIIP and BCW cannot be jointly applied to a CU. When a CU is coded using the CIIP mode, the BCW index of the current CU is set to 2, e.g., equal weights.
[0200] 2.1.9 Bidirectional Optical Flow (BDOF)
[0201] The Bidirectional Optical Flow (BDOF) tool is included in VVC. BDOF (previously known as BIO) was included in the Joint Exploration Model (JEM). Compared with the JEM version, BDOF in VVC is a simpler version, requiring less computation, especially in terms of the number of multiplications and the multiplier size.
[0202] BDOF is used to refine the bidirectional prediction signal of a CU at the 4×4 sub-block level. BDOF is applied to a CU if all of the following conditions are met:
[0203] – The CU is coded using the "true" bidirectional prediction mode, i.e., one of the two reference pictures is before the current picture in the display order, and the other is after the current picture in the display order.
[0204] – The distances from the two reference pictures to the current picture (i.e., the POC differences) are the same.
[0205] – Both of the two reference pictures are short-term reference pictures.
[0206] – The CU is not coded using the affine mode or the SbTMVP Merge mode.
[0207] – The CU has more than 64 luma samples.
[0208] – Both the CU height and the CU width are greater than or equal to 8 luma samples.
[0209] – The BCW weight index indicates equal weights.
[0210] – WP is not enabled for the current CU.
[0211] – The CIIP mode is not used for the current CU.
[0212] BDOF is only applied to the luma component. As its name implies, the BDOF mode is based on the optical flow concept, which assumes that the motion of an object is smooth. For each 4×4 sub-block, the motion refinement (v x ,v y ) is calculated by minimizing the difference between the L0 and L1 predicted samples. Then, the motion refinement is used to adjust the bidirectional predicted sample values in the 4×4 sub-block. The following steps apply to the BDOF process.
[0213] First, the horizontal and vertical gradients of the two prediction signals are calculated by directly computing the difference between two neighboring samples and k = 0, 1.
[0214]
[0215] where I (k) (i, j) is the sample value at the prediction signal coordinates (i, j) in list k (k = 0, 1), and shift1 is calculated based on the luma bit depth bitDepth as shift1 = max(6, bitDepth - 6).
[0216] Then, the autocorrelations and cross-correlations of the gradients S1, S2, S3, S5, and S6 are calculated as follows:
[0217]
[0218] where
[0219]
[0220] where Ω is a 6×6 window around the 4×4 sub-block, and the values of n a and n b are set to be equal to min(1, bitDepth - 11) and min(4, bitDepth - 8), respectively.
[0221] Then, using the cross-correlation and autocorrelation terms, the motion refinement (v x ,vy ):
[0222]
[0223] where th′ BIO = 2 max(5,BD-7) . is the floor function, and
[0224] Based on motion refinement and gradients, the following adjustments are calculated for each sample point in a 4×4 sub-block:
[0225]
[0226] Finally, the BDOF sample points of the CU are calculated by adjusting the bi-predicted sample points as follows:
[0227] pred BDOF (x,y) = (I (0) (x,y) + I (1) (x,y) + b(x,y) + o offset ) >> shift (3 - 30)
[0228] These values are chosen such that the multipliers in the BDOF process do not exceed 15 bits, and the maximum bit-width of the intermediate parameters in the BDOF process remains within 32 bits.
[0229] To derive the gradient values, some predicted sample points I (k) (i,j) outside the current CU boundary in the list k (k = 0, 1) need to be generated. Figure 14 An example extended CU region used in BDOF is shown. As Figure 14 shown, BDOF in VVC uses an extended row / column around the CU boundary. To control the computational complexity of generating predicted sample points outside the boundary, the predicted sample points in the extended region (white positions) are generated by directly obtaining the reference sample points at nearby integer positions (using the floor() operation on the coordinates) without interpolation, and the normal 8-tap motion compensation interpolation filter is used to generate the predicted sample points within the CU (gray positions). These extended sample point values are only used for gradient calculation. For the remaining steps in the BDOF process, if any sample points and gradient values outside the CU boundary are needed, they are filled (i.e., repeated) from their nearest neighbors.
[0230] When the width and / or height of a CU is greater than 16 luma samples, it is divided into sub-blocks with a width and / or height equal to 16 luma samples, and the sub-block boundaries are considered CU boundaries during the BDOF process. The maximum unit size limit of the BDOF process is 16×16. For each sub-block, the BDOF process can be skipped. When the SAD between the initial L0 and L1 prediction samples is less than the threshold, the BDOF process is not applied to the sub-block. The threshold is set to be equal to (8*W*(H>>1), where W represents the sub-block width and H represents the sub-block height. To avoid the additional complexity of SAD calculation, the SAD calculated during the DMVR process between the initial L0 and L1 prediction samples is reused here.
[0231] If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, then bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, i.e., the luma_weight_lx_flag of any one of the two reference pictures is 1, then BDOF is also disabled. When the CU is encoded and decoded using the symmetric MVD mode or the CIIP mode, BDOF is also disabled.
[0232] 2.1.10 Decoder-side Motion Vector Refinement (DMVR)
[0233] To improve the accuracy of the MV in the Merge mode, decoder-side motion vector refinement based on bilateral matching (BM) is applied in VVC. During the bidirectional prediction operation, refined MVs are searched around the initial MVs in the reference picture list L0 and the reference picture list L1. The BM method calculates the distortion between two candidate blocks in the reference picture list L0 and list L1. Figure 15 An example of decoder-side motion vector refinement is shown. As Figure 15 shown, the SAD between the red blocks based on each MV candidate around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bidirectional prediction signal.
[0234] In VVC, the application of DMVR is restricted to CUs encoded and decoded using the following modes and features:
[0235] – CU-level Merge mode with bidirectional prediction MV
[0236] – For the current picture, one reference picture is from the past and the other reference picture is from the future
[0237] – The distances (i.e., POC differences) from the two reference pictures to the current picture are the same
[0238] – Both reference pictures are short-term reference pictures
[0239] – The CU has more than 64 luma samples
[0240] – Both the CU height and the CU width are greater than or equal to 8 luma samples
[0241] – The BCW weight index indicates equal weights
[0242] – WP is not enabled for the current block
[0243] – The CIIP mode is not used for the current block
[0244] The refined MV derived through the DMVR process is used to generate inter - predicted samples and is also used for temporal motion vector prediction in future picture coding. The original MV is used for the de - blocking process and is also used for spatial motion vector prediction in future CU coding.
[0245] Additional features of DMVR are mentioned in the following sub - articles.
[0246] 2.1.10.1 Search Scheme
[0247] In DMVR, the search points are around the initial MV, and the MV offset follows the MV difference mirroring rule. In other words, any point checked by DMVR represented by a candidate MV pair (MV0, MV1) follows the following two equations:
[0248] MV0′ = MV0 + MV_offset (3 - 31)
[0249] MV1′ = MV1 - MV_offset (3 - 32)
[0250] where MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two integer luma samples from the initial MV. The search includes an integer - sample offset search stage and a fractional - sample refinement stage.
[0251] A 25 - point full search is adopted for the integer - sample offset search. First, the SAD of the initial MV pair is calculated. If the SAD of the initial MV pair is less than the threshold, the integer - sample stage of DMVR terminates. Otherwise, the SADs of the remaining 24 points are calculated and checked in raster - scan order. The point with the minimum SAD is selected as the output of the integer - sample offset search stage. To reduce the loss caused by the uncertainty of DMVR refinement, it is recommended to give priority to using the original MV during the DMVR process. The SAD between the reference blocks of the initial MV candidates reduces the SAD value by 1 / 4.
[0252] Integer sample point search is followed by fractional sample point refinement. To save computational complexity, fractional sample point refinement is derived by using the parametric error surface equation instead of performing additional search using SAD comparison. Fractional sample point refinement is conditionally invoked based on the output of the integer sample point search stage. When the integer sample point search stage is terminated at the center with the minimum SAD in the first iteration or the second iteration search, fractional sample point refinement is further applied.
[0253] In sub-pixel offset estimation based on the parametric error surface, the cost at the center position and the costs at four neighboring positions from the center are used to fit a 2-D parabolic error surface equation of the following form:
[0254] E(x,y) = A(x - x min ) 2 + B(y - y min ) 2 + C (3 - 33)
[0255] where (x min , y min ) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost value. By solving the above equation using the cost values of five search points, (x min , y min ) is calculated as follows:
[0256] x min = (E(-1,0) - E(1,0)) / (2(E(-1,0) + E(1,0) - 2E(0,0))) (3 - 34)
[0257] y min = (E(0,-1) - E(0,1)) / (2((E(0,-1) + E(0,1) - 2E(0,0))) (3 - 35)
[0258] x min and y min values are automatically limited between -8 and 8 because all cost values are positive with the minimum being E(0,0). This corresponds to the half-peak shift of the 1 / 16 pixel MV accuracy in VVC. The calculated fraction (x min , y min ) is added to the integer distance refined MV to obtain the sub-pixel refined incremental MV.
[0259] 2.1.10.2 Bilinear Interpolation and Sample Filling
[0260] In VVC, the resolution of the MV is 1 / 16 luma samples. An 8-tap interpolation filter is used to interpolate samples at fractional positions. In DMVR, the search points are centered around the initial fractional pixel MV with integer sample offsets, so interpolation of these samples at fractional positions is required for the DMVR search process. To reduce the computational complexity, a bilinear interpolation filter is used in DMVR to generate the fractional samples for the search process. Another important effect is that, by using the bilinear filter, within a 2-sample search range, DMVR does not access more reference samples compared to the normal motion compensation process. After obtaining the refined MV through the DMVR search process, the normal 8-tap interpolation filter is applied to generate the final prediction. To avoid accessing more reference samples in the normal MC process, samples that are not required for the interpolation process based on the original MV but are required for the interpolation process based on the refined MV are filled from these available samples.
[0261] 2.1.10.3 Maximum DMVR Processing Unit
[0262] When the width and / or height of a CU is greater than 16 luma samples, it is further split into sub-blocks with a width and / or height equal to 16 luma samples. The maximum unit size limit for the DMVR search process is 16×16.
[0263] 2.1.11 Geometric Partitioning Mode (GPM)
[0264] In VVC, geometric partitioning mode is supported for inter prediction. A CU-level flag is used as a Merge mode to signal the geometric partitioning mode, and other Merge modes include the regular Merge mode, MMVD mode, CIIP mode, and sub-block Merge mode. For each possible CU size w×h = 2 m ×2 n , where m,n ∈ {3…6}, excluding 8×64 and 64×8, the geometric partitioning mode supports a total of 64 partitions.
[0265] When using this mode, the CU is divided into two parts by a geometrically positioned line ( Figure 16 ). Figure 16 Examples of GPM partitions grouped by the same angle are shown. The position of the partition line is mathematically derived from the angle and offset parameters of a specific partition. Each part of the geometric partition in the CU uses its own motion for inter prediction; only unidirectional prediction is allowed for each partition, that is, each part has a motion vector and a reference index. Unidirectional prediction motion constraints are applied to ensure that, like normal bidirectional prediction, each CU only requires two motion-compensated predictions. The unidirectional prediction motion for each partition is derived using the process described in 3.4.11.1.
[0266] If the geometric partitioning mode is used for the current CU, further signal the geometric partition index indicating the partition mode (angle and offset) of the geometric partitioning and two Merge indices (one for each partition). Explicitly signal the number of maximum GPM candidate sizes in the SPS, and specify the syntax binarization for the GPM Merge indices. After predicting each part of the geometric partitioning, use hybrid processing with adaptive weights to adjust the sample values along the geometric partitioning edges as described in 3.4.11.2. This is a prediction signal for the entire CU, and like other prediction modes, the transform and quantization processes will be applied to the entire CU. Finally, store the motion field of the CU predicted using the geometric partitioning mode as described in 3.4.11.3.
[0267] 2.1.11.1 Unidirectional Prediction Candidate List Construction
[0268] The unidirectional prediction candidate list is directly derived from the Merge candidate list constructed according to the extended Merge prediction process in 3.4.1. Denote n as the index of the unidirectional predicted motion in the geometric unidirectional prediction candidate list. The LX motion vector of the nth extended Merge candidate (where X equals the parity of n) is used as the nth unidirectional predicted motion vector for the geometric partitioning mode. These motion vectors are marked with "x" in Figure 17 . Figure 17 Shows an example of unidirectional prediction MV selection for the geometric partitioning mode. In the case where the corresponding LX motion vector of the nth extended Merge candidate does not exist, use the L(1-X) motion vector of the same candidate as the unidirectional predicted motion vector for the geometric partitioning mode.
[0269] 2.1.11.2 Blending along Geometric Partitioning Edges
[0270] After predicting each part of the geometric partitioning using its own motion, blend the two prediction signals to derive the samples around the geometric partitioning edges. The blending weight for each position of the CU is derived based on the distance between the individual position and the partitioning edge.
[0271] The distance from the position (x,y) to the partitioning edge can be derived as:
[0272]
[0273] ρ x,j ={0 i % 16 = 8 or (i % 16 ≠ 0 and h ≥ w) ±(j × w) >> 2 otherwise (3-38)
[0274] ρ y,j= {±(j × h) >> 2i % 16 = 8 or (I % 16 ≠ 0 and h ≥ w) 0 otherwise (3 - 39)
[0275] where i, j are indices of the angle and offset of the geometric segmentation that depend on the geometric segmentation index transmitted through the signal. ρ x,j and ρ y,j sign depends on the angle index i.
[0276] The weight of each part of the geometric segmentation is derived as follows:
[0277] wIdxL(x, y) = partIdx? 32 + d(x, y) : 32 - d(x, y) (3 - 40)
[0278]
[0279] w1(x, y) = 1 - w0(x, y) (3 - 42)
[0280] partIdx depends on the angle index i. Figure 18 An example of the weight w0 is shown in. Figure 18 Examples of generating the bending weight by exemplifying the geometric segmentation pattern are shown.
[0281] 2.1.11.3 Motion Field Storage of Geometric Segmentation Pattern
[0282] Mv1 from the first part of the geometric segmentation, Mv2 from the second part of the geometric segmentation, and the combination Mv of Mv1 and Mv2 are stored in the motion field of the geometric segmentation pattern codec CU.
[0283] The type of the stored motion vector at each individual position in the motion field is determined as:
[0284] sType = abs(motionIdx) < 32? 2 : (motionIdx ≤ 0? (1 - partIdx) : partIdx) (3 - 43)
[0285] where motionIdx is equal to d(4x + 2, 4y + 2), which is recalculated by Equation (3 - 36). partIdx depends on the angle index i.
[0286] If sType is equal to 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field, otherwise if sType = 2, then the combination Mv of Mv0 and Mv2 is stored. The combined Mv is generated using the following procedure:
[0287] (1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), then simply combine Mv1 and Mv2 to form a bi - directional prediction motion vector.
[0288] (2) Otherwise, if Mv1 and Mv2 are from the same list, then store only the uni - directional prediction motion Mv2.
[0289] 2.1.12 Combined Inter - and Intra - Prediction (CIIP)
[0290] Figure 19 Examples of the top and left neighboring blocks used in CIIP weight derivation are shown. In VVC, when a CU is coded / decoded in Merge mode, if the CU contains at least 64 luma samples (i.e., the CU width times the CU height is equal to or greater than 64), and if both the CU width and CU height are less than 128 luma samples, then an additional flag is signaled to indicate whether the combined inter / intra - prediction (CIIP) mode is applied to the current CU. As the name implies, CIIP prediction combines an inter - prediction signal with an intra - prediction signal. The same inter - prediction process applied to the regular Merge mode is used to derive the inter - prediction signal P in the CIIP mode inter ; and the intra - prediction signal P is derived after the regular intra - prediction process with planar mode intra . Then, a weighted average is used to combine the intra - and inter - prediction signals, where the weight values are calculated according to the coding / decoding modes of the top and left neighboring blocks as Figure 19 shown as follows:
[0291] – If the top neighbor is available and is intra - coded, then set isIntraTop to 1, otherwise set isIntraTop to 0;
[0292] – If the left neighbor is available and is intra - coded, then set isIntraLeft to 1, otherwise set isIntraLeft to 0;
[0293] – If (isIntraLeft + isIntraTop) equals 2, then set wt to 3;
[0294] – Otherwise, if (isIntraLeft + isIntraTop) equals 1, then set wt to 2;
[0295] – Otherwise, set wt to 1.
[0296] CIIP prediction is formed as follows:
[0297] P CIIP = ((4 - wt)*P inter+wt*P intra +2) >> 2 (3 - 43)
[0298] 2.1.13 Reference Picture Resampling (RPR)
[0299] In HEVC, unless a new sequence using a new SPS starts with an Intra Random Access Point (IRAP) picture, the spatial resolution of a picture cannot be changed. VVC allows the resolution of a picture at a certain position within a sequence to change without encoding IRAP pictures that are always intra-coded. This feature is sometimes referred to as Reference Picture Resampling (RPR) because when the reference picture used for inter prediction has a different resolution from the current picture being decoded, this feature requires resampling of the reference picture. To avoid additional processing steps, the RPR process in VVC is designed to be embedded in the motion compensation process and is performed at the block level. During the motion compensation stage, the scaling ratio is used together with the motion information to locate the reference samples in the reference picture used in the interpolation process.
[0300] In VVC, the scaling ratio is limited to be greater than or equal to 1 / 2 (2x downsampling from the reference picture to the current picture) and less than or equal to 8 (8x upsampling). Three sets of resampling filters with different frequency cut-off values are specified to handle various scaling ratios between the reference picture and the current picture. The three sets of resampling filters are applied to scaling ratios ranging from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8 respectively. Each set of resampling filters has 16 phases for luminance and 32 phases for chrominance, which is the same as the case of the motion compensation interpolation filter. It is worth noting that the filter bank of normal MC interpolation is used in the case where the scaling ratio ranges from 1 / 1.25 to 8. In fact, the normal MC interpolation process is a special case of the resampling process with scaling ratios in the range of 1 / 1.25 to 8. In addition to the regular translational block motion, the affine mode also has three sets of 6-tap interpolation filters for the luminance component to cover different scaling ratios in RPR. The horizontal and vertical scaling ratios are derived based on the picture width and height and the left, right, top, and bottom scaling offsets specified for the reference picture and the current picture.
[0301] To support this feature, the picture resolution and the corresponding consistency window are signaled in the Picture Parameter Set (PPS) instead of in the SPS, while the maximum picture resolution is signaled in the SPS.
[0302] 2.1.14 Other Aspects of Inter Prediction
[0303] To reduce the memory bandwidth, 4×4 sized CUs for inter - frame coding / decoding are not allowed in VVC. For 4×8 / 8×4 CUs for inter - frame coding / decoding, only unidirectional modes are allowed. When the motion information from the Merge mode is bi - directional, it is converted to unidirectional by retaining only the motion information in list 0.
[0304] 3 Inter - frame prediction tools being studied in the Enhanced Compression Model (ECM)
[0305] 3.1 Local Illumination Compensation (LIC)
[0306] LIC is an inter - frame prediction technique that models the local illumination change between the current block and its predicted block as a function of the illumination change between the current block template and the reference block template. The parameters of the function can be represented by a scale α and an offset β, forming a linear equation, i.e., α*p[x]+β to compensate for the illumination change, where p[x] is the reference sample pointed to by the MV at the x - position on the reference picture. Since α and β can be derived based on the current block template and the reference block template, they do not require signaling overhead, except for indicating the use of LIC by signaling the LIC flag for the AMVP mode.
[0307] Local Illumination Compensation is used for unidirectional - predicted inter - frame CUs with the following modifications.
[0308] ● Intra - frame neighboring samples can be used for LIC parameter derivation;
[0309] ● For blocks smaller than 32 luma samples, LIC is disabled;
[0310] ● For both non - sub - block and affine modes, LIC parameter derivation is based on the samples of the modulo - block corresponding to the current CU, rather than based on the partial modulo - block samples corresponding to the top - left 16×16 unit;
[0311] ● The samples of the reference block template are generated by using MC with the block MV without rounding it to integer - pixel precision.
[0312] 3.2 Non - adjacent spatial candidates
[0313] The non - adjacent spatial Merge candidates in Joint Video Exploration Team (JVET) - L0399 are inserted after TMVP in the regular Merge candidate list. The modes of the spatial Merge candidates are as Figure 20 shown. Figure 20 Shows an example of the spatial neighboring blocks used to derive the spatial Merge candidates. The distance between the non - adjacent spatial candidates and the current coded block is based on the width and height of the current coded block. No row - buffer limit is applied.
[0314] 3.3 Template Matching (TM)
[0315] Template matching (TM) is a decoder-side MV derivation method that refines the motion information of the current CU by finding the closest match between a template in the current picture (i.e., the top and / or left neighboring blocks of the current CU) and a block in the reference picture (i.e., of the same size as the template). Figure 21 An example of template matching performed on the search region around the initial MV is shown. As Figure 21 shown, a better MV is searched for around the initial motion search of the current CU within the [-8, +8] pixel search range. The template matching method in JVET-J0021 is used with the following modifications: the search step size is determined based on the AMVR mode, and TM can be cascaded with the bilateral matching process in the Merge mode.
[0316] In the AMVP mode, the MVP candidate is determined based on the template matching error to select the one that achieves the minimum difference between the current block template and the reference block template, and then TM is only performed on this specific MVP candidate for MV refinement. TM refines this MVP candidate starting from the full-pixel MVD accuracy (or 4 pixels in the 4-pixel AMVR mode) within the [-8, +8] pixel search range by using iterative diamond search. The AMVP candidate can be further refined by using cross search with full-pixel MVD accuracy (or 4 pixels in the 4-pixel AMVR mode), and then half-pixel and quarter-pixel are used sequentially according to the AMVR mode specified in Table 1. This search process ensures that the MVP candidate remains at the same MV accuracy indicated by the AMVR mode after the TM process.
[0317] Table 1. Search patterns of AMVR and Merge mode with AMVR
[0318]
[0319] In the Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. As shown in Table 1, TM can be performed all the way to 1 / 8 pixel MVD accuracy or skip those with more than half-pixel MVD accuracy, depending on whether an alternative interpolation filter is used according to the motion information of Merge (used when AMVR is in the half-pixel mode). In addition, when the TM mode is enabled, template matching can be an independent process or an additional MV refinement process between the block-based and sub-block-based bilateral matching (BM) methods, depending on whether BM can be enabled according to its enabling conditions check.
[0320] 3.4 Multi-pass decoder-side motion vector refinement
[0321] Apply multiple-pass decoder-side motion vector refinement. In the first pass, bilateral matching (BM) is applied to the coded block. In the second pass, BM is applied to each 16×16 sub-block within the coded block. In the third pass, the MV in each 8×8 sub-block is refined by applying bidirectional optical flow (BDOF). To perform spatial and temporal motion vector prediction, the refined MVs are stored.
[0322] 3.4.1 First Pass - Block-Based Bilateral Matching MV Refinement
[0323] In the first pass, the refined MVs are derived by applying BM to the coded block. Similar to decoder-side motion vector refinement (DMVR), in the bidirectional prediction operation, the refined MVs are searched around the two initial MVs (MV0 and MV1) in the reference picture lists L0 and L1. Based on the minimum bilateral matching cost between the two reference blocks in L0 and L1, the refined MVs (MV0_pass1 and MV1_pass1) are derived around the starting MVs.
[0324] BM performs a local search to derive the integer sample precision intDeltaMV. The local search adopts a 3×3 square search pattern, traversing the search range [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block dimension, and the maximum values of sHor and sVer are 8.
[0325] The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW*cbH is greater than 64, the mean removed sum of absolute differences (MRSAD) cost function is applied to remove the distorted DC effect between the reference blocks. When the bilCost at the center point of the 3×3 search pattern has the minimum cost, the intDeltaMV local search terminates. Otherwise, the current minimum cost search point becomes the new center point of the 3×3 search pattern, and the search for the minimum cost continues until it reaches the end of the search range.
[0326] Further apply the existing fractional sample refinement to derive the final deltaMV. Then, the refined MVs after the first pass can be derived as:
[0327] ● MV0_pass1 = MV0 + deltaMV
[0328] ● MV1_pass1 = MV1 – deltaMV
[0329] 3.4.2 Second Pass - Sub-Block-Based Bilateral Matching MV Refinement
[0330] In the second pass, refined MVs are derived by applying BM to 16×16 grid sub-blocks. For each sub-block, refined MVs are searched around the two MVs (MV0_pass1 and MV1_pass1) obtained in the first pass in the reference picture lists L0 and L1. Based on the minimum bilateral matching cost between two reference sub-blocks in L0 and L1, the refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived.
[0331] For each sub-block, BM performs a full search to derive the integer-sample accuracy intDeltaMV. The full search has a search range of [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block dimension, and the maximum values of sHor and sVer are 8.
[0332] Figure 22 An example of the diamond region in the search area is shown. The bilateral matching cost is calculated by applying the cost factor to the sum of the absolute transform difference (SATD) costs between two reference sub-blocks as follows: bilCost = satdCost * costFactor. The search area (2*sHor + 1)*(2*sVer + 1) is divided into up to 5 diamond search areas, as Figure 22 shown. Each search area is assigned a costFactor, which is determined by the distance (intDeltaMV) between each search point and the starting MV. Each diamond region is processed in order starting from the center of the search area. In each region, the search points are processed in raster scan order from the upper left corner to the lower right corner of the region. When the minimum bilCost within the current search area is less than the threshold (the threshold is equal to sbW * sbH), the integer-pixel full search terminates; otherwise, the integer-pixel full search continues to the next search area until all search points are checked.
[0333] Existing VVC DMVR fractional-sample refinement is further applied to derive the final deltaMV(sbIdx2). Then, the refined MVs in the second pass can be derived as:
[0334] ● MV0_pass2(sbIdx2) = MV0_pass1 + deltaMV(sbIdx2)
[0335] ● MV1_pass2(sbIdx2) = MV1_pass1 – deltaMV(sbIdx2)
[0336] 3.4.3 Third pass - sub-block based bidirectional optical flow MV refinement
[0337] In the third pass, refined MVs are derived by applying BDOF to 8×8 grid sub-blocks. For each 8×8 sub-block, BDOF refinement is applied to derive scaled Vx and Vy without clipping starting from the refined MVs of the parent and child blocks in the second pass. The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample precision and clipped between -32 and 32.
[0338] The refined MVs for the third pass (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) are derived as follows:
[0339] ● MV0_pass3(sbIdx3) = MV0_pass2(sbIdx2) + bioMv
[0340] ● MV1_pass3(sbIdx3) = MV0_pass2(sbIdx2) – bioMv
[0341] 3.5 OBMC
[0342] When applying Overlapped Block Motion Compensation (OBMC) as described in JVET-L0101, the motion information of neighboring blocks and weighted prediction are used to refine the top and left boundary pixels of the CU.
[0343] The conditions for not applying OBMC are as follows:
[0344] ● When OBMC is disabled at the SPS level
[0345] ● When the current block has an Intra mode or Intra Block Copy (IBC) mode
[0346] ● When the current block applies LIC
[0347] ● When the area of the current luma block is less than or equal to 32
[0348] Sub-block boundary OBMC is performed by applying the same blending to the top, left, bottom, and right sub-block boundary pixels using the motion information of neighboring sub-blocks. Enabled for sub-block based codec tools:
[0349] ● Affine AMVP mode;
[0350] ● Affine Merge mode and Sub-block based Temporal Motion Vector Prediction (SbTMVP);
[0351] ● Sub-block based bilateral matching.
[0352] 3.6 Sample-based BDOF
[0353] In sample-based BDOF, instead of deriving motion refinement (Vx, Vy) on a block basis, it is performed for each sample.
[0354] The coding block is divided into 8×8 sub-blocks. For each sub-block, whether to apply BDOF is determined by checking the SAD between two reference sub-blocks against a threshold. If it is decided to apply BDOF to the sub-block, for each sample point in the sub-block, a 5×5 sliding window is used, and the existing BDOF process is applied to each sliding window to derive Vx and Vy. The derived motion refinement (Vx, Vy) is applied to adjust the bi-predicted sample value of the central sample point of the window.
[0355] 3.7 Interpolation
[0356] The 8-tap interpolation filter used in VVC is replaced by a 12-tap filter. The interpolation filter is derived from a sine function, whose frequency response is cutoff at the Nyquist frequency and truncated by a cosine window function. Table 2 gives the filter coefficients for all 16 phases. Figure 23 An example of the frequency responses of the interpolation filter and the VVC interpolation filter at half-pixel phases is shown. Figure 23 The frequency responses of all interpolation filters and the VVC interpolation filter at half-pixel phases are compared.
[0357] Table 2. Filter Coefficients of the 12-tap Interpolation Filter
[0358]
[0359] 3.8 Multiple Hypothesis Prediction (MHP)
[0360] In the multiple hypothesis inter prediction mode (JVET-M0425), in addition to the regular bi-predicted signal, one or more additional motion-compensated prediction signals are signaled. The resulting overall prediction signal is obtained by weighted sample-wise superposition. Using the bi-predicted signal p bi and the first additional inter prediction signal / hypothesis h3, the resulting prediction signal p3 is obtained as follows:
[0361] p3 = (1 - α)p bi + αh3
[0362] The weighting factor α is specified by the new syntax element add_hyp_weight_idx according to the following mapping:
[0363] add_hyp_weight_idx α 0 1 / 4 1 -1 / 8
[0364] Similar to the above, more than one additional prediction signal can be used. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal.
[0365] p n+1 = (1 - α n+1 )p n + α n+1 h n+1
[0366] The resulting overall prediction signal is the last p obtained n (i.e., the p with the largest index n n ). Within this EE, up to two additional prediction signals can be used (i.e., n is limited to 2).
[0367] The motion parameters of each additional prediction hypothesis can be signaled explicitly by specifying a reference index, a motion vector predictor index, and a motion vector difference, or implicitly by specifying a Merge index. A separate multi-hypothesis Merge flag differentiates between these two signaling modes.
[0368] For the inter-frame AMVP mode, MHP is applied only when non-equal weights in the BCW are selected in the bi-prediction mode.
[0369] A combination of MHP and BDOF is possible, but BDOF is applied only to the bi-prediction signal part of the prediction signal (i.e., the ordinary first two hypotheses).
[0370] 3.9 Adaptive Reordering of Merge Candidates Based on Template Matching (ARMC-TM)
[0371] The Merge candidates are adaptively reordered by template matching (TM). The reordering method is applied to the regular Merge mode, the template matching (TM) Merge mode, and the affine Merge mode (excluding SbTMVP candidates). For the TM Merge mode, the Merge candidates are reordered before the refinement process.
[0372] After constructing the Merge candidate list, the Merge candidates are divided into several subgroups. For the regular Merge mode and the TM Merge mode, the subgroup size is set to 5. For the affine Merge mode, the subgroup size is set to 3. The Merge candidates within each subgroup are reordered in ascending order according to the template matching cost value based on the template. For simplicity, the Merge candidates in the last but not the first subgroup are not reordered.
[0373] The template matching cost of the Merge candidate is measured by the sum of absolute differences (SAD) between the samples of the template of the current block and its corresponding reference samples. The template includes a set of reconstructed samples adjacent to the current block. The reference samples of the template are located by the motion information of the Merge candidate.
[0374] When the Merge candidate uses bidirectional prediction, the reference samples of the template of the Merge candidate are also generated by bidirectional prediction, as Figure 24 shown. Figure 24 An example of the template in the reference picture and the reference samples of the template is shown.
[0375] For a block - based Merge candidate with a sub - block size equal to Wsub×Hsub, the upper template includes several sub - templates of size Wsub×1, and the left - hand template includes several sub - templates of size 1×Hsub. As Figure 25 shown, the motion information of the sub - blocks in the first row and the first column of the current block is used to derive the reference samples of each sub - template. Figure 25 An example of the template and the reference samples of the template of a block with sub - block motion using the motion information of the sub - blocks of the current block is shown.
[0376] 3.10 Geometric Partitioning Mode (GPM) with Merge Motion Vector Difference (MMVD)
[0377] The GPM in VVC is extended by applying motion vector refinement on top of the existing GPM unidirectional MVs. First, for a GPM CU, a signaling flag is used to specify whether to use this mode. If this mode is used, each geometric partition of the GPM CU can further decide whether to signal an MVD. If, after selecting a GPM Merge candidate, an MVD is signaled for a geometric partition, the motion of the partition will be further refined by the signaled MVD information. All other processes are the same as in GPM.
[0378] The MVD is signaled in the form of a distance and direction pair, similar to in MMVD. There are nine candidate distances (1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixel, 3 pixel, 4 pixel, 6 pixel, 8 pixel, 16 pixel) and eight candidate directions (four horizontal / vertical directions and four diagonal directions) in GPM with MMVD (GPM - MMVD). In addition, when pic_fpel_mmvd_enabled_flag equals 1, the MVD is left - shifted by 2 as in MMVD.
[0379] 3.11 Geometric Partitioning Mode (GPM) with Template Matching (TM)
[0380] Apply template matching to GPM. When the GPM mode is enabled for a CU, a CU-level flag is signaled to indicate whether TM is applied to the two geometric partitions. Use TM to refine the motion information for each geometric partition. When TM is selected, according to the partition angle, use the left, upper, or left and upper neighboring samples to construct a template as shown in Table 3. Then, with the half-pixel interpolation filter disabled, use the same Merge mode search pattern to refine the motion by minimizing the difference between the current template and the template in the reference picture.
[0381] Table 3. Templates for the 1st and 2nd geometric partitions, where A represents using upper samples, L represents using left samples, and L+A represents using both left and upper samples.
[0382]
[0383] The GPM candidate list is constructed as follows:
[0384] 1. The interleaved list 0 MV candidates and list 1 MV candidates are directly derived from the regular Merge candidate list, where the list 0 MV candidates have a higher priority than the list 1 MV candidates. Apply an adaptive threshold pruning method based on the current CU size to remove redundant MV candidates.
[0385] 2. The interleaved list 1 MV candidates and list 0 MV candidates are further directly derived from the regular Merge candidate list, where the list 1 MV candidates have a higher priority than the list 0 MV candidates. The same adaptive threshold pruning method is also applied to remove redundant MV candidates.
[0386] 3. Zero MV candidates will be filled until the GPM candidate list is full.
[0387] 4. GPM-MMVD and GPM-TM can only be enabled for one GPM CU. This is achieved by first signaling the GPM-MMVD syntax. When both GPM-MMVD control flags are equal to false (i.e., GPM-MMVD is disabled for the two GPM partitions), the GPM-TM flag is signaled to indicate whether template matching is applied to the two GPM partitions. Otherwise (at least one GPM-MMVD flag is equal to true), the value of the GPM-TM flag is presumed to be false.
[0388] 4. Technical problems solved by the disclosed technical solution
[0389] An example design for illumination compensation has the following problems. First, for each pixel in the current block, the prediction is derived only based on the corresponding pixel from the motion-compensated block. This ignores the relevant pixels in the current block and the neighboring pixels in the reference block, thus motivating an improved LIC method that takes into account all the information from neighboring pixels in the reference block.
[0390] Second, the LIC parameters are derived based on a template. However, the distribution of pixels in the template may be different from that in the current block. This makes the DC prediction inaccurate. This motivates the proposal of an improved DC prediction method.
[0391] Third, only one linear model is applied to the LIC block, which cannot be optimal for all scenarios. This motivates multi-mode LIC.
[0392] 5. List of Solutions and Embodiments
[0393] To solve the above problems and some other problems not mentioned, the methods outlined below are disclosed. These embodiments should be regarded as examples for explaining general concepts and should not be interpreted narrowly. In addition, these embodiments can be applied individually or in any combination.
[0394] The disclosed methods are applicable to other scenarios that employ illumination compensation, such as inter-view prediction in multi-view coding and decoding.
[0395] The term "block" can represent a coding block (CB), a coding unit (CU), a prediction unit (PU), a prediction block (PB), etc.
[0396] Example 1
[0397] An improved method for local illumination compensation (LIC) is disclosed. The following method is disclosed:
[0398] In one example, the new multi-parameter LIC (MPLIC) model can be applied as follows:
[0399] Y pred = α0Y0 + α1Y1 + α2Y2.. + α N-1 Y N-1 + β
[0400] where Y pred is the predicted sample, and Y0, Y1, Y2,... Y N-1 are reference samples derived based on at least one motion vector (MV). α0, α1, α2,... α N-1 , β are model parameters.
[0401] Example 2
[0402] In one example, Y0 can be the reference pixel pointed to by the MV for the current pixel, and Y1, Y2,... YN-1 It can be a neighboring pixel of Y0.
[0403] Example 3
[0404] In one example, Y0 can be a reference sample for the current sample X0 derived based on at least one MV. i (i = 1, 2, …, N - 1) can be derived through a predefined offset relative to Y0. In one example, if the current block is unidirectionally predicted, Y0 can be a reference sample derived based on one MV. In one example, if the current block is bidirectionally predicted and only one model can be derived, Y0 can be a reference sample derived based on two MVs. In one example, if the current block is bidirectionally predicted and two models can be derived for two directions, Y0 can be a reference sample derived based on one MV.
[0405] Example 4
[0406] In one example, Y0 can be a reference pixel obtained through a block vector or a template matching tool, etc.
[0407] Example 5
[0408] In one example, the MPLIC method can be used as an additional mode of LIC.
[0409] Example 6
[0410] In one example, the MPLIC method can replace the LIC method disclosed in the background art.
[0411] Example 7
[0412] In one example, the selection of N and neighboring positions Y1, Y2, … Y n can depend on the following factors: For example, the selection may depend on codec information such as block size, quantization parameter (QP) value, etc. For example, different selections can be made at the sequence / picture / strip / slice / CTU / CU level.
[0413] Example 8
[0414] In one example, the MPLIC method can be used for blocks with specific codec modes, such as AMVP or affine AMVP or Merge or sub - block Merge.
[0415] Example 9
[0416] In one example, the MPLIC method can be used under certain conditions, such as for certain sequences or frames or block sizes or QP values.
[0417] Example 10
[0418] In one example, the above method can be applied to all color components or a subset thereof. In one example, for different color components, the selection of N and the neighbors may be different. In one example, the above method can be applied only to the luminance component.
[0419] Example 11
[0420] In one example, the MPLIC method can be applied only to uni - directional prediction. For example, the MPLIC method can be applied to both uni - directional prediction and bi - directional prediction.
[0421] Example 12
[0422] In one example, the MPLIC method can be used only when the motion vector is not fractional precision. For example, when the motion vector is integer precision, the MPLIC method can completely replace the existing LIC method. For example, at least for the AMVR mode, the MPLIC method can be applied only when the MV precision is full - pixel. For example, when the MPLIC method is applied, the MV can be signaled only in at least full - pixel form. For example, when the MPLIC method is applied, the MV can be rounded to full - pixel.
[0423] Example 13
[0424] In one example, at least two different models can be employed. One model can process samples along rows, and a second model can then process samples along columns, and vice versa. In one example, the number of models to be applied can depend on factors such as the precision of the motion vector.
[0425] Example 14
[0426] In one example, for the Merge mode, the use of the MPLIC model is inherited from its use among the neighbors.
[0427] Example 15
[0428] In one example, for each Merge candidate (such as X) using the two - parameter LIC model disclosed in the background art, an additional Merge candidate is inserted into the MV candidate list. In one example, the new Merge candidate can inherit all the information of the X candidate and replace the LIC model with a new multi - parameter LIC model. For example, when constructing a new Merge candidate using the new multi - parameter LIC model, only a subset of the candidates in the MV candidate list is considered. In one example, a block can be encoded and decoded in a specific mode, such as regular Merge, MMVD, affine Merge, sub - block Merge, template - matching Merge, etc. The MV candidate list can be a regular Merge list, an affine Merge list, a template - matching Merge list, etc.
[0429] Example 16
[0430] In one example, α0, α1, α2, … αn, β are model parameters derived by minimizing the following error function based on the current block template and the reference block template, and they do not require signaling overhead except for signaling the MPLIC flag for the AMVP / affine AMVP mode to indicate the use of MPLIC.
[0431] In one example, the error function can be:
[0432]
[0433] where X0 is a sample point in the current template, Y0 is the reference sample point of X0, and Y i (i = 1, 2, …, N - 1) has a predefined offset relative to Y0, ||T|| is the number of template sample points, and λ n is the regularization parameter, and o n is predefined. For example, o0 = 1, and o i (i = 1, 2, …, N - 1) = 0.
[0434] Alternatively, the error function can be
[0435]
[0436] Example 17
[0437] In one example, an improved method for DC prediction is disclosed. The following method is disclosed:
[0438] In one example, a mapping function is defined. For each sample point in the current block, the corresponding virtual sample point is derived using the mapping function, thereby obtaining a virtual block. In one example, a mapping function is formed between the reference block pixels and the reference template. This mapping function is used to map the sample point values in the current block to the sample point values in the current template adjacent to the current block. (See example Figure 26 ). Figure 26 An example of virtual block generation for improved DC prediction is shown. In one example, the DC of the virtual block is used as the DC prediction of the current block.
[0439] Example 18
[0440] In one example, the mapping function can be constructed based on criteria such as the minimum SAD, sum of squared errors (SSE), etc. between the sample points in the reference block and the sample points in the reference template. In one example, if more than one reference template sample point yields the same minimum cost for the reference block sample point, the sample point closest in spatial domain distance is selected.
[0441] Example 19
[0442] In one example, the use of the method can depend on encoding / decoding information, such as block size, QP, etc. For example, other options and conditions listed in other disclosures.
[0443] Example 20
[0444] In one example, a multi-model LIC is disclosed. The following method is disclosed. In one example, reference samples can be classified into different subsets, and different LIC models can be applied to each subset.
[0445] In one example, the LIC model for a subset can be a previously disclosed two-parameter linear model or multi-parameter linear model, or a combination thereof. In one example, the classification of reference samples can be based on a predefined threshold, such as the average of sample values, etc. In one example, the classification can be based on an implicit method, such as clustering. In one example, a signal transmission threshold can be explicitly passed. In one example, choosing between an implicit threshold and an explicitly communicated threshold can have additional flexibility. In one example, when used in conjunction with a geometric partitioning mode (GPM), the classification of samples can be based on the partitioning of templates obtained through partitioning in the GPM. In one example, the use of the method can depend on encoding / decoding, such as block size, QP, etc. For example, other options and conditions listed in other disclosures.
[0446] Example 21
[0447] In one example, for each of the methods disclosed above, a syntax element can be signaled in the bitstream (e.g., a flag) to specify whether the disclosed prediction mode is ultimately selected for encoding / decoding the current block. Additionally, the syntax element can depend on whether the block is AMVP encoded or Merge encoded.
[0448] Example 22
[0449] Additionally, the syntax element can depend on whether the block is unidirectionally predicted or bidirectionally predicted. In one example, if the current block is bidirectionally predicted, the syntax element can be presumed to be false.
[0450] Example 23
[0451] Additionally, the syntax element can depend on the type of Merge prediction used.
[0452] Example 24
[0453] In one example, the syntax element is signaled only for all or a subset of the Y, U, V planes.
[0454] Example 25
[0455] In addition, syntax elements can be signaled conditionally.
[0456] Example 26
[0457] Whether / how to signal a syntax element can depend on the dimensions of the current block. For example, a syntax element can be signaled only if the current block is larger than a predefined size. For example, a syntax element can be signaled only if the sum of the width and height of the current block is greater than a predefined threshold. In one example, greater than can be replaced by "less than", "not greater than", or "not less than".
[0458] Example 27
[0459] Whether / how to signal a syntax element can depend on the QP selected for encoding / decoding the current block. For example, a syntax element is signaled only if the QP of the current block is greater than a predefined threshold. In one example, greater than can be replaced by "less than", "not greater than", or "not less than".
[0460] Example 28
[0461] In one example, a syntax element can be binarized into a fixed-length code, (truncated) unary code, exponential Golomb code, etc. In one example, a syntax element can be encoded / decoded using at least one coding context in arithmetic coding. In one example, a syntax element can be encoded / decoded using bypass coding.
[0462] Example 29
[0463] Whether and / or how to apply the methods disclosed above can be signaled at the sequence level, group of pictures level, picture level, slice level, and / or slice group level, for example, in the sequence header, picture header, SPS, video parameter set (VPS), dependent parameter set (DPS), decoding capability information (DCI), PPS, adaptive parameter set (APS), slice header, and / or slice group header.
[0464] Example 30
[0465] Whether and / or how to apply the methods disclosed above can be signaled at the PB, transform block (TB), CB, PU, TU, CU, virtual pipeline data unit (VPDU), CTU, CTU row, slice, tile, sub-picture, and / or other types of regions containing more than one sample or pixel.
[0466] Example 31
[0467] Whether and / or how to apply the methods disclosed above can depend on coding information such as block size, color format, single / double tree segmentation, color component, slice / picture type.
[0468] 6. References
[0469] [1] ITU-T and ISO / IEC, "High Efficiency Video Coding", Rec. ITU-T H.265|ISO / IEC 23008-2 (current version).
[0470] [2] B. Bross, J. Chen, S. Liu, and Y.-K. Wang, "Versatile Video Coding (Draft 10)", JVET-2001 document, 19th JVET meeting: by teleconference, June 22 - July 1, 2020.
[0471] [3] A. Browne, J. Chen, Y. Ye, S. Kim, "Algorithm description of the Versatile Video Coding and Test Model 14 (VTM 14)", JVET-W2002, September 2021.
[0472] [4] M. Coban, F. Le Léannec, J. Strom, "Algorithm description of the Enhanced Compression Model 2 (ECM 2)", JVET-W2025, September 2021.
[0473] [5] VTM software: https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM.git
[0474] [6] ECM software: https: / / vcgit.hhi.fraunhofer.de / ecm / ECM.git
[0475] Figure 27 is a block diagram showing an example video processing system 4000 in which the various techniques disclosed herein can be implemented. Various specific implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8-bit or 10-bit multi-component pixel values), or may be received in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.
[0476] System 4000 may include a codec component 4004 that may implement various codec or encoding methods described in this document. The codec component 4004 may reduce the average bit rate of a video from the input 4002 to the output of the codec component 4004 to produce a coded representation of the video. Thus, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of the codec component 4004 may be stored or transmitted via a connected communication as represented by component 4006. The bitstream (or coded) representation of the stored or transmitted video received at the input 4002 may be used by component 4008 to generate pixel values or a displayable video to be sent to the display interface 4010. The process of generating a user-viewable video from the bitstream representation is sometimes referred to as video decompression. Further, although certain video processing operations are referred to as "codec" operations or tools, it should be understood that encoding tools or operations are used at the encoder, while the corresponding decoding tools or operations that reverse the encoding results will be performed by the decoder.
[0477] Examples of a peripheral bus interface or a display interface may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), DisplayPort, etc. Examples of a storage interface include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE) interface, etc. The techniques described in this document may be embodied in various electronic devices such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0478] Figure 28 is a block diagram of an example video processing apparatus 4100. The apparatus 4100 may be used to implement one or more methods described herein. The apparatus 4100 may be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The (multiple) processors 4102 may be configured to implement one or more methods described in this document. The memory (multiple memories) 4104 may be used to store data and code for implementing the methods and techniques described herein. The video processing circuitry 4106 may be used to implement some of the techniques described in this document in hardware circuitry. In some embodiments, the video processing circuitry 4106 may be at least partially included in the processor 4102, such as a graphics coprocessor.
[0479] Figure 29is a flowchart of an example method 4200 for video processing. Method 4200 includes determining to apply multi-parameter local illumination compensation (MPLIC) to visual media data at step 4202. At step 4204, a conversion between the visual media data and a bitstream is performed based on the MPLIC. According to an example, the conversion step 4204 may include encoding at an encoder or decoding at a decoder.
[0480] It should be noted that method 4200 may be implemented in a device for processing video data, the device including a processor and a non-transitory memory having instructions thereon, such as video encoder 4400, video decoder 4500, and / or encoder 4600. In such a case, the instructions, when executed by the processor, cause the processor to perform method 4200. Additionally, method 4200 may be executed by a non-transitory computer-readable medium that includes a computer program product for use in a video codec device. The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium such that, when executed by a processor, cause the video codec device to perform method 4200.
[0481] Figure 30 is a block diagram showing an example video codec system 4300 that may utilize the techniques of the present disclosure. Video codec system 4300 may include a source device 4310 and a destination device 4320. Source device 4310 generates encoded video data, which may be referred to as a video encoding device. Destination device 4320 may decode the encoded video data generated by source device 4310, which may be referred to as a video decoding device.
[0482] Source device 4310 may include a video source 4312, a video encoder 4314, and an input / output (I / O) interface 4316. Video source 4312 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data or a combination of such sources. The video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a codec representation of a picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be directly transmitted to destination device 4320 via I / O interface 4316 over network 4330. The encoded video data may also be stored on a storage medium / server 4340 for access by destination device 4320.
[0483] The target device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. The I / O interface 4326 may include a receiver and / or a modem. The I / O interface 4326 may obtain encoded video data from a source device 4310 or a storage medium / server 4340. The video decoder 4324 may decode the encoded video data. The display device 4322 may display the decoded video data to a user. The display device 4322 may be integrated with the target device 4320 or may be external to the target device 4320 that may be configured to interface with an external display device.
[0484] The video encoder 4314 and the video decoder 4324 may operate according to video compression standards, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or additional standards.
[0485] Figure 31 is a block diagram showing an example of a video encoder 4400, which may be Figure 30 the video encoder 4314 in the system 4300 shown in. The video encoder 4400 may be configured to perform any or all of the techniques of the present disclosure. The video encoder 4400 includes a plurality of functional components. The techniques described in the present disclosure may be shared among various components of the video encoder 4400. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.
[0486] The functional components of the video encoder 4400 may include a splitting unit 4401, a prediction unit 4402, a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a buffer 4413, and an entropy coding unit 4414. The prediction unit may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra prediction unit 4406.
[0487] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in the IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0488] In addition, some components, such as the motion estimation unit 4404 and the motion compensation unit 4405, may be highly integrated but are shown separately in the example of the video encoder 4400 for purposes of explanation.
[0489] The splitting unit 4401 can split an image into one or more video blocks. The video encoder 4400 and the video decoder 4500 can support various video block sizes.
[0490] The mode selection unit 4403 can select, for example, one of the coding / decoding modes (intra or inter) based on an error result, and provide the resulting intra or inter coded block to the residual generation unit 4407 to generate residual block data and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 4403 can select a combination of intra and inter prediction (CIIP) mode, where the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 4403 can also select the resolution of the motion vector for the block (e.g., sub-pixel or integer pixel accuracy).
[0491] To perform inter prediction on a current video block, the motion estimation unit 4404 can generate motion information for the current video block by comparing one or more reference frames from the cache 4413 with the current video block. The motion compensation unit 4405 can determine a predicted video block for the current video block based on the motion information and decoded samples of a picture from the cache 4413 other than the picture associated with the current video block.
[0492] The motion estimation unit 4404 and the motion compensation unit 4405 can perform different operations on the current video block, e.g., depending on whether the current video block is in an I-slice, a P-slice, or a B-slice.
[0493] In some examples, the motion estimation unit 4404 can perform uni-directional prediction for the current video block, and the motion estimation unit 4404 can search the reference pictures in list 0 or list 1 for a reference video block for the current video block. Then the motion estimation unit 4404 can generate a reference index indicating the reference picture in list 0 or list 1 that contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 4404 can output the reference index, the prediction direction indicator, and the motion vector as the motion information for the current video block. The motion compensation unit 4405 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0494] In other examples, the motion estimation unit 4404 may perform bidirectional prediction for a current video block. The motion estimation unit 4404 may search for a reference video block of the current video block in the reference pictures in list 0 and may also search for another reference video block of the current video block in the reference pictures in list 1. Then, the motion estimation unit 4404 may generate a reference index indicating the reference pictures containing the reference video blocks in list 0 and list 1 and a motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 4404 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 4405 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.
[0495] In some examples, the motion estimation unit 4404 may output a complete set of motion information for the decoding process of the decoder. In some examples, the motion estimation unit 4404 may not output a complete set of motion information of the current video. Instead, the motion estimation unit 4404 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is similar enough to the motion information of neighboring video blocks.
[0496] In one example, the motion estimation unit 4404 may indicate in a syntax structure associated with the current video block a value indicating to the video decoder 4500 that the current video block has the same motion information as another video block.
[0497] In another example, the motion estimation unit 4404 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0498] As discussed above, the video encoder 4400 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 4400 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0499] The intra prediction unit 4406 may perform intra prediction on a current video block. When the intra prediction unit 4406 performs intra prediction on the current video block, the intra prediction unit 4406 may generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data of the current video block may include a predicted video block and various syntax elements.
[0500] The residual generation unit 4407 may generate residual data for a current video block by subtracting a (plurality of) predicted video blocks of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0501] In other examples, there may be no residual data for the current video block, for example, in skip mode, and the residual generation unit 4407 may not perform the subtraction operation.
[0502] The transform processing unit 4408 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0503] After the transform processing unit 4408 generates the transform coefficient video blocks associated with the current video block, the quantization unit 4409 may quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0504] The inverse quantization unit 4410 and the inverse transform unit 4411 may apply inverse quantization and inverse transform to the transform coefficient video blocks, respectively, to reconstruct the residual video blocks from the transform coefficient video blocks. The reconstruction unit 4412 may add the reconstructed residual video blocks to the corresponding samples of one or more predicted video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block for storage in the buffer 4413.
[0505] After the reconstruction unit 4412 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.
[0506] The entropy encoding unit 4414 may receive data from other functional components of the video encoder 4400. When the entropy encoding unit 4414 receives data, the entropy encoding unit 4414 may perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.
[0507] Figure 32 is a block diagram illustrating an example of a video decoder 4500, which may be Figure 30 the video decoder 4324 in the system 4300 shown in. The video decoder 4500 may be configured to perform any or all of the techniques of the present disclosure. In the example shown, the video decoder 4500 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video decoder 4500. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.
[0508] In the example shown, video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, video decoder 4500 may perform a decoding process that is generally the reverse of the encoding process described with respect to video encoder 4400.
[0509] Entropy decoding unit 4501 may retrieve an encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., coded video data blocks). Entropy decoding unit 4501 may decode the entropy-coded video data, and based on the entropy-decoded video data, motion compensation unit 4502 may determine motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. For example, motion compensation unit 4502 may determine such information by performing AMVP and Merge modes.
[0510] Motion compensation unit 4502 may generate a motion-compensated block, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter to be used with sub-pixel precision may be included in a syntax element.
[0511] Motion compensation unit 4502 may compute interpolation of sub-integer pixels of a reference block using the interpolation filter used by video encoder 4400 during encoding of a video block. Motion compensation unit 4502 may determine the interpolation filter used by video encoder 4400 based on received syntax information and use the interpolation filter to generate a prediction block.
[0512] Motion compensation unit 4502 may use some syntax information to determine the size of blocks used to encode (multiple) frames and / or (multiple) slices of an encoded video sequence, partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.
[0513] Intra prediction unit 4503 may form a prediction block from spatially adjacent blocks using, for example, an intra prediction mode received in the bitstream. Inverse quantization unit 4504 inverse quantizes, i.e., de-quantizes, the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 4501. Inverse transform unit 4505 applies an inverse transform.
[0514] The reconstruction unit 4506 can add the residual block to the corresponding prediction block generated by the motion compensation unit 4502 or the intra prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block effect artifacts. Then the decoded video block is stored in the buffer 4507, which provides reference blocks for subsequent motion compensation / intra prediction and also generates the decoded video for presentation on a display device.
[0515] Figure 33 is a schematic diagram of an example encoder 4600. The encoder 4600 is suitable for implementing the technologies of VVC. The encoder 4600 includes three loop filters, namely, a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Different from the DF 4602 that uses a predefined filter, the SAO 4604 and the ALF 4606 utilize the original samples of the current picture. By adding offsets and applying finite impulse response (FIR) filters respectively, and using the coding / decoding side information of the signal transmission offsets and filter coefficients, the mean square error between the original samples and the reconstructed samples is reduced. The ALF 4606 is located at the last processing stage of each picture and can be regarded as a tool for trying to capture and repair the artifacts created by the previous stages.
[0516] The encoder 4600 also includes an intra prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive the input video. The intra prediction component 4608 is configured to perform intra prediction, while the ME / MC component 4610 is configured to perform inter prediction using the reference pictures obtained from the reference picture buffer 4612. The residual blocks from the inter prediction or intra prediction are fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are fed into an entropy coding / decoding component 4618. The entropy coding / decoding component 4618 performs entropy coding / decoding on the prediction results and the quantized transform coefficients and sends them to a video decoder (not shown). The quantized components output from the quantization component 4616 can be fed into an inverse quantization component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 is capable of outputting the images to the DF 4602, the SAO 4604, and the ALF 4606 for filtering before these images are stored in the reference picture buffer 4612.
[0517] Next, a list of preferred solutions for some examples is provided.
[0518] The following solutions show examples of the technologies discussed herein.
[0519] 1. A method for processing video data, comprising: determining to apply multi-parameter local illumination compensation (MPLIC) to visual media data; and performing a conversion between the visual media data and a bitstream based on the MPLIC.
[0520] 2. The method according to solution 1, wherein the MPLIC is applied as follows:
[0521] Y pred = α0Y0 + α1Y1 + α2Y2.. + α N-1 Y N-1 + β
[0522] where Y pred is a predicted sample, Y0, Y1, Y2,... Y N-1 reference samples are derived based on at least one MV, and α0, α1, α2,... αN-1, β are model parameters.
[0523] 3. The method according to any one of solutions 1-2, wherein Y0 is a reference pixel pointed to by the MV for the current pixel, and Y1, Y2,... Y N-1 are neighboring pixels of Y0.
[0524] 4. The method according to any one of solutions 1-3, wherein Y0 is a reference sample for the current sample X0 derived based on at least one MV, and wherein Y i (i = 1, 2,..., N-1) is derived with a predefined offset relative to Y0.
[0525] 5. The method according to any one of solutions 1-4, wherein when the current block uses uni-directional prediction, Y0 is a reference sample derived based on one MV, wherein when the current block is bi-directional prediction and only one model is derived, Y0 is a reference sample derived based on two MVs, or wherein when the current block is bi-directional prediction and two models are derived for two directions, Y0 is a reference sample derived based on one MV.
[0526] 6. The method according to any one of solutions 1-5, wherein Y0 is a reference pixel obtained by a block vector or a template matching tool.
[0527] 7. The method according to any one of solutions 1-6, wherein the selection of N and the neighboring positions Y1, Y2,... Y n depends on factors including codec information, and different selections are adopted at the sequence level, picture level, slice level, tile level, coding tree unit (CTU) level, coding unit (CU) level, or a combination thereof.
[0528] 8. The method according to any one of Solutions 1-7, wherein MPLIC is used for blocks that employ Advanced Motion Vector Prediction (AMVP), Affine AMVP, or Merge mode or sub-block Merge mode.
[0529] 9. The method according to any one of Solutions 1-8, wherein MPLIC is used for certain sequences, frames, block sizes, quantization parameter (QP) values, or combinations thereof.
[0530] 10. The method according to any one of Solutions 1-9, wherein MPLIC is applied to luminance, blue-difference chrominance, red-difference chrominance, or combinations thereof.
[0531] 11. The method according to any one of Solutions 1-10, wherein MPLIC is applied only to uni-directional prediction or to both uni-directional prediction and bi-directional prediction.
[0532] 12. The method according to any one of Solutions 1-11, wherein MPLIC is used only when the motion vector is not fractional precision.
[0533] 13. The method according to any one of Solutions 1-12, wherein MPLIC includes two models, and wherein the first model processes samples along rows and the second model processes samples along columns.
[0534] 14. The method according to any one of Solutions 1-13, wherein when the current block uses the Merge mode, the use of MPLIC is inherited from neighboring blocks.
[0535] 15. The method according to any one of Solutions 1-14, wherein α0, α1, α2, … α n and β are model parameters derived by minimizing an error function based on the current block template and the reference block template.
[0536] 16. The method according to any one of Solutions 1-15, wherein the error function is:
[0537]
[0538] where X0 is a sample in the current template, Y0 is the reference sample of X0, Y i (i = 1, 2, …, N-1) has a predefined offset relative to Y0, ||T|| is the number of template samples, λ n is the regularization parameter, and o n is a predefined value.
[0539] 17. The method according to any one of Solutions 1-16, wherein the error function is:
[0540]
[0541] where X0 is a sample point in the current template, Y0 is the reference sample point of X0, and Y i (i = 1, 2, …, N - 1) has a predefined offset relative to Y0, and ||T|| is the number of template sample points.
[0542] 18. The method according to any one of Solutions 1 - 17, wherein DC (Direct Current) prediction is employed, and wherein for each sample point in the current block, a mapping function is used to derive a corresponding virtual sample point to generate a virtual block.
[0543] 19. The method according to any one of Solutions 1 - 18, wherein the mapping function is formed between the reference block pixels and the reference template, wherein the mapping function is used to map the sample point values in the current block to the sample point values in the current template adjacent to the current block, and wherein the DC of the virtual block is regarded as the DC prediction of the current block.
[0544] 20. The method according to any one of Solutions 1 - 19, wherein the mapping function is constructed based on the criterion between the sample points in the reference block and the sample points in the reference template, and wherein when more than one reference template sample point generates the same minimum cost for the reference block sample point, the sample point closest in spatial domain distance is selected.
[0545] 21. The method according to any one of Solutions 1 - 20, wherein the use of the mapping function depends on the codec information.
[0546] 22. The method according to any one of Solutions 1 - 21, wherein MPLIC employs reference sample points classified into different subsets, and wherein different local illumination compensation (LIC) models are applied to each subset.
[0547] 23. The method according to any one of Solutions 1 - 22, wherein the LIC model for the subset is a two - parameter linear model or a multi - parameter linear model, wherein the classification of the reference sample points is based on a predefined threshold, wherein the classification is based on clustering, wherein the classification is explicitly transmitted through a signal, or wherein when used in conjunction with a geometric partitioning mode (GPM), the classification of the sample points is based on the partitioning of the template obtained by GPM partitioning.
[0548] 24. The method according to any one of Solutions 1 - 23, wherein the bitstream includes one or more syntax elements indicating whether a prediction mode is selected for the current block.
[0549] 25. The method according to any one of Solutions 1-24, wherein the syntax element depends on: whether the current block is encoded / decoded by AMVP, whether the current block is encoded / decoded by Merge, whether the current block is unidirectionally predicted, whether the current block is bidirectionally predicted, type Merge prediction, or a combination thereof.
[0550] 26. The method according to any one of Solutions 1-25, wherein the syntax element is conditionally signaled depending on the dimension of the current block to signal the syntax element, depending on the quantization parameter (QP) used for encoding / decoding the current block, context encoding, binarization, or a combination thereof to signal the syntax element.
[0551] 27. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to execute the method according to any one of Solutions 1 to 26.
[0552] 28. A non-transitory computer-readable medium, comprising a computer program product for use in a video encoding / decoding device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium such that, when executed by a processor, cause the video encoding / decoding device to execute the method according to any one of Solutions 1 to 26.
[0553] 29. A non-transitory computer-readable recording medium that stores a bitstream of a video generated by a method executed by a video processing device, wherein the method comprises: determining to apply multi-parameter local illumination compensation (MPLIC) to visual media data; and generating the bitstream based on the determination.
[0554] 30. A method for storing a video bitstream, comprising: determining to apply multi-parameter local illumination compensation (MPLIC) to visual media data; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0555] 31. The method, apparatus, or system described in this document.
[0556] The following solutions show further examples of the techniques discussed herein.
[0557] 1. A method for processing video data, comprising: determining to apply multi-parameter local illumination compensation (MPLIC) to visual media data; and performing a conversion between the visual media data and a bitstream based on the MPLIC.
[0558] 2. The method according to Solution 1, wherein the MPLIC is applied as follows: Y pred = α0Y0 + α1Y1 + α2Y2.. + α N-1 Y N-1 + β where Y pred is a predicted sample point, Y0, Y1, Y2,... Y N-1 are reference sample points derived based on at least one motion vector (MV), and α0, α1, α2,... αN-1, β are model parameters.
[0559] 3. The method according to any one of Solutions 1-2, wherein Y0 is a reference pixel pointed to by the MV for the current pixel, and Y1, Y2,... Y N-1 are neighboring pixels of Y0.
[0560] 4. The method according to any one of Solutions 1-3, wherein Y0 is a reference sample point for the current sample point X0 derived based on at least one MV, and wherein Y i (i = 1, 2,..., N-1) is derived with a predefined offset relative to Y0.
[0561] 5. The method according to any one of Solutions 1-4, wherein when the current block uses uni-directional prediction, Y0 is a reference sample point derived based on one MV, wherein when the current block is bi-directional prediction and only one model is derived, Y0 is a reference sample point derived based on two MVs, or wherein when the current block is bi-directional prediction and two models are derived for two directions, Y0 is a reference sample point derived based on one MV.
[0562] 6. The method according to any one of Solutions 1-5, wherein Y0 is a reference pixel obtained by a block vector or a template matching tool.
[0563] 7. The method according to any one of Solutions 1-6, wherein MPLIC is used as an additional local illumination compensation (LIC) mode or in place of LIC.
[0564] 8. The method according to any one of Solutions 1-7, wherein the selection of N and the neighboring positions Y1, Y2,... Y n depends on factors including codec information, the codec information including the block size of the quantization parameter (QP) value, wherein different selections are made at the sequence level, picture level, slice level, tile level, coding tree unit (CTU) level, coding unit (CU) level, or a combination thereof.
[0565] 9. The method according to any one of Solutions 1-8, wherein MPLIC is used for blocks employing Advanced Motion Vector Prediction (AMVP), Affine AMVP, or Merge mode or sub-block Merge mode.
[0566] 10. The method according to any one of Solutions 1-9, wherein MPLIC is used for certain sequences, frames, block sizes, Quantization Parameter (QP) values, or combinations thereof.
[0567] 11. The method according to any one of Solutions 1-10, wherein MPLIC is applied to all color components or a subset of color components, or wherein the selection of N and neighbors is different for different color components, or wherein MPLIC is applied only to the luminance component.
[0568] 12. The method according to any one of Solutions 1-11, wherein MPLIC is applied only to uni-directional prediction or to both uni-directional and bi-directional prediction.
[0569] 13. The method according to any one of Solutions 1-12, wherein MPLIC is used only when the motion vector is not fractional precision, or wherein when the motion vector is integer precision, MPLIC completely replaces LIC, or wherein at least for the Adaptive Motion Vector Resolution (AMVR) mode, MPLIC is applied only when the MV precision is full pixel, or wherein when MPLIC is applied, the MV can only be signaled in at least full pixel form, or when MPLIC is applied, the MV is rounded to full pixel.
[0570] 14. The method according to any one of Solutions 1-13, wherein MPLIC includes two models, and wherein the first model processes samples along rows and the second model processes samples along columns, or wherein the number of models to be applied depends on the precision of the motion vector.
[0571] 15. The method according to any one of Solutions 1-14, wherein when the current block uses Merge mode, the use of MPLIC is inherited from neighboring blocks.
[0572] 16. The method according to any one of Solutions 1 - 15, wherein for each Merge candidate X using a two-parameter LIC model, an additional Merge candidate is inserted into the MV candidate list, or wherein the additional Merge candidate inherits all information of the Merge candidate X and replaces the LIC model with a multi-parameter LIC model, or wherein only a subset of candidates in the MV candidate list is considered to construct an additional Merge candidate using a multi-parameter LIC model, or wherein a block is encoded or decoded using regular Merge, motion vector difference (MMVD), affine Merge, sub-block Merge, template matching Merge, or a combination thereof, or wherein the MV candidate list is a Merge list, an affine Merge list, a template matching Merge list, or a combination thereof.
[0573] 17. The method according to any one of Solutions 1 - 16, wherein α0, α1, α2, … α n and β are model parameters derived by minimizing an error function based on a current block template and a reference block template, and wherein no signaling overhead is used except for signaling an MPLIC flag to indicate the use of MPLIC for the AMVP or affine AMVP mode.
[0574] 18. The method according to any one of Solutions 1 - 17, wherein the error function is: where X0 is a sample point in the current template, Y0 is the reference sample point of X0, Y i (i = 1, 2, …, N - 1) has a predefined offset relative to Y0, ||T|| is the number of template sample points, λ n is a regularization parameter, and o n is a predefined value or o0 = 1, o i (i = 1, 2, …, N - 1) = 0.
[0575] 19. The method according to any one of Solutions 1 - 18, wherein the error function is: E = σ ||T|| (α0Y0 + α1Y1 + α2Y2.. + α N-1 Y N-1 + β - X0) 2 where X0 is a sample point in the current template, Y0 is the reference sample point of X0, Y i (i = 1, 2, …, N - 1) has a predefined offset relative to Y0, and ||T|| is the number of template sample points.
[0576] 20. The method according to any one of Solutions 1-19, wherein DC (Direct Current) prediction is employed, and wherein for each sample point in the current block, a mapping function is used to derive a corresponding virtual sample point to generate a virtual block.
[0577] 21. The method according to any one of Solutions 1-20, wherein the mapping function is formed between the pixels of the reference block and the reference template, and the mapping function is used to map the sample point value in the current block to the sample point value in the current template adjacent to the current block, or the DC of the virtual block is regarded as the DC prediction of the current block.
[0578] 22. The method according to any one of Solutions 1-21, wherein the mapping function is constructed based on the criterion between the sample points in the reference block and the sample points in the reference template, or when more than one reference template sample point generates the same minimum cost for the reference block sample point, the sample point closest in spatial domain distance is selected.
[0579] 23. The method according to any one of Solutions 1-22, wherein the use of the mapping function depends on the codec information including the block size or quantization parameter (QP) size.
[0580] 24. The method according to any one of Solutions 1-23, wherein MPLIC employs reference sample points classified into different subsets, and different local illumination compensation (LIC) models are applied to each subset.
[0581] 25. The method according to any one of Solutions 1-24, wherein the LIC model for the subset is a two-parameter linear model or a multi-parameter linear model, the classification of the reference sample points is based on a predefined threshold, the classification is based on clustering, the threshold is explicitly transmitted by the signal, additional flexibility in selecting between the implicit threshold and the explicitly communicated threshold is adopted, when used in conjunction with the geometric partitioning mode (GPM), the classification of the sample points is based on the partitioning of the template obtained by the GPM partitioning, or the use of MPLIC depends on the codec information including the block size QP size.
[0582] 26. The method according to any one of Solutions 1-25, wherein the bitstream includes one or more syntax elements indicating whether a prediction mode is selected for the current block.
[0583] 27. The method according to any one of Solutions 1-26, wherein the syntax element depends on: whether the current block is AMVP decoded, whether the current block is Merge decoded, whether the current block is unidirectionally predicted, whether the current block is bidirectionally predicted, type Merge prediction, or a combination thereof, or wherein when the current block is bidirectionally predicted, the syntax element is presumed to be false.
[0584] 28. The method according to any one of Solutions 1-27, wherein the syntax element is conditionally signaled, the syntax element is signaled only for a subset of the Y, U, and V planes, the syntax element is signaled depending on the dimensions of the current block, the syntax element is signaled depending on the quantization parameter (QP) used to decode the current block, the syntax element is context decoded, the syntax element is binarized into a fixed length code, a truncated unary code, an exponential Golomb code, the syntax element is decoded with at least one context in arithmetic coding including bypass coding, or a combination thereof.
[0585] 29. The method according to any one of Solutions 1-28, wherein the syntax element is signaled only when the current block is larger than a predefined size, or wherein the syntax element is signaled only when the sum of the width and height of the current block is greater than, less than, not greater than, or not less than a predefined threshold, or wherein the syntax element is signaled only when the QP of the current block is greater than, less than, not greater than, or not less than a predefined threshold, or a combination thereof.
[0586] 30. The method according to any one of Solutions 1-29, wherein the use of the method is signaled at the sequence level, picture group level, picture level, slice level, or slice group level, including in the sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header, or slice group header, or wherein the use of the method is signaled at a prediction block (PB), transform block (TB), decoded block (CB), prediction unit (PU), transform unit (TU), decoded unit (CU), virtual pipeline decoding unit (VPDU), decoded tree unit (CTU), CTU row, slice, picture, sub-picture, or other region containing more than one sample or pixel.
[0587] 31. The method according to any one of Solutions 1-30, wherein the application of the method depends on decoding information including block size, color format, single tree segmentation, double tree segmentation, color component, slice type, or picture type.
[0588] 32. The method according to any one of Solutions 1-31, wherein the conversion includes encoding the visual media data into the bitstream.
[0589] 33. The method according to any one of Solutions 1-31, wherein the conversion includes decoding the visual media data from the bitstream.
[0590] 34. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of Solutions 1 to 33.
[0591] 35. A non-transitory computer-readable medium, comprising a computer program product for use by a video codec device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium such that, when executed by a processor, cause the video codec device to perform the method according to any one of Solutions 1 to 33.
[0592] 36. A non-transitory computer-readable recording medium that stores a bitstream of a video generated by a method executed by a video processing device, wherein the method includes: determining to apply multi-parameter local illumination compensation (MPLIC) to visual media data; and generating the bitstream based on the determination.
[0593] 37. A method for storing a video bitstream, the method includes: determining to apply multi-parameter local illumination compensation (MPLIC) to visual media data; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0594] In the solutions described herein, the encoder can conform to the format rules by generating an encoded / decoded representation according to the format rules. In the solutions described herein, the decoder can use the format rules to parse the syntax elements in the encoded / decoded representation to generate a decoded video in the case of knowing the presence or absence of the syntax elements according to the format rules.
[0595] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. The bitstream representation of the current video block may correspond, for example, to bits that are co-located within the bitstream or are distributed at different positions, as defined by the syntax. For example, a macroblock may be encoded based on the transform and coding / decoding error residual values, and may also use bits in the header and other fields in the bitstream. Additionally, during the conversion, the decoder may parse the bitstream based on this determination, knowing whether certain fields are present or not, as described in the above solution. Similarly, the encoder may determine whether to include certain syntax fields and may generate the coded / decoded representation accordingly by including or excluding the syntax fields in the coded / decoded representation.
[0596] The disclosed solutions, examples, embodiments, modules, and functional operations described in this document, as well as other solutions, examples, embodiments, modules, and functional operations, may be implemented in digital electronic circuitry or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in a combination of one or more of them. The disclosed embodiments, as well as other embodiments, may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatuses, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to the hardware, the apparatus may also include code that creates an execution environment for the computer program being considered, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to a suitable receiver device.
[0597] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a part of a file that contains other programs or data (e.g., one or more scripts in a markup language document), in a single file dedicated to the program being considered, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[0598] The processes and logical flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[0599] By way of example, processors suitable for the execution of a computer program include both general and special purpose microprocessors, as well as any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data (e.g., magnetic disks, magneto-optical disks, or optical disks) from which it receives data or to which it transfers data, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including by way of example semiconductor memory devices, such as erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and compact disc read only memory (CD ROM) and digital versatile disc read only memory (DVD-ROM) discs. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0600] Although this patent document contains many details, these should not be construed as limiting the scope of any subject matter or what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular technology. Certain features described in the context of separate embodiments in this patent document may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Additionally, although the above may describe features as acting in certain combinations and even initially so claimed, in some cases, one or more features from the claimed combination may be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.
[0601] Similarly, although operations are depicted in the figures in a particular order, this should not be construed as requiring that such operations be performed in the particular order shown or in a sequential order, or that all of the illustrated operations be performed to achieve the desired result. Additionally, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0602] Only a few specific embodiments and examples are described, and other specific embodiments, enhancements, and variations can be made based on what is described and illustrated in this patent document.
[0603] When there is no intermediate component between a first component and a second component other than a wire, trace, or another medium, the first component is directly coupled to the second component. When there is an intermediate component between the first component and the second component other than a wire, trace, or another medium, the first component is indirectly coupled to the second component. The term "coupled" and its variants include both direct coupling and indirect coupling. Unless otherwise stated, the use of the term "about" means a range that includes the subsequent number ±10%.
[0604] Although several embodiments are provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. The examples should be considered illustrative rather than restrictive, and are not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or some features may be omitted or not implemented.
[0605] Moreover, without departing from the scope of the present disclosure, the techniques, systems, subsystems, and methods described and illustrated as separate or discrete in various embodiments may be combined or integrated with other systems, modules, techniques, or methods. Other items shown or discussed as being coupled may be directly connected or may be indirectly coupled or communicate through some interface, device, or intermediate component, whether electrically, mechanically, or otherwise. Those skilled in the art can identify other examples of alterations, substitutions, and changes and can make them without departing from the spirit and scope disclosed herein.
Claims
1. A method for processing video data, comprising: Determining to apply multi-parameter local illumination compensation (MPLIC) to visual media data; And Performing a conversion between the visual media data and a bitstream based on the MPLIC.
2. The method according to claim 1, wherein The MPLIC is applied as follows: Y pred = α0Y0 + α1Y1 + α2Y2.. + α N-1 Y N-1 + β where Y pred is a prediction sample point, Y0, Y1, Y2, … Y N-1 are reference sample points derived based on at least one motion vector (MV), and α0, α1, α2, … αN-1, β are model parameters.
3. The method according to any one of claims 1-2, wherein Y0 is the reference pixel pointed to by MV for the current pixel, and Y1, Y2, … Y N-1 are neighboring pixels of Y0.
4. The method according to any one of claims 1 to 3, wherein Y0 is a reference sample for the current sample X0 derived based on at least one MV, and where Y i (i = 1, 2, …, N-1) is derived with a predefined offset relative to Y0.
5. The method according to any one of claims 1-4, wherein, When the current block uses uni-directional prediction, Y0 is a reference sample point derived based on one MV, where when the current block is bi-directional prediction and only one model is derived, Y0 is a reference sample point derived based on two MVs, or where when the current block is bi-directional prediction and two models are derived for two directions, Y0 is a reference sample point derived based on one MV.
6. The method according to any one of claims 1-5, wherein Y0 is a reference pixel obtained by a block vector or a template matching tool.
7. The method according to any one of claims 1-6, wherein MPLIC is used as an additional local illumination compensation (LIC) mode or in place of LIC.
8. The method according to any one of claims 1-7, wherein N and neighboring positions Y1, Y2, … Y n The selection depends on factors including coding and decoding information, the coding and decoding information including the block size of quantization parameter (QP) values, where different selections are made at the sequence level, picture level, slice level, tile level, coding tree unit (CTU) level, coding unit (CU) level, or a combination thereof.
9. The method according to any one of claims 1-8, wherein MPLIC is used for blocks adopting advanced motion vector prediction (AMVP), affine AMVP, or Merge mode or sub-block Merge mode.
10. The method according to any one of claims 1-9, wherein, MPLIC is used for certain sequences, frames, block sizes, quantization parameter (QP) values, or combinations thereof.
11. The method according to any one of claims 1 to 10, wherein, MPLIC is applied to all color components or a subset of color components, or where the selection of N and neighbors is different for different color components, or where MPLIC is only applied to the luminance component.
12. The method according to any one of claims 1-11, wherein, MPLIC is only applied to uni-directional prediction or for both uni-directional prediction and bi-directional prediction.
13. The method according to any one of claims 1-12, wherein, MPLIC is only used when the motion vector is not fractional precision, or where when the motion vector is integer precision, MPLIC completely replaces LIC, or where at least for the adaptive motion vector resolution (AMVR) mode, MPLIC is only applied when the MV precision is full pixel, or where when MPLIC is applied, the MV can only be signaled in at least full pixel form, or when MPLIC is applied, the MV is rounded to full pixel.
14. The method according to any one of claims 1-13, wherein, MPLIC includes two models, and where the first model processes samples along rows and the second model processes samples along columns, or where the number of models to be applied depends on the precision of the motion vector.
15. The method according to any one of claims 1-14, wherein When the current block uses the Merge mode, the use of MPLIC is inherited from neighboring blocks.
16. The method according to any one of claims 1-15, wherein For each Merge candidate X using a two-parameter LIC model, an additional Merge candidate is inserted into the MV candidate list, or where the additional Merge candidate inherits all information of the Merge candidate X and replaces the LIC model with a multi-parameter LIC model, or where only a subset of candidates in the MV candidate list is considered to use the multi-parameter LIC model to construct the additional Merge candidate, or where the block is encoded and decoded using regular Merge, motion vector difference (MMVD), affine Merge, sub-block Merge, template matching Merge, or combinations thereof, or where the MV candidate list is a Merge list, an affine Merge list, a template matching Merge list, or combinations thereof.
17. The method according to any one of claims 1-16, wherein, α0, α1, α2, … α n and β are model parameters derived by minimizing an error function based on a current block template and a reference block template, and no signaling overhead is used other than signaling an MPLIC flag for the AMVP or affine AMVP mode to indicate the use of MPLIC.
18. The method according to any one of claims 1-17, wherein, The error function is: where X0 is a sample point in the current template, Y0 is the reference sample point of X0, Y i (i = 1, 2, …, N - 1) has a predefined offset relative to Y0, ||T|| is the number of template sample points, λ n is the regularization parameter, and o n is a predefined value or o0 = 1, o i (i = 1, 2, …, N - 1) = 0.
19. The method according to any one of claims 1-18, wherein, The error function is: where X0 is a sample point in the current template, Y0 is the reference sample point of X0, and Y i (i = 1, 2, …, N-1) has a predefined offset relative to Y0, and ||T|| is the number of template sample points.
20. The method according to any one of claims 1-19, wherein DC (Direct Current) prediction is adopted, and for each sample point in the current block, a mapping function is used to derive the corresponding virtual sample point to generate a virtual block.
21. The method according to any one of claims 1-20, wherein, The mapping function is formed between the pixels of the reference block and the reference template, where the mapping function is used to map the sample point value in the current block to the sample point value in the current template adjacent to the current block, or where the DC of the virtual block is regarded as the DC prediction of the current block.
22. The method according to any one of claims 1-21, wherein, The mapping function is constructed based on the criterion between the sample points in the reference block and the sample points in the reference template, or when more than one reference template sample point generates the same minimum cost for the reference block sample point, the sample point closest in spatial domain distance is selected.
23. The method according to any one of claims 1-22, wherein, The use of the mapping function depends on the codec information including the block size or quantization parameter (QP) size.
24. The method according to any one of claims 1-23, wherein, MPLIC (Motion Picture Lossless Intra Coding) adopts reference sample points classified into different subsets, and different local illumination compensation (LIC) models are applied to each subset.
25. The method according to any one of claims 1-24, wherein, The LIC model for the subset is a two-parameter linear model or a multi-parameter linear model, where the classification of the reference sample points is based on a predefined threshold, where the classification is based on clustering, where the threshold is explicitly signaled, where additional flexibility in choosing between an implicit threshold and an explicitly communicated threshold is adopted, where when used in conjunction with a geometric partitioning mode (GPM), the classification of the sample points is based on the partitioning of the template obtained by GPM partitioning, or where the use of MPLIC depends on the codec information including the block size QP size.
26. The method according to any one of claims 1-25, wherein, The bitstream includes one or more syntax elements indicating whether a prediction mode is selected for the current block.
27. The method according to any one of claims 1-26, wherein, The syntax element depends on: whether the current block is coded by AMVP (Advanced Motion Vector Prediction), whether the current block is coded by Merge, whether the current block is unidirectional predicted, whether the current block is bidirectional predicted, type Merge prediction or their combination, or when the current block is bidirectional predicted, it is presumed that the syntax element is false.
28. The method according to any one of claims 1-27, wherein, The syntax element is conditionally signaled, the syntax element is signaled only for a subset of the Y, U, and V planes, the syntax element is signaled depending on the dimension of the current block, the syntax element is signaled depending on the quantization parameter (QP) used to code the current block, the syntax element is context-coded, the syntax element is binarized into a fixed-length code, a truncated unary code, an exponential Golomb code, the syntax element is coded with at least one context in arithmetic coding including bypass coding, or their combination.
29. The method according to any one of claims 1-28, wherein, The syntax element is signaled only when the current block is larger than a predefined size, or where the syntax element is signaled only when the sum of the width and height of the current block is greater than, less than, not greater than, or not less than a predefined threshold, or where the syntax element is signaled only when the QP of the current block is greater than, less than, not greater than, or not less than a predefined threshold, or their combination.
30. The method according to any one of claims 1-29, wherein, The use of the method is signaled at the sequence level, picture group level, picture level, slice level, or tile group level, including in the sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header, or tile group header, or where the use of the method is signaled at a prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline decoding unit (VPDU), codec tree unit (CTU), CTU row, slice, tile, sub-picture, or other region containing more than one sample or pixel.
31. The method according to any one of claims 1-30, wherein, The application of the method depends on codec information including block size, color format, single-tree segmentation, dual-tree segmentation, color component, slice type, or picture type.
32. The method according to any one of claims 1 to 31, wherein, The conversion includes encoding the visual media data into the bitstream.
33. The method according to any one of claims 1 - 31, wherein, The conversion includes decoding the visual media data from the bitstream.
34. An apparatus for processing video data, comprising: A processor; and a non-transitory memory having instructions thereon, where the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 33.
35. A non-transitory computer-readable medium comprising a computer program product for use in a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium such that, when executed by a processor, cause the video codec device to perform the method according to any one of claims 1 to 33.
36. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method executed by a video processing device, where the method comprises: determining to apply multi-parameter local illumination compensation (MPLIC) to visual media data; and generating the bitstream based on the determination.
37. A method for storing a bitstream of a video, comprising: determining to apply multi-parameter local illumination compensation (MPLIC) to visual media data; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.