Using virtual candidate prediction and weighted prediction in video processing
By introducing local illumination compensation and generalized bidirectional prediction technology into video encoding and decoding, the encoding and decoding process of video blocks is optimized, the shortcomings of compression rate and parallel implementation in existing technologies are solved, and more efficient video encoding and decoding is achieved.
Patent Information
- Application Number
- CN202080009266.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-01-17
- Filing Date
- 2020-01-17
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2040-01-17
AI Technical Summary
Existing video coding and decoding technologies have deficiencies in compression rate and parallel implementation, and need to be improved to provide better coding and decoding solutions.
The local illumination compensation (LIC) technology is introduced to optimize the encoding and decoding process of video blocks by constructing a motion candidate list and applying a weight set for generalized bidirectional prediction.
It improves the compression rate of video encoding and decoding, reduces the complexity of encoding and decoding, and supports more efficient parallel processing.
Smart Images

Figure CN113302919B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is intended to timely claim, under applicable patent law and / or the rules of the Paris Convention, International Patent Application No. PCT / CN2019 / 072154, filed on January 17, 2019. The entire disclosure of the aforementioned application is incorporated by reference as a part of the disclosure of this patent. Technical Field
[0003] This patent document relates to the field of video encoding and decoding. Background Art
[0004] Currently, efforts are underway to improve the performance of current video codec technologies to provide better compression rates or to provide video encoding and decoding schemes that allow for lower complexity or parallelized implementations. Several new video codec tools have recently been proposed by industry experts and are currently being tested to determine their effectiveness. Summary of the Invention
[0005] This document provides techniques for incorporating local illumination compensation in embodiments of a video encoder or decoder.
[0006] In one example aspect, a method of video processing is disclosed. The method includes constructing a motion candidate list having at least one motion candidate for conversion between a current block of video and a bitstream representation of the video; and determining whether a generalized bi-prediction (GBI) processing tool is enabled based on a candidate type of a first motion candidate. The GBI processing tool includes deriving a final prediction based on applying equal or unequal weights to predictions derived from different reference lists according to a weight set. The method also includes performing the conversion based on the determination.
[0007] In another example aspect, a method of video processing is disclosed. The method includes determining an operational state of local illumination compensation (LIC) prediction, generalized bidirectional (GBI) prediction, or weighted prediction for a conversion between a current block of video and a bitstream representation of the video. The current block is associated with a prediction technique that uses at least one virtual motion candidate. The method also includes performing the conversion based on the determination.
[0008] In another example aspect, a method of video processing is disclosed. The method includes, for a conversion between a current block of video and a bitstream representation of the video, determining a state of a filtering operation or one or more parameters of the filtering operation for the current block based on characteristics of local illumination compensation (LIC) prediction, generalized bidirectional (GBI) prediction, weighted prediction, or affine motion prediction. The method also includes performing the conversion based on the determination.
[0009] In another example aspect, a method of video processing is disclosed. The method includes, in converting between a video block and a bitstream representation of the video block, determining that the video block is a boundary block of a coding tree unit (CTU) in which the video block is located and therefore enabling a local illumination compensation (LIC) codec for the video block, deriving parameters for local illumination compensation (LIC) for the video block based on the determination that the LIC codec is enabled for the video block, and performing the conversion by adjusting pixel values of the video block using the LIC.
[0010] In another example aspect, a method of video processing is disclosed. The method includes, in converting between a video block and a bitstream representation of the video block, determining that the video block is an internal block of a codec tree unit (CTU) in which the video block is located and therefore disabling a local illumination compensation (LIC) codec tool for the video block, inheriting parameters of the LIC of the video block, and performing the conversion by adjusting pixel values of the video block using the LIC.
[0011] In yet another example aspect, another method of video processing is disclosed. The method includes, in converting between a video block and a bitstream representation of the video block, determining that both local illumination compensation and intra block copy codecs are enabled for use with a current block, and performing the conversion by performing local illumination compensation (LIC) and intra block copy operations on the video block.
[0012] In yet another example aspect, another method of video processing is disclosed, comprising performing conversion by performing local illumination compensation (LIC) and intra block copy operations on a video block, and performing conversion between a current block and a corresponding bitstream representation of the current block.
[0013] In yet another example aspect, another method of video processing is disclosed. The method includes determining local illumination compensation (LIC) parameters for the video block using at least some samples of neighboring blocks of the video block during conversion between a video block of video and a bitstream representation of the video, and performing conversion between the video block and the bitstream representation by performing the LIC using the determined parameters.
[0014] In yet another example aspect, another method of video processing is disclosed, the method including performing conversion between a video and a bitstream representation of the video, wherein the video is represented as a video frame including video blocks, and local illumination compensation (LIC) is enabled only for the video blocks using a geometric prediction structure including a triangle prediction mode.
[0015] In yet another example aspect, another method of video processing is disclosed. The method includes performing a conversion between a video and a bitstream representation of the video, wherein the video is represented as a video frame including video blocks, and in the conversion to its corresponding bitstream representation, performing local illumination compensation (LIC) on less than all pixels of a current block.
[0016] In yet another example aspect, another method of video processing is disclosed. The method includes determining, in converting between a video block and a bitstream representation of the video block, that both local illumination compensation (LIC) and generalized bi-prediction (GBi) or multi-hypothesis inter prediction codecs are enabled for use with a current block, and performing the conversion by performing LIC and GBi or multi-hypothesis inter prediction operations on the video block.
[0017] In yet another example aspect, another method of video processing is disclosed. The method includes determining, in converting between a video block and a bitstream representation of the video block, that both local illumination compensation (LIC) and combined inter-intra prediction (CIIP) codecs are enabled for use with a current block, and performing the conversion by performing LIC and CIIP operations on the video block.
[0018] In yet another example aspect, another method of video processing is disclosed, the method comprising determining that a video block is associated with a prediction technique comprising pairwise prediction or combined bi-prediction; determining an operational state of local illumination compensation (LIC), generalized bi-prediction (GBI), or weighted prediction, enabling or disabling the operational state based on the determination that the video block is associated with the prediction technique; and performing further processing on the video block according to the operational state of LIC, GBI, or weighted prediction.
[0019] In yet another example aspect, another method of video processing is disclosed, comprising determining a characteristic related to local illumination compensation (LIC), generalized bidirectional prediction (GBI), or weighted prediction; and applying a deblocking filter to a video block based on the determination of the characteristic.
[0020] In yet another example aspect, another method of video processing is disclosed, comprising determining a characteristic regarding an affine mode, affine parameters, or affine type associated with a video block; and applying a deblocking filter to the video block based on the determination of the characteristic.
[0021] In yet another example aspect, another video processing method is disclosed. The method includes storing data associated with a combined inter-intra prediction (CIIP) flag in a history-based motion vector prediction (HMVP) table; storing motion information in the HMVP table; and performing further processing on a video block using the data stored in the HMVP table and the motion information stored in the HMVP table.
[0022] In yet another example aspect, another method of video processing is disclosed. The method includes storing data associated with a local illumination compensation (LIC) flag in a history-based motion vector prediction (HMVP) table; storing motion information in the HMVP table; and performing further processing on a video block using the data stored in the HMVP table and the motion information stored in the HMVP table.
[0023] In yet another example aspect, another method of video processing is disclosed, including determining to perform local illumination compensation (LIC); dividing a video block into video processing data units (VPDUs) based on the determination to perform LIC; and performing LIC on the VPDUs sequentially or in parallel.
[0024] In yet another representative aspect, the various techniques described herein may be embodied as a computer program product stored on a non-transitory computer-readable medium, including program code for executing the methods described herein.
[0025] In yet another representative aspect, a video decoder device may implement the method as described herein.
[0026] The details of one or more embodiments are set forth in the accompanying drawings, the accompanying figures, and the following description. Other features will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 An example of the derivation process of Merge candidate list construction is shown.
[0028] Figure 2 Example locations of spatial merge candidates are shown.
[0029] Figure 3 An example of candidate pairs considered for redundancy checking of spatial merge candidates is shown.
[0030] Figure 4Example locations of the second PU (Prediction Unit) for N×2N and 2N×N partitions are shown.
[0031] Figure 5 It is a diagram of the motion vector scaling of the time domain Merge candidate.
[0032] Figure 6 An example of candidate positions of the time-domain merge candidates C0 and C1 is shown.
[0033] Figure 7 An example of a combined bi-predictive Merge candidate is shown.
[0034] Figure 8 An example of a process of deriving motion vector prediction candidates is shown.
[0035] Figure 9 is an example illustration of motion vector scaling of spatial motion vector candidates.
[0036] Figure 10 An example of an advanced temporal motion vector predictor (ATMVP) for a coding unit (CU) is shown.
[0037] Figure 11 An example of one CU having four sub-blocks (AD) and its neighboring blocks (ad) is shown.
[0038] Figure 12 An example of a planar motion vector prediction process is shown.
[0039] Figure 13 This is a flowchart of an example of encoding with different motion vector (MV) precisions.
[0040] Figure 14 is an example illustration of a sub-block in which OBMC is applied.
[0041] Figure 15 An example of adjacent sample points used to derive IC parameters is shown.
[0042] Figure 16 It is a diagram of dividing the coding unit (CU) into two triangular prediction units.
[0043] Figure 17 An example of the positions of neighboring blocks is shown.
[0044] Figure 18 An example is shown in which the CU applies the first weighting factor group.
[0045] Figure 19 An example of a motion vector storage implementation is shown.
[0046] Figure 20 An example of a simplified affine motion model is shown.
[0047] Figure 21 An example of an affine MVF (Motion Vector Field) for each sub-block is shown.
[0048] Figure 22 Examples of a 4-parameter affine model and a 6-parameter affine model are shown.
[0049] Figure 23 An example of a motion vector predictor (MVP) for the AF_INTER mode is shown.
[0050] Figures 24A-24B Examples of candidates for the AF_MERGE mode are shown.
[0051] Figure 25 Candidate positions of the affine merge mode are shown.
[0052] Figure 26 An example process of bilateral matching is shown.
[0053] Figure 27 An example process of template matching is shown.
[0054] Figure 28 An implementation of unilateral motion estimation (ME) in frame rate upconversion (FRUC) is shown.
[0055] Figure 29 An embodiment of the Ultimate Motion Vector Expression (UMVE) search process is shown.
[0056] Figure 30 An example of a UMVE search point is shown.
[0057] Figure 31 An example of a distance index and distance offset mapping is shown.
[0058] Figure 32 An example of an optical flow trajectory is shown.
[0059] Figure 33A-Figure 33BExamples of Bi-directional Optical flow (BIO) without block extension are shown: a) access locations outside the block; b) padding is used to avoid extra memory access and computation.
[0060] Figure 34 An example of using Decoder-side Motion Vector Refinement (DMVR) based on bilateral template matching is shown.
[0061] Figure 35 An example of adjacent samples used in a bilateral filter is shown.
[0062] Figure 36 An example of a window covering two samples used in the weight calculation is shown.
[0063] Figure 37 An example of a decoding flow using the proposed history-based motion vector prediction (HMVP) method is shown.
[0064] Figure 38 An example of updating a table in the proposed HMVP method is shown.
[0065] Figure 39 is a block diagram of a hardware platform used to implement the video encoding or decoding techniques described in this document.
[0066] Figure 40 An example of a hardware platform for implementing the methods and techniques described in this document is shown.
[0067] Figure 41 is a flowchart of an example method for video processing.
[0068] Figure 42 is a flow chart of an example method of video processing according to the present disclosure.
[0069] Figure 43 is a flow chart of another example method of video processing according to the present disclosure.
[0070] Figure 44 is a flowchart of yet another example method of video processing according to the present disclosure. DETAILED DESCRIPTION
[0071] This document provides several technologies that can be embodied in digital video encoders and decoders. For clear understanding, section headings are used in this document and do not limit the scope of the technologies and embodiments disclosed in each section to only that section.
[0072] 1. Overview
[0073] This patent document relates to video codec technology. Specifically, it relates to local illumination compensation (LIC) in video codecs. The patent application can be applied to existing video codec standards, such as HEVC, or to upcoming standards, such as the Versatile Video Codec (VVC). The patent application can also be applied to future video codec standards or codecs.
[0074] 2. Examples of video encoding / decoding technologies
[0075] Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture that uses temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the VVC standard with the goal of a 50% bitrate reduction compared to HEVC.
[0076] 2.1. Inter-frame prediction in HEVC / H.265
[0077] Each inter-predicted PU has motion parameters for one or two reference picture lists. The motion parameters include motion vectors and reference picture indices. The use of one of the two reference picture lists can also be signaled using inter_pred_idc. The motion vector can be explicitly encoded as a delta relative to the prediction value.
[0078] When a CU is coded in skip mode, a PU is associated with the CU and there are no significant residual coefficients, no coded motion vector deviations or reference picture indices. By specifying the Merge mode, the motion parameters of the current PU are obtained from neighboring PUs including spatial and temporal candidates. Merge mode can be applied to any inter-predicted PU, not just for skip mode. An alternative to Merge mode is explicit transmission of motion parameters, where the motion vector (more accurately, the motion vector difference compared to the motion vector prediction value), the corresponding reference picture index for each reference picture list and the use of the reference picture list are explicitly signaled per PU. In this disclosure, this mode is referred to as Advanced Motion Vector Prediction (AMVP).
[0079] When signaling indicates that one of the two reference picture lists is to be used, a PU is generated from a block of samples. This is called "unidirectional prediction." Unidirectional prediction can be used for both P slices and B slices.
[0080] When signaling indicates that two reference picture lists are to be used, a PU is generated from two sample blocks. This is called "bi-prediction." Bi-prediction is only available for B slices.
[0081] The following text provides detailed information about the inter prediction modes specified in HEVC. The description starts with the Merge mode.
[0082] 2.1.1.Merge mode
[0083] 2.1.1.1. Merge Mode Candidate Derivation
[0084] When predicting a PU using Merge mode, an index pointing to an entry in the Merge candidate list is parsed from the bitstream and used to retrieve motion information. The construction of this list is specified in the HEVC standard and can be summarized according to the following sequence of steps:
[0085] Step 1: Initial candidate derivation
[0086] Step 1.1: Spatial Candidate Derivation
[0087] Step 1.2: Redundancy check of spatial candidates
[0088] Step 1.3: Time Domain Candidate Derivation
[0089] Step 2: Additional candidate insertions
[0090] Step 2.1: Create bidirectional prediction candidates
[0091] Step 2.2: Insert zero motion candidates
[0092] These steps are also schematically depicted in Figure 1 For spatial Merge candidate derivation, a maximum of four Merge candidates are selected from the candidates located at five different positions. For temporal Merge candidate derivation, a maximum of one Merge candidate is selected from the two candidates. Since a constant number of candidates for each PU is assumed at the decoder, additional candidates are generated when the number of candidates obtained from step 1 does not reach the maximum number of Merge candidates (MaxNumMergeCand) signaled in the slice header. Since the number of candidates is constant, the index of the best Merge candidate is encoded using truncated unary binarization (TU). If the size of the CU is equal to 8, all PUs of the current CU share a single Merge candidate list, which is the same as the Merge candidate list of the 2N×2N prediction unit.
[0093] Hereinafter, operations associated with the aforementioned steps will be described in detail.
[0094] 2.1.1.2. Spatial Candidate Derivation
[0095] In the derivation of spatial Merge candidates, Figure 2 Up to four Merge candidates are selected from the candidates at the positions depicted in . The order of derivation is A1, B1, B0, A0, and B2. Position B2 is considered only when any PU at position A1, B1, B0, A0 is unavailable (for example, because it belongs to another strip or slice) or is intra-coded. After adding the candidate at position A1, a redundancy check is performed on the addition of the remaining candidates, which ensures that candidates with the same motion information are excluded from the list, thereby improving coding efficiency. Figure 3 An example of candidate pairs considered for redundancy check of spatial merge candidates is shown. In order to reduce computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the candidate pairs with Figure 3 The arrows in the list link pairs, and candidates are added to the list only if the corresponding candidates used for redundancy checking do not have the same motion information. Another source of duplicate motion information is a "second PU" associated with a partition other than 2N×2N. As an example, Figure 4 The second PU is depicted in the N×2N and 2N×N cases, respectively. When the current PU is partitioned into N×2N, the candidate at position A1 is not considered for list construction. In fact, adding this candidate would result in two prediction units having the same motion information, which is redundant for having only one PU in the codec unit. Similarly, position B1 is not considered when the current PU is partitioned into 2N×N.
[0096] 2.1.1.3. Time Domain Candidate Derivation
[0097] In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal merge candidate, the scaled motion vector is derived based on the co-located PU belonging to the picture with the smallest POC (Picture Order Count) difference with the current picture in a given reference picture list. The reference picture list to be used for derivation of the co-located PU is explicitly signaled in the slice header. Figure 5 It is a diagram of the motion vector scaling of the time domain Merge candidate. Figure 5 As shown by the dotted line in the figure, the scaled motion vector of the temporal Merge candidate is obtained, which is scaled from the motion vector of the collocated PU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the collocated picture and the collocated picture. The reference picture index of the temporal Merge candidate is set equal to zero. The actual implementation of the scaling process is described in the HEVC specification. For B slices, two motion vectors (one for reference picture list 0 and the other for reference picture list 1) are obtained and combined to generate a bidirectional prediction Merge candidate.
[0098] In the collocated PU(Y) belonging to the reference frame, the position of the temporal candidate is selected between candidates C0 and C1, such as Figure 6 If the PU at position C0 is not available, is intra-coded, or is outside the current CTU row, position C1 is used. Otherwise, position C0 is used in the derivation of the time-domain merge candidate.
[0099] 2.1.1.4. Additional candidate insertion
[0100] In addition to spatial and temporal Merge candidates, there are two additional types of Merge candidates: combined bi-directional prediction Merge candidate and zero Merge candidate. Combined bi-directional prediction Merge candidate is generated by utilizing spatial and temporal Merge candidates. Combined bi-directional prediction Merge candidate is only used for B slices. Combined bi-directional prediction candidate is generated by combining the first reference picture list motion parameters of an initial candidate with the second reference picture list motion parameters of another initial candidate. If these two tuples provide different motion hypotheses, they will form a new bi-directional prediction candidate. As an example, Figure 7Shown is the case when two candidates in the original list (on the left) (with mvL0 and refIdxL0 or mvL1 and refIdxL1) are used to create a combined bi-predictive Merge candidate that is added to the final list (on the right). There are many rules for combining that are considered to generate these additional Merge candidates.
[0101] Zero-motion candidates are inserted to fill the remaining entries in the Merge candidate list and thus reach the MaxNumMergeCand capacity. These candidates have zero spatial displacement and a reference picture index that starts at zero and increases each time a new zero-motion candidate is added to the list. The number of reference frames used by these candidates is 1 and 2 for unidirectional and bidirectional prediction, respectively. Finally, no redundancy check is performed on these candidates.
[0102] 2.1.1.5. Motion Estimation Regions for Parallel Processing
[0103] To speed up the encoding process, motion estimation can be performed in parallel, thereby deriving motion vectors for all prediction units within a given region at the same time. Deriving Merge candidates from spatial neighbors may interfere with parallel processing because a prediction unit cannot derive motion parameters from neighboring PUs until its associated motion estimation is completed. To ease the trade-off between coding efficiency and processing latency, HEVC defines a Motion Estimation Region (MER), and uses the "log2_parallel_merge_level_minus2" syntax element to signal the size of the MER in the picture parameter set. When defining a MER, Merge candidates that fall into the same region are marked as unavailable and are therefore not considered in the list construction.
[0104] 2.1.2.AMVP
[0105] AMVP exploits the spatio-temporal correlation of motion vectors with neighboring PUs, which is used for explicit transmission of motion parameters. For each reference picture list, a motion vector candidate list is constructed by first checking the availability of neighboring PU positions in the left and upper time domains, removing redundant candidates, and adding zero vectors to make the candidate list a constant length. The encoder can then select the best prediction value from the candidate list and send the corresponding index indicating the selected candidate. Similar to Merge index signaling, the index of the best motion vector candidate is encoded using truncated unary. In this case, the maximum value to be encoded is 2 (see Figure 8 ). In the following sections, details about the derivation process of motion vector prediction candidates will be provided.
[0106] 2.1.2.1. Derivation of AMVP Candidates
[0107] In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. For spatial motion vector candidate derivation, two motion vector candidates are finally derived based on the motion vector of each PU located at five different positions, such as Figure 8 Depicted.
[0108] For temporal motion vector candidate derivation, a motion vector candidate is selected from two candidates derived based on two different collocated positions. After generating the first list of first spatio-temporal candidates, duplicate motion vector candidates in the list are removed. If the number of possible candidates is greater than two, motion vector candidates whose reference picture index within the associated reference picture list is greater than 1 are removed from the list. If the number of spatio-temporal motion vector candidates is less than two, an additional zero motion vector candidate is added to the list.
[0109] 2.1.2.2. Spatial Motion Vector Candidates
[0110] In the derivation of spatial motion vector candidates, Figure 2 Consider up to two candidates out of the five possible candidates for PU derivation at the depicted position, and those positions are the same as the positions of the motion merge. The derivation order for the left side of the current PU is defined as A0, A1, and scaled A0, scaled A1. The derivation order for the upper side of the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. Therefore, for each side, there are four cases that can be used as motion vector candidates, two of which do not require spatial scaling and two use spatial scaling. The four different cases are summarized as follows:
[0111] No airspace scaling
[0112] -(1) Same reference picture list and same reference picture index (same POC)
[0113] -(2) Different reference picture lists, but same reference pictures (same POC)
[0114] Airspace scaling
[0115] -(3) Same reference picture list, but different reference pictures (different POC)
[0116] -(4) Different reference picture lists, and different reference pictures (different POCs)
[0117] First, the case without spatial scaling is checked, followed by spatial scaling. When the POC differs between the reference pictures of the neighboring PU and the reference picture of the current PU, spatial scaling is considered regardless of the reference picture list. If all PUs of the left candidate are unavailable or intra-coded, scaling of the upper motion vector is allowed to facilitate the parallel derivation of the left and upper MV candidates. Otherwise, spatial scaling is not allowed for the upper motion vector.
[0118] like Figure 9 As depicted, in the spatial scaling process, the motion vectors of neighboring PUs are scaled in a similar manner to temporal scaling. The main difference is that the reference picture list and the index of the current PU are given as input; the actual scaling process is the same as the temporal scaling process.
[0119] 2.1.2.3. Temporal Motion Vector Candidates
[0120] Except for the reference picture index derivation, all the processes for deriving the temporal Merge candidate are the same as those for deriving the spatial motion vector candidate (see Figure 6 ). The reference picture index is signaled to the decoder.
[0121] 2.2. New inter-frame prediction method in JEM
[0122] 2.2.1 Motion Vector Prediction Based on Sub-CU
[0123] In JEM with QTBT (QuadTrees plus Binary Trees), each CU can have at most one motion parameter set for each prediction direction. By dividing the large CU into sub-CUs and deriving the motion information of all sub-CUs of the large CU, two sub-CU level motion vector prediction methods are considered in the encoder. The optional temporal motion vector prediction (ATMVP) method allows each CU to extract multiple motion information sets from multiple blocks that are smaller than the current CU in the collocated reference picture. In the Spatial-Temporal Motion Vector Prediction (STMVP) method, the motion vector of the sub-CU is recursively derived by using the temporal motion vector predictor and the spatial neighboring motion vectors.
[0124] In order to maintain a more accurate motion field for sub-CU motion prediction, motion compression of reference frames is currently disabled.
[0125] 2.2.1.1. Optional temporal motion vector prediction
[0126] In the optional temporal motion vector prediction (ATMVP) method, the temporal motion vector prediction (TMVP) is modified by extracting multiple sets of motion information (including motion vectors and reference indices) from blocks smaller than the current CU. Figure 10 As shown, a sub-CU is a square N×N block (N is set to 4 by default).
[0127] ATMVP predicts the motion vectors of sub-CUs within a CU in two steps. The first step is to use the so-called time domain vector to identify the corresponding block in the reference picture. The reference picture is also called the motion source picture. The second step is to divide the current CU into sub-CUs and obtain the motion vector and reference index of each sub-CU from the block corresponding to each sub-CU, such as Figure 10 shown.
[0128] In the first step, the reference picture and the corresponding block are determined by the motion information of the spatially adjacent blocks of the current CU. To avoid repeated scanning of adjacent blocks, the first merge candidate in the merge candidate list of the current CU is used. The first available motion vector and its associated reference index are set to the temporal vector and the index of the motion source picture. In this way, in ATMVP, the corresponding block can be identified more accurately than in TMVP, where the corresponding block (sometimes called a collocated block) is always in the lower right or center position relative to the current CU.
[0129] In the second step, the corresponding blocks of the sub-CU are identified by the time domain vector in the motion source picture by adding the time domain vector to the coordinates of the current CU. For each sub-CU, the motion information of its corresponding block (the minimum motion grid covering the center sample point) is used to derive the motion information of the sub-CU. After identifying the motion information of the corresponding N×N block, it is converted into the motion vector and reference index of the current sub-CU in the same way as the TMVP of HEVC, where motion scaling and other processes apply. For example, the decoder checks whether the low latency condition is met (that is, the POC of all reference pictures of the current picture is less than the POC of the current picture), and may use the motion vector MV x (motion vector corresponding to reference picture list X) to predict the motion vector MV of each sub-CU y (For example, where X equals 0 or 1, and Y equals 1-X).
[0130] 2.2.1.2. Spatial-Temporal Motion Vector Prediction (STMVP)
[0131] In this method, the motion vector of a sub-CU is recursively derived in raster scan order. Figure 11Let us consider an 8×8 CU, which contains four 4×4 sub-CUs: A, B, C, and D. The neighboring 4×4 blocks in the current frame are labeled a, b, c, and d.
[0132] The motion derivation of sub-CU A starts by identifying its two spatial neighbors. The first neighbor is the N×N block (block c) on the upper side of sub-CU A. If this block c is not available or is intra-coded, check the other N×N blocks on the upper side of sub-CU A (from left to right, starting from block c). The second neighbor is the block (block b) on the left side of sub-CU A. If block b is not available or is intra-coded, check the other blocks on the left side of sub-CU A (from top to bottom, starting from block b). The motion information obtained from the adjacent blocks of each list is scaled to the first reference frame of the given list. Next, the temporal motion vector prediction value (TMVP) of sub-block A is derived by following the same process of TMVP derivation as specified by HEVC. The motion information of the collocated block at position D is extracted and scaled accordingly. Finally, after retrieving and scaling the motion information, all available motion vectors (up to 3) are averaged separately for each reference list. The average motion vector is assigned as the motion vector of the current sub-CU.
[0133] 2.2.1.3. Sub-CU Motion Prediction Mode Signaling
[0134] Sub-CU mode is enabled as an additional Merge candidate, and no additional syntax elements are required to signal these modes. Two additional Merge candidates are added to the Merge candidate list of each CU to represent ATMVP mode and STMVP mode. If the sequence parameter set indicates that ATMVP and STMVP are enabled, up to seven Merge candidates can be used. The encoding and decoding logic of the additional Merge candidates is the same as the encoding logic of the Merge candidates in HM, which means that for each CU in a P slice or B slice, two RD checks may be required for the two additional Merge candidates.
[0135] In JEM, all bins of the Merge index are context-coded by CABAC, whereas in HEVC, only the first bin is context-coded, and the remaining bins are context-bypass coded.
[0136] 2.2.2. Pairwise Average Candidates
[0137] Pairwise average candidates are generated by averaging predefined candidate pairs in the current Merge candidate list, and the predefined pairs are defined as {(0,1),(0,2),(1,2),(0,3),(1,3),(2,3)}, where these numbers represent the Merge index of the Merge candidate list. The average motion vector is calculated separately for each reference list. If two motion vectors are available in one list, they are averaged even if they point to different reference pictures; if only one motion vector is available, it is used directly; if no motion vector is available, the list remains invalid. Pairwise average candidates replace the combined candidates in the HEVC standard.
[0138] The complexity analysis of the pairwise averaging candidates is summarized in Table 1. For the worst case of additional computation for averaging (last column in Table 1), each pair requires 4 additions and 4 shifts (MVx and MVy in L0 and L1), and each pair requires 4 reference index comparisons (refIdx0 is valid in L0 and L1, and refIdx1 is valid). There are 6 pairs, for a total of 24 additions, 24 shifts, and 24 reference index comparisons. The combined candidate pairs in the HEVC standard use 2 reference index comparisons per pair (refIdx0 is valid in L0, and refIdx1 is valid in L1), for a total of 12 pairs, for a total of 24 reference index comparisons.
[0139] Table 1: Operational analysis of pairwise average candidates
[0140]
[0141] 2.2.3. Planar Motion Vector Prediction
[0142] In JVET-K0135, planar motion vector prediction is proposed.
[0143] In order to generate a smooth fine-grained motion field, Figure 12 A brief description of the planar motion vector prediction process is given.
[0144] Planar motion vector prediction is achieved by averaging horizontal and vertical linear interpolations on a 4×4 block basis, as shown below.
[0145] P(x,y)=(H×P h (x,y)+W×P v (x,y)+H×W) / (2×H×W) Equation (1)
[0146] W and H represent the width and height of the block. (x, y) are the coordinates of the current sub-block relative to the top-left sub-block. All distances are expressed as pixel distances divided by 4. P(x, y) is the motion vector of the current sub-block.
[0147] The horizontal prediction P at position (x,y) h (x,y) and vertical prediction P v (x,y) is calculated as follows:
[0148] P h (x,y)=(W-1-x)×L(-1,y)+(x+1)×R(W,y) Equation (2)
[0149] P v (x,y)=(H-1-y)×A(x,-1)+(y+1)×B(x,H) Equation (3)
[0150] Where L(-1, y) and R(W, y) are the motion vectors of the 4×4 blocks to the left and right of the current block. A(x, -1) and B(x, H) are the motion vectors of the 4×4 blocks above and below the current block.
[0151] The reference motion information of the left column and the upper row neighboring blocks are derived from the spatial domain neighboring blocks of the current block.
[0152] The reference motion information of the neighboring blocks in the right column and bottom row is derived as follows.
[0153] (1) Derivation of motion information of the adjacent 4×4 block in the lower right temporal domain
[0154] (2) Using the derived motion information of the lower right adjacent 4×4 block and the motion information of the upper right adjacent 4×4 block, the motion vector of the right column adjacent 4×4 block is calculated as described in the following equation (4).
[0155] (3) Using the derived motion information of the lower right adjacent 4×4 block and the motion information of the lower left adjacent 4×4 block, the motion vector of the bottom row adjacent 4×4 block is calculated as described in equation (5).
[0156] R(W,y)=((Hy-1)×AR+(y+1)×BR) / H Equation (4)
[0157] B(x,H)=((Wx-1)×BL+(x+1)×BR) / W Equation (5)
[0158] Where AR is the motion vector of the upper right spatially adjacent 4×4 block, BR is the motion vector of the lower right temporally adjacent 4×4 block, and BL is the motion vector of the lower left spatially adjacent 4×4 block.
[0159] The motion information obtained from neighboring blocks of each list is scaled to the first reference picture of the given list.
[0160] 2.2.4 Adaptive Motion Vector Difference Resolution
[0161] In HEVC, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the PU's motion vector and the predicted motion vector) is signaled in units of quarter luma samples. In JEM, Locally Adaptive Motion Vector Resolution (LAMVR) is introduced. In JEM, MVD can be encoded and decoded in units of quarter luma samples, integer luma samples, or four luma samples. The MVD resolution is controlled at the codec unit (CU) level, and the MVD resolution flag is conditionally signaled for each CU with at least one non-zero MVD component.
[0162] For a CU with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter luma sample MV precision is used in the CU. When the first flag (equal to 1) indicates that quarter luma sample MV precision is not used, another flag is signaled to indicate whether integer luma sample MV precision or quad luma sample MV precision is used.
[0163] When the first MVD resolution flag of a CU is zero or is not coded for the CU (meaning all MVDs in the CU are zero), a quarter luma sample MV resolution is used for the CU. When the CU uses integer luma sample MV precision or four luma sample MV precision, the MVP in the CU's AMVP candidate list is rounded to the corresponding precision.
[0164] In the encoder, CU-level RD check is used to determine which MVD resolution to use for the CU. That is, for each MVD resolution, three CU-level RD checks are performed. In order to speed up the encoder, the following encoding scheme is applied in JEM:
[0165] During RD check of a CU with normal quarter luma sample MVD resolution, the motion information of the current CU (integer luma sample precision) is stored. The stored motion information (after rounding) is used as the starting point for further small-scale motion vector refinement during RD check of the same CU with integer luma sample and 4 luma sample MVD resolution, so that the time-consuming motion estimation process is not repeated three times.
[0166] · Conditionally call the RD check for the CU with 4-luma sample MVD resolution. For a CU, when the RD cost of the integer luma sample MVD resolution is much greater than the RD cost of the quarter-luma sample MVD resolution, skip the RD check for the 4-luma sample MVD resolution of the CU.
[0167] The encoding process is as Figure 13 shown. First, test the 1 / 4 pixel MV, calculate the RD cost and denote it as RDCost0, then test the integer MV, and the RD cost is denoted as RDCost1. If RDCost1 < th * RDCost0 (where th is a positive value), then test the 4-pixel MV; otherwise, skip the 4-pixel MV. Basically, when checking the integer or 4-pixel MV, the motion information and RD cost of the 1 / 4 pixel MV are already known, and this motion information and RD cost can be reused to accelerate the encoding process of the integer or 4-pixel MV.
[0168] 2.2.5. Higher Motion Vector Storage Precision
[0169] In HEVC, the motion vector precision is a quarter pixel (for 4:2:0 video, a quarter luma sample and an eighth chroma sample). In JEM, the precision of the internal motion vector storage and Merge candidates is increased to 1 / 16 pixel. The higher motion vector precision (1 / 16 pixel) is used for motion compensated inter prediction of CUs encoded / decoded in skip / Merge mode. As described in Section 2.2.2, for CUs encoded / decoded in normal AMVP mode, integer pixel or quarter pixel motion is used.
[0170] The SHVC upsampling interpolation filter with the same filter length and normalization factor as the HEVC motion compensation interpolation filter is used as the motion compensation interpolation filter for additional fractional pixel positions. In JEM, the chroma component motion vector precision is 1 / 32 sample, and an additional interpolation filter for the 1 / 32 pixel fractional position is derived by using the average of two filters at two neighboring 1 / 16 pixel fractional positions.
[0171] 2.2.6. Overlapped Block Motion Compensation
[0172] Overlapped Block Motion Compensation (OBMC) was previously used in H.263. In JEM, unlike H.263, OBMC can be turned on and off using CU-level syntax. When OBMC is used in JEM, OBMC is performed on all motion compensation (MC) block boundaries except the right and bottom boundaries of the CU. In addition, it is applied to both luminance and chrominance components. In JEM, an MC block corresponds to a codec block. When a CU is encoded or decoded in sub-CU mode (including Sub-CU Merge, Affine and FRUC (Frame Rate Up Conversion) modes), each sub-block of the CU is an MC block. In order to handle CU boundaries in a unified manner, OBMC is performed on all MC block boundaries at the sub-block level, where the sub-block size is set to be equal to 4×4, as shown in Figure 14 shown.
[0173] When OBMC is applied to the current sub-block, in addition to the current motion vector, the motion vectors of four adjacent adjacent sub-blocks (if available and different from the current motion vector) are also used to derive the prediction block for the current sub-block. These multiple prediction blocks based on multiple motion vectors are combined to generate the final prediction signal for the current sub-block.
[0174] The prediction block based on the motion vector of the neighboring sub-block is denoted as P N , where N represents the index of the adjacent upper, lower, left, and right sub-blocks, and the prediction block based on the motion vector of the current sub-block is represented as P C When P N When the OBMC is based on the motion information of the adjacent sub-block containing the same motion information as the current sub-block, the OBMC is not N Otherwise, P N Each sample point is added to P C In the same point in the same N Four rows / columns are added to P C The weighting factors {1 / 4, 1 / 8, 1 / 16, 1 / 32} are used for P N , and weighting factors {3 / 4, 7 / 8, 15 / 16, 31 / 32} are used for P C The exception is small MC blocks (ie when the height or width of the codec block is equal to 4 or the CU is coded in sub-CU mode), for which only P N Two rows / columns are added to P C In this case, the weighting factors {1 / 4, 1 / 8} are used for P N , and weighting factors {3 / 4, 7 / 8} are used for P C For P generated based on the motion vector of the vertically (horizontally) adjacent sub-blockN , P N The samples in the same row (column) of P are added to P with the same weighting factor. C .
[0175] In JEM, for CUs with a size less than or equal to 256 luma samples, a CU-level flag is signaled to indicate whether OBMC is applied to the current CU. For CUs with a size greater than 256 luma samples or not encoded in AMVP mode, OBMC is applied by default. At the encoder, when OBMC is applied to a CU, its impact is taken into account during the motion estimation stage. The prediction signal formed by OBMC using the motion information of the top and left neighboring blocks is used to compensate for the upper and left boundaries of the original signal of the current CU, and then the normal motion estimation process is applied.
[0176] 2.2.7. Local lighting compensation
[0177] Local Illumination Compensation (LIC) is based on a linear model of illumination variations using a scaling factor a and an offset b, and is adaptively enabled or disabled for each codec unit (CU) in inter-mode codecs.
[0178] When LIC is applied to a CU, the least square error method is used to derive parameters a and b by using the neighboring samples of the current CU and its corresponding reference samples. More specifically, as Figure 15 As shown, the IC parameters are derived and applied to each prediction direction separately using the subsampling (2:1 subsampling) neighboring samples and corresponding samples of the CU in the reference picture (identified by the motion information of the current CU or sub-CU).
[0179] When the CU is encoded or decoded in Merge mode, the LIC flag is copied from the adjacent block in a manner similar to the motion information copying in Merge mode; otherwise, the LIC flag is signaled for the CU to indicate whether LIC is applied.
[0180] When LIC is enabled for a picture, an additional CU-level RD check is required to determine whether LIC is applied to the CU. When LIC is enabled for a CU, the Mean-Removed Sum of Absolute Difference (MR-SAD) and the Mean-Removed Sum of Absolute Hadamard-Transformed Difference (MR-SATD) (instead of SAD and SATD) are used for integer-pixel motion search and fractional-pixel motion search, respectively.
[0181] To reduce coding complexity, the following coding scheme is applied in JEM. When there is no significant illumination change between the current picture and its reference pictures, LIC is disabled for the entire picture. To identify this situation, the encoder computes the histogram of the current picture and each of its reference pictures. If the histogram difference between the current picture and each of its reference pictures is less than a given threshold, LIC is disabled for the current picture; otherwise, LIC is enabled for the current picture.
[0182] 2.2.8. Hybrid Intra- and Inter-frame Prediction
[0183] In JVET-L0100, multi-hypothesis prediction is proposed, where hybrid intra-frame and inter-frame prediction is a way to generate multiple hypotheses.
[0184] When multi-hypothesis prediction is applied to improve intra mode, multi-hypothesis prediction combines an intra prediction and a merge index prediction. In a merge CU, a flag is signaled for the merge mode to select the intra mode from the intra candidate list when the flag is true. For the luma component, the intra candidate list is derived from 4 intra prediction modes including DC mode, planar mode, horizontal mode and vertical mode, and the size of the intra candidate list can be 3 or 4, depending on the block shape. When the CU width is greater than twice the CU height, the horizontal mode is not included in the intra mode list, and when the CU height is greater than twice the CU width, the vertical mode is removed from the intra mode list. Weighted averaging is used to combine one intra prediction mode selected by the intra mode index and one merge index prediction selected by the merge index. For chroma components, DM is always applied without additional signaling. The weights used for combined prediction are described as follows. When DC mode or planar mode is selected, or the CB width or height is less than 4, equal weights are applied. For those CBs whose width and height are greater than or equal to 4, when horizontal / vertical mode is selected, a CB is first divided vertically / horizontally into four equal-area regions. Each weight set, denoted as (w_intra i ,w_inter i), where i is 1 to 4, and (w_intra1, w_inter1) = (6, 2), (w_intra2, w_inter2) = (5, 3), (w_intra3, w_inter3) = (3, 5), and (w_intra4, w_inter4) = (2, 6) will be applied to the corresponding areas. (w_intra1, w_inter1) is used for the area closest to the reference sample, and (w_intra4, w_inter4) is used for the area farthest from the reference sample. The combined prediction can then be calculated by adding the two weighted predictions and shifting them right by 3 bits. In addition, the intra prediction mode of the intra hypothesis of the predicted value can be saved for reference by subsequent neighboring CUs.
[0185] 2.2.9. Triangle Prediction Unit Mode
[0186] The concept of triangle prediction unit mode is to introduce a new triangle partitioning for motion compensation prediction. Figure 16 As shown in the figure, it divides the CU into two triangular prediction units (PUs) along the diagonal or anti-diagonal direction. Each triangular prediction unit in the CU is inter-predicted using its own unidirectional prediction motion vector and reference frame index, which are derived from the unidirectional prediction candidate list. After predicting the triangular prediction unit, an adaptive weighting process is performed on the diagonal edges. Then, the transformation and quantization process are applied to the entire CU. Note that this mode is only applicable to Skip mode and Merge mode.
[0187] One-way prediction candidate list
[0188] The unidirectional prediction candidate list includes five unidirectional prediction motion vector candidates. It is derived from seven adjacent blocks, including five spatial adjacent blocks (1 to 5) and two temporal collocated blocks (6 to 7), as shown in Figure 2. Figure 17 As shown in the figure, the motion vectors of seven adjacent blocks are collected and placed into a unidirectional prediction candidate list in the order of unidirectional prediction motion vector, L0 motion vector of bidirectional prediction motion vector, L1 motion vector of bidirectional prediction motion vector, and average motion vector of L0 and L1 motion vector of bidirectional prediction motion vector. If the number of candidates is less than five, a zero motion vector is added to the list.
[0189] Adaptive weighting process
[0190] After predicting each triangle prediction unit, an adaptive weighting process is applied to the diagonal edge between the two triangle prediction units to derive the final prediction for the entire CU. The two weighting factor groups are listed below:
[0191] The first weighting factor group: {7 / 8, 6 / 8, 4 / 8, 2 / 8, 1 / 8} and {7 / 8, 4 / 8, 1 / 8} for luma and chroma samples respectively;
[0192] The second weighting factor group: {7 / 8, 6 / 8, 5 / 8, 4 / 8, 3 / 8, 2 / 8, 1 / 8} and {6 / 8, 4 / 8, 2 / 8} for luma and chroma samples respectively.
[0193] A weighting factor set is selected based on the comparison of the motion vectors of the two triangle prediction units. When the reference pictures of the two triangle prediction units are different from each other or the difference between their motion vectors is greater than 16 pixels, the second weighting factor set is used. Otherwise, the first weighting factor set is used. For example Figure 18 shown.
[0194] Motion vector storage
[0195] Motion vector of triangle prediction unit ( Figure 19 Mv1 and Mv2 in 4×4 grid are stored in 4×4 grid. For each 4×4 grid, whether to store unidirectional prediction motion vector or bidirectional prediction motion vector depends on the position of the 4×4 grid in the CU. Figure 19 As shown, unidirectional prediction motion vectors Mv1 or Mv2 are stored for the 4×4 grid located in the non-weighted area. On the other hand, bidirectional prediction motion vectors are stored for the 4×4 grid located in the weighted area. The bidirectional prediction motion vectors are derived from Mv1 and Mv2 according to the following rules:
[0196] 1. In the case where Mv1 and Mv2 have motion vectors from different directions (L0 or L1), Mv1 and Mv2 are simply combined to form a bidirectional prediction motion vector.
[0197] 2. In the case that both Mv1 and Mv2 come from the same L0 (or L1) direction,
[0198] 2.a. If the reference picture of Mv2 is the same as the picture in the L1 (or L0) reference picture list, Mv2 is scaled to that picture. Mv1 and the scaled Mv2 are combined to form a bidirectional prediction motion vector.
[0199] 2.b. If the reference picture of Mv1 is the same as a picture in the L1 (or L0) reference picture list, Mv1 is scaled to that picture. The scaled Mv1 and Mv2 are combined to form a bidirectional prediction motion vector.
[0200] 2.c. Otherwise, only store Mv1 for the weighted region.
[0201] 2.2.10. Affine Motion Compensated Prediction
[0202] In HEVC, only the translational motion model is used for motion compensation prediction (MCP). In the real world, there are many kinds of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In JEM, a simplified affine transformation motion compensation prediction is applied. Figure 20 As shown, the affine motion field of a block is described by two control point motion vectors.
[0203] The motion vector field (MVF) of a block is described by the following equation:
[0204]
[0205] Where (v 0x ,v 0y ) is the motion vector of the upper left control point, (v 1x ,v 1y ) is the motion vector of the upper right control point.
[0206] To further simplify motion compensated prediction, sub-block based affine transformation prediction is applied. The sub-block size M×N is derived from equation (7), where MvPre is the fractional precision of the motion vector (1 / 16 in JEM), (v 2x ,v 2y ) is the motion vector of the lower left control point calculated according to equation (6).
[0207]
[0208] After M and N are derived from equation (7), they should be adjusted downward if necessary to be divisors of w and h, respectively.
[0209] In order to derive the motion vector of each M×N sub-block, as Figure 21 As shown, the motion vector of the center sample of each sub-block is calculated according to Equation 1 and rounded to 1 / 16 fractional precision. Then, the motion compensation interpolation filter mentioned in Section 2.2.5 is applied to generate the prediction of each sub-block using the derived motion vector.
[0210] After MCP, the high-precision motion vector of each sub-block is rounded and saved with the same precision as the normal motion vector.
[0211] 2.2.10.1.AF_INTER Mode
[0212] In JEM, there are two affine motion modes: AF_INTER mode and AF_MERGE mode. For CUs with width and height greater than 8, AF_INTER mode can be applied. A CU-level affine flag is signaled in the bitstream to indicate whether AF_INTER mode is used. In this mode, adjacent blocks are used to construct a pair of motion vectors {(v0, v1)|v0={v A , v B , v c},v1={v D , v E}} candidate list. Figure 23 As shown, v0 is selected from the motion vector of block A, block B or block C. The motion vector from the adjacent block is scaled according to the reference list and the relationship between the POC of the reference of the adjacent block, the POC of the reference of the current CU and the POC of the current CU. And the method of selecting v1 from adjacent blocks D and E is similar. If the number of candidate lists is less than 2, the list is filled with motion vector pairs composed by copying each AMVP candidate. When the candidate list is greater than 2, the candidates are first sorted according to the consistency of the adjacent motion vectors (the similarity of the two motion vectors in a pair of candidates), and only the first two candidates are retained. The RD cost check is used to determine which motion vector pair candidate is selected as the control point motion vector prediction (CPMVP) of the current CU. And the index indicating the position of the CPMVP in the candidate list is signaled in the bitstream. After determining the CPMVP of the current affine CU, affine motion estimation is applied and the control point motion vector (CPMV) is found. Then, the difference between the CPMV and the CPMVP is signaled in the bitstream.
[0213] In AF_INTER mode, when using 4 / 6 parameter affine mode, 2 / 3 control points are required, and therefore 2 / 3 MVDs need to be encoded and decoded for these control points, such as Figure 22 In JVET-K0337, it is proposed to derive MV as follows, that is, predict mvd1 and mvd2 from mvd0.
[0214]
[0215]
[0216]
[0217] in mvd iand mv1 are the predicted motion vector, motion vector difference and motion vector of the upper left pixel (i=0), upper right pixel (i=1) or lower left pixel (i=2), respectively. Figure 22 Note that the addition of two motion vectors (e.g., mvA(xA, yA) and mvB(xB, yB)) is equal to the separate summation of the two components, i.e., newMV = mvA + mvB, where the two components of newMV are set to (xA + xB) and (yA + yB), respectively.
[0218] 2.2.10.2. Fast Affine ME Algorithm in AF_INTER Mode
[0219] In affine mode, the MVs of two or three control points need to be jointly determined. Direct joint search for multiple MVs is computationally complex. A fast affine ME algorithm is proposed and applied to VTM / BMS.
[0220] The fast affine ME algorithm is described for a 4-parameter affine model, and the idea can be extended to a 6-parameter affine model.
[0221]
[0222]
[0223] Replacing (a-1) with a', the motion vector can be rewritten as:
[0224]
[0225] Assuming that the motion vectors of the two control points (0,0) and (0,w) are known, we can derive the affine parameters from equation (13):
[0226]
[0227] The motion vector can be rewritten in vector form as:
[0228]
[0229] in
[0230]
[0231]
[0232] P = (x, y) is the pixel position.
[0233] At the encoder, the MVD of AF_INTER is derived iteratively. i (P) is the MV derived in the i-th iteration at position P, and dMV Ci It is expressed as MV in the i-th iteration C Updated deviation. Then, in the (i+1)th iteration,
[0234]
[0235]
[0236] Pic ref Indicated as a reference picture, and Pic cur Represents the current picture, and represents Q=P+MV i (P). Assuming we use MSE as the matching criterion, we need to minimize:
[0237]
[0238] Assumptions Small enough, we can rewrite it approximately using the first-order Taylor expansion as follows.
[0239]
[0240] in, Indicates E i+1 (P)=Pic cur (P)-Pic ref (Q),
[0241]
[0242] We can deduce this by setting the derivative of the error function to zero. Then, according to
[0243]
[0244]
[0245]
[0246]
[0247] To calculate the deviation MV of the control points (0,0) and (0,w).
[0248] Assuming that such an MVD derivation process is iterated n times, the final MVD is calculated as follows:
[0249]
[0250]
[0251]
[0252]
[0253] Using JVET-K0337, the deviation MV of the control point (0,0) represented by mvd0 is predicted from the deviation MV of the control point (0,w) represented by mvd1. Now, only mvd1 is actually encoded.
[0254] 2.2.10.3.AF_MERGE Mode
[0255] When applying CU in AF_MERGE mode, it obtains the first block encoded and decoded in affine mode from the valid adjacent reconstructed blocks. And the selection order of candidate blocks is from left, top, top right, bottom left to top left, such as Figure 24A As shown, the motion vectors v2, v3, and v4 of the upper left, upper right, and lower left corners of the CU containing block A are derived. The motion vector v0 of the upper left corner of the current CU is calculated based on v2, v3, and v4. Then, the motion vector v1 of the upper right corner of the current CU is calculated.
[0256] After deriving the CPMV v0 and v1 of the current CU, the MVF of the current CU is generated according to the simplified affine motion model equation 1. In order to identify whether the current CU is coded or decoded in AF_MERGE mode, the affine flag is signaled in the bitstream when there is at least one neighboring block coded or decoded in affine mode.
[0257] Figure 24A and Figure 24B Example showing AF_MERGE candidates
[0258] In JVET-L0366, which is planned to be adopted into VTM 3.0, the affine merge candidate list is constructed by the following steps:
[0259] (1) Insert inherited affine candidates
[0260] An inherited affine candidate is one that is derived from the affine motion model of its effectively adjacent affine codec blocks. In a common basis, such as Figure 25 As shown, the scanning order of candidate positions is: A1, B1, B0, A0 and B2.
[0261] After a candidate is derived, a full pruning process is performed to check whether the same candidate has been inserted into the list. If there is an identical candidate, the derived candidate is discarded.
[0262] (2) Inserting constructed affine candidates
[0263] If the number of candidates in the affine Merge candidate list is less than MaxNumAffineCand (set to 5 in this disclosure), the constructed affine candidate is inserted into the candidate list. The constructed affine candidate refers to a candidate constructed by combining the neighboring motion information of each control point.
[0264] The motion information of the control point is first obtained from Figure 25 The specified spatial and temporal neighbors are derived as shown. CPk (k = 1, 2, 3, 4) represents the kth control point. A0, A1, A2, B0, B1, B2, and B3 are the spatial locations of the predicted CPk (k = 1, 2, 3); T is the temporal location of the predicted CP4.
[0265] The coordinates of CP1, CP2, CP3, and CP4 are (0,0), (W,0), (H,0), and (W,H), respectively, where W and H are the width and height of the current block.
[0266] The motion information for each control point is obtained according to the following priority order:
[0267] For CP1, the priority is B2->B3->A2. If B2 is available, B2 is used. Otherwise, if B2 is not available, B3 is used. If both B2 and B3 are unavailable, A2 is used. If all three candidates are unavailable, motion information for CP1 cannot be obtained.
[0268] For CP2, the checking priority is B1->B0.
[0269] For CP3, the checking priority is A1->A0.
[0270] For CP4, use T.
[0271] Second, affine merge candidates are constructed using combinations of control points.
[0272] Constructing a 6-parameter affine candidate requires motion information from three control points. The three control points can be selected from one of the following four combinations: {CP1, CP2, CP4}, {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4}. The combinations {CP1, CP2, CP3}, {CP2, CP3, CP4}, and {CP1, CP3, CP4} are converted into a 6-parameter motion model represented by the top-left, top-right, and bottom-left control points.
[0273] Constructing a 4-parameter affine candidate requires motion information for two control points. These two control points can be selected from one of the following six combinations: {CP1, CP4}, {CP2, CP3}, {CP1, CP2}, {CP2, CP4}, {CP1, CP3}, {CP3, CP4}. The combinations {CP1, CP4}, {CP2, CP3}, {CP2, CP4}, {CP1, CP3}, {CP3, CP4} are converted into a 4-parameter motion model represented by the top left and top right control points.
[0274] The combinations of constructed affine candidates are inserted into the candidate list in the following order:
[0275] {CP1,CP2,CP3}, {CP1,CP2,CP4}, {CP1,CP3,CP4}, {CP2,CP3,CP4}, {CP1,CP2}, {CP1,CP3}, {CP2,CP3}, {CP1,CP4}, {CP2,CP4}, {CP3,CP4}.
[0276] For a combined reference list X (X is 0 or 1), the reference index with the highest usage in the control point is selected as the reference index of list X, and the motion vector pointing to the difference reference picture will be scaled.
[0277] After a candidate is derived, a full pruning process is performed to check whether the same candidate has been inserted into the list. If there is an identical candidate, the derived candidate will be discarded.
[0278] (3) Fill with zero motion vector
[0279] If the number of candidates in the affine merge candidate list is less than 5, a zero motion vector with a zero reference index is inserted into the candidate list until the list is full.
[0280] 2.2.11. Pattern matching motion vector derivation
[0281] The Pattern Matched Motion Vector Derivation (PMMVD) mode is a special Merge mode based on the Frame Rate Up Conversion (FRUC) technology. In this mode, the decoder derives the motion information of the block instead of signaling the motion information of the block.
[0282] When the CU's Merge flag is true, the FRUC flag is signaled to the CU. When the FRUC flag is false, the Merge index is signaled and the normal Merge mode is used. When the FRUC flag is true, an additional FRUC mode flag is signaled to indicate which method (bilateral matching or template matching) will be used to derive the motion information of the block.
[0283] On the encoder side, the decision on whether to use the FRUC Merge mode for a CU is based on the RD cost selection made for normal merge candidates. That is, two matching modes (bilateral matching and template matching) are tested for a CU using RD cost selection. The matching mode that results in the minimum cost is further compared with the other CU modes. If the FRUC matching mode is the most effective mode, the FRUC flag is set to true for the CU, and the relevant matching mode is used.
[0284] The motion inference process in FRUC Merge mode is a two-step process: first, a CU-level motion search is performed, followed by sub-CU-level motion refinement. At the CU level, an initial motion vector for the entire CU is derived based on bilateral matching or template matching. First, a list of MV candidates is generated, and the candidate that results in the minimum matching cost is selected as the starting point for further CU-level refinement. A local search based on bilateral matching or template matching is then performed near the starting point, and the MV that results in the minimum matching cost is selected as the MV for the entire CU. Subsequently, the motion information is further refined at the sub-CU level, using the derived CU motion vector as the starting point.
[0285] For example, the following derivation process is performed for the W×H CU motion information derivation. In the first stage, the MV of the entire W×H CU is derived. In the second stage, the CU is further divided into M×M sub-CUs. The value of M is calculated as shown in (16), and D is a predefined partition depth, which is set to 3 by default in JEM. The MV of each sub-CU is then derived.
[0286]
[0287] like Figure 26 As shown, bilateral matching is used to derive the motion information of the current CU by finding the closest match between two blocks along the motion trajectory of the current CU in two different reference pictures. Under the assumption of continuous motion trajectory, the motion vectors MV0 and MV1 pointing to the two reference blocks should be proportional to the temporal distance between the current picture and the two reference pictures (i.e., TD0 and TD1). As a special case, when the current picture is temporally between the two reference pictures and the temporal distance from the current picture to the two reference pictures is the same, bilateral matching becomes mirror-based bidirectional MV.
[0288] like Figure 27As shown, template matching is used to derive the motion information of the current CU by finding the closest match between the template in the current picture (the top and / or left adjacent blocks of the current CU) and the block in the reference picture (the same size as the template). In addition to the aforementioned FRUC Merge mode, template matching is also applied to AMVP mode. In JEM, as done in HEVC, AMVP has two candidates. New candidates are derived through the template matching method. If the candidate newly derived by template matching is different from the first existing AMVP candidate, it is inserted at the beginning of the AMVP candidate list, and then the list size is set to 2 (meaning removing the second existing AMVP candidate). When applied to AMVP mode, only CU-level search is applied.
[0289] 2.2.11.1. CU-level MV candidate set
[0290] MV candidates set at the CU level include:
[0291] (i) If the current CU is in AMVP mode, it is the original AMVP candidate,
[0292] (ii) All Merge candidates,
[0293] (iii) Several MVs in the interpolated MV field introduced in Section 2.2.11.3,
[0294] (iv) Top and left neighboring motion vectors
[0295] When bilateral matching is used, each valid MV of the Merge candidate is used as input to generate an MV pair assuming bilateral matching. For example, in reference list A, one valid MV of the Merge candidate is (MVa, refa). Then, the reference picture refb of its paired bilateral MV is found in the other reference list B, so that refa and refb are on different sides of the current picture in the temporal domain. If such refb is not available in reference list B, refb is determined to be a reference different from refa, and its temporal distance to the current picture is the minimum in list B. After determining refb, MVb is derived by scaling MVa based on the temporal distance between the current picture and refa and refb.
[0296] The four MVs from the interpolated MV field are also added to the candidate list at the CU level. More specifically, the interpolated MVs at positions (0, 0), (W / 2, 0), (0, H / 2), and (W / 2, H / 2) of the current CU are added.
[0297] When FRUC is applied in AMVP mode, the original AMVP candidates are also added to the MV candidate set at the CU level.
[0298] At the CU level, for AMVP CU, up to 15 MVs are added to the candidate list, and for Merge CU, up to 13 MVs are added to the candidate list.
[0299] 2.2.11.2. Sub-CU level MV candidate set
[0300] MV candidates set at the sub-CU level include:
[0301] (i) MV determined from CU-level search,
[0302] (ii) top, left, upper left and upper right adjacent MVs,
[0303] (iii) a scaled version of the collocated MV from the reference picture,
[0304] (iv) Up to 4 ATMVP candidates,
[0305] (v) A maximum of 4 STMVP candidates.
[0306] The scaled MV from the reference picture is derived as follows: All reference pictures in the two lists are traversed. The MV at the collocated position of the sub-CU in the reference picture is scaled to the reference of the starting CU level MV.
[0307] ATMVP and STMVP candidates are limited to the top four.
[0308] At the sub-CU level, up to 17 MVs are added to the candidate list.
[0309] 2.2.11.3. Generation of interpolated MV fields
[0310] Before encoding or decoding a frame, an interpolated motion field is generated for the entire picture based on unilateral ME. The motion field can then be used as a CU-level or sub-CU-level MV candidate later.
[0311] First, the motion fields of each reference picture in both reference lists are traversed at the 4×4 block level. For each 4×4 block, if the motion associated with the block passes through a 4×4 block in the current picture (e.g. Figure 28 As shown) and the block is not assigned any interpolated motion, the motion of the reference block is scaled to the current picture according to the temporal distances TD0 and TD1 (in the same way as the MV scaling of TMVP in HEVC), and the scaled motion is assigned to the block in the current frame. If no scaled MV is assigned to the 4×4 block, the motion of the block is marked as unavailable in the interpolated motion field.
[0312] 2.2.11.4. Interpolation and Matching Cost
[0313] When the motion vector points to a fractional sample position, motion compensation interpolation is required. To reduce complexity, bilinear interpolation is used for bilateral matching and template matching instead of conventional 8-tap HEVC interpolation.
[0314] The calculation of the matching cost is slightly different at different steps. When selecting a candidate from the candidate set at the CU level, the matching cost is the absolute sum difference (SAD) of the bilateral match or template match. After determining the starting MV, the matching cost C of the bilateral match of the sub-CU level search is calculated as follows:
[0315]
[0316] where w is a weighting factor set to 4 based on experience, MV and MV s Indicates the current MV and the starting MV respectively. SAD is still used as the matching cost of template matching in sub-CU level search.
[0317] In FRUC mode, MV is derived by using only luma samples. The derived motion will be used for both luma and chroma for MC inter-frame prediction. After the MV is determined, the final MC is performed using an 8-tap interpolation filter for luma and a 4-tap interpolation filter for chroma.
[0318] 2.2.11.5.MV Refinement
[0319] MV refinement is a pattern-based MV search with bilateral matching cost or template matching cost as the criterion. In JEM, two search modes are supported - Unrestricted Center-Biased Diamond Search (UCBDS) and Adaptive Cross Search, which are used for MV refinement at CU level and sub-CU level respectively. For both CU and sub-CU level MV refinement, MV is searched directly with quarter luma sample MV precision, and then followed by eighth luma sample MV refinement. The search range for MV refinement for CU and sub-CU steps is set to equal to 8 luma samples.
[0320] 2.2.11.6. Prediction Direction Selection in Template Matching FRUC Merge Mode
[0321] In bilateral matching Merge mode, bidirectional prediction is always applied because the motion information of the CU is derived based on the closest match between two blocks along the motion trajectory of the current CU in two different reference pictures. There is no such restriction for template matching Merge mode. In template matching Merge mode, the encoder can choose from unidirectional prediction in list 0, unidirectional prediction in list 1, or bidirectional prediction for the CU. The selection is based on the template matching cost as follows:
[0322] If costBi<=factor*min(cost0,cost1)
[0323] Bidirectional prediction is used;
[0324] Otherwise, if cost0 <= cost1
[0325] Then use the one-way prediction in list 0;
[0326] otherwise,
[0327] Use the one-way prediction in Listing 1;
[0328] Where cost0 is the SAD of template matching for list 0, cost1 is the SAD of template matching for list 1, and costBi is the SAD of template matching for bidirectional prediction. The value of factor is equal to 1.25, which means that the selection process is biased towards bidirectional prediction.
[0329] Inter prediction direction selection is only applied to the template matching process at the CU level.
[0330] 2.2.12. Generalized Bidirectional Prediction
[0331] In traditional bidirectional prediction, the prediction values from L0 and L1 are averaged with equal weight 0.5 to generate the final prediction value. The prediction value generation formula is shown in Equation (32).
[0332] P TraditionalBiPred =(P L0 +P L1 +RoundingOffset)>>shiftNum Equation (32)
[0333] In equation (32), P TraditionalBiPred is the final prediction value of the traditional two-way prediction, P L0 and P L1 are the predicted values from L0 and L1 respectively, and RoundingOffset and shiftNum are used to normalize the final predicted value.
[0334] Generalized Bi-prediction (GBI) is proposed to allow different weights to be applied to the prediction values from L0 and L1. The prediction value is generated as shown in Equation (33).
[0335] P GBi =((1-w1)*P L0 +w1*P L1 +RoundingOffset GBi )>>shiftNum GBi Equation (33)
[0336] In equation (33), P GBi is the final predicted value of GBi, (1-w1) and w1 are the selected GBI weights applied to the predicted values of L0 and L1 respectively. GBi and shiftNum GBi It is used to normalize the final prediction value in GBi.
[0337] The supported weights for w1 are {-1 / 4, 3 / 8, 1 / 2, 5 / 8, 5 / 4}. One equal-weight set and four unequal-weight sets are supported. For the equal-weight case, the process for generating the final prediction value is exactly the same as in the traditional bidirectional prediction mode. For true bidirectional prediction under random access (RA) conditions, the number of candidate weight sets is reduced to three.
[0338] For Advanced Motion Vector Prediction (AMVP) mode, if the CU is bi-predictive, the weight selection in the GBI is explicitly signaled at the CU level. For Merge mode, the weight selection is inherited from the Merge candidate. In this proposal, the GBI supports weighted averaging of DMVR-generated templates and the final prediction value of BMS-1.0.
[0339] 2.2.13. Multi-hypothesis inter-frame prediction
[0340] In the multi-hypothesis inter-frame prediction mode, in addition to the traditional unidirectional / bidirectional prediction signal, one or more additional prediction signals are signaled. The resulting overall prediction signal is obtained by weighted superposition on a sample-by-sample basis. uni / bi and the first additional inter-frame prediction signal / hypothesis h3, the resulting prediction signal p3 is obtained as follows:
[0341] p3=(1-α)p uni / bi +αh3 Equation (34)
[0342] The changes to the prediction unit syntax are shown in bold in the table below:
[0343]
[0344] The weighting factor α is specified by the syntax element add_hyp_weight_idx according to the following mapping:
[0345] add_hyp_weight_idx α 0 1 / 4 1 -1 / 8
[0346] Note that for the additional prediction signal, the concept of prediction list 0 / list 1 is abolished and a combined list is used instead. This combined list is generated by alternately inserting reference frames from list 0 and list 1 with increasing reference indices, omitting reference frames that have already been inserted, thus avoiding double entries.
[0347] Similar to the above, more than one additional prediction signal may be used. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal.
[0348] p n+1 =(1α n+1 )p n +α n+1 h n+1 Equation (35)
[0349] The resulting overall prediction signal is obtained as the final p n (ie p n has the largest index n).
[0350] Note that also for inter-frame prediction blocks using MERGE mode (rather than skip mode), additional inter-frame prediction signals can be specified. Also note that in the case of MERGE, not only the uni / bi-directional prediction parameters, but also the additional prediction parameters of the selected Merge candidate can be used for the current block.
[0351] 2.2.14. Multi-hypothesis prediction of one-way prediction in AMVP model
[0352] When multi-hypothesis prediction is applied to improve unidirectional prediction in AMVP mode, a flag is signaled to enable or disable multi-hypothesis prediction with inter_dir equal to 1 or 2, where 1, 2, and 3 represent list 0, list 1, and bidirectional prediction, respectively. Additionally, when the flag is true, a merge index is signaled. This way, multi-hypothesis prediction transforms unidirectional prediction into bidirectional prediction, where one motion is derived using the original syntax elements in AMVP mode, while the other motion is derived using the merge scheme. The final prediction combines these two predictions into a bidirectional prediction using a 1:1 weighting. A merge candidate list is first derived from the merge mode, excluding sub-CU candidates (e.g., affine, alternative temporal motion vector prediction (ATMVP)). Next, it is split into two separate lists: one for list 0 (L0), which contains all L0 motion from the candidates, and another for list 1 (L1), which contains all L1 motion. After removing redundancy and filling gaps, two merge lists are generated for L0 and L1, respectively. There are two constraints when applying multi-hypothesis prediction to improve AMVP mode. First, multi-hypothesis prediction is enabled for those CUs whose luma codec block (CB) area is greater than or equal to 64. Second, multi-hypothesis prediction is only applicable to L1 when in low-latency B pictures.
[0353] 2.2.15. Multi-hypothesis prediction in skip / merge mode
[0354] When multi-hypothesis prediction is applied to skip or merge mode, whether multi-hypothesis prediction is enabled is explicitly signaled. In addition to the original prediction, an additional merge index prediction is selected. Therefore, each candidate of multi-hypothesis prediction implies a merge candidate pair, where one merge candidate is used for the first merge index prediction and the other merge candidate is used for the second merge index prediction. However, in each pair, the merge candidate used for the second merge index prediction is implicitly derived as the subsequent merge candidate (i.e., the already signaled merge index plus 1) without signaling any additional merge index. After removing redundancy by excluding pairs containing similar merge candidates and filling in the gaps, a candidate list for multi-hypothesis prediction is formed. Then, the motion from the pair of two merge candidates is obtained to generate the final prediction, where a 5:3 weight is applied to the first merge index prediction and the second merge index prediction, respectively. In addition, a merge or skip CU with multi-hypothesis prediction enabled can save motion information of additional hypotheses in addition to the motion information of the existing hypotheses for reference by subsequent adjacent CUs. Note that sub-CU candidates (e.g., affine, ATMVP) are excluded from the candidate list, and multi-hypothesis prediction is not applicable to skip mode for low-latency B pictures. In addition, when multi-hypothesis prediction is applied to Merge or Skip mode, for those CUs whose CU width or CU height is less than 16, or those CUs whose CU width and CU height are both equal to 16, a bilinear interpolation filter is used for multi-hypothesis motion compensation. Therefore, the worst-case bandwidth (access samples required per sample) of each Merge or Skip CU with multi-hypothesis prediction enabled is calculated in Table 1, and each number is less than half of the worst-case bandwidth of each 4×4 CU with multi-hypothesis prediction disabled.
[0355] 2.2.16. Final motion vector expression
[0356] The final motion vector representation (UMVE) is used in Skip or Merge mode using the proposed motion vector representation method.
[0357] UMVE reuses the same Merge candidates as those used in VVC. Among the Merge candidates, candidates can be selected and further extended by the proposed motion vector representation method.
[0358] UMVE provides a new motion vector representation with simplified signaling, including the starting point, motion magnitude, and motion direction.
[0359] Figure 29 An example of the UMVE search process is shown.
[0360] Figure 30An example of a UMVE search point is shown.
[0361] The proposed technique uses the Merge candidate list as is, but for the extension of UMVE, only candidates of the default Merge type (MRG_TYPE_DEFAULT_N) are considered.
[0362] The base candidate index defines the starting point. The base candidate index indicates the best candidate among the candidates in Table 2, as shown below.
[0363] Table 2. Example base candidate IDX
[0364] Basic Candidate IDX 0 1 2 3 Nth MVP First MVP Second MVP Third MVP Fourth MVP
[0365] If the number of basic candidates is equal to 1, the basic candidate IDX is not signaled.
[0366] The distance index is the motion amplitude information. The distance index indicates the predefined distance from the starting point information. The predefined distance is shown in Table 3 as follows:
[0367] Table 3. Example distances IDX
[0368]
[0369] The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent four directions as shown in Table 4 below.
[0370] Table 4. Example Direction IDX
[0371] Direction IDX 00 01 10 11 x-axis + – N / A N / A y-axis N / A N / A + –
[0372] The UMVE flag is signaled immediately after the Skip and Merge flags are sent. If the Skip and Merge flags are true, the UMVE flag is parsed. If the UMVE flag is 1, the UMVE syntax is parsed. However, if it is not 1, the AFFINE flag is parsed. If the AFFINE flag is 1, this is AFFINE mode. However, if it is not 1, the Skip / Merge index is parsed for VTM Skip / Merge mode.
[0373] No additional line buffering is required for UMVE candidates, as software skip / merge candidates are used directly as base candidates. Using the input UMVE index, the MV complement is determined before motion compensation. No long line buffer is required for this purpose.
[0374] 2.2.17. Affine Merge Mode with Prediction Offset
[0375] UMVE is extended to the affine merge mode, hereinafter also referred to as the UMVE affine mode. The proposed method selects the first available affine merge candidate as the base predictor. It then applies a motion vector offset to the motion vector value of each control point from the base predictor. If no affine merge candidate is available, the proposed method is not used.
[0376] The inter prediction direction of the selected basic prediction value and the reference index for each direction are used without change.
[0377] In the current embodiment, the affine model of the current block is assumed to be a 4-parameter model, and only two control points need to be derived. Therefore, only the first two control points of the basic prediction value will be used as the control point prediction value.
[0378] For each control point, the zero_MVD flag is used to indicate whether the control point of the current block has the same MV value as the corresponding control point prediction value. If the zero_MVD flag is true, no additional signaling is required for the control point. Otherwise, the distance index and offset direction index are signaled for the control point.
[0379] A distance offset table of size 5 is used, as shown in the following table. The distance index is signaled to indicate which distance offset to use. The mapping of distance index and distance offset value is shown in Table 5.
[0380] Table 5 - Example distance offset table
[0381] Distance DX 0 1 2 3 4 Distance Offset 1 / 2 pixel 1 pixel 2 pixels 4 pixels 8 pixels
[0382] Figure 31 An example of a distance index and distance offset mapping is shown.
[0383] The direction index can represent four directions as shown below, where only the x or y direction may have MV differences, but not both directions.
[0384] Offset direction IDX 00 01 10 11 x-dir-factor +1 –1 0 0 y-dir-factor 0 0 +1 –1
[0385] If inter prediction is unidirectional, the signaled distance offset is applied to the offset direction of the prediction value of each control point. The result will be the MV value of each control point.
[0386] For example, when the basic prediction value is unidirectional and the motion vector value of the control point is MVP (v px ,v py ). When the distance offset and direction index are signaled, the motion vector of the corresponding control point of the current block will be calculated as follows.
[0387] MV(v x ,v y )=MVP(vpx ,v py )+MV(x-dir-factor*distance-offset,y-dir-factor*distance-offset) Equation (36)
[0388] If inter prediction is bidirectional, the signaled distance offset is applied in the signaled offset direction of the L0 motion vector of the control point predictor; and the same distance offset with the opposite direction is applied to the L1 motion vector of the control point predictor. The result will be the MV value of each control point in each inter prediction direction.
[0389] For example, when the basic prediction value is unidirectional and the motion vector value of the control point on L0 is MVP L0 (v 0px ,v 0py ), and the motion vector of the control point on L1 is MVP L1 (v 1px ,v 1py ). When the distance offset and direction index are signaled, the motion vector of the corresponding control point of the current block will be calculated as follows.
[0390] MV L0 (v 0x ,v 0y )=MVP L0 (v 0px ,v 0py )+MV(x-dir-factor*distance-offset,y-dir-factor*distance-offset) Equation (37)
[0391] MV L1 (v 0x ,v 0y )=MVP L1 (v 0px ,v 0py )+MV(-x-dir-factor*distance-offset,-y-dir-factor*distance-offset) Equation (38)
[0392] 2.2.18. Bidirectional Optical Flow
[0393] Bidirectional Optical Flow (BIO) is a sample-by-sample motion refinement that is performed on top of block-by-block motion compensation for bidirectional prediction. Sample-level motion refinement does not use signaling.
[0394] Hypothesis I(k) is the luminance value from reference k (k=0,1) after block motion compensation, and I (k) The horizontal and vertical components of the gradient. Assuming that the optical flow is valid, the motion vector field (v x ,v y ) is given by the following equation:
[0395]
[0396] Combine the optical flow equation with Hermite interpolation to obtain the motion trajectory of each sample point, and finally get the function value I (k) and derivatives The only third-order polynomial that matches is the BIO prediction:
[0397]
[0398] Here, τ0 and τ1 represent the distance to the reference frame, as Figure 31 As shown. The distances τ0 and τ1 are calculated based on the POC of Ref0 and Ref1: τ0 = POC(current) - POC(Ref0), τ1 = POC(Refl) - POC(current). If the two predictions come from the same temporal direction (either both from the past or both from the future), the signals are different (i.e., τ0·τ1 < 0). In this case, BIO is only applied when the predictions are not from the same moment (i.e., τ0≠τ1), both reference regions have non-zero motion (MVx0, MVy0, MVx1, MVy1≠0), and the block motion vector is proportional to the temporal distance (MVx0 / MVx1 = MVy0 / MVy1 = -τ0 / τ1).
[0399] By minimizing the difference between point A and point B ( Figure 32 The motion vector field (v x , v y ). The model only uses the first linear term of the local Taylor expansion for Δ:
[0400]
[0401] All values in equation (40) depend on the sample position (i′, j′), which has been ignored from the signs so far. Assuming that the motion is consistent in the local surrounding area, we minimize Δ within a (2M+1)×(2M+1) square window Ω centered at the current prediction point (i, j), where M is equal to 2:
[0402]
[0403] For this optimization problem, JEM uses a simplified approach, first minimizing in the vertical direction and then minimizing in the horizontal direction. This yields:
[0404]
[0405]
[0406] in,
[0407]
[0408]
[0409]
[0410] To avoid division by zero or very small values, regularization parameters r and m are introduced into equations (43) and (44).
[0411] r=500·4 d-8 Equation (46)
[0412] m=700·4 d-8 Equation (47)
[0413] Here d is the bit depth of the video samples.
[0414] In order to keep the memory access of BIO the same as that of conventional bi-predictive motion compensation, all predictions and gradient values I (k) , are only calculated for the positions inside the current block. In equation (45), the (2M+1)×(2M+1) square window Ω centered at the current prediction point on the boundary of the prediction block needs to access the positions outside the block (such as Figure 33A In JEM, the I outside the block (k) , The value of is set equal to the nearest available value inside the block. This can be implemented as padding, such as Figure 33B shown.
[0415] Using BIO, it is possible to refine the motion field for each sample. To reduce the computational complexity, a block-based BIO design is used in JEM. The motion refinement is calculated based on a 4×4 block. In the block-based BIO, the s in Equation 30 is aggregated over all samples in the 4×4 block. n value, then s n The aggregated value of is used to derive the BIO motion vector offset of the 4×4 block. More specifically, the following formula is used for block-based BIO derivation:
[0416]
[0417] where b k represents the sample set of the kth 4×4 block belonging to the prediction block. n ((s n,bk )>>4) instead to derive the associated motion vector offset.
[0418] In some cases, the MV refinement of BIO may be unreliable due to noise or irregular motion. Therefore, in BIO, the amplitude of MV refinement is clipped to a threshold thBIO. The threshold is determined based on whether the reference pictures of the current picture are all from one direction. If all the reference pictures of the current picture are from one direction, the threshold is set to 12×2 14-d ; otherwise, it is set to 12×2 13-d .
[0419] The gradient of BIO is calculated simultaneously with the motion compensated interpolation using an operation consistent with the HEVC motion compensation process (2D separable FIR). The input to this 2D separable FIR is the same reference frame sample as the motion compensation process and the fractional position (fracX, fracY) according to the fractional part of the block motion vector. In the case of , the signal is first interpolated vertically using BIOfilterS corresponding to the fractional position fracy with a de-scaling offset d-8, and then a gradient filter BIOfilterG corresponding to the fractional position fracX with a de-scaling offset 18-d is applied in the horizontal direction. In the case of , the gradient filter is first applied vertically using BIOfilterG corresponding to the fractional position fracY with a descaling offset d-8, and then the signal displacement is performed in the horizontal direction using BIOfilterS corresponding to the fractional position fracX with a descaling offset 18-d. The length of the interpolation filters BIOfilterG for gradient calculation and BIOfilterF for signal displacement is short (6 taps) in order to maintain a reasonable complexity. Table 6 shows the filters used for gradient calculation for different fractional positions of block motion vectors in BIO. Table 7 shows the interpolation filters used for prediction signal generation in BIO.
[0420] Table 6: Example filters for gradient calculation in BIO
[0421] Fractional pixel position Gradient interpolation filter (BIOfilterG) 0 {8,-39,-3,46,-17,5} 1 / 16 {8,-32,-13,50,-18,5} 1 / 8 {7,-27,-20,54,-19,5} 3 / 16 {6,-21,-29,57,-18,5} 1 / 4 {4,-17,-36,60,-15,4} 5 / 16 {3,-9,-44,61,-15,4} 3 / 8 {1,-4,-48,61,-13,3} 7 / 16 {0,1,-54,60,-9,2} 1 / 2 {-1,4,-57,57,-4,1}
[0422] Table 7: Example interpolation filters for prediction signal generation in BIO
[0423] Fractional pixel position Interpolation filter for prediction signal (BIOfilterS) 0 {0,0,64,0,0,0} 1 / 16 {1,-3,64,4,-2,0} 1 / 8 {1,-6,62,9,-3,1} 3 / 16 {2,-8,60,14,-5,1} 1 / 4 {2,-9,57,19,-7,2} 5 / 16 {3,-10,53,24,-8,2} 3 / 8 {3,-11,50,29,-9,2} 7 / 16 {3,-11,44,35,-10,3} 1 / 2 {3,-10,35,44,-11,3}
[0424] In JEM, BIO is applied to all bidirectionally predicted blocks when the two predictions come from different reference pictures. When LIC is enabled for a CU, BIO is disabled.
[0425] In JEM, OBMC is applied to a block after the normal MC process. To reduce computational complexity, BIO is not applied during the OBMC process. This means that BIO is only applied to the MC process of a block when using its own MV, and is not applied to the MC process when using the MV of a neighboring block during the OBMC process.
[0426] 2.2.19. Decoder-side motion vector refinement
[0427] In bidirectional prediction, two prediction blocks, formed using motion vectors (MVs) from list 0 and list 1, are combined to form a single prediction signal for a block region. In the decoder-side motion vector refinement (DMVR) method, the two motion vectors of the bidirectional prediction are further refined through a bilateral template matching process. Bilateral template matching is applied in the decoder to perform a distortion-based search between the bilateral template and the reconstructed samples in the reference picture to obtain the refined MVs without sending additional motion information.
[0428] In DMVR, the bilateral template is generated as a weighted combination (i.e., average) of two predicted blocks from the initial MV0 of list 0 and MV1 of list 1, respectively, as Figure 34 As shown. The template matching operation involves calculating a cost metric between the generated template and the sample area (around the initial prediction block) in the reference picture. For each of the two reference pictures, the MV that produces the minimum template cost is considered to be the updated MV in the list to replace the original MV. In JEM, 9 MV candidates are searched for each list. The 9 candidate MVs include the original MV and 8 surrounding MVs, which are offset by one luminance sample relative to the original MV in the horizontal or vertical direction or both. Finally, two new MVs (i.e., MV0′ and MV1′), as shown in Figure 33, are used to generate the final bidirectional prediction result. The sum of absolute differences (SAD) is used as the cost metric. Please note that when calculating the cost of a prediction block generated by a surrounding MV, the rounded MV (rounded to integer pixels) is actually used to obtain the prediction block, not the actual MV.
[0429] DMVR is applied to the Merge mode of bi-prediction, where one MV comes from the past reference picture and the other comes from the future reference picture, without transmitting additional syntax elements. In JEM, DMVR is not applied when LIC, affine motion, FRUC or sub-CU Merge candidate is enabled for the CU.
[0430] 3. Related Tools
[0431] 3.1.1. Diffusion filter
[0432] In JVET-L0157, a diffusion filter is proposed, where the intra / inter prediction signal of a CU can be further modified by a diffusion filter.
[0433] 3.1.1.1 Uniform Diffusion Filter
[0434] The uniform diffusion filter is constructed by dividing the prediction signal by a given value h I or h IV This is achieved by convolution with a fixed mask (defined below).
[0435] In addition to the prediction signal itself, a row of reconstructed samples to the left and above the block is used as input for the filtered signal, where the use of these reconstructed samples in inter blocks can be avoided.
[0436] Let pred be the prediction signal on a given block obtained by intra-frame or motion compensated prediction. In order to handle the boundary points of the filter, the prediction signal needs to be expanded to the prediction signal pred ext This extended prediction can be formed in two ways: either, as an intermediate step, a row of reconstructed samples to the left and above the block is added to the prediction signal and the resulting signal is then mirrored in all directions. Alternatively, only the prediction signal itself is mirrored in all directions. The latter extension is used for inter blocks. In this case, only the prediction signal itself consists of the extended prediction signal pred ext input.
[0437] If you want to use the filter h I , then it is proposed to use h I *pred replaces the predicted signal pred, using the aforementioned boundary extension. Here, the filter mask h I Given as follows:
[0438]
[0439] If you want to use the filter h IV , then it is proposed to use h IV *pred replaces the predicted signal pred. Here, the filter h IV Given as follows:
[0440] h IV =h I *h I *h I *h I Equation (50)
[0441] 3.1.1.2. Directional diffusion filter
[0442] Instead of using a signal adaptive diffusion filter, a directional filter is used, a horizontal filter h hor and vertical filter h ver , which still has a fixed mask. More precisely, the one in the previous section corresponds to the mask h I The uniform diffusion filtering is simply restricted to be applied only in the vertical direction or in the horizontal direction.
[0443]
[0444] is applied to the prediction signal to implement a straight filter, and by using the transposed mask to implement a horizontal filter.
[0445] 3.1.2. Bilateral filter
[0446] A bilateral filter was proposed in JVET-L0406 and is always applied to luma blocks with non-zero transform coefficients and a slice quantization parameter greater than 17. Therefore, no signaling is required to indicate the use of the bilateral filter. The bilateral filter, if applied, is performed on the decoded samples immediately after the inverse transform. In addition, the filter parameters (i.e., weights) are explicitly derived from codec information.
[0447] The filtering process is defined as:
[0448]
[0449] Among them, P 0,0 is the intensity of the current sample point and P′ 0,0 is the modified intensity of the current sample point, P k,0 and W k are the intensity and weighting parameters of the kth adjacent sample point respectively. Figure 35 An example of a current sample point and its four neighboring sample points (ie, K=4) is depicted in FIG.
[0450] More specifically, the weight W associated with the kth neighboring sample point k (x) is defined as follows:
[0451] W k (x) = Distance k ×Range k(x) Equation (52)
[0452] in
[0453]
[0454]
[0455] σ d Depends on the codec mode and codec block size. The described filtering process is applied to intra-codec blocks and to inter-codec blocks when the TU is further partitioned to enable parallel processing.
[0456] In order to better capture the statistical characteristics of the video signal and improve the performance of the filter, the weight function generated by equation (52) is composed of the parameters σ listed in Table 8 d Adjust the parameter σ d Depends on the codec mode and block segmentation parameters (minimum size).
[0457] Table 8 Example values of σd for different block sizes and coding modes
[0458] Min(block width, block height) Intra-frame mode Interframe mode 4 82 62 8 72 52 other 52 32
[0459] In order to further improve the coding performance, for inter-frame coding blocks when TU is not divided, the intensity difference between the current sample and one of its adjacent samples is replaced by the representative intensity difference between two windows covering the current sample and the adjacent samples. Therefore, the filtering process equation is modified as follows:
[0460]
[0461] Among them, P k,m and P 0,m Indicated by P k,0 and P 0,0 The mth sample value in the window centered at . In this proposal, the window size is set to 3×3. Figure 36 The coverage P is depicted in 2,0 and P 0,0 Example of two windows.
[0462] 3.1.3. Intra-frame Block Copy
[0463] Decoder:
[0464] In some cases, the currently (partially) decoded picture is considered a reference picture. The current picture is placed at the last position in reference picture list 0. Therefore, for slices that use the current picture as the only reference picture, the slice type is considered to be a P slice. The bitstream syntax in this method follows the same syntax structure used for inter codecs, and the decoding process is unified with inter codecs. The only significant difference is that the block vector (which is a motion vector pointing to the current picture) always uses integer pixel resolution.
[0465] The changes from the block-level CPR_flag method are:
[0466] 1. In this mode the encoder searches for block width and height that are both less than or equal to 16.
[0467] 2. When the luma block vector is odd, chroma interpolation is enabled.
[0468] 3. When the SPS flag is turned on, enable the adaptive motion vector resolution (AMVP) of CPR mode. In this case, when using AMVR, at the block level, the block vector can switch between 1 pixel integer resolution and 4 pixel integer resolution.
[0469] Encoder side:
[0470] The encoder performs RD checks on blocks with a width or height no greater than 16. For non-Merge mode, a block vector search is first performed using a hash-based search. If no valid candidate is found in the hash search, a local search based on block matching is performed.
[0471] In hash-based search, hash key matching (32-bit CRC) between the current block and reference blocks is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on 4×4 blocks. For larger current block sizes, a hash key match with a reference block occurs when all of its 4×4 blocks match the hash keys in the corresponding reference positions. If multiple reference blocks are found to match the current block with the same hash key, the block vector cost of each candidate is calculated and the one with the lowest cost is selected.
[0472] In block matching search, the search range is set to 64 pixels to the left and top of the current block.
[0473] 3.1.4. History-based Motion Vector Prediction
[0474] PCT Application No. PCT / CN2018 / 093987, filed on July 2, 2018, and entitled “MOTION VECTOR PREDICTION BASED ON LOOK-UPTABLES,” the contents of which are incorporated herein by reference, describes one or more lookup tables in which at least one motion candidate is stored to predict motion information of a block.
[0475] A history-based MVP (HMVP) method is proposed, where the HMVP candidate is defined as the motion information of the previously coded block. During the encoding / decoding process, a table with multiple HMVP candidates is maintained. When a new slice is encountered, the table is cleared. Whenever there is an inter-frame coded block, the associated motion information is added to the last entry of the table as a new HMVP candidate. The entire encoding and decoding process is Figure 37 In one example, the table size is set to L (eg, L=16 or 6, or 44), which indicates that a maximum of L HMVP candidates can be added to the table.
[0476] (1) In one embodiment, if there are more than L HMVP candidates from previously coded blocks, a First-In-First-Out (FIFO) rule is applied so that the table always contains the latest L previously coded motion candidates. Figure 38 An example of applying the FIFO rule to remove HMVP candidates and add new candidates to the table used in the proposed method is depicted.
[0477] (2) In another embodiment, whenever a new motion candidate is added (such as when the current block is inter-coded and in non-affine mode), a redundancy check process is first applied to identify whether there is the same or similar motion candidate in the LUT.
[0478] 3.2.JVET-M0101
[0479] 3.2.1. Add HEVC-style weighted prediction (WP)
[0480] To implement HEVC-style weighted prediction, the following syntax changes (shown in bold) may be made to the PPS and slice header:
[0481]
[0482]
[0483]
[0484]
[0485]
[0486] 3.2.2. Calling in a mutually exclusive manner (bi-prediction with weighted averaging (BWA, also known as GBI) and WP
[0487] Although both BWA and WP apply weights to the motion compensated prediction signal, the specific operation is different. In the case of bidirectional prediction, WP applies linear weights and constant offsets to each prediction signal, depending on the (weight, offset) parameters associated with the ref_idx used to generate the prediction signal. Although there are some range restrictions on these (weight, offset) values, the allowed range is relatively large and there are no normalization restrictions on the parameter values. WP parameters are signaled at the picture level (in the slice header). WP has been shown to provide very large codec gains for fading sequences, and it is generally desirable to enable WP for such content.
[0488] BWA uses a CU-level index to indicate how to combine the two prediction signals in the case of bidirectional prediction. BWA applies a weighted average, i.e., the two weights add up to 1. BWA provides a codec gain of about 0.66% for the CTC RA configuration, but its efficiency is much lower than WP (JVET-L0201) for fading content.
[0489] Method #1 :
[0490] The PPS syntax element weight_bipred_flag is checked to determine whether gbi_idx is signaled at the current CU. The syntax signaling is modified in bold as follows:
[0491]
[0492] If the current slice refers to a PPS with weight_bipred_flag set to 1, the first option is to disable BWA for all reference pictures of the current slice.
[0493] Method #2 :
[0494] BWA is disabled only when both reference pictures used in bidirectional prediction have weighted prediction turned on, i.e., the (weight, offset) parameters of these reference pictures have non-default values. This allows bidirectional prediction CUs using reference pictures with default WP parameters (i.e., WP is not called for these CUs) to still use BWA. The syntax signaling is modified in bold as follows:
[0495]
[0496] 4. Question
[0497] In LIC, two parameters including the scaling parameter a and the offset b need to be derived by using adjacent reconstructed pixels, which may cause a delay problem.
[0498] Hybrid intra-frame and inter-frame prediction, diffusion filter, bilateral filter, OBMC and LIC need to further modify the inter-frame prediction signal in different ways, and they all have delay problems.
[0499] 5. Some Example Embodiments and Techniques
[0500] The following detailed inventions should be considered as examples to explain the general concept. These inventions should not be interpreted narrowly. In addition, these inventions can be combined in any way.
[0501] Example 1:
[0502] In some embodiments, LIC is performed only on blocks located at the CTU boundary (referred to as boundary blocks), and the LIC parameters can be derived using adjacent reconstructed samples of the current CTU. In this case, LIC is always disabled for blocks not located at the CTU boundary (referred to as intra blocks) without being signaled.
[0503] In some embodiments, the selection of adjacent reconstructed samples outside the current CTU may depend on the position of the block relative to the CTU covering the current block.
[0504] In some embodiments, for blocks at the left boundary of a CTU, only the left reconstructed neighboring samples of the CU may be used to derive the LIC parameters.
[0505] In some embodiments, for a block at the upper boundary of a CTU, only the upper reconstructed neighboring samples of the block may be used to derive the LIC parameters.
[0506] In some embodiments, for a block at the top left corner of a CTU, the LIC parameters may be derived using the left and / or top reconstructed neighboring samples of the block.
[0507] In some embodiments, for a block at the upper right corner of a CTU, the LIC parameters are derived using the upper right and / or upper reconstructed neighboring samples of the block.
[0508] In some embodiments, for a block at the bottom left corner of a CTU, the LIC parameters are derived using the bottom left and / or left reconstructed neighboring samples of the block.
[0509] Example 2:
[0510] In some embodiments, LIC parameter sets may be derived only for blocks located at a CTU boundary, and internal blocks of the CTU may inherit from one or more of these LIC parameter sets.
[0511] In some embodiments, a LIC parameter lookup table is maintained for each CTU, and each derived LIC parameter set is inserted into the LIC parameter table. The LIC parameter lookup table can be maintained using the method described in PCT Application No. PCT / CN2018 / 093987, filed on July 2, 2018, entitled "Motion Vector Prediction Based on Look-up Tables," the contents of which are incorporated herein by reference.
[0512] In some embodiments, such a lookup table is maintained for each reference picture, and when deriving LIC parameters, LIC parameters are derived for all reference pictures.
[0513] In some embodiments, such a lookup table is cleared at the beginning of each CTU or CTU row or slice or slice group or picture.
[0514] In some embodiments, for intra blocks coded in AMVP mode or affine inter mode, if the LIC flag is true, the LIC parameter set used is explicitly signaled. For example, an index is signaled to indicate which lookup table entry is used for each reference picture. In some embodiments, for unidirectionally predicted blocks, the LIC flag will be false if there is no valid entry for the reference picture in the lookup table. In some embodiments, for bidirectionally predicted blocks, the LIC flag will be false if there is no valid entry for any reference picture in the lookup table. In some embodiments, for bidirectionally predicted blocks, the LIC flag may be true if there is a valid entry for at least one reference picture in the lookup table. For each reference picture with valid LIC parameters, a LIC parameter index is signaled.
[0515] In some embodiments, more than one lookup table is maintained. In some embodiments, when encoding / decoding the current CTU, it only uses parameters from the lookup table generated by some previously encoded / decoded CTU, and the lookup table generated by the current CTU is used for some subsequent CTUs. In some embodiments, each lookup table can correspond to a reference picture / a reference picture list / a specific range of block sizes / a specific codec mode / a specific block shape, etc.
[0516] In some embodiments, if the intra block is coded in Merge or Affine Merge mode, for spatial and / or temporal Merge candidates, the LIC flag and LIC parameters are inherited from the corresponding neighboring blocks.
[0517] In some embodiments, the LIC flag is inherited and the LIC parameters may be explicitly signaled. In some embodiments, the LIC flag is explicitly signaled and the LIC parameters may be inherited. In some embodiments, both the LIC flag and the LIC parameters are explicitly signaled. In some embodiments, the difference between the LIP parameters of a block and the inherited parameters is signaled.
[0518] In some embodiments, for boundary blocks coded in AMVP mode or affine inter mode, if the LIC flag is true, the LIC parameters used are implicitly derived using the same method as in 2.2.7.
[0519] In some embodiments, if the boundary block is coded in Merge or Affine Merge mode, for spatial and / or temporal Merge candidates, the LIC flag and LIC parameters are inherited from the corresponding neighboring blocks.
[0520] In some embodiments, the LIC flag is inherited and the LIC parameters may be explicitly signaled. In some embodiments, the LIC flag is explicitly signaled and the LIC parameters may be inherited. In some embodiments, both the LIC flag and the LIC parameters are explicitly signaled.
[0521] In some embodiments, if a block is coded in Merge mode, LIC is always disabled for combined Merge candidates or averaged Merge candidates.
[0522] In some embodiments, if the LIC flag of either of the two merge candidates used to generate the combined merge candidate or the average merge candidate is true, the LIC flag is set to true. In some embodiments, for an internal block, if both merge candidates used to generate the combined merge candidate or the average merge candidate use LIC, the LIC parameters may be inherited from either one of them. In some embodiments, the LIC flag is inherited and the LIC parameters may be explicitly signaled. In some embodiments, the LIC flag is explicitly signaled and the LIC parameters may be inherited. In some embodiments, both the LIC flag and the LIC parameters are explicitly signaled.
[0523] In some embodiments, if a block is encoded or decoded in Merge mode from an HMVP Merge candidate, the LIC is always disabled. In some embodiments, if a block is encoded or decoded in Merge mode from an HMVP Merge candidate, the LIC is always enabled. In some embodiments, if the LIC flag of the HMVP Merge candidate is true, the LIC flag is set to true. For example, the LIC flag can be explicitly signaled. In some embodiments, if the LIC flag of the HMVP Merge candidate is true, the LIC flag is set to true and the LIP parameters are inherited from the HMVP Merge candidate. In some embodiments, the LIC parameters are explicitly signaled. In some embodiments, the LIC parameters are implicitly derived.
[0524] Example 3:
[0525] LIC can be used with the intra block copy (IBC, or current picture reference) mode.
[0526] In some embodiments, if a block is encoded in intra block copy mode, an indication of LIC usage (e.g., a LIC flag) may also be signaled. In some embodiments, in Merge mode, if an IBC-encoded block inherits motion information from a neighboring block, it may also inherit the LIC flag. In some embodiments, it may also inherit LIC parameters.
[0527] In some embodiments, if the LIC flag is true, the LIC parameters are implicitly derived using the same method as in 2.2.7. In some embodiments, the methods described in Example 1 and / or Example 2 may be used. In some embodiments, if LIC is enabled for encoding and decoding a block, an indication of IBC usage may also be signaled.
[0528] Example 4:
[0529] In some embodiments, hybrid intra and inter prediction (also called combined intra-inter prediction), diffusion filters, bilateral filters, transform domain filtering methods, OBMC, LIC, or any other tool that modifies the inter prediction signal or modifies the reconstructed blocks from motion compensation (e.g., causing latency issues) is used exclusively.
[0530] In some embodiments, if LIC is enabled, all other tools are implicitly disabled. If the LIC flag (explicitly signaled or implicitly derived) is true, the on / off flags of other tools (if any) are not signaled but are implicitly derived to be off.
[0531] In some embodiments, if hybrid intra and inter prediction is enabled, all other tools are implicitly disabled. If the (explicitly signaled or implicitly derived) hybrid intra and inter prediction flag is true, the on / off flags of other tools (if any) are not signaled but are implicitly derived to be off.
[0532] In some embodiments, if the diffusion filter is enabled, all other tools are implicitly disabled. If the diffusion filter flag (explicitly signaled or implicitly derived) is true, the on / off flags of other tools (if any) are not signaled, but are implicitly derived to be off.
[0533] In some embodiments, if OBMC is enabled, all other tools are implicitly disabled. If the OBMC flag (explicitly signaled or implicitly derived) is true, the on / off flags of other tools (if any) are not signaled but are implicitly derived to be off.
[0534] In some embodiments, if the bilateral filter is enabled, all other tools are implicitly disabled. If the (explicitly signaled or implicitly derived) bilateral filter flag is true, the on / off flags of other tools (if any) are not signaled, but are implicitly derived to be off.
[0535] In some embodiments, different tools can be checked sequentially. Furthermore, this checking process terminates when a decision is made to enable one of the tools. In some embodiments, the checking order is LIC → diffusion filter → hybrid intra- and inter-frame prediction → OBMC → bilateral filter. In some embodiments, the checking order is LIC → diffusion filter → bilateral filter → hybrid intra- and inter-frame prediction → OBMC.
[0536] In some embodiments, the order of checking is LIC → OBMC → diffusion filter → bilateral filter → hybrid intra-frame and inter-frame prediction. In some embodiments, the order of checking is LIC → OBMC → diffusion filter → hybrid intra-frame and inter-frame prediction → bilateral filter. In some embodiments, the order can be adaptively changed based on previous codec information and / or based on codec information of the current block (e.g., block size / reference picture / MV information / low delay check flag / slice / picture / slice type), such as based on the mode of neighboring blocks.
[0537] In some embodiments, for any of the above tools (e.g., CIIP, LIC, diffusion filter, bilateral filter, transform domain filtering method), whether to automatically disable it and / or use a different way to decode the current block can depend on the encoding and decoding mode of adjacent or non-adjacent rows or columns.
[0538] In some embodiments, for any of the above tools (eg, CIIP, LIC, diffusion filter, bilateral filter, transform domain filtering methods), whether to automatically disable it may depend on all codec modes of adjacent or non-adjacent rows or columns.
[0539] In some embodiments, for any of the above tools (e.g., CIIP, LIC, diffusion filter, bilateral filter, transform domain filtering method), whether to enable it may depend on at least one of the samples in adjacent or non-adjacent rows or columns not being encoded or decoded in a particular mode.
[0540] In some embodiments, adjacent or non-adjacent rows may include upper rows and / or upper right rows. In some embodiments, adjacent or non-adjacent columns may include left columns and / or lower left columns. In some embodiments, the codec mode / specific mode may include intra mode and / or combined intra-inter mode (CIIP). In some embodiments, if any adjacent / non-adjacent block in an adjacent or non-adjacent row or column is coded with intra and / or CIIP, this method is disabled. In some embodiments, if all adjacent / non-adjacent blocks in an adjacent or non-adjacent row or column are coded with intra and / or CIIP, this method is disabled. In some embodiments, different methods may include short tap filters / fewer adjacent samples for LIC parameter derivation, reference sample filling, etc.
[0541] In some embodiments, for any of the above tools (e.g., CIIP, LIC, diffusion filter, bilateral filter, transform domain filtering method), the tool is disabled when at least one upper adjacent / non-adjacent upper adjacent block in an adjacent or non-adjacent row is intra-coded and / or coded in a combined intra-inter mode.
[0542] In some embodiments, for any of the above tools (e.g., CIIP, LIC, diffusion filter, bilateral filter, transform domain filtering method), the tool is disabled when at least one left adjacent / non-adjacent left adjacent block in an adjacent or non-adjacent column is intra-coded and / or coded in combined intra-inter mode.
[0543] In some embodiments, for any of the above tools (e.g., CIIP, LIC, diffusion filter, bilateral filter, transform domain filtering methods), the tool is disabled when all blocks in adjacent or non-adjacent rows / columns are intra-coded and / or coded in combined intra-inter mode.
[0544] In some embodiments, for a LIC codec block, neighboring samples coded in intra mode and / or mixed intra and inter mode are excluded from the derivation of LIC parameters.
[0545] In some embodiments, for a LIC codec block, neighboring samples coded in non-intra mode and / or non-CIIP mode may be included in the derivation of LIC parameters.
[0546] Example 5:
[0547] In some embodiments, when a tool is disabled for a block (such as when all neighboring samples are intra-coded), such codec tools may still be applied, but in a different manner.
[0548] In some embodiments, a short tap filter may be applied. In some embodiments, padding of reference samples may be applied. In some embodiments, adjacent rows / columns and / or non-adjacent rows / columns of a block in a reference picture (via motion compensation if necessary) may be used as replacements for adjacent rows / columns and / or non-adjacent rows / columns of the current block.
[0549] Example 6:
[0550] In some embodiments, LIC is used exclusively with GBI or multi-hypothesis inter prediction.
[0551] In some embodiments, GBI information may be signaled after the LIC information. In some embodiments, whether GBI information is signaled may depend on the signaled / inferred LIC information. Furthermore, whether GBI information is signaled may depend on the LIC information and weighted prediction information associated with at least one reference picture of the current block or the current slice / slice group / picture containing the current block. In some embodiments, this information may be signaled in the SPS / PPS / slice header / slice group header / slice / CTU / CU.
[0552] In some embodiments, LIC information may be signaled after GBI information. In some embodiments, whether LIC information is signaled may depend on signaled / inferred GBI information. Furthermore, whether LIC information is signaled may depend on GBI information and weighted prediction information associated with at least one reference picture of the current block or the current slice / slice group / picture containing the current block. In some embodiments, this information may be signaled in an SPS / PPS / slice header / slice group header / slice / CTU / CU.
[0553] In some embodiments, if the LIC flag is true, the syntax elements required for GBI or multi-hypothesis inter prediction are not signaled.
[0554] In some embodiments, if the GBI flag is true (eg, unequal weights are applied to two or more reference pictures), syntax elements required for LIC or multi-hypothesis inter prediction are not signaled.
[0555] In some embodiments, if GBI is enabled for blocks of two or more reference pictures with unequal weights, the syntax elements required for LIC or multi-hypothesis inter prediction are not signaled.
[0556] In some embodiments, LIC is used exclusively with sub-block techniques (such as affine mode). In some embodiments, LIC is used exclusively with triangle prediction mode. In some embodiments, LIC is always disabled when triangle prediction mode is enabled for a block. In some embodiments, the LIC flag of a TPM Merge candidate can be inherited from spatial or temporal blocks or other types of motion candidates (e.g., HMVP candidates). In some embodiments, the LIC flag of a TPM Merge candidate can be inherited from some spatial or temporal blocks (e.g., only A1, B1). In some embodiments, the LIC flag of a TPM Merge candidate can always be set to false.
[0557] In some embodiments, if multi-hypothesis inter prediction is enabled, the syntax elements required for LIC or GBI are not signaled.
[0558] In some embodiments, different tools may be checked in a specific order. Furthermore, this checking process terminates when a decision is made to enable one of the tools. In some embodiments, the checking order is LIC -> GBI -> multi-hypothesis inter prediction. In some embodiments, the checking order is LIC -> multi-hypothesis inter prediction -> GBI. In some embodiments, the checking order is GBI -> LIC -> multi-hypothesis inter prediction.
[0559] In some embodiments, the check order is GBI->Multi-Hypothesis Inter Prediction->LIC. In some embodiments, the order can be adaptively changed based on previous codec information and / or based on codec information of the current block (e.g., block size / reference picture / MV information / low delay check flag / slice / picture / slice type), such as based on the mode of neighboring blocks.
[0560] In some embodiments, the above method may be applied only when the current block is uni-predicted.
[0561] Example 7:
[0562] In some embodiments, LIC is used exclusively with combined inter-intra prediction (CIIP). In some embodiments, CIIP information may be signaled after the LIC information. In some embodiments, whether CIIP information is signaled may depend on the signaled / inferred LIC information. Furthermore, whether CIIP information is signaled may depend on the LIC information and weighted prediction information associated with at least one reference picture of the current block or the current slice / slice group / picture containing the current block. In some embodiments, this information may be signaled in the SPS / PPS / slice header / slice group header / slice / CTU / CU.
[0563] In some embodiments, the LIC information may be signaled after the CIIP information.
[0564] In some embodiments, whether LIC information is signaled may depend on signaled / inferred CIIP information. Furthermore, whether LIC information is signaled may depend on CIIP information and weighted prediction information associated with at least one reference picture of the current block or the current slice / slice group / picture containing the current block. In some embodiments, this information may be signaled in an SPS / PPS / slice header / slice group header / slice / CTU / CU.
[0565] In some embodiments, if the LIC flag is true, then the syntax elements required for CIIP are not signaled. In some embodiments, if the CIIP flag is true, then the syntax elements required for LIC are not signaled. In some embodiments, if CIIP is enabled, then the syntax elements required for LIC are not signaled.
[0566] Example 8:
[0567] In some embodiments, weighted prediction cannot be applied when one or more of the following codec tools are applied (except weighted prediction):
[0568] a.BIO (also known as BDOF)
[0569] b.CIIP
[0570] c. Affine prediction
[0571] d. Overlapped Block Motion Compensation (OBMC)
[0572] e. Decoder-side motion vector refinement (DMVR)
[0573] In some embodiments, if codec tools and weighted prediction are mutually exclusive, then when weighted prediction is applied, information (e.g., a flag) indicating whether the codec tools are used is not signaled. Such information (e.g., a flag) can be inferred to be zero. In some embodiments, this information can still be expressively signaled in the SPS / PPS / slice header / slice group header / slice / CTU / CU.
[0574] Example 9:
[0575] In some embodiments, LIC may be used exclusively with weighted prediction at the block level.
[0576] In some embodiments, when the current block is bi-directionally predicted, the signaling of LIC information may depend on weighted_bipred_flag.
[0577] In some embodiments, when the current block is uni-directionally predicted, the signaling of LIC information may depend on weighted_pred_flag.
[0578] In some embodiments, the signaling of LIC information may depend on weighted prediction parameters associated with one or all reference pictures associated with the current block.
[0579] In some embodiments, if weighted prediction is enabled for some or all reference pictures of a block, LIC may be disabled for the block and syntax elements related to LIC are not signaled.
[0580] In some embodiments, even if weighted prediction is enabled for some or all reference pictures of a block, LIC may still be applied and weighted prediction may be disabled for the block. In some embodiments, LIC may be applied to reference pictures to which weighted prediction is not applied and may be disabled on reference pictures to which weighted prediction is applied.
[0581] In some embodiments, LIC may be used together with weighted prediction for a block.
[0582] Example 10:
[0583] In some embodiments, LIC may be used exclusively with weighted prediction at picture level / slice level / slice group level / CTU group level.
[0584] In some embodiments, if weighted prediction is enabled for some or all reference pictures of a picture / slice / slice group / CTU group, LIC is disabled and all related syntax elements are not signaled.
[0585] In some embodiments, if LIC is enabled for a picture / slice / slice group / CTU group, weighted prediction is disabled for all its reference pictures and all related syntax elements are not signaled.
[0586] In some embodiments, LIC may be used together with weighted prediction.
[0587] Example 11:
[0588] In some embodiments, LIC can be used exclusively with CPR mode. In some embodiments, when CPR mode is enabled for a block, signaling of an indication of LIC use and / or side information can be skipped. In some embodiments, when LIC mode is enabled for a block, signaling of an indication of CPR use and / or side information can be skipped.
[0589] Example 12:
[0590] In some embodiments, LIC or / and GBI or / and weighted prediction may be disabled in pairwise prediction or combined bi-prediction or other types of virtual / artificial candidates (eg, zero motion vector candidates).
[0591] In some embodiments, if one of the two candidates involved in the pairwise prediction or combined bidirectional prediction adopts LIC prediction and neither of the two candidates adopts weighted prediction or GBI, LIC may be enabled for the pairwise or combined bidirectional Merge candidate.
[0592] In some embodiments, if both candidates involved in the pairwise prediction or combined bidirectional prediction adopt LIC prediction, LIC may be enabled for the pairwise or combined bidirectional Merge candidate.
[0593] In some embodiments, if one of the two candidates involved in the pairwise prediction or combined bidirectional prediction adopts weighted prediction and neither of the two candidates adopts LIC prediction or GBI, weighted prediction can be enabled for the pairwise or combined bidirectional Merge candidate.
[0594] In some embodiments, if both candidates involved in the pairwise prediction or combined bidirectional prediction adopt weighted prediction, weighted prediction may be enabled for the pairwise or combined bidirectional Merge candidate.
[0595] In some embodiments, if one of the two candidates involved in the pairwise prediction or combined bidirectional prediction adopts GBI and neither of the two candidates adopts LIC prediction or weighted prediction, GBI may be enabled for the pairwise or combined bidirectional Merge candidate.
[0596] In some embodiments, if two candidates involved in a pairwise prediction or combined bidirectional prediction adopt GBI and the GBI index is the same for the two candidates, GBI may be enabled for the pairwise or combined bidirectional Merge candidate.
[0597] Example 13:
[0598] In some embodiments, LIC prediction or / and weighted prediction or / and GBI may be considered in the deblocking filter.
[0599] In some embodiments, the deblocking filtering process may be applied even if two blocks near a boundary have the same motion but different LIC parameters / weighted prediction parameters / GBI indices.
[0600] In some embodiments, if two blocks (near a boundary) have different GBI indices, a stronger boundary filter strength may be assigned.
[0601] In some embodiments, if two blocks have different LIC flags, a stronger boundary filtering strength may be assigned.
[0602] In some embodiments, if two blocks adopt different LIC parameters, a stronger boundary filtering strength may be assigned.
[0603] In some embodiments, if one block adopts weighted prediction while the other block does not adopt weighted prediction, or the two blocks adopt different weighting factors (eg, predicted from reference pictures with different weighting factors), a stronger boundary filtering strength may be assigned.
[0604] In some embodiments, if two blocks are near a boundary, one coded in LIC mode and the other not coded in LIC mode, deblocking filtering may be applied even if the motion vectors are the same.
[0605] In some embodiments, if two blocks are near a boundary, one block is coded in weighted prediction mode and the other block is not coded in weighted prediction mode, deblocking filtering may be applied even if the motion vectors are the same.
[0606] In some embodiments, if two blocks near a boundary, one block is coded with GBI enabled (e.g., unequal weight) and the other block is not coded with GBI enabled, deblocking filtering may be applied even if the motion vectors are the same.
[0607] Example 14:
[0608] In some embodiments, affine mode / affine parameters / affine type may be considered in deblocking filtering.
[0609] In some embodiments, whether and / or how to apply a deblocking filter may depend on the affine mode / affine parameters.
[0610] In some embodiments, if two blocks (near the boundary) have different affine patterns, a stronger boundary filter strength may be assigned.
[0611] In some embodiments, if two blocks (near the boundary) have different affine parameters, a stronger boundary filter strength may be assigned.
[0612] In some embodiments, if two blocks are near a boundary, one coded in affine mode and the other not coded in affine mode, deblocking filtering may be applied even if the motion vectors are the same.
[0613] Example 15:
[0614] Examples 1-14 may also be applicable to other types of filtering processes.
[0615] Example 16:
[0616] In some embodiments, the CIIP flag may be stored in a history-based motion vector prediction (HMVP) table along with the motion information.
[0617] In some embodiments, when comparing two candidate motion information, the CIIP flag is considered in the comparison.
[0618] In some embodiments, when comparing two candidate motion information, the CIIP flag is not considered in the comparison.
[0619] In some embodiments, when a Merge candidate comes from an entry in the HMVP table, the CIIP flag of the entry is also copied to the Merge candidate.
[0620] Example 17:
[0621] In some embodiments, the LIC flag may be stored in a history-based motion vector prediction (HMVP) table along with the motion information.
[0622] In some embodiments, when comparing two candidate motion information, the LIC flag is considered in the comparison.
[0623] In some embodiments, when comparing two candidate motion information, the LIC flag is not considered in the comparison.
[0624] In some embodiments, when a Merge candidate is from an entry in the HMVP table, the LIC flag of the table is also copied to the Merge candidate.
[0625] Example 18:
[0626] In some embodiments, when performing LIC, the block should be divided into VPDU (such as 64*64 or 32*32 or 16*16) process blocks, and each block performs the LIC process sequentially (which means that the LIC process of one process block may depend on the LIC process of another process block) or performs the LIC process in parallel (which means that the LIC process of one process block does not depend on the LIC process of another process block).
[0627] Example 19:
[0628] In some embodiments, LIC can be used with multi-hypothesis prediction (as described in 2.2.13, 2.2.14, 2.2.15).
[0629] In some embodiments, the LIC flag is explicitly signaled for both multi-hypothesis AMVP and Merge modes (e.g., as described in 2.2.14). In some embodiments, the explicitly signaled LIC flag is applied to both AMVP and Merge modes. In some embodiments, the explicitly signaled LIC flag is applied only to AMVP mode, while the LIC flag for Merge mode is inherited from the corresponding Merge candidate. Different LIC flags can be used for AMVP and Merge modes. Furthermore, different LIC parameters can be derived / inherited for AMVP and Merge modes. In some embodiments, the LIC is always disabled for Merge mode. In some embodiments, the LIC flag is not signaled and is always disabled for AMVP mode. However, for Merge mode, the LIC flag and / or LIC parameters can be inherited or derived.
[0630] In some embodiments, the LIC flag is inherited from the corresponding merge candidate in multi-hypothesis merge mode (e.g., as described in 2.2.15). In some embodiments, the LIC flag is inherited for each of the two selected merge candidates, so different LIC flags can be inherited for the two selected merge candidates. Furthermore, different LIC parameters can be derived / inherited for the two selected merge candidates. In some embodiments, the LIC flag is inherited only for the first selected merge candidate, while LIC is always disabled for the second selected merge candidate.
[0631] In some embodiments, the LIC flag is explicitly signaled for the multi-hypothesis inter prediction mode (e.g., as described in 2.2.13). In some embodiments, if a block is predicted in Merge mode (or UMVE mode) and additional motion information, the explicitly signaled LIC flag may apply to both the Merge mode (or UMVE mode) and the additional motion information. In some embodiments, if a block is predicted in Merge mode (or UMVE mode) and additional motion information, the explicitly signaled LIC flag may apply to the additional motion information. For Merge mode, the LIC flag and / or LIC parameters may be inherited or derived. In some embodiments, LIC is always disabled for Merge mode.
[0632] In some embodiments, if a block is predicted in Merge mode (or UMVE mode) with additional motion information, the LIC flag is not signaled and is disabled for the additional motion information. For Merge mode, the LIC flag and / or LIC parameters may be inherited or derived. In some embodiments, if a block is predicted in AMVP mode with additional motion information, the explicitly signaled LIC flag may apply to both AMVP mode and the additional motion information. In some embodiments, the LIC is always disabled for the additional motion information. In some embodiments, different LIC parameters may be derived / inherited for Merge mode (or UMVE mode) / AMVP mode and the additional motion information.
[0633] In some embodiments, when multiple hypotheses are applied to a block, illumination compensation may be applied to some prediction signals, but not all prediction signals. In some embodiments, when multiple hypotheses are applied to a block, one or more flags may be signaled / derived to indicate the use of illumination compensation for the prediction signal.
[0634] Example 20:
[0635] In some embodiments, the LIC flag may be inherited from the base merge candidate in UMVE mode. In some embodiments, the LIC parameters are implicitly derived as described in 2.2.7. In some embodiments, for boundary blocks encoded and decoded in UMVE mode, the LIC parameters are implicitly derived as described in 2.2.7. In some embodiments, for intra blocks encoded and decoded in UMVE mode, the LIC parameters are inherited from the base merge candidate. In some embodiments, the LIC flag may be explicitly signaled in UMVE mode.
[0636] Example 21:
[0637] The above proposed method or LIC can be applied under certain conditions, such as block size, slice / picture / slice type or motion information.
[0638] In some embodiments, when the block size contains less than M*H samples, e.g., 16 or 32 or 64 luma samples, the proposed method or LIC is not allowed.
[0639] In some embodiments, when the minimum size of the width or / and height of the block is less than or not greater than X, the proposed method or LIC is not allowed. In one example, X is set to 8.
[0640] In some embodiments, when the width of the block > th1 or >= th1 and / or the height of the block > th2 or >= th2, the proposed method or LIC is not allowed. In one example, th1 and / or th2 are set to 8.
[0641] In some embodiments, when the width of the block < th1 or <= th1 and / or the height of the block < th2 or <= th2, the proposed method or LIC is not allowed. In one example, th1 and / or th2 are set to 8.
[0642] In some embodiments, LIC is disabled for the affine inter prediction mode or / and the affine Merge mode.
[0643] In some embodiments, LIC is disabled for sub-block coding / decoding tools (such as ATMVP or / and STMVP or / and the planar motion vector prediction mode).
[0644] In some embodiments, LIC is only applied to certain components. For example, LIC is applied to the luma component. Alternatively, LIC is applied to the chroma component.
[0645] In some embodiments, if the LIC flag is true, then BIO or / and DMVR is disabled.
[0646] In some embodiments, LIC is disabled for bi-predicted blocks.
[0647] In some embodiments, LIC is disabled for intra blocks coded in the AMVP mode.
[0648] In some embodiments, LIC is only allowed for uni-predicted blocks.
[0649] Example 22:
[0650] In some embodiments, the selection of neighboring samples for deriving LIC parameters may depend on coding / decoding information, block shape, etc.
[0651] In some embodiments, if the width >= height or the width > height, then only the upper neighboring pixels are used to derive LIC parameters.
[0652] In some embodiments, if width < height, only the left neighboring pixels are used to derive LIC parameters.
[0653] Example 23:
[0654] In some embodiments, only one LIC flag is signaled for a block with a geometric partitioning structure (such as triangle prediction mode). In this case, all partitions (all PUs) of the block share the same value of the LIC enable flag.
[0655] In some embodiments, for some PUs, LIC may always be disabled regardless of whether the LIC flag is signaled.
[0656] In some embodiments, if the block is partitioned from the top right corner to the bottom left corner, a single LIC parameter set is derived and used for both PUs. In some embodiments, LIC is always disabled for the bottom PU. In some embodiments, if the top PU is coded in Merge mode, the LIC flag is not signaled.
[0657] In some embodiments, if the block is partitioned from the top left corner to the bottom right corner, LIC parameters are derived for each PU. In some embodiments, the top neighboring samples of the block are used to derive the LIC parameters for the top PU, and the left neighboring samples of the block are used to derive the LIC parameters for the left PU. In some embodiments, one LIC parameter set is derived and used for both PUs.
[0658] In some embodiments, if both PUs are coded in Merge mode, the LIC flag is not signaled and may be inherited from the Merge candidate. LIC parameters may be derived or inherited.
[0659] In some embodiments, if one PU is coded in AMVP mode and another PU is coded in Merge mode, the signaled LIC flag may apply only to the PU coded in AMVP mode. The PU coded in Merge mode inherits the LIC flag and / or LIC parameters. In some embodiments, the LIC flag is not signaled and is disabled for the PU coded in AMVP mode. The LIC flag and / or LIC parameters are inherited by the PU coded in Merge mode. In some embodiments, the LIC is disabled for the PU coded in Merge mode.
[0660] In some embodiments, if both PUs are coded in AMVP mode, one LIC flag may be signaled for each PU.
[0661] In some embodiments, a PU may utilize reconstructed samples from another PU within the current block that have already been reconstructed to derive LIC parameters.
[0662] Example 24:
[0663] In some embodiments, LIC may be performed on a portion of pixels rather than the entire block.
[0664] In some embodiments, LIC may be performed only on pixels around the block boundary, while not on other pixels within the block.
[0665] In some embodiments, LIC is performed on the top W×N rows or N×H columns, where N is an integer and W and H represent the width and height of the block. For example, N is equal to 4.
[0666] In some embodiments, LIC is performed on the top left (Wm)×(Hn) region, where m and n are integers, and W and H represent the width and height of the block. For example, m and n are equal to 4.
[0667] Example 25:
[0668] In some embodiments, the LIC flag may be used to update the HMVP table.
[0669] In some embodiments, the HMVP candidate may include a LIC flag in addition to the motion vector and other information.
[0670] In some embodiments, the LIC flag is not used to update the HMVP table. For example, a default LIC flag value is set for each HMVP candidate.
[0671] Example 26:
[0672] In some embodiments, if a block is coded with LIC, the updating of the HMVP table may be skipped. In some embodiments, the LIC coded block may also be used to update the HMVP table.
[0673] Example 27:
[0674] In some embodiments, candidate motions may be reordered according to the LIC flags.
[0675] In some embodiments, LIC-enabled candidates may be placed before all or part of LIC-disabled candidates in a Merge / AMVP or other motion candidate list.
[0676] In some embodiments, LIC-disabled candidates may be placed before all or part of LIC-enabled candidates in a Merge / AMVP or other motion candidate list.
[0677] In some embodiments, the Merge / AMVP or other motion candidate list construction process may be different for LIC codec blocks and non-LIC codec blocks.
[0678] In some embodiments, for LIC coded blocks, the Merge / AMVP or other motion candidate lists may not contain motion candidates derived from non-LIC coded spatial / temporal neighboring blocks or non-neighboring blocks or HMVP candidates with LIC flag equal to false.
[0679] In some embodiments, for non-LIC coded blocks, the Merge / AMVP or other motion candidate lists may not contain motion candidates derived from LIC coded spatial / temporal neighboring blocks or non-neighboring blocks or HMVP candidates with LIC flag equal to true.
[0680] Example 28:
[0681] In some embodiments, whether the above methods are enabled or disabled may be signaled in the SPS / PPS / VPS / sequence header / picture header / slice header / slice group header / CTU group, etc. In some embodiments, which method to use may be signaled in the SPS / PPS / VPS / sequence header / picture header / slice header / slice group header / CTU group, etc. In some embodiments, whether the above methods are enabled or disabled and / or which method to apply may depend on block size, video processing data unit (VPDU), picture type, low delay check flag, codec information of the current block (such as reference picture, unidirectional prediction or bidirectional prediction), or a previously coded block.
[0682] 6. Additional Embodiments
[0683] Method #1 :
[0684] The PPS syntax element weight_bipred_flag is checked to determine whether the LIC_enabled_flag is signaled at the current CU. The syntax signaling is modified in bold as follows:
[0685]
[0686] The first option is to disable LIC for all reference pictures of the current slice / slice / slice group, e.g.
[0687] 1. If the current slice references a PPS with weighted_bipred_flag set to 1 and the current block is bidirectionally predicted;
[0688] 2. If the current slice references a PPS with weighted_pred_flag set to 1 and the current block is uni-directionally predicted.
[0689] Method #2 :
[0690] For the bidirectional prediction case:
[0691] LIC is disabled only when both reference pictures used in bidirectional prediction have weighted prediction turned on, i.e., the (weight, offset) parameters of these reference pictures have non-default values. This allows bidirectional prediction CUs using reference pictures with default WP parameters (i.e., WP is not called for these CUs) to still use LIC.
[0692] For the one-way prediction case:
[0693] LIC is disabled only when weighted prediction is enabled for the reference pictures used in unidirectional prediction, i.e., the (weight, offset) parameters of these reference pictures have non-default values. This allows unidirectional prediction CUs using reference pictures with default WP parameters (i.e., WP is not invoked for these CUs) to still use LIC. The syntax signaling is modified in bold as follows:
[0694]
[0695] In the above table, X represents the reference picture list (X is 0 or 1).
[0696] For the above example, the following might apply:
[0697] 1. The sps_LIC_enabled_flag that controls the use of LIC per sequence may be replaced by an indication of the use of LIC per picture / view / slice / slice group / CTU row / region / multiple CTUs / CTUs.
[0698] 2. Control of the signaling of LIC_enabled_flag may depend on the block size.
[0699] 3. The control of the signaling of LIC_enabled_flag may depend on the gbi index.
[0700] 4. The control of the signaling of the gbi index may also depend on the LIC_enabled_flag.
[0701] Figure 39 is a block diagram illustrating an example of the architecture of a computer system or other control device 2600 that can be used to implement various portions of the presently disclosed technology. Figure 39In the embodiment of the present invention, computer system 2600 includes one or more processors 2605 and memory 2610 connected via interconnect 2625. Interconnect 2625 can represent any one or more separate physical buses, point-to-point connections, or both connected through appropriate bridges, adapters, or controllers. Thus, interconnect 2625 can include, for example, a system bus, a Peripheral Component Interconnect (PCI) bus, a HyperTransport or Industry Standard Architecture (ISA) bus, a Small Computer System Interface (SCSI) bus, a Universal Serial Bus (USB), an IIC (I2C) bus, or an Institute of Electrical and Electronics Engineers (IEEE) standard 674 bus, sometimes referred to as "FireWire."
[0702] The processor(s) 2605 may include a central processing unit (CPU) to control the overall operation of, for example, a host computer. In some embodiments, the processor(s) 2605 achieve this by running software or firmware stored in the memory 2610. The processor(s) 2605 may be or include one or more programmable general-purpose or special-purpose microprocessors, digital signal processors (DSPs), programmable controllers, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), etc., or a combination of these devices.
[0703] Memory 2610 may be or include the main memory of a computer system. Memory 2610 represents any suitable form of random access memory (RAM), read-only memory (ROM), flash memory, or the like, or a combination of these devices. When used, memory 2610 may, among other things, contain a set of machine instructions that, when executed by processor 2605, causes processor 2605 to perform operations to implement embodiments of the presently disclosed technology.
[0704] Also, an (optional) network adapter 2615 is connected to the processor(s) 2605 via the interconnect 2625. The network adapter 2615 provides the computer system 2600 with the ability to communicate with remote devices (such as storage clients and / or other storage servers) and may be, for example, an Ethernet adapter or a Fibre Channel adapter.
[0705] Figure 40A block diagram of an example embodiment of a device 2700 that can be used to implement various parts of the technology disclosed herein is shown. The mobile device 2700 can be a laptop, a smartphone, a tablet, a camcorder, or other type of device capable of processing video. The mobile device 2700 includes a processor or controller 2701 that processes data, and a memory 2702 that communicates with the processor 2701 to store and / or buffer data. For example, the processor 2701 may include a central processing unit (CPU) or a microcontroller unit (MCU). In some embodiments, the processor 2701 may include a field-programmable gate array (FPGA). In some embodiments, the mobile device 2700 includes a graphics processing unit (GPU), a video processing unit (VPU), and / or a wireless communication unit, or communicates with a GPU, VPU, and / or wireless communication unit for various visual and / or communication data processing functions of the smartphone device. For example, the memory 2702 may include and store processor-executable code that, when executed by the processor 2701, configures the mobile device 2700 to perform various operations, such as receiving information, commands, and / or data, processing the information and data, and sending or providing the processed information / data to another device, such as an actuator or an external display. To support the various functions of the mobile device 2700, the memory 2702 may store information and data, such as instructions, software, values, images, and other data processed or referenced by the processor 2701. For example, various types of random access memory (RAM) devices, read-only memory (RAM) devices, flash memory devices, and other suitable storage media may be used to implement the storage functions of the memory 2702. In some embodiments, the mobile device 2700 includes an input / output (I / O) unit 2703 to connect the processor 2701 and / or the memory 2702 to other modules, units, or devices. For example, the I / O unit 2703 can interface the processor 2701 and the memory 2702 using various types of wireless interfaces compatible with typical data communication standards, such as between one or more computers in the cloud and a user device. In some embodiments, the mobile device 2700 can interface with other devices using wired connections via the I / O unit 2703. The mobile device 2700 can also interface with other external interfaces, such as data storage devices and / or visual or audio display devices 2704, to retrieve and transmit data and information that can be processed by the processor, stored in the memory, or presented on the display device 2704 or an output unit of an external device.For example, display device 2704 may display video frames modified based on MVP according to the disclosed techniques.
[0706] Figure 41 4100 is a flowchart representation of a method for video processing. The method 4100 includes determining (4102) that the video block is a boundary block of a codec tree unit (CTU) in which the video block is located and thus enabling a local illumination compensation (LIC) codec for the video block, deriving (4104) parameters for local illumination compensation (LIC) for the video block based on the determination that the LIC codec is enabled for the video block, and performing (4106) the conversion by adjusting pixel values of the video block using the LIC. For example, Section 5, Item 1 discloses some example variations and embodiments of the method 4100.
[0707] The following clause-based list describes certain features and aspects of the disclosed technology listed in Section 5.
[0708] 1. A video processing method, comprising:
[0709] In converting between the video block and a bitstream representation of the video block, determining that the video block is a boundary block of a codec tree unit (CTU) in which the video block is located, and thus enabling a local illumination compensation (LIC) codec tool for the video block;
[0710] deriving parameters of local illumination compensation (LIC) for the video block based on determining that the LIC codec is enabled for the video block; and
[0711] The conversion is performed by adjusting the pixel values of the video block using the LIC.
[0712] 2. The method of clause 1, wherein the conversion generates video blocks from a bitstream representation.
[0713] 3. The method of clause 1 , wherein the conversion generates a bitstream representation from the video blocks.
[0714] 4. A method according to any of clauses 1 to 3, wherein the derivation uses samples of neighbouring blocks depending on the position of the current block within the CTU.
[0715] 5. A method according to clause 4, wherein (1) when the current block is located at the left boundary of a CTU, the derivation uses only left reconstructed neighboring samples, or (2) when the current block is located at the upper boundary of a CTU, the derivation uses only upper reconstructed neighboring samples, or (3) when the current block is located at the upper left corner of a CTU, the derivation uses left and / or upper reconstructed neighboring samples, or (4) when the current block is located at the upper right corner of a CTU, the derivation uses right and / or upper reconstructed neighboring samples, or (5) when the current block is located at the lower left corner of a CTU, the derivation uses left and / or lower reconstructed neighboring samples.
[0716] 6. A video processing method, comprising:
[0717] In converting between the video block and a bitstream representation of the video block, determining that the video block is an internal block of a codec tree unit (CTU) in which the video block is located, and thus disabling a local illumination compensation (LIC) codec tool for the video block;
[0718] Inherit the parameters of the LIC of the video block; and
[0719] The conversion is performed by adjusting the pixel values of the video block using the LIC.
[0720] 7. The method of clause 6, wherein the inheritance comprises:
[0721] Maintaining a lookup table (LUT) of LIC parameters previously used for other blocks; and
[0722] LIC parameters for the video block are derived based on one or more LIC parameters from a lookup table.
[0723] 8. The method of clause 7, wherein the other blocks comprise blocks from a CTU or blocks from a reference picture for the current block.
[0724] 9. A method according to any of clauses 7 to 9, wherein the lookup table is cleared and rebuilt for each CTU or CTU row or slice or slice group or picture of the video block.
[0725] 10. A method according to clause 7, wherein the current block uses Advanced Motion Vector Prediction (AMVP) mode or Affine Inter mode, and the bitstream representation is configured to include a flag indicating which LIC parameters are used for the current block.
[0726] 11. The method of clause 10, wherein the flag indicates an index of an entry in the LUT.
[0727] 12. A method according to any of clauses 7 to 11, wherein a plurality of LUTs are used.
[0728] 13. A method according to clause 12, wherein at least one LUT is maintained for each reference picture or each reference picture list used in the transformation of the current block.
[0729] 14. The method of clause 7, wherein the current block uses Merge mode or Affine Merge mode for the spatial or temporal Merge candidate motion vector, and wherein the inheritance comprises inheriting one or both of a LIC flag and a LIC parameter of a neighboring block.
[0730] 15. The method of clause 7, wherein the current block uses a merge mode or an affine merge mode for a spatial or temporal merge candidate motion vector, and wherein the inheritance comprises configuring the bitstream with one or more difference values, and calculating the LIC flag or the LIC parameter based on the one or more difference values and one or both of the LIC flag and the LIC parameter of a neighboring block.
[0731] 16. A method according to clause 1, wherein the current block is encoded using Advanced Motion Vector Prediction (AMVP) or Affine Inter mode, and wherein the parameters are derived using least squares error minimization, in which a sub-set of neighboring samples is used for error minimization.
[0732] 17. A video processing method, comprising:
[0733] In converting between the video block and the bitstream representation of the video block, determining that both local illumination compensation and intra block copy codecs are enabled for use with the current block; and
[0734] The conversion is performed by performing local illumination compensation (LIC) and intra block copy operations on the video blocks.
[0735] 18. The method of clause 17, wherein a LIC flag in the bitstream is configured to indicate that LIC is enabled for the current block.
[0736] 19. The method of clause 18, wherein the converting generates video blocks from the bitstream representation.
[0737] 20. The method of clause 18, wherein the converting generates a bitstream representation from the video blocks.
[0738] 21. The method of clause 17, wherein the current block uses Merge mode, and wherein the current block inherits the LIC flag value of the adjacent block.
[0739] 22. The method of clause 17, wherein the current block uses Merge mode, and wherein the current block inherits LIC parameters of a neighboring block.
[0740] 23. A video processing method, comprising:
[0741] During conversion between bitstreams of a video including multiple pictures having multiple blocks, a local illumination compensation (LIC) mode for a current block of the video is determined based on LIC mode rules; and conversion between the current block and a corresponding bitstream representation of the current block is performed.
[0742] 24. The method of clause 23, wherein the LIC mode rule specifies the use of a LIC mode that does not include inter prediction, diffusion filter, bilateral filter, overlapped block motion compensation, or tools that modify the inter prediction signal of the current block.
[0743] 25. The method of clause 24, wherein the LIC mode rules are explicitly signaled in the bitstream.
[0744] 26. The method of clause 23, wherein the LIC mode rule specifies that the LIC mode is to be used only if the current block does not use a generalized bi-predictive (GBI) codec mode.
[0745] 27. A method according to any of clauses 23 or 26, wherein the LIC mode rules are explicitly signalled in the bitstream and the GBI indication is omitted from the bitstream.
[0746] 28. The method of clause 23, wherein the LIC mode rule specifies that the LIC mode is to be used only if the current block does not use the multiple hypothesis coding mode.
[0747] 29. A method according to any of clauses 23 or 28, wherein the LIC mode rules are explicitly signalled in the bitstream and the multi-hypothesis codec mode is implicitly disabled.
[0748] 30. The method of any of clauses 28 to 29, wherein the LIC mode rule specifies that LIC is to be applied to a selected subset of the prediction signals of the multiple hypothesis mode.
[0749] 31. The method of clause 23, wherein the LIC mode rule specifies inheritance of the LIC flag from a base Merge candidate of the current block, the base Merge candidate also using the Final Motion Vector Expression (UMVE) mode.
[0750] 32. The method of clause 31, wherein the current block inherits its LIC flag from a base Merge candidate of the UMVE mode.
[0751] 33. A method according to clause 32, wherein the current block inherits using the calculation in section 2.2.7.
[0752] 34. A method according to any of clauses 1 to 33, wherein LIC is used only if the current block also satisfies a condition related to the size or slice type or slice type or picture type of the current block or the type of motion information associated with the current block.
[0753] 35. A method according to clause 34, wherein the condition excludes block sizes smaller than M*H samples, where M and H are pre-specified integer values.
[0754] 36. The method of clause 35, wherein M*H is equal to 16 or 32 or 64.
[0755] 37. The method of clause 34, wherein the condition specifies that the width or height of the current block is less than or not greater than X, where X is an integer.
[0756] 38. The method of clause 37, wherein X=8.
[0757] 39. A method according to clause 34, wherein the condition specifies that the current block corresponds to a luma sample.
[0758] 40. The method of clause 34, wherein the condition specifies that the current block corresponds to a chroma sample.
[0759] 41. A video processing method, comprising:
[0760] During conversion between a video block of video and a bitstream representation of the video, determining local illumination compensation (LIC) parameters for the video block using at least some samples of neighboring blocks of the video block; and
[0761] Conversion between video blocks and bitstream representations is performed by performing LIC using the determined parameters.
[0762] 42. The method of clause 41, wherein the converting generates video blocks from the bitstream representation.
[0763] 43. The method of clause 41, wherein the converting generates a bitstream representation from the video blocks.
[0764] 44. A method according to any of clauses 41 to 43, wherein identifying at least some samples of the neighbouring blocks depends on codec information of the current block or on the shape of the current block.
[0765] 45. A method according to clause 44, wherein the width of the current block is greater than or equal to the height of the current block, and thus only upper neighboring pixels are used to derive the LIC parameters.
[0766] 46. A video processing method, comprising:
[0767] Conversion between video and a bitstream representation of the video is performed, wherein the video is represented as video frames comprising video blocks and local illumination compensation (LIC) is enabled only for video blocks using a geometric prediction structure comprising a triangle prediction mode.
[0768] 47. The method of clause 46, wherein the converting generates video blocks from the bitstream representation.
[0769] 48. The method of clause 46, wherein the converting generates a bitstream representation from the video blocks.
[0770] 49. A method according to any of clauses 46 to 48, wherein, for a current block with LIC enabled, all prediction unit partitions share the same LIC flag value.
[0771] 50. A method according to any of clauses 46 to 48, wherein, for a current block partitioned from the upper right corner to the lower left corner, a single LIC parameter set is used for both partitions of the current block.
[0772] 51. A method according to any of clauses 46 to 48, wherein, for a current block partitioned from the upper left corner to the lower right corner, a single LIC parameter set is used for each partition of the current block.
[0773] 52. A video processing method, comprising:
[0774] A conversion is performed between a video and a bitstream representation of the video, wherein the video is represented as a video frame comprising video blocks, and in the conversion to its corresponding bitstream representation, local illumination compensation (LIC) is performed on less than all pixels of a current block.
[0775] 53. The method of clause 52, wherein the converting generates video blocks from the bitstream representation.
[0776] 54. The method of clause 52, wherein the converting generates a bitstream representation from the video blocks.
[0777] 55. A method according to any of clauses 52 to 54, wherein LIC is performed only on pixels on the boundary of the current block.
[0778] 56. A method according to any of clauses 52 to 54, wherein LIC is performed only on pixels in a top portion of the current block.
[0779] 57. A method according to any of clauses 52 to 54, wherein LIC is performed only on pixels in the right part of the current block.
[0780] 58. A method for video processing, comprising:
[0781] In converting between a video block and a bitstream representation of the video block, determining that both local illumination compensation (LIC) and generalized bi-prediction (GBi) or multi-hypothesis inter prediction codecs are enabled for use with the current block; and
[0782] The conversion is performed by performing LIC and GBi or multi-hypothesis inter prediction operations on the video blocks.
[0783] 59. A method as described in clause 58, wherein information associated with the GBi codec tool is signalled subsequent to information associated with the LIC codec tool.
[0784] 60. A method as recited in clause 58, wherein information associated with the LIC codec is signalled subsequent to information associated with the GBi codec.
[0785] 61. A method according to clause 59 or 60, wherein the information is signaled in a sequence parameter set (SPS), a picture parameter set (PPS), a slice header, a slice group header, a slice, a codec unit (CU) or a codec tree unit (CTU).
[0786] 62. The method of clause 58, wherein the LIC flag is true, and wherein the bitstream representation does not include one or more syntax elements associated with a GBi or multi-hypothesis inter prediction codec.
[0787] 63. The method of clause 58, wherein the GBi flag is true, and wherein the bitstream representation does not include one or more syntax elements associated with a LIC or multiple hypothesis inter prediction codec.
[0788] 64. The method of clause 58, wherein a multi-hypothesis inter prediction (LIC) flag is true, and wherein the bitstream representation does not include one or more syntax elements associated with the LIC or GBi codec.
[0789] 65. A method for video processing, comprising:
[0790] In converting between a video block and a bitstream representation of the video block, determining that both local illumination compensation (LIC) and combined inter-intra prediction (CIIP) codecs are enabled for use with the current block; and
[0791] The conversion is performed by performing LIC and CIIP operations on the video blocks.
[0792] 66. A method as described in clause 65, wherein information associated with the CIIP codec tool is signaled after information associated with the LIC codec tool.
[0793] 67. A method as recited in clause 65, wherein information associated with the LIC codec tool is signaled subsequent to information associated with the CIIP codec tool.
[0794] 68. A method according to clause 66 or 67, wherein the information is signalled in a sequence parameter set (SPS), a picture parameter set (PPS), a slice header, a slice group header, a slice, a codec unit (CU) or a codec tree unit (CTU).
[0795] 69. The method of clause 65, wherein the LIC flag is true, and wherein the bitstream representation does not include one or more syntax elements associated with the CIIP codec tool.
[0796] 70. The method of clause 65, wherein the CIIP flag is true, and wherein the bitstream representation does not include one or more syntax elements associated with the LIC codec tool.
[0797] 71. A method according to any of clauses 1 to 70, wherein the bitstream is configured with a bit field that controls the operation of the conversion.
[0798] 72. The method of clause 71, wherein the bit field is included at a sequence level, a picture level, a video level, a sequence header level, a picture header level, a slice header level, a slice level, or a codec tree unit group level.
[0799] 73. A video encoding apparatus comprising a processor configured to implement the method according to any one or more of clauses 1 to 72.
[0800] 74. A video decoding apparatus comprising a processor configured to implement the method according to any one or more of clauses 1 to 72.
[0801] 75. A computer-readable medium having stored thereon code which, when executed, causes a processor to perform the method of any one or more of clauses 1 to 72.
[0802] Figure 424200 is a flow chart of a method 4200 for video processing. The method 4200 includes constructing (4210) a motion candidate list having at least one motion candidate for conversion between a current block of video and a bitstream representation of the video. The method 4200 includes determining (4220) whether a generalized bidirectional prediction (GBI) processing tool is enabled based on a candidate type of a first motion candidate. The GBI processing tool includes deriving a final prediction based on applying equal or unequal weights to predictions derived from different reference lists according to a weight set. The method 4200 also includes performing (4230) the conversion based on the determination.
[0803] In some embodiments, the motion candidate list includes a merge candidate list. In some embodiments, the first motion candidate is a pairwise average merge candidate or a zero motion vector merge candidate. In some embodiments, the GBI processing tool is disabled. In some embodiments, equal weight is applied to each prediction derived from a different reference picture list. In some embodiments, the pairwise average merge candidate is generated by averaging predefined existing candidate pairs in the merge candidate list.
[0804] Figure 43 4300 is a flowchart representation of a method 4300 for video processing. The method 4300 includes determining (4310) an operational state of local illumination compensation (LIC) prediction, generalized bidirectional (GBI) prediction, or weighted prediction for a conversion between a current block of video and a bitstream representation of the video. The current block is associated with a prediction technique that uses at least one virtual motion candidate. The method 4300 also includes performing (4320) the conversion based on the determination.
[0805] In some embodiments, the prediction technique comprises at least one of pairwise prediction or combined bi-prediction.In some embodiments, the at least one virtual motion candidate comprises a zero motion vector candidate.
[0806] In some embodiments, the operating state indicates that at least one of LIC prediction, GBI prediction, or weighted prediction is disabled. In some embodiments, two motion candidates are associated with prediction techniques. In some embodiments, the operating state of LIC prediction indicates that LIC prediction is enabled when one of the two motion candidates uses LIC prediction and neither motion candidate uses GBI prediction or weighted prediction. In some embodiments, the operating state of LIC prediction indicates that LIC prediction is enabled when both motion candidates use LIC prediction. In some embodiments, the operating state of weighted prediction indicates that weighted prediction is enabled when one of the two motion candidates uses weighted prediction and neither motion candidate uses LIC prediction or GBI prediction. In some embodiments, the operating state of weighted prediction indicates that weighted prediction is enabled when both motion candidates use weighted prediction. In some embodiments, the operating state of GBI prediction indicates that GBI prediction is enabled when one of the two motion candidates uses GBI prediction and neither motion candidate uses LIC prediction or weighted prediction. In some embodiments, the operating state of GBI prediction indicates that GBI prediction is enabled when both motion candidates use GBI prediction and use the same GBI index. Item 12 in the previous section also describes additional variations and implementations of method 4300.
[0807] Figure 44 4400 is a flowchart representation of a method 4400 for video processing. The method 4400 includes, for converting between a current block of video and a bitstream representation of the video, determining (4410) a state of a filtering operation or one or more parameters of a filtering operation on the current block based on characteristics of a local illumination compensation (LIC) prediction, a generalized bidirectional (GBI) prediction, a weighted prediction, or an affine motion prediction for the current block. The method 4400 includes performing (4420) the conversion based on the determination.
[0808] In some embodiments, the characteristics of the LIC prediction, GBI prediction, or weighted prediction include at least a flag for the LIC prediction, parameters for the weighted prediction, or an index for the GBI prediction. In some embodiments, the filtering operation includes a deblocking filter. In some embodiments, one or more parameters of the filtering operation include a boundary filter strength for the deblocking filter. In some embodiments, the boundary filter strength is increased when the GBI index of the current block differs from the GBI index of the neighboring block. In some embodiments, the boundary filter strength is increased when the LIC flag of the current block differs from the LIC flag of the neighboring block. In some embodiments, the boundary filter strength is increased when parameters of the LIC prediction of the current block differ from parameters of the LIC prediction of the neighboring block. In some embodiments, the boundary filter strength is increased when the current block uses a different weighting factor for weighted prediction than the neighboring block. In some embodiments, the boundary filter strength is increased when only one of the current block and the neighboring block uses weighted prediction. In some embodiments, the boundary filter strength is increased when parameters of the affine prediction of the current block differ from parameters of the affine prediction of the neighboring block. In some embodiments, the boundary filter strength is increased when only one of the current block and the neighboring block uses affine motion prediction.
[0809] In some embodiments, the current block and the neighboring blocks have the same motion vector. In some embodiments, the state of the filtering operation indicates that a deblocking filter is applied when characteristics of the LIC prediction, GBI prediction, or weighted prediction of the current block differ from characteristics of the LIC prediction, GBI prediction, or weighted prediction of the neighboring blocks. In some embodiments, the state of the filtering operation indicates that a deblocking filter is applied when only one of the current block and the neighboring blocks employs LIC prediction. In some embodiments, the state of the filtering operation indicates that a deblocking filter is applied when only one of the current block and the neighboring blocks employs weighted prediction. In some embodiments, the state of the filtering operation indicates that a deblocking filter is applied when only one of the current block and the neighboring blocks employs GBI prediction. In some embodiments, the state of the filtering operation indicates that a deblocking filter is applied when only one of the current block and the neighboring blocks employs affine prediction.
[0810] In some embodiments, the conversion generates the current block from the bitstream representation. In some embodiments, the conversion generates the bitstream representation from the current block. Items 13-15 in the previous sections also describe additional variations and implementations of method 4400.
[0811] The disclosed and other embodiments, modules, and functional operations described in this document may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in a combination of one or more thereof. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium, for execution by a data processing apparatus or to control the operation of the data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter that implements a machine-readable propagated signal, or a combination of one or more thereof. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an operating environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.
[0812] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one or more computers located at one site or distributed across multiple sites and interconnected by a communications network.
[0813] The processes and logic flows described herein can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0814] By way of example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated into, special-purpose logic circuitry.
[0815] Although this patent document contains many details, these should not be construed as limitations on any invention or the scope of what is claimed, but rather as descriptions of features specific to particular embodiments of particular inventions. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments, alone or in any suitable subcombination. Furthermore, although the features described above may be described as functioning in certain combinations, and even initially claimed to be so protected, in some cases one or more features in the combination may be deleted from the claimed combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.
[0816] Similarly, while operations may be depicted in a particular order in the drawings, this should not be understood as requiring that these operations be performed in the particular order shown, or in sequential order, or that all illustrated operations be performed, in order to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0817] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A video processing method, comprising: Constructing a motion candidate list having at least one motion candidate for conversion between a current block of a video and a bitstream of the video; selecting a motion candidate from the motion candidate list; as well as performing the conversion based on the selected motion candidate; wherein, for the selected candidate motion, whether to disable a generalized bidirectional prediction processing tool is based on a candidate type of the selected candidate motion, wherein the generalized bidirectional prediction processing tool indicates applying different weights at a codec unit CU level to two prediction blocks generated from reference picture list 0 and reference picture list 1 to derive a final prediction block for the current block, The method further comprises: In response to the selected motion candidate being the first Merge candidate, disabling the generalized bidirectional prediction processing tool for the selected motion candidate, The motion vector of the first Merge candidate corresponding to the first reference list is generated based on the motion vectors of two existing motion candidates added to the Merge candidate list corresponding to the first reference list, and Herein, if the motion vectors of the two existing motion candidates corresponding to the first reference list are available, the motion vector of the first Merge candidate is generated by averaging the motion vectors of the two existing motion candidates corresponding to the first reference list, and if only the motion vector of one existing motion candidate corresponding to the first reference list is available, the motion vector of the first Merge candidate corresponding to the first reference list is generated by using the available motion vector of the existing motion candidate corresponding to the first reference list.
2. The method according to claim 1, wherein The motion candidate list includes a Merge candidate list.
3. The method according to any one of claims 1 to 2, wherein Equal weights at the CU level are applied to the two prediction blocks generated from reference picture list 0 and reference picture list 1 to derive the final prediction block for the current block.
4. The method according to any one of claims 1 to 2, wherein In response to the selected motion candidate being a second Merge candidate, disabling a generalized bi-prediction processing tool for the selected motion candidate, wherein the motion vector of the second Merge candidate is equal to zero.
5. The method according to any one of claims 1-2, wherein the converting generates the current block from a bitstream.
6. The method according to any one of claims 1-2, wherein the converting generates a bitstream based on the current block.
7. The method according to claim 1, further comprising: determining whether a generalized bi-prediction (GBI) processing tool is enabled based on a candidate type of a first motion candidate in the motion candidate list, wherein the GBI processing tool includes deriving a final prediction based on applying equal or unequal weights to predictions derived from different reference lists according to a weight set; and The converting is performed based on the determination.
8. The method according to claim 7, wherein: The motion candidate list includes a Merge candidate list.
9. The method according to claim 8, wherein The first motion candidate is a pairwise average Merge candidate or a zero motion vector Merge candidate.
10. The method according to claim 9, wherein: The GBI processing tool is disabled.
11. The method according to claim 10, wherein: Equal weight is applied to each prediction derived from a different reference picture list.
12. The method according to any one of claims 9 to 11, wherein The pairwise average Merge candidate is generated by averaging predefined existing candidate pairs in the Merge candidate list.
13. The method according to claim 1, further comprising: determining, for the conversion between the current block of the video and the bitstream of the video, an operating state of local illumination compensation (LIC) prediction, generalized bidirectional (GBI) prediction, or weighted prediction, wherein the current block is associated with a prediction technique using at least one virtual motion candidate; and The converting is performed based on the determination.
14. The method according to claim 13, wherein The prediction technique includes at least one of pairwise prediction or combined bidirectional prediction.
15. The method according to any one of claims 13 to 14, wherein The at least one virtual motion candidate includes a zero motion vector candidate.
16. The method according to any one of claims 13 to 14, wherein The operating state indicates that at least one of the LIC prediction, the GBI prediction, or the weighted prediction is disabled.
17. The method according to any one of claims 13 to 14, wherein Two motion candidates are associated with the prediction technique.
18. The method according to claim 17, wherein The operation state of the LIC prediction indicates that the LIC prediction is enabled if one of the two motion candidates adopts the LIC prediction and neither of the two motion candidates adopts the GBI prediction or the weighted prediction.
19. The method according to claim 17, wherein The operation state of the LIC prediction indicates that the LIC prediction is enabled if both of the two motion candidates adopt the LIC prediction.
20. The method according to claim 17, wherein The operation state of the weighted prediction indicates that the weighted prediction is enabled if one of the two motion candidates adopts the weighted prediction and neither of the two motion candidates adopts the LIC prediction or the GBI prediction.
21. The method according to claim 17, wherein The operation state of the weighted prediction indicates that the weighted prediction is enabled if both of the two motion candidates adopt the weighted prediction.
22. The method according to claim 17, wherein The operation state of the GBI prediction indicates that the GBI prediction is enabled if one of the two motion candidates adopts the GBI prediction and neither of the two motion candidates adopts the LIC prediction or the weighted prediction.
23. The method according to claim 17, wherein The operation state of the GBI prediction indicates that the GBI prediction is enabled if both the two motion candidates adopt the GBI prediction and use the same GBI index.
24. The method according to any one of claims 7 to 11, 13-14, wherein The conversion generates the current block from the bitstream.
25. The method according to any one of claims 7 to 11, 13-14, wherein The conversion generates the bitstream from the current block.
26. A video processing apparatus comprising a processor configured to implement the method according to any one of claims 7 to 25.
27. A computer readable medium having stored thereon code which, when executed, causes a processor to implement the method of any one of claims 7 to 25.
28. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein: The instructions, when executed by the processor, cause the processor to: Constructing a motion candidate list having at least one motion candidate for conversion between a current block of a video and a bitstream of the video; selecting a motion candidate from the motion candidate list; as well as performing the conversion based on the selected motion candidate; wherein, for the selected candidate motion, whether to disable a generalized bidirectional prediction processing tool is based on a candidate type of the selected candidate motion, wherein the generalized bidirectional prediction processing tool indicates applying different weights at a codec unit CU level to two prediction blocks generated from reference picture list 0 and reference picture list 1 to derive a final prediction block for the current block, The instructions, when executed by the processor, further cause the processor to: In response to the selected motion candidate being the first Merge candidate, disabling the generalized bidirectional prediction processing tool for the selected motion candidate, The motion vector of the first Merge candidate corresponding to the first reference list is generated based on the motion vectors of two existing motion candidates added to the Merge candidate list corresponding to the first reference list, and Herein, if the motion vectors of the two existing motion candidates corresponding to the first reference list are available, the motion vector of the first Merge candidate is generated by averaging the motion vectors of the two existing motion candidates corresponding to the first reference list, and if only the motion vector of one existing motion candidate corresponding to the first reference list is available, the motion vector of the first Merge candidate corresponding to the first reference list is generated by using the available motion vector of the existing motion candidate corresponding to the first reference list.
29. The apparatus according to claim 28, wherein The motion candidate list includes a Merge candidate list.
30. The device according to any one of claims 28-29, wherein Equal weights at the CU level are applied to the two prediction blocks generated from reference picture list 0 and reference picture list 1 to derive the final prediction block for the current block.
31. The device according to any one of claims 28-29, wherein In response to the selected motion candidate being a second Merge candidate, disabling a generalized bi-prediction processing tool for the selected motion candidate, wherein the motion vector of the second Merge candidate is equal to zero.
32. The device according to any one of claims 28-29, wherein The conversion generates a current block from a bitstream.
33. The apparatus of any one of claims 28-29, wherein the converting generates a bitstream based on the current block.
34. A non-transitory computer-readable storage medium storing instructions that cause a processor to: Constructing a motion candidate list having at least one motion candidate for conversion between a current block of a video and a bitstream of the video; selecting a motion candidate from the motion candidate list; and performing the conversion based on the selected motion candidate; in, for the selected candidate motion, whether to disable a generalized bidirectional prediction processing tool based on the candidate type of the selected candidate motion, wherein the generalized bidirectional prediction processing tool indicates applying different weights at a codec unit (CU) level to two prediction blocks generated from reference picture list 0 and reference picture list 1 to derive a final prediction block for the current block, The instructions further cause the processor to: In response to the selected motion candidate being the first Merge candidate, disabling the generalized bidirectional prediction processing tool for the selected motion candidate, The motion vector of the first Merge candidate corresponding to the first reference list is generated based on the motion vectors of two existing motion candidates added to the Merge candidate list corresponding to the first reference list, and Herein, if the motion vectors of the two existing motion candidates corresponding to the first reference list are available, the motion vector of the first Merge candidate is generated by averaging the motion vectors of the two existing motion candidates corresponding to the first reference list, and if only the motion vector of one existing motion candidate corresponding to the first reference list is available, the motion vector of the first Merge candidate corresponding to the first reference list is generated by using the available motion vector of the existing motion candidate corresponding to the first reference list.
35. The medium of claim 34, wherein The motion candidate list includes a Merge candidate list.
36. The medium according to any one of claims 34-35, wherein In response to the selected motion candidate being a second Merge candidate, disabling a generalized bi-prediction processing tool for the selected motion candidate, wherein the motion vector of the second Merge candidate is equal to zero.
37. A method for storing a bitstream of a video, the method comprising: constructing a motion candidate list having at least one motion candidate; selecting a motion candidate from the motion candidate list; as well as generating a bitstream based on the selected motion candidate; storing the generated bit stream in a non-transitory computer-readable recording medium; wherein, for the selected candidate motion, whether to disable a generalized bidirectional prediction processing tool is based on a candidate type of the selected candidate motion, wherein the generalized bidirectional prediction processing tool indicates applying different weights at a codec unit CU level to two prediction blocks generated from reference picture list 0 and reference picture list 1 to derive a final prediction block for the current block, The method further comprises: In response to the selected motion candidate being the first Merge candidate, disabling the generalized bidirectional prediction processing tool for the selected motion candidate, The motion vector of the first Merge candidate corresponding to the first reference list is generated based on the motion vectors of two existing motion candidates added to the Merge candidate list corresponding to the first reference list, and Herein, if the motion vectors of the two existing motion candidates corresponding to the first reference list are available, the motion vector of the first Merge candidate is generated by averaging the motion vectors of the two existing motion candidates corresponding to the first reference list, and if only the motion vector of one existing motion candidate corresponding to the first reference list is available, the motion vector of the first Merge candidate corresponding to the first reference list is generated by using the available motion vector of the existing motion candidate corresponding to the first reference list.
Citation Information
Patent Citations
Method and Apparatus of Adaptive Bi-Prediction for Video Coding
US20180184117A1
Bidirectional Prediction In Video Compression
US20180332298A1
Method and apparatus for video coding
US20200186818A1