Affine prediction improvements for video coding
Patent Information
- Application Number
- CN202180041405.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-08-22
- Filing Date
- 2021-06-08
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2041-06-08
Smart Images

Figure CN115918080B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This is the national phase of International Patent Application No. PCT / CN2021 / 098811, filed on June 8, 2021, claiming priority and benefits to International Patent Application No. PCT / CN2020 / 094839, filed on June 8, 2020; U.S. Provisional Patent Application No. 63 / 046,634, filed on June 30, 2020; and International Patent Application No. PCT / CN2020 / 110663, filed on August 22, 2020. The entire disclosure of the foregoing applications is incorporated herein by reference as part of the disclosure of this application. Technical Field
[0003] This patent document relates to image and video processing. Background Technology
[0004] Digital video accounts for the largest share of bandwidth usage in the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention
[0005] This document discloses techniques for image and video encoders and decoders that can be used to perform image or video encoding or decoding.
[0006] In one example aspect, a video processing method is disclosed. The method includes: for a conversion between video blocks and a bitstream of a video, determining the gradient of a prediction vector at the sub-block level of the video block according to a rule, wherein the rule specifies the use of the same gradient value assigned to all samples within the sub-blocks of the video block; and performing the conversion based on the determination.
[0007] In another example, a video processing method is disclosed. The method includes: for a conversion between a current video block and a video bitstream, determining a prediction refinement using optical flow techniques after performing a bidirectional prediction technique to obtain motion information of the current video block; and performing the conversion based on the determination.
[0008] In another example, a video processing method is disclosed. The method includes: for a conversion between a current video block and a bitstream of the video, determining whether to apply encoding / decoding tools related to affine prediction to the current video block based on whether the current video block satisfies a condition; and performing the conversion based on the determination, wherein the condition relates to one or more control point motion vectors of the current video block, or a first size of the current video block, or a second size of a sub-block of the current video block, or one or more motion vectors derived for one or more sub-blocks of the current video block, or a prediction mode of the current video block.
[0009] In another example, a video processing method is disclosed. The method includes: for a conversion between a current video block and a bitstream of the video, determining, according to rules, whether and how to apply motion compensation to a sub-block at a sub-block index: (1) whether an affine mode is applied to the current video block, wherein the bitstream includes an affine motion indicator indicating whether an affine mode is applied, (2) the color component to which the current video block belongs, and (3) the color format of the video; and performing the conversion based on the determination.
[0010] In another example, a video processing method is disclosed. The method includes: for a conversion between a current video block and a bitstream of video, determining, according to rules, whether and how to compute predictions of sub-blocks of the current video block based on the following: (1) whether an affine mode is applied to the current video block, wherein the bitstream includes an affine motion indicator indicating whether an affine mode is applied, (2) the color component to which the current video block belongs, and (3) the color format of the video; and performing the conversion based on the determination.
[0011] In another example, a video processing method is disclosed. The method includes: for a conversion between a current video block and a bitstream of video, determining a value of a first syntax element in the bitstream according to a rule, wherein the value of the first syntax element indicates whether an affine mode is applied to the current video block, and wherein the rule specifies that the value of the first syntax element is based on: (1) a second syntax element in the bitstream indicating whether motion compensation based on an affine model is used to generate a predicted sample of the current video block, or (2) a third syntax element in the bitstream indicating whether motion compensation based on a 6-parameter affine model is enabled for a codec layer video sequence (CLVS); and performing the conversion based on the determination.
[0012] In another example, a video processing method is disclosed. The method includes: for a conversion between a current video block and a bitstream of the video, determining whether to enable an interleaving prediction tool for the current video block based on a relationship between a first motion vector of a first sub-block of a first style of the current video block and a second motion vector of a second sub-block of a second style of the current video block; and performing the conversion based on the determination.
[0013] In another example, a different video processing method is disclosed. This method includes: a transformation between video blocks and the codec representation of the video; determining the gradient of the prediction vector at the sub-block level for the video block, based on a rule specifying that the same gradient value is used for all samples in each sub-block; and performing the transformation based on the determination.
[0014] In another example, a different video processing method is disclosed. This method includes: for a conversion between a current video block and a bitstream representation of the video, determining, according to rules, a prediction refinement using motion information employing optical flow techniques; and performing the conversion based on the determination, wherein the rules specify that prediction refinement is performed after performing bidirectional prediction to obtain motion information.
[0015] In another example, a different video processing method is disclosed. This method includes: a conversion between a current video block and a bitstream representation of the video, determining that an encoding / decoding tool is enabled for the current video block because the current video block satisfies a condition; and performing the conversion based on the determination, wherein the condition relates to the control point motion vector of the current video block, the size of the current video block, or the prediction mode of the current video block.
[0016] In another example, a different video processing method is disclosed. This method includes: for a conversion between video blocks and a encoded / decoded representation of the video, determining, according to rules, whether and how to compute predictions for sub-blocks of the video blocks; and performing the conversion based on the determination, wherein the rules are based on one or more of whether an affine mode is enabled for the video block, or the color component to which the video block belongs, or the color format of the video.
[0017] In yet another example, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the methods described above.
[0018] In yet another example, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the methods described above.
[0019] In yet another example, a computer-readable medium on which code is stored is disclosed. This code embodies one of the methods described herein in the form of processor-executable code.
[0020] These and other features are described throughout this document. Attached Figure Description
[0021] Figure 1 An example derivation process for constructing the Merge candidate list is shown.
[0022] Figure 2 Example locations of spatial Merge candidates are shown.
[0023] Figure 3 Candidate pairs are shown that take into account for redundancy checks for spatial Merge candidates.
[0024] Figures 4A-4B Examples of the locations of the second PU divided into N×2N and 2N×N segments are shown.
[0025] Figure 5 This is a diagram illustrating motion vector scaling for time merge candidates.
[0026] Figure 6 Example candidate positions C0 and C1 for time merging are shown.
[0027] Figure 7 An example of combined bidirectional prediction of Merge candidates is shown.
[0028] Figure 8 An example of the derivation process for motion vector prediction candidates is shown.
[0029] Figure 9 A diagram illustrating the scaling of motion vector candidates for spatial motion vectors is shown.
[0030] Figure 10 An example of ATMVP motion prediction for CU is shown.
[0031] Figure 11 An example of a CU with four sub-blocks (AD) and their adjacent blocks (ad) is shown.
[0032] Figure 12 This is an example illustration of a sub-block applying OBMC.
[0033] Figure 13 An example of adjacent samples used to derive IC parameters is shown.
[0034] Figure 14 An example of a simplified affine motion model is shown.
[0035] Figure 15 An example of an affine MVF for each sub-block is shown.
[0036] Figure 16 An example of an MVP for AF_INTER is shown.
[0037] Figure 17 Example candidates for AF_MERGE are shown.
[0038] Figure 18 An example of a bilateral mapping is shown.
[0039] Figure 19 An example of template mapping is shown.
[0040] Figure 20 An example of a one-sided ME in FRUC is shown.
[0041] Figure 21 An example of an optical flow trajectory is shown.
[0042] Figures 22A-22BAn example of BIO with / o block extension is shown.
[0043] Figure 23 An example of an interpolation sample used in BIO is shown.
[0044] Figure 24 An example of the proposed DMVR based on bilateral template matching is shown.
[0045] Figure 25 An example of sub-block MV VSB and pixel Δv(i,j) is shown.
[0046] Figure 26 An example of phase-change level filtering is shown.
[0047] Figure 27 An example illustration of applying an 8-tap horizontal filter is shown.
[0048] Figure 28 An example of non-uniform phase vertical filtering is shown.
[0049] Figure 29 An example of the interleaved prediction process is shown.
[0050] Figure 30 An example of weighted values in a sub-block is shown.
[0051] Figure 31 An example partitioning style is shown.
[0052] Figure 32 An example partitioning style is shown.
[0053] Figure 33 This is a block diagram illustrating a video encoding / decoding system according to some embodiments of the present disclosure.
[0054] Figure 34 This is a block diagram of an example hardware platform used for video processing.
[0055] Figure 35 This is a flowchart of an example method for video processing.
[0056] Figure 36 This is a block diagram illustrating an example video codec system.
[0057] Figure 37 This is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0058] Figure 38 This is a block diagram illustrating a decoder according to some embodiments of the present disclosure.
[0059] Figure 39A and Figure 39BAn example of cache bandwidth limitations for interleaved affine prediction is shown.
[0060] Figure 40 Examples of sub-block rows in style 0 and style 1 are shown.
[0061] Figure 41 An example of selecting SB0 (black solid block) from SB1 (green dashed block) is shown.
[0062] Figure 42 and Figure 43 The positions of sample-style 0 and 1 are shown respectively.
[0063] Figure 44 The sub-blocks in style 1 are shown.
[0064] Figure 45 The sub-blocks in style 1 are shown.
[0065] Figure 46 An array of 4x4 weights is shown.
[0066] Figure 47 An 8x8 array of weights is shown.
[0067] Figures 48 to 54 This is a flowchart of an example method for video processing. Detailed Implementation
[0068] The use of section headings in this document is for ease of understanding and does not limit the applicability of the techniques and embodiments disclosed in each section to that section. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs.
[0069] introduce
[0070] This document relates to video codec technology. Specifically, it concerns motion compensation in video codecs. It can be applied to existing video codec standards, such as HEVC, or standards yet to be finalized (Multi-Functional Video Codec). It may also be applicable to future video codec standards or video codecs.
[0071] 2. Preliminary Discussion
[0072] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Visual. These two organizations jointly developed H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Codec (AVC), as well as the H.265 / HEVC standard. Since H.262, video codec standards have been based on a hybrid video codec architecture, utilizing time prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and applied them to reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Group (JVET) was established between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) to develop the VVC standard, which aims to reduce the bit rate by 50% compared to HEVC.
[0073] The latest version of the VVC draft, namely Universal Video Codec (Draft 2), can be found at the following URL:
[0074] http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 11_Ljubljana / wg11 / JVET-K1001-v7.zip
[0075] The latest reference software for VVC is called VTM, which can be found at the following website:
[0076] https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / tags / VTM-2.1
[0077] 2.1. Inter-frame prediction in HEVC / H.265
[0078] Each inter-frame predicted PU has motion parameters for one or two lists of reference images. The motion parameters include motion vectors and reference image indices. The use of one of the two reference image lists can also be signaled using `inter_pred_idc`. The motion vectors can be explicitly encoded as increments relative to the predictor.
[0079] When a CU is encoded and decoded in skip mode, a PU is associated with the CU and has no significant residual coefficients, no encoded motion vector increments, or a reference picture index. A Merge mode is specified, thereby obtaining the motion parameters of the current PU from neighboring PUs (including spatial and temporal candidates). The Merge mode can be applied to any PU for inter-frame prediction, not just skip mode. An alternative to the Merge mode is explicit transmission of motion parameters, where each PU explicitly signals the motion vector (more precisely, the motion vector difference compared to the motion vector predictor), the corresponding reference picture index for each reference picture list, and the reference picture list used. This mode is named Advanced Motion Vector Prediction (AMVP) in this disclosure.
[0080] When signaling indicates that one of two lists of reference images should be used, a PU is generated from a sample block. This is called "one-way prediction". One-way prediction can be used for both P-bands and B-bands.
[0081] When signaling indicates that two lists of reference images should be used, PUs are generated from two sample blocks. This is called "bidirectional prediction". Bidirectional prediction is only applicable to B-strips.
[0082] The following text provides details about the inter-frame prediction modes specified in HEVC. The description will begin with Merge mode.
[0083] 2.1.1. Merge Mode
[0084] 2.1.1.1. Candidate Derivation of Merge Pattern
[0085] When predicting a PU using the Merge pattern, the indexes pointing to entries in the Merge candidate list are parsed from the bitstream and used to retrieve motion information. The construction of this list is specified in the HEVC standard and can be summarized in the following steps:
[0086] Step 1: Initial Candidate Derivation
[0087] Step 1.1: Spatial Candidate Derivation
[0088] Step 1.2: Redundancy check of spatial candidates
[0089] Step 1.3: Derivation of Time Candidates
[0090] Step 2: Additional candidate insertion
[0091] Step 2.1: Creation of bidirectional prediction candidates
[0092] Step 2.2: Insert zero-motion candidates
[0093] exist Figure 1These steps are also schematically depicted. For spatial merge candidate derivation, up to four merge candidates are selected from candidates located at five distinct positions. For temporal merge candidate derivation, up to one merge candidate is selected from two candidates. Because a constant number of candidates is assumed at the decoder, additional candidates are generated when the number of candidates obtained from step 1 does not reach the maximum number (MaxNumMergeCand) of merge candidates signaled in the stripe header. Since the number of candidates is constant, the index of the best merge candidate is encoded using truncated univariate binarization (TU). If the size of the CU is equal to 8, all PUs of the current CU share a single merge candidate list, which is the same as the merge candidate list of the 2N×2N prediction unit.
[0094] The operations associated with the above steps are described in detail below.
[0095] Figure 1 An example derivation process for constructing the Merge candidate list is shown.
[0096] 2.1.1.2. Spatial Candidate Derivation
[0097] In the derivation of the space Merge candidate, in the position of Figure 2 At most four merge candidates are selected from the candidates at the indicated positions. The derivation order is A1, B1, B0, A0, and B2. Position B2 is considered only if any PU at positions A1, B1, B0, or A0 is unavailable (e.g., because it belongs to another slice or tile) or is intra-frame encoded / decoded. After adding the candidate at position A1, a redundancy check is performed on the addition of the remaining candidates. This redundancy check ensures that candidates with the same motion information are excluded from the list, thereby improving encoding / decoding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the above redundancy check. Instead, only those with the same motion information are considered. Figure 3 Pairs linked by arrows are added to the list only if the corresponding candidate used for redundancy checking does not have the same motion information. Another source of duplicate motion information is a "second PU" associated with a segmentation different from 2Nx2N. As an example, Figure 4A and Figure 4B The second prediction unit (PU) is depicted in the N×2N and 2N×N cases, respectively. When the current PU is divided into N×2N units, the candidate at position A1 is not considered for list construction. In fact, adding this candidate would result in two prediction units with the same motion information, which is redundant for an encoding / decoding unit with only one PU. Similarly, when the current PU is divided into 2N×N units, position B1 is not considered.
[0098] Figure 2 Example locations of spatial Merge candidates are shown.
[0099] Figure 3 Candidate pairs are shown that take into account for redundancy checks for spatial Merge candidates.
[0100] Figures 4A-4B Examples of the locations of the second PU divided into N×2N and 2N×N segments are shown.
[0101] 2.1.1.3. Derivation of Time Candidates
[0102] In this step, only one candidate is added to the list. Specifically, in the derivation of the merge candidate at this time, the scaling motion vector is derived based on the co-located PU belonging to the image with the smallest POC difference to the current image within the given list of reference images. The list of reference images to be used for deriving the co-located PU is explicitly signaled in the strip header. Figure 5 As shown by the dashed lines, a scaled motion vector for the temporal merge candidate is obtained, which is scaled from the motion vector of the juxtaposed PU using the POC distances tb and td, where tb is defined as the POC difference between the reference image and the current image, and td is defined as the POC difference between the reference image and the juxtaposed image. The reference image index for the temporal merge candidate is set to zero. The actual implementation of the scaling process is described in the HEVC specification. For the B-strip, two motion vectors are obtained, one for reference image list 0 and the other for reference image list 1, and combined to produce a bidirectional prediction merge candidate.
[0103] Figure 5 This is a diagram illustrating motion vector scaling for time merge candidates.
[0104] In the juxtaposed PU(Y) belonging to the reference frame, the position of the time candidate is selected between candidate C0 and C1, such as... Figure 6 As shown. If the PU at position C0 is unavailable, intra-frame encoded, or outside the current CTU line, then position C1 is used. Otherwise, position C0 is used in the derivation of the time merge candidate.
[0105] Figure 6 Example candidate positions C0 and C1 for time merging are shown.
[0106] 2.1.1.4. Additional Candidate Insertion
[0107] In addition to spatial and temporal merge candidates, there are two additional types of merge candidates: combined bidirectional prediction merge candidates and zero merge candidates. Combined bidirectional prediction merge candidates are generated by utilizing spatial and temporal merge candidates. These combined bidirectional prediction merge candidates are only used for B-strips. They are generated by combining the motion parameters of the first reference image list of the initial candidate with the motion parameters of the second reference image list of another. If these two tuples provide different motion hypotheses, they will form a new bidirectional prediction candidate. As an example, Figure 7 This describes the case where two candidates from the original list (on the left) (with mvL0 and refIdxL0 or mvL1 and refIdxL1) are used to create bidirectional predictive merge candidates that are added to the final list (on the right). There are many rules regarding the combinations considered to generate these additional merge candidates.
[0108] Figure 7 An example of combined bidirectional prediction of Merge candidates is shown.
[0109] Zero-motion candidates are inserted to populate the remaining entries in the Merge candidate list, thus reaching the MaxNumMergeCand capacity. These candidates have a null spatial displacement and a reference image index, which starts at zero and increments each time a new zero-motion candidate is added to the list. For unidirectional and bidirectional prediction, the number of reference frames used for these candidates is 1 and 2, respectively. Finally, no redundancy checks are performed on these candidates.
[0110] 2.1.1.5. Motion estimation region for parallel processing
[0111] To accelerate the encoding process, motion estimation can be performed in parallel, thereby simultaneously deriving the motion vectors of all prediction units within a given region. Deriving merge candidates from spatial neighbors can interfere with parallel processing because a prediction unit cannot derive motion parameters from neighboring PUs until its associated motion estimation is complete. To mitigate the trade-off between encoding / decoding efficiency and processing latency, HEVC defines a Motion Estimation Region (MER), the size of which is signaled using the image parameter set of the "log2_parallel_merge_level_minus2" syntax element. When defining the MER, merge candidates falling within the same region are marked as unavailable and therefore not considered in the list construction.
[0112] 2.1.2. AMVP
[0113] AMVP utilizes the spatiotemporal correlation of motion vectors with neighboring PUs for explicit transmission of motion parameters. For each list of reference images, a motion vector candidate list is constructed by first checking the availability of the positions of the left and top temporally adjacent PUs, removing redundant candidates, and adding zero vectors to make the candidate list a constant length. The encoder can then select the best predictor from the candidate list and send the corresponding index indicating the selected candidate. Similar to the Merge index signaling, the index of the best motion vector candidate is encoded using a truncated unary. The maximum value to be encoded in this case is 2 (see...). Figure 8 The following sections provide details of the derivation process for the motion vector prediction candidates.
[0114] 2.1.2.1. Derivation of AMVP Candidates
[0115] Figure 8 The derivation process of motion vector prediction candidates is summarized.
[0116] In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. For the derivation of spatial motion vector candidates, two motion vector candidates are ultimately derived based on the motion vectors of each PU located at five different positions, as follows: Figure 2 As shown.
[0117] For temporal motion vector candidate derivation, a motion vector candidate is selected from two candidates derived based on two different juxtaposition positions. After generating the first spatiotemporal candidate list, duplicate motion vector candidates are removed from the list. If the number of potential candidates is greater than 2, motion vector candidates whose reference image index is greater than 1 in the associated reference image list are removed from the list. If the number of spatiotemporal motion vector candidates is less than 2, additional zero motion vector candidates are added to the list.
[0118] 2.1.2.2. Candidate Spatial Motion Vectors
[0119] In the derivation of spatial motion vector candidates, at most two candidates are considered from five potential candidates. This is based on the fact that... Figure 2 The PUs at the indicated positions are the same as those where motion is merged. The derivation order to the left of the current PU is defined as A0, A1, and scaled A0, scaled A1. The derivation order to the top of the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. Therefore, for each side, there are four cases that can be used as motion vector candidates, two of which do not require spatial scaling and two of which do. The four different cases are summarized below.
[0120] No spatial scaling
[0121] –(1) Same list of reference images, and same reference image indices (same POC)
[0122] –(2) Different lists of reference images, but the same reference image (same POC)
[0123] • Spatial scaling
[0124] –(3) Same list of reference images, but different reference images (different POCs)
[0125] –(4) Different lists of reference images, and different reference images (different POCs)
[0126] First, the case without spatial scaling is checked, then spatial scaling is checked. Spatial scaling is considered, regardless of the list of reference images, when the POC differs between the reference images of adjacent PUs and the current PU. Scaling of the upper motion vector is allowed if all PUs of the left-hand candidate are unavailable or intra-frame encoded / decoded to aid in the parallel derivation of the left-hand and upper MV candidates. Otherwise, spatial scaling of the upper motion vector is not allowed.
[0127] During spatial scaling, the motion vectors of adjacent PUs are scaled in a manner similar to time scaling, such as... Figure 9 As shown. The main difference is that the current PU's reference image list and index are used as input; the actual scaling process is the same as the time scaling process.
[0128] 2.1.2.3. Candidate Time Motion Vectors
[0129] Except for the derivation of the reference image index, all the procedures used to derive the temporal Merge candidate are the same as those used to derive the spatial motion vector candidate (see [link]). Figure 6 The reference image index is signaled to the decoder.
[0130] 2.2. A new inter-frame prediction method in JEM
[0131] 2.2.1. Motion Vector Prediction Based on Sub-CU
[0132] In JEM employing QTBT, each CU can have at most one set of motion parameters for each prediction direction. Two sub-CU-level motion vector prediction methods are considered in the encoder by dividing the large CU into sub-CUs and deriving the motion information of all sub-CUs of the large CU. The Alternate Temporal Motion Vector Prediction (ATMVP) method allows each CU to obtain multiple sets of motion information from multiple blocks smaller than the current CU in the juxtaposed reference image. In the Spatiotemporal Motion Vector Prediction (STMVP) method, the motion vectors of the sub-CUs are recursively derived using temporal motion vector predictors and spatially adjacent motion vectors.
[0133] To preserve a more accurate motion field for sub-CU motion prediction, motion compression of the reference frame is currently disabled.
[0134] Figure 10 An example of ATMVP motion prediction for CU is shown.
[0135] 2.2.1.1. Prediction of Alternate Time Motion Vectors
[0136] In the Alternate Temporal Motion Vector Prediction (ATMVP) method, the Temporal Motion Vector Prediction (TMVP) is modified by extracting multiple sets of motion information (including motion vectors and reference indices) from blocks smaller than the current CU. For example... Figure 10 As shown, the sub-CU is a square N×N block (N is set to 4 by default).
[0137] ATMVP predicts the motion vectors of sub-CUs within a CU in two steps. The first step is to identify corresponding blocks in a reference image using so-called time vectors. This reference image is called the motion source image. The second step is to split the current CU into sub-CUs and obtain the motion vector and reference index of each sub-CU from the blocks corresponding to each sub-CU, such as... Figure 10 As shown.
[0138] In the first step, reference images and corresponding blocks are determined using motion information of spatially adjacent blocks in the current CU. To avoid repeated scanning of adjacent blocks, the first merge candidate in the current CU's merge candidate list is used. The first available motion vector and its associated reference index are set to the index of the time vector and the motion source image. In this way, in ATMVP, corresponding blocks can be identified more accurately than in TMVP, where the corresponding block (sometimes called the juxtaposed block) is always located at the lower right or center position relative to the current CU.
[0139] In the second step, the corresponding block of the sub-CU is identified from the time vector in the motion source image by adding the time vector to the coordinates of the current CU. For each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the center sample) is used to derive the sub-CU's motion information. After identifying the motion information of the corresponding N×N blocks, it is converted into the motion vector and reference index of the current sub-CU in the same way as the TMVP of HEVC, where motion scaling and other processes are applied. For example, the decoder checks whether a low-latency condition is met (i.e., the POC of all reference images of the current image is less than the POC of the current image), and may use the motion vector MV. x (The motion vector corresponding to the reference image list X) is used to predict the motion vector MV of each sub-CU. y (X equals 0 or 1, Y equals 1-X).
[0140] 2.2.1.2. Spatiotemporal motion vector prediction
[0141] In this method, the motion vector of the sub-CU is recursively derived according to the raster scan sequence. Figure 11 This illustrates the concept. Let's consider an 8×8 CU containing four 4×4 sub-CUs A, B, C, and D. Adjacent 4×4 blocks in the current frame are labeled a, b, c, and d.
[0142] Motion derivation for sub-CU A begins with identifying its two spatial neighbors. The first neighbor is the N×N block (block c) above sub-CU A. If block c is unavailable or intra-coded, the other N×N blocks above sub-CU A are checked (from left to right, starting with block c). The second neighbor is the block to the left of sub-CU A (block b). If block b is unavailable or intra-coded, the other blocks to the left of sub-CU A are checked (from top to bottom, starting with block b). Motion information obtained from adjacent blocks in each list is scaled to the first reference frame for the given list. Next, the temporal motion vector predictor (TMVP) for sub-block A is derived by following the same procedure as the TMVP derivation specified in HEVC. Motion information for the juxtaposed block at position D is extracted and scaled accordingly. Finally, after retrieving and scaling the motion information, all available motion vectors (up to three) are averaged for each reference list. The averaged motion vector is designated as the motion vector for the current sub-CU.
[0143] Figure 11 An example of a CU with four sub-blocks (AD) and their adjacent blocks (ad) is shown.
[0144] 2.2.1.3. Sub-CU Motion Prediction Mode Signaling
[0145] Sub-CU modes are enabled as additional Merge candidates, and no additional syntax elements are required to signal these modes. Two additional Merge candidates are added to the Merge candidate list for each CU to represent the ATMVP and STMVP modes. A maximum of seven Merge candidates are used if the sequence parameter set indicates that ATMVP and STMVP are enabled. The encoding logic for the additional Merge candidates is the same as that for the Merge candidates in the HM, meaning that for each CU in a P or B stripe, two RD checks are required for the two additional Merge candidates.
[0146] In JEM, all bits (bins) of the Merge index are context-coded using CABAC. In HEVC, however, only the first bit is context-coded; the remaining bits are context-bypass encoded.
[0147] 2.2.2. Adaptive Motion Vector Difference Resolution
[0148] In HEVC, when the `use_integer_mv_flag` in the strip header is equal to 0, motion vector difference (MVD) (between the motion vector of the PU and the predicted motion vector) is signaled in units of quarter-luminance samples. In JEM, Local Adaptive Motion Vector Resolution (LAMVR) is introduced. In JEM, MVD can be encoded and decoded in units of quarter-luminance samples, integer luminance samples, or four luminance samples. MVD resolution is controlled at the codec unit (CU) level, and an MVD resolution flag is conditionally signaled for each CU with at least one non-zero MVD component.
[0149] For a CU with at least one non-zero MVD component, signaling informs a first flag to indicate whether quarter-luminance sample MV precision is used in the CU. When the first flag (equal to 1) indicates that quarter-luminance sample MV precision is not used, signaling informs another flag to indicate whether integer luminance sample MV precision or four-luminance sample MV precision is used.
[0150] When the first MVD resolution flag of the CU is zero, or when no encoding or decoding is performed for the CU (meaning all MVDs in the CU are zero), a quarter-luminance sample MV resolution is used for the CU. When the CU uses integer luminance sample MV precision or four-luminance sample MV precision, the MVPs in the CU's AMVP candidate list are rounded to the corresponding precision.
[0151] In the encoder, CU-level RD checks are used to determine which MVD resolution will be used for the CU. That is, CU-level RD checks are performed three times for each MVD resolution. To speed up the encoder, the following encoding scheme is applied in JEM.
[0152] • During the RD check of a CU with a normal quarter-luminance sample MVD resolution, the motion information of the current CU (integer luminance sample precision) is stored. The stored motion information (rounded) is used as the starting point for further small-scale motion vector refinement during the RD check of the same CU with integer luminance sample and four-luminance sample MVD resolutions, so that the time-consuming motion estimation process is not repeated three times.
[0153] • Conditionally invoke the RD check for CUs with 4-luminance sample MVD resolution. For a CU, skip the RD check for the CU's 4-luminance sample MVD resolution if the RD cost is much greater than the quarter-luminance sample MVD resolution.
[0154] 2.2.3. Higher motion vector storage accuracy
[0155] In HEVC, the motion vector precision is one-quarter of a pixel (one-quarter of the luminance sample and one-eighth of the chrominance sample for a 4:2:0 video). In JEM, the precision of internal motion vector storage and merge candidates is increased to 1 / 16 pixel. This higher motion vector precision (1 / 16 pixel) is used for motion-compensated inter-frame prediction in CUs encoded in skip / merge mode. For CUs encoded in normal AMVP mode, integer pixel or one-quarter pixel motion is used, as described in Section 2.2.2.
[0156] The SHVC upsampling interpolation filter has the same filter length and normalization factor as the HEVC motion-compensated interpolation filter and is used as the motion-compensated interpolation filter for additional fractional pixel locations. In JEM, the chroma component motion vector accuracy is 1 / 32 sample, and the additional interpolation filter for the 1 / 32 pixel fractional location is derived by averaging the filters for two adjacent 1 / 16 pixel fractional locations.
[0157] 2.2.4. Overlapping Block Motion Compensation
[0158] Overlapping Block Motion Compensation (OBMC) was previously used in H.263. In JEM, unlike H.263, OBMC can be turned on and off using syntax at the CU level. When using OBMC in JEM, OBMC is performed on all motion compensation (MC) block boundaries except for the right and bottom boundaries of the CU. Furthermore, it applies to both luma and chroma components. In JEM, MC blocks correspond to codec blocks. When encoding and decoding a CU in subCU modes (including subCU Merge, affine, and FRUC modes), each sub-block of the CU is an MC block. To handle CU boundaries in a uniform manner, OBMC is performed at the sub-block level for all MC block boundaries, where the sub-block size is set to equal to 4×4, such as... Figure 12 As shown.
[0159] When OBMC is applied to the current sub-block, in addition to the current motion vector, the motion vectors of four adjacent connected sub-blocks (if available and different from the current motion vector) are used to derive the prediction block for the current sub-block. These multiple prediction blocks based on multiple motion vectors are combined to produce the final prediction signal for the current sub-block.
[0160] The prediction block based on the motion vectors of neighboring sub-blocks is represented as P. N , where N represents the indices of the adjacent upper, lower, left, and right sub-blocks, and the predicted block based on the motion vector of the current sub-block is represented as P. C When P N When using motion information from neighboring sub-blocks that contain the same motion information as the current sub-block, do not start from P NExecute OBMC. Otherwise, P N Each sample is added to P C The same samples in, that is, P N Add four rows / columns to P C The weighting factors {1 / 4, 1 / 8, 1 / 16, 1 / 32} are used for P. N Weighting factors {3 / 4, 7 / 8, 15 / 16, 31 / 32} are used for P C An exception is small MC blocks (i.e., when the height or width of the codec block is equal to 4 or the CU is encoded / decoded in sub-CU mode). For small MC blocks, only P... N Two rows / columns are added to P C In this case, the weighting factors {1 / 4, 1 / 8} are used for P. N Weighting factors {3 / 4, 7 / 8} are used for P C For P generated based on the motion vectors of vertically (horizontally) adjacent sub-blocks N , will P N Samples in the same row (column) are added to P with the same weighting factor. C .
[0161] In JEM, for CUs with a size of 256 lumen samples or less, signaling informs the CU level flag to indicate whether OBMC should be applied to the current CU. For CUs with a size greater than 256 lumen samples or not encoded / decoded in AMVP mode, OBMC is applied by default. At the encoder, when OBMC is applied to the CU, its impact is considered during the motion estimation phase. OBMC uses a predicted signal formed by motion information from the top and left adjacent blocks to compensate for the top and left boundaries of the original signal of the current CU, and then the normal motion estimation process is applied.
[0162] 2.2.5. Local lighting compensation
[0163] Local Illumination Compensation (LIC) is based on a linear model for illumination variations, using a scaling factor a and an offset b. Furthermore, this LIC is adaptively enabled or disabled for each inter-frame mode codec's codec unit (CU).
[0164] When LIC is applied to CU, the least squares error method is used to derive parameters a and b by using the neighboring samples of the current CU and their corresponding reference samples. More specifically, as... Figure 13 As shown, neighboring and corresponding samples (identified by motion information of the current CU or sub-CU) of the CU in the reference image are used for subsampling (2:1 subsampling). IC parameters are derived and applied to each prediction direction respectively.
[0165] When encoding and decoding a CU in Merge mode, the LIC flag is copied from the adjacent block in a manner similar to motion information copying in Merge mode; otherwise, signaling informs the CU of the LIC flag to indicate whether LIC is applied.
[0166] When LIC is enabled for an image, additional CU-level RD checks are required to determine whether LIC is applied to the CU. When LIC is enabled for a CU, the sum of absolute differences with the mean removed (MR-SAD) and the sum of absolute Hadamard transform differences with the mean removed (MR-SATD) are used instead of SAD and SATD for integer pixel motion search and fractional pixel motion search, respectively.
[0167] To reduce coding complexity, the following coding scheme was applied in JEM.
[0168] • When there is no significant lighting change between the current image and its reference images, LIC is disabled for the entire image. To identify this situation, histograms of the current image and each reference image of the current image are calculated at the encoder. If the histogram difference between the current image and each reference image of the current image is less than a given threshold, LIC is disabled for the current image; otherwise, LIC is enabled for the current image.
[0169] 2.2.6. Affine Motion Compensation Prediction
[0170] In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). However, in the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. JEM applies a simplified affine transformation motion compensation prediction. For example... Figure 14 As shown, the affine motion field of the block is described by two control point motion vectors.
[0171] The motion vector field (MVF) of the block is described by the following equation:
[0172]
[0173] Where (v 0x v 0y ) is the motion vector of the top left control point, and (v 1x v 1y ) is the motion vector of the upper right control point.
[0174] To further simplify motion compensation prediction, a sub-block-based affine transformation prediction was applied. Equation 2 derives the sub-block size M×N, where MvPre is the fractional precision of the motion vector (1 / 16 in JEM), (v... 2x v 2y ) is the motion vector of the lower left control point, calculated according to Equation 1.
[0175]
[0176] As derived from Equation 2, M and N should be adjusted downwards if necessary, so that they become the divisors of w and h, respectively.
[0177] To derive the motion vector for each M×N sub-block, such as Figure 15 As shown, the motion vector of the center sample of each sub-block is calculated according to Equation 1 and rounded to 1 / 16 fractional precision. Then, the motion-compensated interpolation filter mentioned in Section 2.2.3 is applied to generate predictions for each sub-block with derived motion vectors.
[0178] After MCP, the high-precision motion vector of each sub-block is rounded and saved with the same precision as the normal motion vector.
[0179] In JEM, there are two affine motion modes: AF_INTER mode and AF_MERGE mode. AF_INTER mode can be applied to CUs with both width and height greater than 8. Signaling in the bitstream informs the CU-level affine flag to indicate whether AF_INTER mode is used. In this mode, a motion vector pair {(v0,v1)|v0={v...} is constructed using neighboring blocks. A ,v B ,v c},v1={v D ,v E The candidate list. For example... Figure 16 As shown, v0 is selected from the motion vectors of blocks A, B, or C. The motion vectors from neighboring blocks are scaled based on the relationship between the reference list and the POC of the references of neighboring blocks, the POC of the reference of the current CU, and the POC of the current CU. The method for selecting v1 from neighboring blocks D and E is similar. If the number of candidates in the candidate list is less than two, the list is populated by copying the motion vector pairs consisting of each AMVP candidate. When the candidate list is greater than two, the candidates are first sorted based on the consistency of adjacent motion vectors (the similarity between the two motion vectors in a candidate pair), and only the top two candidates are retained. An RD cost check is used to determine which motion vector pair candidate is selected as the control point motion vector prediction (CPMVP) for the current CU. The index indicating the position of the CPMVP in the candidate list is signaled in the bitstream. After determining the CPMVP of the current affine CU, affine motion estimation is applied and the control point motion vector (CPMV) is found. The difference between the CPMV and CPMVP is then signaled in the bitstream.
[0180] When the CU is applied in AF_MERGE mode, it obtains the first block encoded / decoded in affine mode from the valid neighbor reconstructed blocks. The selection order of candidate blocks is from left, top, top right, bottom left to top left, as follows: Figure 17 As shown. If the adjacent lower left block A is encoded and decoded in affine mode, as... Figure 17 As shown, the motion vectors v2, v3, and v4 of the upper left, upper right, and lower left corners of the CU containing block A are derived. Then, the motion vector v0 of the upper left corner of the current CU is calculated based on v2, v3, and v4. Next, the motion vector v1 of the upper right corner of the current CU is calculated.
[0181] After deriving the CPMVv0 and v1 of the current CU, the MVF of the current CU is generated according to Equation 1 of the simplified affine motion model. In order to identify whether the current CU is encoded or decoded in AF_MERGE mode, when at least one neighboring block is encoded or decoded in affine mode, the affine flag is signaled in the bitstream.
[0182] 2.2.7. Derivation of Motion Vectors for Pattern Matching
[0183] Pattern Matched Motion Vector Derivation (PMMVD) mode is a special merge mode based on Frame Rate Upconversion (FRUC) technology. In this mode, the motion information of the block is not notified by signaling, but is deduced on the decoder side.
[0184] When the Merge flag of the CU is true, the signaling notifies the CU of the FRUC flag. When the FRUC flag is false, the signaling notifies the Merge index and uses the regular Merge mode. When the FRUC flag is true, the signaling notifies additional FRUC mode flags to indicate which method (bilateral matching or template matching) will be used to derive the block's motion information.
[0185] On the encoder side, the decision on whether to use the FRUC Merge mode for the CU is based on the RD cost selection made for the normal Merge candidates. That is, the two matching modes of the CU (bilateral matching and template matching) are checked using RD cost selection. The CU mode that results in the minimum cost is then further compared with other CU modes. If the FRUC matching mode is the most efficient, the FRUC flag of the CU is set to true, and the relevant matching mode is used.
[0186] The motion derivation process in the FRUC Merge mode has two steps. First, a CU-level motion search is performed, followed by sub-CU-level motion refinement. At the CU level, the initial motion vector for the entire CU is derived based on bilateral matching or template matching. First, a candidate MV list is generated, and the candidate resulting in the minimum matching cost is selected as the starting point for further CU-level refinement. Then, a local search based on bilateral matching or template matching is performed near the starting point, and the MV result with the minimum matching cost is taken as the MV for the entire CU. Subsequently, starting from the derived CU motion vector, the motion information is further refined at the sub-CU level.
[0187] For example, for the derivation of motion information for a W×HCU, the following derivation process is performed. In the first stage, the MV of the entire W×HCU is derived. In the second stage, the CU is further divided into M×M sub-CUs. The value of M is calculated according to (16), and D is a predefined splitting depth, which is set to 3 by default in JEM. Then the MV of each sub-CU is derived.
[0188]
[0189] like Figure 18 As shown, bilateral matching is used to deduce the motion information of the current CU by finding the closest match between two blocks along the current CU's motion trajectory in two different reference images. Under the assumption of continuous motion trajectories, the motion vectors MV0 and MV1 pointing to the two reference blocks should be proportional to the temporal distances between the current image and the two reference images, i.e., TD0 and TD1. As a special case, when the current image is temporally located between the two reference images and the temporal distances from the current image to the two reference images are the same, bilateral matching becomes a mirror-based bidirectional MV.
[0190] like Figure 19 As shown, template matching is used to deduce the motion information of the current CU by finding the closest match between a template in the current image (the top and / or left adjacent block of the current CU) and a block in the reference image (of the same size as the template). Besides the FRUC Merge pattern mentioned earlier, template matching also applies to the AMVP pattern. In JEM, as in HEVC, there are two AMVP candidates. New candidates are derived using the template matching method. If a newly derived candidate differs from the first existing AMVP candidate, it is inserted at the beginning of the AMVP candidate list, and the list size is set to 2 (meaning the second existing AMVP candidate is removed). When applied to the AMVP pattern, only the CU-level search is performed.
[0191] 2.2.7.1. CU-level MV candidate set
[0192] The MV candidate set at the CU level includes:
[0193] (i) If the current CU is in AMVP mode, the original AMVP candidates
[0194] (ii) All Merge candidates
[0195] (iii) Interpolating several MVs in the MV field, which will be introduced in Section 2.2.7.3.
[0196] (iv) Top and left adjacent motion vectors
[0197] When using bilateral matching, each valid MV of the Merge candidate is used as input to generate MV pairs with the bilateral matching hypothesis. For example, a valid MV of the Merge candidate is (MVa, refa) in reference list A, and then the reference image refb of its paired bilateral MV is found in another reference list B, such that refa and refb are on different sides of the current image in time. If such a refb is not available in reference list B, then refb is determined as a reference different from refa, and its temporal distance to the current image is the smallest reference in list B. After determining refb, MVb is derived by scaling MVa based on the temporal distance between the current image and refa and refb.
[0198] Four MVs from the interpolated MV field are also added to the CU-level candidate list. More specifically, the interpolated MVs are added at the current CU positions (0, 0), (W / 2, 0), (0, H / 2), and (W / 2, H / 2).
[0199] When FRUC is applied in AMVP mode, the original AMVP candidates are also added to the CU-level MV candidate set.
[0200] At the CU level, up to 15 MVs for AMVP CU and up to 13 MVs for Merge CU were added to the candidate list.
[0201] 2.2.7.2. Sub-CU Level MV Candidate Set
[0202] The candidate set of MVs at the sub-CU level includes:
[0203] (i) Based on the MV determined by the CU level search,
[0204] (ii) The top, left, top-left and top-right adjacent MVs
[0205] (iii) A scaled version of the juxtaposed MV from the reference image.
[0206] (iv) A maximum of 4 ATMVP candidates
[0207] (v) A maximum of 4 STMVP candidates
[0208] The derivation of the scaled MV from the reference images is as follows. Iterate through all reference images in both lists. The MV at the juxtaposition position of a sub-CU in the reference image is scaled to the reference of the MV at the starting CU level.
[0209] The ATMVP and STMVP candidates are limited to the top four.
[0210] At the sub-CU level, up to 17 MVs are added to the candidate list.
[0211] 2.2.7.3. Generation of Interpolated MV Fields
[0212] Before encoding and decoding the frames, an interpolated motion field is generated for the entire image based on a one-sided ME. The motion field can then be used later as a CU-level or sub-CU-level MV candidate.
[0213] First, traverse the motion field of each reference image in both reference lists at the 4×4 block level. For each 4×4 block, if the motion associated with that block passes through 4×4 blocks in the current image (e.g., ... Figure 20 (As shown) If no interpolated motion is assigned to the block, the motion of the reference block is scaled to the current image based on the time distances TD0 and TD1 (the same way as the MV scaling in TMVP in HEVC), and the scaled motion is assigned to the block in the current frame. If no scaled MV is assigned to the 4×4 block, the motion of that block is marked as unavailable in the interpolated motion field.
[0214] 2.2.7.4. Interpolation and Matching Costs
[0215] When the motion vector points to the position of a fractional sample, motion-compensated interpolation is required. To reduce complexity, bilinear interpolation is used instead of the conventional 8-tap HEVC interpolation for bilateral matching and template matching.
[0216] The calculation of matching cost varies slightly at different steps. When selecting candidates from the candidate set at the CU level, the matching cost is the absolute sum of differences (SAD) between bilateral matches or template matches. After determining the starting MV, the matching cost C for bilateral matches searched at the sub-CU level is calculated as follows:
[0217]
[0218] Where w is a weighting factor set to 4 based on experience, and MV and MV s These indicate the current MV and the starting MV, respectively. SAD is still used as the matching cost for template matching in sub-CU level searches.
[0219] In FRUC mode, the motion signature (MV) is derived solely using luma samples. The derived motion is used for luma and chroma predictions in inter-frame MC. After determining the MV, the final MC is performed using an 8-tap interpolation filter for luma and a 4-tap interpolation filter for chroma.
[0220] 2.2.7.5. MV Refinement
[0221] MV refinement is a style-based MV search based on either bilateral matching cost or template matching cost. JEM supports two search styles—Unrestricted Center-Oblique Diamond Search (UCBDS) and Adaptive Cross Search—for MV refinement at the CU level and sub-CU level, respectively. For both CU and sub-CU level MV refinement, the MV is searched directly with 1 / 4 luminance sample MV precision, followed by 1 / 8 luminance sample MV refinement. The search range for MV refinement at both the CU and sub-CU step sizes is set to equal 8 luminance samples.
[0222] 2.2.7.6. Selection of Prediction Direction in Template Matching FRUC Merge Mode
[0223] In the bilateral matching merge mode, bidirectional prediction is always applied because the motion information of the CU is derived based on the closest match between two blocks along the current CU's motion trajectory in two different reference images. There is no such restriction for the template matching merge mode. In template matching merge mode, the encoder can choose between unidirectional prediction from list 0, unidirectional prediction from list 1, or bidirectional prediction of the CU. The choice is based on the template matching cost, as follows:
[0224] If costBi <= factor * min(cost0, cost1)
[0225] Use two-way forecasting;
[0226] Otherwise, if cost0 <= cost1
[0227] Use one-way prediction from list 0;
[0228] otherwise,
[0229] Use one-way prediction from List 1;
[0230] Where cost0 is the SAD matching of template 0, cost1 is the SAD matching of template 1, and costBi is the SAD matching of template 1 for bidirectional prediction. The value of factor is 1.25, which means that the selection process is biased towards bidirectional prediction.
[0231] Inter-frame prediction direction selection is only applicable to CU-level template matching processes.
[0232] 2.2.8. Improved Generalized Two-Way Forecasting
[0233] VTM-3.0 adopts the generalized bidirectional prediction improvement (GBi) proposed in JVET-L0646.
[0234] GBi was proposed in JVET-C0047. JVET-K0248 improved the gain-complexity tradeoff of GBi and was adopted by BMS2.1. In bidirectional prediction mode, BMS2.1 GBi applies unequal weights to predictors from L0 and L1. In inter-frame prediction mode, multiple weight pairs including equal weight pairs (1 / 2, 1 / 2) are evaluated based on rate-distortion optimization (RDO), and the GBi index signaling of the selected weight pairs is communicated to the decoder. In Merge mode, the GBi index is inherited from the neighboring CU. In BMS2.1 GBi, predictor generation in bidirectional prediction mode is shown in Equation (1).
[0235] P GBi =(w0*P L0 +w1*P L1 +RoundingOffset GBi >>shiftNum GBi ,
[0236] Where P GBi This is the final predictor for GBi. w0 and w1 are the selected GBi weight pairs and are applied to the predictors for list 0 (L0) and list 1 (L1), respectively. RoundingOffset GBi and shiftNum GBi Used for the final predictor in normalized GBi. The supported w1 weight set is {-1 / 4, 3 / 8, 1 / 2, 5 / 8, 5 / 4}, where five weights correspond to one equal weight pair and four unequal weight pairs. The blend gain, i.e., the sum of w1 and w0, is fixed at 1.0. Therefore, the corresponding w0 weight set is {5 / 4, 5 / 8, 1 / 2, 3 / 8, -1 / 4}. Weight pairs are selected at the CU level.
[0237] For non-low-latency images, the weight set size is reduced from five to three, where the w1 weight set is {3 / 8, 1 / 2, 5 / 8} and the w0 weight set is {5 / 8, 1 / 2, 3 / 8}. This reduction in weight set size for non-low-latency images is applied to all GBi tests in BMS2.1 GBi and this contribution.
[0238] In JVET-L0646, a combined solution based on JVET-L0197 and JVET-L0296 is proposed to further improve GBi performance. Specifically, the following modifications are applied to the existing GBi design in BMS2.1.
[0239] 2.2.8.1. GBi Encoder Error Fixes
[0240] To reduce GBi encoding time, the current encoder design stores unidirectional predicted motion vectors estimated from GBi weights equal to 4 / 8 and reuses them for unidirectional prediction searches with other GBi weights. This fast encoding method works for both translational and affine motion models. In VTM2.0, 6-parameter and 4-parameter affine models are used. When the GBi weights are equal to 4 / 8, the BMS2.1 encoder does not distinguish between 4-parameter and 6-parameter affine models when storing unidirectional predicted affine vectors (MVs). Therefore, after encoding with GBi weights of 4 / 8, a 4-parameter affine MV can be overwritten by a 6-parameter affine MV. The stored 6-parameter affine MV can be used for 4-parameter affine MEs with other GBi weights, or the stored 4-parameter affine MV can be used for 6-parameter affine MEs. The proposed GBi encoder bug fix is to separate the storage of 4-parameter and 6-parameter affine MVs. When the GBi weights are equal to 4 / 8, the encoder stores these affine MVs based on the affine model type and reuses the corresponding affine MVs based on the affine model type for other GBi weights.
[0241] 2.2.8.2. GBi Encoder Acceleration
[0242] Five encoder acceleration methods are proposed to reduce encoding time when GBi is enabled.
[0243] (1) Affine motion estimation that conditionally skips certain GBi weights
[0244] In BMS2.1, affine ME, including 4-parameter and 6-parameter affine MEs, is performed on all GBi weights. For unequal GBi weights (weights not equal to 4 / 8), we recommend conditionally skipping affine ME. Specifically, affine ME is performed on other GBi weights if and only if the affine mode is selected as the current best mode and it is not the affine merge mode after evaluating the 4 / 8 GBi weights. If the current image is not a low-latency image, the bidirectional prediction ME for the translation model is skipped when performing affine ME for unequal GBi weights. If the affine mode is not selected as the current best mode, or if the affine merge is selected as the current best mode, then affine ME is skipped for all other GBi weights.
[0245] (2) In 1-pixel and 4-pixel MVD precision coding, reduce the number of weights used for RD cost checking of low-latency images.
[0246] For low-latency images, five weights are used for RD cost checking across all MVD accuracies (including 1 / 4 pixel, 1 pixel, and 4 pixel). The encoder will first check the RD cost for 1 / 4 pixel MVD accuracy. We recommend skipping a portion of the GBi weights used for RD cost checking at 1 pixel and 4 pixel MVD accuracy. We sort unequal weights based on their RD cost at 1 / 4 pixel MVD accuracy. During encoding at 1 pixel and 4 pixel MVD accuracy, only the top two weights with the lowest RD cost, along with the GBi weight 4 / 8, will be evaluated. Therefore, for 1 pixel and 4 pixel MVD accuracy of low-latency images, a maximum of three weights will be evaluated.
[0247] (3) When the reference images L0 and L1 are the same, skip the bidirectional prediction search conditionally.
[0248] For some images in the RA, the same image may appear in two reference image lists (list-0 and list-1). For example, for the random access codec configuration in CTC, the reference image structure for the first group of images (GOP) is as follows.
[0249] POC:16,TL:0,[L0: 0 [L1:] 0 ]
[0250] POC:8,TL:1,[L0: 0 16 [L1:] 16 0]
[0251] POC:4,TL:2,[L0:0 8 [L1:] 8 16]
[0252] POC:2,TL:3,[L0:0 4 [L1:] 4 8]
[0253] POC:1,TL:4,[L0:0 2 [L1:] 2 4]
[0254] POC:3,TL:4,[L0:2 0][L1:48]
[0255] POC:6, TL:3, [L0:4 0][L1:816]
[0256] POC:5, TL:4, [L0:4 0][L1:68]
[0257] POC:7, TL:4, [L0:6 4][L1:816]
[0258] POC:12,TL:2,[L0: 8 0][L1:16 8 ]
[0259] POC:10,TL:3,[L0:8 0][L1:12 16]
[0260] POC:9, TL:4, [L0:8 0][L1:10 12]
[0261] POC:11,TL:4,[L0:10 8][L1:12 16]
[0262] POC:14,TL:3,[L0: 12 8][L1: 12 16]
[0263] POC:13,TL:4,[L0:12 8][L1:14 16]
[0264] POC:15,TL:4,[L0: 14 12][L1:16 14 ]
[0265] We can see that images 16, 8, 4, 2, 1, 12, 14, and 15 have the same reference images (multiple) in both lists. For bidirectional prediction of these images, the L0 and L1 reference images may be the same. We propose that the encoder skips bidirectional prediction MEs with unequal GBi weights when 1) the two reference images in bidirectional prediction are the same, 2) the temporal layer is greater than 1, and 3) the MVD accuracy is 1 / 4 pixel. For affine bidirectional prediction MEs, this fast skipping method only applies to 4-parameter affine MEs.
[0266] (4) Skip RD cost checks with unequal GBi weights based on the time layer and the POC distance between the reference image and the current image.
[0267] When the time layer is equal to 4 (the highest time layer in RA) or the POC distance between the reference image (list-0 or list-1) and the current image is equal to 1 and the encoding / decoding QP is greater than 32, we recommend skipping RD cost evaluations for those with unequal GBi weights.
[0268] (5) During ME, floating-point calculations used for unequal GBi weights will be changed to fixed-point calculations.
[0269] For existing bidirectional predictive search, the encoder fixes the MV of one list and refines the MV of the other list. The objective is modified before ME to reduce computational complexity. For example, if the MV of list-1 is fixed and the encoder wants to refine the MV of list-0, the objective for refining the MV of list-0 is modified using equation (5). O is the original signal, P1 is the predicted signal of list-1, and w is the GBi weight of list-1.
[0270] T=((O<<3)-w*P1)*(1 / (8-w)) (5)
[0271] The term (1 / (8-w)) is stored with floating-point precision, which increases computational complexity. We suggest changing equation (5) to fixed-point, as in equation (6).
[0272] T=(O*a1-P1*a2+round)>>N (6)
[0273] Where a1 and a2 are scaling factors, which are calculated as follows:
[0274] γ=(1<<N) / (8-w); a1=γ<<3; a2=γ*w; round=1<<(N-1)
[0275] 2.2.8.3. GBi CU Size Limitation
[0276] In this approach, GBi is disabled for small CUs. In inter-frame prediction mode, if bidirectional prediction is used and the CU area is less than 128 lumen samples, GBi is disabled without any signaling.
[0277] 2.2.9. Bidirectional optical flow
[0278] 2.2.9.1. Theoretical Analysis
[0279] In BIO, motion compensation is first performed to generate the first prediction for the current block (in each prediction direction). The first prediction is used to derive the spatial gradient, temporal gradient, and optical flow for each sub-block / pixel within the block, and then they are used to generate the second prediction, which is the final prediction for the sub-block / pixel. A detailed description follows.
[0280] Bidirectional optical flow (BIO) is a per-sample motion refinement performed on top of block-by-block motion compensation used for bidirectional prediction. Sample-level motion refinement does not use signaling.
[0281] Let I (k) The brightness value from reference k (k = 0, 1) after block motion compensation, and They are I (k) The horizontal and vertical components of the gradient. Assuming optical flow is effective, the motion vector field (v...)x ,v y The following equation gives the result.
[0282]
[0283] Combining this optical flow equation with Hermite interpolation of the motion trajectory for each sample yields a unique third-order polynomial, which terminates at the function value I. (k) and derivative The polynomial matches. The value of the polynomial at t=0 is the BIO prediction:
[0284]
[0285] Here, τ0 and τ1 represent the distances to the reference frame, such as... Figure 21 As shown. The distances τ0 and τ1 are calculated based on the POCs of Ref0 and Ref1: τ0 = POC(current) - POC(Ref0), τ1 = POC(Ref1) - POC(current). If the two predictions originate from the same time direction (both from the past or both from the future), they have different signs (i.e., τ0·τ1 < 0). In this case, BIO is applied only when the predictions do not originate from the same time (i.e., τ0 ≠ τ1), both reference regions have non-zero motion (MVx0, MVy0, MVx1, MVy1 ≠ 0), and the block motion vector is proportional to the time distance (MVx0 / MVx1 = MVy0 / MVy1 = -τ0 / τ1).
[0286] By minimizing the difference Δ between the values at points A and B Figure 37 The motion vector field (v) is determined by the intersection of the motion trajectory on the plane and the reference frame plane. x ,v y The model uses only the first linear term of the local Taylor expansion of Δ:
[0287]
[0288] All values in equation (9) depend on the sample location (i′,j′), which has been omitted from the notation so far. Assuming the motion is consistent in the local surrounding region, we minimize Δ within a (2M+1)×(2M+1) square window centered at the current prediction point (i,j), where M equals 2:
[0289]
[0290] For this optimization problem, JEM uses a simplified approach, first minimizing in the vertical direction and then minimizing in the horizontal direction. This leads to...
[0291]
[0292]
[0293] in,
[0294]
[0295]
[0296]
[0297] To avoid division by zero or very small values, regularization parameters r and m are introduced in equations (11) and (12).
[0298] r = 500·4 d-8 (14)
[0299] m = 700·4 d-8 (15)
[0300] Here, d is the bit depth of the video sample.
[0301] To maintain the same memory access as regular bidirectional predictive motion compensation, all predictions and gradient values are calculated only for the position within the current block. (k) , In equation (13), a 2M+1)×(2M+1) square window centered on the current prediction point on the boundary of the prediction block needs to access locations outside the block (e.g., Figure 22A (As shown) In JEM, I outside the block (k) , The value is set to equal to the nearest available value within the block. For example, this can be implemented as padding, such as... Figure 22B As shown.
[0302] Figures 22A-22B BIO without block extension: Figure 22A The access location outside the block is shown.
[0303] Figure 22B An example of padding used to avoid additional memory access and computation is shown.
[0304] Using BIO, the motion field of each sample can be refined. To reduce computational complexity, a block-based BIO design is used in JEM. Motion refinement is calculated based on 4×4 blocks. In block-based BIO, s in equation (13) of all samples in the 4×4 are aggregated. n The value, then the aggregated s n The values are used to derive the BIO motion vector offset for a 4×4 block. More specifically, the following equation is used for block-based BIO derivation:
[0305]
[0306]
[0307]
[0308] Where b k Let represent the sample set belonging to the k-th 4x4 block of the prediction block. Then, in equations (11) and (12), s... n Replace with ((s) n,bk )>>4) to derive the associated motion vector offset.
[0309] In some cases, the MV clique (regiment) of BIO may be unreliable due to noise or irregular motion. Therefore, in BIO, the amplitude of the MV clique is capped to a threshold value thBIO. The threshold value is determined based on whether all reference images of the current image come from the same direction. If all reference images of the current image come from the same direction, the threshold value is set to 12×2. 14-d Otherwise, it is set to 12×2 13-d .
[0310] Using the same operation as the HEVC motion compensation process (2D separable FIR), the gradient and motion compensation interpolation of BIO are calculated simultaneously. The input to this 2D separable FIR is the same reference frame sample as the motion compensation process and the fractional positions (fracX, fracY) based on the fractional part of the block motion vector. In the horizontal gradient... In the case of the signal, vertical interpolation is first performed using BIOfilters, which correspond to the fractional position fracY with a descaling offset of d-8. Then, a gradient filter BIOfilterG is applied in the horizontal direction, which corresponds to the fractional position fracX with a descaling offset of 18-d. In the vertical gradient... In this case, the gradient filter is first applied vertically using BIOfilterG, corresponding to the fractional position fracY with a descaling offset of d-8. Then, signal translation is performed horizontally using BIOfilterS, corresponding to the fractional position fracX with a descaling offset of 18-d. To maintain reasonable complexity, the interpolation filters for gradient computation BIOfilterG and signal translation BIOfilterF are relatively short (6 taps). Table 1 shows the filters used in BIO for gradient computation at different fractional positions of the block motion vector. Table 2 shows the interpolation filters used in BIO for predictive signal generation.
[0311] Table 1: Filters for gradient calculation in BIO
[0312]
[0313]
[0314] Table 2: Interpolation filters used for predictive signal generation in BIO
[0315] 0 {0,0,64,0,0,0} 1 / 16 {1,-3,64,4,-2,0} 1 / 8 {1,-6,62,9,-3,1} 3 / 16 {2,-8,60,14,-5,1} 1 / 4 {2,-9,57,19,-7,2} 5 / 16 {3,-10,53,24,-8,2} 3 / 8 {3,-11,50,29,-9,2} 7 / 16 {3,-11,44,35,-10,3} 1 / 2 {3,-10,35,44,-11,3}
[0316] In JEM, when two predictions come from different reference images, BIO is applied to all bidirectional prediction blocks. BIO is disabled when LIC is enabled for CU.
[0317] In JEM, OBMC is applied to a block after the normal MC process. To reduce computational complexity, BIO is not applied during OBMC. This means that when using the block's own MV, BIO is only applied to the block's MC process, but when using the MV of an adjacent block during OBMC, BIO is not applied to the MC process.
[0318] BIO in VTM-3.0 proposed in JVET-L0256 (2.2.9.2)
[0319] Step 1: Determine if BIO is applicable (W and H are the width and height of the current block).
[0320] BIO is not applicable if
[0321] Affine codec
[0322] ATMVP encoding and decoding
[0323] ·(iPOC-iPOC0)*(iPOC-iPOC1)>=0
[0324] • H = 4 or (W = 4 and H = 8)
[0325] • Use weighted prediction
[0326] The GBi weight is not (1, 1).
[0327] If BIO is not used,
[0328] The total SAD between the two reference blocks (denoted as R0 and R1) is less than the threshold.
[0329]
[0330] Step 2: Data Preparation
[0331] For a WxH block, interpolate (W+2)x(H+2) samples.
[0332] The internal WxH samples are interpolated using an 8-tap interpolation filter, and the four outer rows of the samples are interpolated using a bilinear filter, similar to normal motion compensation. Figure 23 (The black circle in the middle)
[0333] For each location, the gradient is computed on the two reference blocks (denoted as R0 and R1).
[0334] Gx0(x,y)=(R0(x+1,y)-R0(x-1,y))>>4
[0335] Gy0(x,y)=(R0(x,y+1)-R0(x,y-1))>>4
[0336] Gx1(x,y)=(R1(x+1,y)-R1(x-1,y))>>4
[0337] Gy1(x,y)=(R1(x,y+1)-R1(x,y-1))>>4
[0338] For each location, the internal value is calculated as follows:
[0339] T1=(R0(x,y)>>6)-(R1(x,y)>>6), T2=(Gx0(x,y)+Gx1(x,y))>>3, T3=(Gy0(x,y)+Gy1(x,y))>>3
[0340] B1(x,y)=T2*T2, B2(x,y)=T2*T3, B3(x,y)=-T1*T2, B5(x,y)=T3*T3, B6(x,y)=-T1*T3
[0341] Figure 23 An example of an interpolation sample used in BIO is shown.
[0342] Step 3: Calculate the prediction for each block
[0343] If the SAD between two 4×4 reference blocks is less than the threshold, then skip the BIO of the 4×4 block.
[0344] Calculate Vx and Vy.
[0345] Calculate the final prediction for each position in the 4×4 block.
[0346] b(x,y)=(Vx(Gx 0 (x,y)-Gx 1 (x,y))+Vy(Gy 0 (x,y)-Gy 1 (x,y)+1)>>1
[0347] P(x,y)=(R 0 (x,y)+R 1 (x,y)+b(x,y)+offset)>>shift
[0348] b(x,y) is called the correction term.
[0349] BIO in VTM-3.0
[0350] 8.3.4 Inter-frame block decoding process
[0351] - If predFlagL0 and predFlagL1 are equal to 1, DiffPicOrderCnt(currPic, refPicList0[refIdx0])*DiffPicOrderCnt(currPic, refPicList1[refIdx1])<0, MotionModelIdc[xCb][yCb] is equal to 0 and MergeModeList[merge_idx[xCb][yCb]] is not equal to SbCol, then set the value of bioAvailableFlag to TRUE.
[0352] - Otherwise, set the value of bioAvailableFlag to FALSE.
[0353] …
[0354] - If bioAvailableFlag equals TRUE, then apply the following:
[0355] - The variable shift is set to equal MAX(2, 14-bitDepth).
[0356] The variables cuLevelAbsDiffThres and subCuLevelAbsDiffThres are set to equal (1 << (bitDepth - 8 + shift)) * cbWidth * cbHeight and 1 << (bitDepth – 3 + shift). The variable cuLevelSumAbsoluteDiff is set to 0.
[0357] - For xSbIdx = 0..(cbWidth>>2)-1 and ySbIdx = 0..(cbHeight>>2)-1, the derivation of the current sub-block's variables subCuLevelSumAbsoluteDiff[xSbIdx][ySbIdx] and bidirectional optical flow utilization flag bioUtilizationFlag[xSbIdx][ySbIdx] is as follows:
[0358] subCuLevelSumAbsoluteDiff[xSbIdx][ySbIdx]=∑ i ∑ j Abs
[0359] (predSamplesL0L[(xSbIdx<<2)+1+i][(ySbIdx<<2)+1+j]-predSamplesL1L[(xSbIdx<<2)+1+i][(ySbIdx<<2)+1+j]), where i, j=0..3
[0360] bioUtilizationFlag[xSbIdx][ySbIdx]=subCuLevelSumAbsoluteDiff[xSbIdx][ySbIdx]>=
[0361] subCuLevelAbsDiffThres
[0362] cuLevelSumAbsoluteDiff+=subCuLevelSumAbsoluteDiff[xSbIdx][ySbIdx]
[0363] - If cuLevelSumAbsoluteDiff is less than cuLevelAbsDiffThres, then set bioAvailableFlag to FALSE.
[0364] - If bioAvailableFlag equals TRUE, then by calling the bidirectional optical flow sample prediction procedure specified in Clause 8.3.4.5, the predicted sample predSamplesL[xL+xSb][yL+ySb] within the current luminance codec subblock is derived using the luminance codec subblock width sbWidth, luminance codec subblock height sbHeight, sample arrays predSamplesL0L and predSamplesL1L, and variables predFlagL0, predFlagL1, refIdxL0, and refIdxL1, where xL = 0..sbWidth-1 and yL = 0..sbHeight-1.
[0365] 8.3.4.3 Fractional Sample Interpolation Process
[0366] 8.3.4.3.1 Overview
[0367] The input to this process is:
[0368] -Luminance position (xSb, ySb), specifies the top-left sample of the current codec sub-block relative to the top-left luminance sample of the current image.
[0369] - The variable sbWidth specifies the width of the current codec sub-block in the luminance sample.
[0370] - The variable sbHeight specifies the height of the current codec sub-block in the luminance sample.
[0371] - The brightness motion vector mvLX is given in units of 1 / 16 brightness samples.
[0372] - The chromaticity motion vector mvCLX is given in units of 1 / 32 chromaticity samples.
[0373] - Selected reference image sample arrays refPicLXL and refPicLXCb and refPicLXCr.
[0374] -Bidirectional optical flow enable flag bioAvailableFlag.
[0375] The output of this process is:
[0376] - When bioAvailableFlag is FALSE, predict the (sbWidth)x(sbHeight) array predSamplesLXL of the luminance sample values, or when bioAvailableFlag is TRUE, predict the (sbWidth+2)x(sbHeight+2) array predSamplesLXL of the luminance sample values.
[0377] - Two (sbWidth / 2)x(sbHeight / 2) arrays, predSamplesLXCb and predSamplesLXCr, are used to predict the chromaticity sample values.
[0378] Let (xIntL, yIntL) be the brightness position given in units of the full sample, and (xFracL, yFracL) be the offset given in units of 1 / 16 of the sample. These variables are used only in this clause to specify the fractional sample positions within the reference sample arrays refPicLXL, refPicLXCb, and refPicLXCr.
[0379] When bioAvailableFlag equals TRUE, the derivation of the corresponding predicted luminance sample value predSamplesLXL[xL][yL] for each luminance sample position (xL = -1..sbWidth, yL = -1..sbHeight) within the predicted luminance sample array predSamplesLXL is as follows:
[0380] The derivation of variables xIntL, yIntL, xFracL, and yFracL is as follows:
[0381] xIntL=xSb-1+(mvLX[0]>>4)+xL
[0382] yIntL=ySb-1+(mvLX[1]>>4)+yL
[0383] xFracL=mvLX[0]&15
[0384] yFracL=mvLX[1]&15
[0385] The value of -bilinearFiltEnabledFlag is derived as follows:
[0386] - If xL equals -1 or sbWidth, or yL equals -1 or sbHeight, then set the value of bilinearFiltEnabledFlag to TRUE.
[0387] - Otherwise, set the value of bilinearFiltEnabledFlag to FALSE.
[0388] The predicted brightness sample values predSamplesLXL[xL][yL] are derived by calling the procedure specified in Clause 8.3.4.3.2 with (xIntL, yIntL), (xFracL, yFracL), refPicLXL, and bilinearFiltEnabledFlag as inputs.
[0389] When bioAvailableFlag equals FALSE, the derivation of the corresponding predicted luminance sample value predSamplesLXL[xL][yL] for each luminance sample position (xL = 0..sbWidth-1, yL = 0..sbHeight-1) in the predicted luminance sample array predSamplesLXL is as follows:
[0390] The derivation of variables xIntL, yIntL, xFracL, and yFracL is as follows:
[0391] xIntL=xSb+(mvLX[0]>>4)+xL
[0392] yIntL=ySb+(mvLX[1]>>4)+yL
[0393] xFracL=mvLX[0]&15
[0394] yFracL=mvLX[1]&15
[0395] The variable bilinearFiltEnabledFlag is set to FALSE.
[0396] - The predicted brightness sample values predSamplesLXL[xL][yL] are derived by calling the procedure specified in Clause 8.3.4.3.2 with (xIntL, yIntL), (xFracL, yFracL), refPicLXL, and bilinearFiltEnabledFlag as inputs.
[0397] …
[0398] 8.3.4.5 Bidirectional Optical Flow Prediction Process
[0399] The input to this process is:
[0400] - Two variables, nCbW and nCbH, specify the width and height of the current codec block.
[0401] - Two (nCbW+2)x(nCbH+2) brightness prediction sample arrays, predSamplesL0 and predSamplesL1,
[0402] - The prediction list uses the flags predFlagL0 and predFlagL1.
[0403] -Refer to indices refIdxL0 and refIdxL1,
[0404] - The bidirectional optical flow utilization flag bioUtilizationFlag[xSbIdx][ySbIdx], where xSbIdx = 0..(nCbW>>2)-1, ySbIdx = 0..(nCbH>>2)-1
[0405] The output of this process is an array pbSamples of (nCbW)x(nCbH) luminance prediction sample values.
[0406] The variable bitDepth is set to BitDepthY.
[0407] The variable shift2 is set to equal Max(3, 15-bitDepth), and the variable offset2 is set to equal 1<<(shift2-1).
[0408] The variable mvRefineThres is set to equal 1 << (13-bit Depth).
[0409] For xSbIdx=0..(nCbW>>2)-1 and ySbIdx=0..(nCbH>>2)-1,
[0410] - If bioUtilizationFlag[xSbIdx][ySbIdx] is FALSE, then for x = xSb..xSb+3 and y = ySb..ySb+3, the derivation of the predicted sample value of the current prediction unit is as follows:
[0411] pbSamples[x][y]=Clip3(0, (1< <bitDepth)-1,
[0412] (predSamplesL0[x][y]+predSamplesL1[x][y]+offset2)>>shift2)
[0413] Otherwise, the derivation of the predicted sample value for the current prediction unit is as follows:
[0414] - The derivation of the position (xSb, ySb) of the top-left sample of the current sub-block relative to the top-left samples of the predicted sample arrays predSamplesL0 and predSampleL1 is as follows:
[0415] xSb=(xSbIdx<<2)+1
[0416] ySb=(ySbIdx<<2)+1
[0417] - For x = xSb – 1..xSb + 4, y = ySb – 1..ySb + 4, apply the following:
[0418] The derivation of the position (hx, vy) of each corresponding sample (x, y) in the predicted sample array is as follows:
[0419] hx = Clip3(1, nCbW, x)
[0420] vy = Clip3(1, nCbH, y)
[0421] The derivation of variables gradientHL0[x][y], gradientVL0[x][y], gradientHL1[x][y], and gradientVL1[x][y] is as follows:
[0422] gradientHL0[x][y]=(predSamplesL0[hx+1][vy]–predSampleL0[hx-1][vy])>>4
[0423] gradientVL0[x][y]=(predSampleL0[hx][vy+1]-predSampleL0[hx][vy-1])>>4
[0424] gradientHL1[x][y]=(predSampleL1[hx+1][vy]–predSampleL1[hx-1][vy])>>4
[0425] gradientVL1[x][y]=(predSampleL1[hx][vy+1]–predSampleL1[hx][vy-1])>>4
[0426] - The derivation of variables temp, tempX and tempY is as follows:
[0427] temp[x][y]=(predSampleL0[hx][vy]>>6)-(predSampleL1[hx][vy]>>6)
[0428] tempX[x][y]=(gradientHL0[x][y]+gradientHL1[x][y])>>3
[0429] tempY[x][y]=(gradientVL0[x][y]+gradientVL1[x][y])>>3
[0430] - The derivation of variables sGx2, sGy2, sGxGy, sGxdI and sGydI is as follows:
[0431] sGx2=∑ x ∑ y
[0432] (tempX[xSb+x][ySb+y]*tempX[xSb+x][ySb+y]), where x, y=-1..4
[0433] sGy2=∑ x ∑ y
[0434] (tempY[xSb+x][ySb+y]*tempY[xSb+x][ySb+y]), where x, y=-1..4
[0435] sGxGy=∑ x ∑ y
[0436] (tempX[xSb+x][ySb+y]*tempY[xSb+x][ySb+y]), where x, y=-1..4
[0437] sGxdI=∑ x ∑ y (-tempX[xSb+x][ySb+y]*temp[xSb+x][ySb+y]), where x, y=-1..4
[0438] sGydI=∑ x ∑ y (-tempY[xSb+x][ySb+y]*temp[xSb+x][ySb+y]), where x, y=-1..4
[0439] -The refined derivation of the horizontal and vertical motion of the current sub-block is as follows:
[0440] vx=sGx2>0? Clip3(-mvRefineThres, mvRefineThres, -(sGxdI<<3)>>Floor(Log2(sGx2))): 0
[0441] vy=sGy2>0? Clip3(-mvRefineThres,mvRefineThres,((sGydI<<3)-((vx*sGxGym)<<12+
[0442] vx*sGxGys)>>1)>>Floor(Log2(sGy2))): 0
[0443] sGxGym=sGxGy>>12;
[0444] sGxGys=sGxGy&((1<<12)-1)
[0445] For x = xSb-1..xSb+2 and y = ySb-1..ySb+2, apply the following:
[0446] sampleEnh=Round((vx*(gradientHL1[x+1][y+1]-gradientHL0[x+1][y+1]))>>1)
[0447] +Round((vy*(gradientVL1[x+1][y+1]-gradientVL0[x+1][y+1]))>>1)
[0448] pbSamples[x][y]=Clip3(0, (1< <bitDepth)-1,(predSamplesL0[x+1][y+1]
[0449] +predSamplesL1[x+1][y+1]+sampleEnh+offset2)>>shift2)
[0450] 2.2.10. Decoder-side motion vector refinement
[0451] DMVR is a decoder-side motion vector derivation (DMVD).
[0452] In bidirectional prediction, for the prediction of a block region, two prediction blocks formed using motion vectors (MV) from list 0 and MV from list 1 are combined to form a single prediction signal. In the decoder-side motion vector refinement (DMVR) method, the two bidirectional prediction motion vectors are further refined through a bilateral template matching process. The bilateral template matching applied in the decoder performs a distortion-based search between the bilateral templates and reconstructed samples in the reference image to obtain the refined MV without transmitting additional motion information.
[0453] In DMVR, such as Figure 24 As shown, bilateral templates are generated from the initial MV0 of list 0 and MV1 of list 1 as a weighted combination (i.e., average) of the two prediction blocks. The template matching operation involves calculating a cost metric between the generated template and the sample regions in the reference images (around the initial prediction block). For each of the two reference images, the MV that produces the minimum template cost is considered the updated MV for that list, replacing the original MV. In JEM, nine MV candidates are searched for each list. The nine candidate MVs include the original MV and eight surrounding MVs, one of which has a brightness sample offset to the original MV in the horizontal or vertical direction, or both. Finally, as... Figure 24 The two new MVs shown, MV0′ and MV1′, are used to generate the final bidirectional prediction result. The sum of absolute differences (SAD) is used as the cost metric. Note that when calculating the cost of the prediction block generated by a surrounding MV, the predicted block is actually obtained using a rounded MV (to an integer pixel) instead of the true MV.
[0454] DMVR is applied to bidirectional prediction merge patterns, where one MV comes from a past reference picture and the other from a future reference picture, without transferring additional syntax elements. In JEM, DMVR is not applied when LIC, affine motion, FRUC, or subCU merge candidates are enabled for a CU.
[0455] Figure 24An example of the proposed DMVR based on bilateral template matching is shown.
[0456] 2.2.11.JVET-N0236
[0457] This contribution proposes a method for refining sub-block-based affine motion compensation predictions using optical flow. After performing sub-block-based affine motion compensation, the prediction samples are refined by adding a difference derived from the optical flow equation; this is called Optical Flow Prediction Refinement (PROF). This method can achieve pixel-level granularity inter-frame prediction without increasing memory access bandwidth.
[0458] To achieve finer motion compensation granularity, this contribution proposes a method for refining sub-block-based affine motion compensation predictions using optical flow. After performing sub-block-based affine motion compensation, the brightness prediction samples are refined by adding a difference derived from the optical flow equation. The proposed PROF (Prediction Refinement Using Optical Flow) is described in the following four steps.
[0459] Step 1) Perform sub-block-based affine motion compensation to generate sub-block predictions.
[0460] Step 2) Calculate the spatial gradient g of the sub-block prediction at each sample location using a 3-tap filter [-1, 0, 1]. x (i,j) and g y (i,j).
[0461] g x (i,j)=I(i+1,j)-I(i-1,j)
[0462] g y (i,j)=I(i,j+1)-I(i,j-1)
[0463] Sub-block prediction extends each side by one pixel for gradient calculation. To reduce memory bandwidth and complexity, pixels on the extended boundaries are copied from the nearest integer pixel position in the reference image. This avoids additional interpolation of the filled regions.
[0464] Step 3) Calculate and refine the brightness prediction using the optical flow equation.
[0465] ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j)
[0466] Among them, such as Figure 25As shown, Δv(i,j) is the difference between the pixel MV (represented as v(i,j)) calculated for sample position (i,j) and the sub-block MV of the sub-block to which pixel (i,j) belongs.
[0467] Figure 25 An example of sub-block MV VSB and pixel Δv(i,j) is shown (arrow 2502).
[0468] Because the affine model parameters and the pixel position relative to the sub-block center do not change between sub-blocks, Δv(i,j) can be calculated for the first sub-block and reused for other sub-blocks in the same CU. Let x and y be the horizontal and vertical offsets from the pixel position to the sub-block center, Δv(x,y) can be derived using the following equation.
[0469]
[0470] For a four-parameter affine model
[0471]
[0472] For a 6-parameter affine model
[0473]
[0474] Where (v 0x ,v 0y ), (v 1x ,v 1y ), (v 2x ,v 2y ) are the motion vectors of the control points at the top left, top right, and bottom left corners, and w and h are the width and height of the CU.
[0475] Step 4) Finally, refine the brightness prediction and add it to the sub-block prediction I(i,j). The final prediction I' is generated by the following equation.
[0476] I′(i,j)=I(i,j)+ΔI(i,j)
[0477] 2.2.12.JVET-N0510
[0478] Phase change affine sub-block motion compensation (PAMC)
[0479] To better approximate the affine motion model in the affine sub-blocks, a phase-change MC is applied to the sub-blocks. In the proposed method, the affine codec block is also divided into 4×4 sub-blocks, and the sub-block MV is derived for each sub-block as in VTM4.0. The MC for each sub-block is divided into two stages. The first stage is to filter a (4+L-1)×(4+L-1) reference block window with (4+L-1) row horizontal filtering, where L is the filter tap length of the interpolation filter. However, unlike translation MC, in the proposed phase-change affine sub-block MC, the filtering phase is different for each sample row. For each sample row, the MVx derivation is as follows.
[0480] MVx = (subblockMVx << 7 + dMvVerX × (rowIdx - L / 2 – 2)) >> 7 (Equation 1)
[0481] The filtered phase for each sample row is derived from MVx. subblockMVx is the x-component of the MV of the derived subblock MV, as done in VTM4.0. rowIdx is the sample row index. dMvVerX is (cuBottomLeftCPMVx - cuTopLeftCPMVx) << (7 - log2LumaCbHeight), where cuBottomLeftCPMVx is the x-component of the MV of the lower left control point of the CU, cuTopLeftCPMVx is the x-component of the MV of the upper left control point of the CU, and LumaCbHeight is log2 of the height of the luminance codec block (CB).
[0482] After horizontal filtering, 4×(4+L-1) horizontally filtered samples are generated. Figure 1 The concept of the proposed horizontal filtering is illustrated. Gray dots represent samples from the reference block window, and orange dots represent samples for horizontal filtering. The blue tubes representing 8×1 samples indicate the application of an 8-tap horizontal filter, such as... Figure 26 and Figure 27 As shown. Each sample row requires four horizontal filters. The filter phase is the same across sample rows. However, the filter phase is different across different rows. This generates skewed 4×11 samples.
[0483] In the second stage, samples filtered by 4×(4+L-1) levels ( Figure 1 The orange samples in the sample column are further vertically filtered. For each sample column, the MVy derivation is as follows.
[0484] MVy=(subblockMVy<<7+dMvHorY×(columnIdx-2))>>7 (Equation 2)
[0485] The filtered phase for each sample column is derived from MVy. subblockMVy is the y-component of the MV of the derived subblock MV, as done in VTM4.0. columnIdx is the sample column index. dMvHorY is (cuTopRightCPMVy – cuTopLeftCPMVy) << (7 – log2LumaCbWidth), where cuTopRightCPMVy is the y-component of the MV of the upper right control point of the CU, cuTopLeftCPMVy is the y-component of the MV of the upper left control point of the CU, and log2LumaCbWidth is log2 of the width of the luminance CB.
[0486] After vertical filtering, 4×4 affine sub-block prediction samples are generated. Figure 28 The concept of the proposed vertical filtering is illustrated. Light orange dots represent horizontally filtered samples from the first stage. Red dots represent vertically filtered samples used as the final prediction samples.
[0487] In this scheme, the interpolation filter set used is the same as that in VTM4.0. The only difference is that the horizontal filter phase on a sample row is different, and the vertical filter phase on a sample column is different. As for the number of filtering operations per affine sub-block in the proposed method, it is the same as that in VTM4.0.
[0488] 2.2.13. JVET-K0102 Interleaving Prediction
[0489] like Figure 29 The interleaved prediction shown is proposed to achieve finer-grained MV without excessively increasing complexity.
[0490] First, the coded block is divided into sub-blocks with two different partitioning styles. The first partitioning style is the same as the partitioning style in VTM, such as... Figure 31 The second partitioning style also divides the codec block into 4x4 sub-blocks, but with a 2x2 offset, as shown in style 0. Figure 32 The style shown is 1.
[0491] Secondly, AMC uses these two partitioning styles to generate two auxiliary predictions. The MV of each sub-block in the partitioning style is derived from the CPMV using an affine model.
[0492] The final prediction is calculated as a weighted sum of the two auxiliary predictions, using the following formula:
[0493]
[0494] like Figure 30 As shown, auxiliary prediction samples located at the center of the sub-block are associated with a weighting value of 3, while auxiliary prediction samples located at the boundary of the sub-block are associated with a weighting value of 1.
[0495] Figure 29 An example of the interleaved prediction process is shown.
[0496] Figure 30 An example of weighted values in a sub-block is shown.
[0497] 3. The technical problem solved by the publicly disclosed technical solutions
[0498] 1. In JVET-N0236, optical flow is used for affine prediction, but it is far from optimal.
[0499] 2. When applying interleaved prediction, the required bandwidth may increase.
[0500] 3. For affine mode, how to apply motion compensation to the chromaticity components with 4x4 sub-blocks remains an open question.
[0501] 4. Interleaved affine prediction increases cache bandwidth requirements. For AMC of style 1 in interleaved affine prediction, reference sample lines need to be extracted from the cache buffer, resulting in higher cache bandwidth.
[0502] 4. Example Solutions and Implementation Examples
[0503] To address these issues, we propose different forms for deriving refined prediction samples with optical flow. Furthermore, we suggest using information from neighboring (e.g., adjacent or non-adjacent) blocks (such as reconstructed samples or motion information), and / or the gradient of a sub-block and its prediction block, to obtain the final prediction block for the current sub-block.
[0504] The following items should be considered as examples to illustrate general concepts. These items should not be interpreted in a narrow sense. Furthermore, these items can be combined in any way.
[0505] Let Ref0 and Ref1 represent the reference images of the current image from list 0 and list 1, respectively, and let τ0 = POC(current) - POC(Ref0) and τ1 = POC(Ref1) - POC(current). Let refblk0 and refblk1 represent the reference blocks of the current block from Ref0 and Ref1, respectively. For a sub-block in the current block, its MV (Modular Value) pointing to the corresponding reference sub-block in refblk0 is determined by (v x v y ) represents the current sub-block's MV, which is represented by (mvL0) and (ref Ref 0 and Ref 1 respectively). x mvL0 y ) and (mvL1 x mvL1 y )express.
[0506] In the following discussion, SatShift(x, n) is defined as
[0507]
[0508] Shift(x, n) is defined as Shift(x, n) = (x + offset 0) >> n.
[0509] In one example, offset0 and / or offset1 are set to (1 < 0.05).<n)> >1 or (1<<(n-1)). In another example, offset0 and / or offset1 are set to 0.
[0510] In another example, offset0 = offset1 = ((1<<n)> >1)-1 or ((1<<(n-1)))-1.
[0511] Clip3(min, max, x) is defined as
[0512]
[0513] In the following discussion, an operation between two motion vectors means that the operation will be applied to both components of the motion vector. For example, MV3 = MV1 + MV2 is equivalent to MV3. x =MV1 x +MV2 x and MV3 y =MV1 y +MV2 y Alternatively, this operation can be applied only to the horizontal or vertical components of the two motion vectors.
[0514] In the following discussion, such as Figure 2 As shown, the left adjacent block, the lower left adjacent block, the upper adjacent block, the upper right adjacent block, and the upper left adjacent block are represented as blocks A1, A0, B1, B0, and B2.
[0515] Example of cache bandwidth constraints in affine mode interleaving prediction
[0516] The reference sample line of the cinematographic element (MC) of a child block in style 0 can be reused in the MC of a spatially adjacent child block in style 1. For example, as Figure 39A As shown, the sub-block in style 0 (with light yellow) and its corresponding sub-block in style 1 (with light blue). Under the suggested cache bandwidth limit, as Figure 39B As shown, the number of reference samples (denoted by N) extracted from the cache buffer of the sub-block in style 0 is:
[0517]
[0518] Where num filter_tap This indicates the number of filter taps, such as 6, 7, or 8, and the Margin is a default non-negative number such as 0 or 1.
[0519] To ensure that the reference sample rows required for the corresponding sub-block in Style 1 are within the N rows required for the sub-block in Style 1, the MV_y of the sub-block in Style 1 is constrained. When the MV_y of the sub-block in Style 1 exceeds the constraint, interleaved affine prediction is not used.
[0520] like Figure 40 As shown, in the interleaved prediction of affine patterns, A0, A1, A2, and A3 represent sub-block rows of style 0, and B0, B1, B2, B3, and B4 represent sub-block rows of style 1.
[0521] For the B0 sub-block line, if MV y (A0)-MV y (B0) > Margin, interleaving prediction without using affine mode;
[0522] For sub-block line B4, if MV y (B4)-MV y (A3) > Margin, interleaved prediction without using affine mode;
[0523] For sub-block rows B1, B2, and B3, which sub-block row in style 0 shares a cache buffer with them is determined by the affine model. Taking sub-block row B1 as an example, if ΔMV y =MV y (B1)-MV y If (A0) > 0, then the predicted sample of the sub-block in row B1 is closer to the sub-block in row A1 than the sub-block in row A0. We can use an affine model to calculate ΔMV. y ΔMV y = dΔx + eΔy. Here, we set... and Where `subblock_width` and `subblock_height` represent the width and height of the subblock. When ΔMV y >0&&MV y (A1)-MV y (B1) > Margin or ΔMV y ≤0&&MV y (B1)-MV y When (A0) > Margin, affine mode interleaving prediction is not used. Furthermore, sub-blocks B2 and B3 are similar.
[0524] In the description above, MV y(Ax) represents the vertical component of the motion vector of the sub-block in mode 0, and MV y (Bx) represents the vertical component of the motion vector of the corresponding sub-block in mode 1.
[0525] 1. It is suggested that gradients (e.g., G(x, y)) can be derived at the sub-block level, where all samples within the sub-block are assigned the same gradient value.
[0526] a. In one example, all samples within a sub-block share the same motion displacement and gradient information. This allows for refinement to be performed at the sub-block level. For sub-block (x... s y s Only one refinement is derived as (Vx(x)). s y s )×Gx(x s y s )+Vy(x s ,y s )×Gy(x s y s )).
[0527] i. In addition, alternatively, the motion displacement V(x, y) is derived at the sub-block level, such as 2x2 or 4x1.
[0528] b. For example, sub-blocks can be defined in different shapes, such as 2x2, 2x1, or 4x1.
[0529] c. For example, the gradient of a sub-block can be calculated using the average sample value of its sub-blocks or a subset of samples within a sub-block.
[0530] d. For example, the gradient Gx(x) of the sub-block s ,y s The gradient can be calculated as the average gradient of each sample or the average gradient of a subset of samples within a sub-block.
[0531] e. In one example, the sub-block size can be predefined or dynamically derived from decoded information (e.g., based on the block size).
[0532] i. Alternatively, the sub-block size indication can be signaled in the bitstream.
[0533] f. In one example, the size of the sub-block used for gradient calculation and the size of the sub-block used for motion vector / motion vector difference can be the same.
[0534] i. Alternatively, the size of the sub-block used for gradient calculation can be smaller than the size of the block used for motion vector / motion vector difference.
[0535] ii. Alternatively, the size of the sub-block used for gradient calculation can be larger than the size of the block used for motion vector / motion vector difference.
[0536] 2. Instead of performing PROF for each list of reference images, PROF (or other prediction refinement encoding / decoding methods that rely on optical flow information) can be applied after performing bidirectional prediction to reduce complexity.
[0537] a. In one example, a provisional prediction signal can first be generated (e.g., using an average / weighted average) based on at least two sets of decoded motion information, and further refinement can be applied based on the generated provisional prediction signal.
[0538] i. In one example, the gradient calculation is derived from the generated temporary prediction signal.
[0539] b. In one example, suppose two predictors from two reference lists of samples located at (x, y) are denoted as P0(x, y) and P1(x, y), respectively. The optical flow from the two lists is denoted as V. 0 (x, y) and V 1 (x, y). Intermediate bidirectional prediction P b (x, y) is derived as follows:
[0540] P b (x, y)=α×P0(x, y)+β×P1(x, y)+o
[0541] Where α and β are weighting factors for the two predictors, and o is the offset. The spatial gradient can be represented using P. b (x, y) derivation. The final bidirectional prediction derivation is as follows:
[0542] P'(x, y) = P b (x, y) + Gx(x, y) × (α × V) 0 x (x, y) + β×V 1 x (x, y))+Gy(x, y)×(α×V 0 y (x, y) + β×V 1 y (x, y))
[0543] 3. It is suggested that the application of encoding / decoding tools related to block affine prediction can be determined by whether certain conditions are met, where these conditions depend on encoded / decoded information such as the block's CPMV and / or the width of the block or sub-block, and / or the height of the block or sub-block, and / or one or more MVs derived for one or more sub-blocks, and / or whether the block is bidirectionally or unidirectionally predicted. In the following discussion, the width and height of the block are represented by W and H, respectively.
[0544] a. The block can be encoded or decoded in at least one of the following modes: affine Merge mode, affine inter-frame mode, affine skip mode, affine Merge mode with motion vector difference (affine MMVD, also known as affine UMVE), or any other mode using affine prediction.
[0545] b. In one example, the encoding / decoding tool could be affine prediction.
[0546] c. In one example, the encoding / decoding tool could be interleaving prediction.
[0547] d. In one example, it is required that if the condition is not met, the encoding / decoding tool cannot be applied to blocks in a consistent bitstream.
[0548] i. Alternatively, when the conditions are not met, another codec tool can be used to replace the current codec tool.
[0549] (i) In one example, the above method is applied only when the current block is encoded or decoded in affine Merge mode.
[0550] e. In one example, if the condition is not met, the codec tool is turned off in the block at the encoder and the decoder.
[0551] i. In one example, if the condition is not met, no signaling notification is given indicating whether the syntax element of the codec tool should be applied in the block.
[0552] (i) When there is no signaling notification syntax element, the codec tool is inferred to be off.
[0553] f. In one example, the codec tool is affine prediction, and applies normal (non-affine) prediction when the conditions are not met.
[0554] i. In one example, one or more CPMVs can be used as the motion vectors of the current block without applying an affine motion model.
[0555] g. In one example, the codec tool uses interleaved prediction, and if the conditions are not met, it applies normal affine prediction (without interleaved prediction).
[0556] i. In one example, when the condition is not met, the predicted sample of the second partitioning pattern in the interleaved prediction is set to be equal to the predicted sample of the first partitioning pattern at the corresponding position.
[0557] h. In one example, the bounding rectangle (or bounding box) is a rectangle derived based on information. Conditions depend on the bounding rectangle for determination. For example, a bounding box or bounding rectangle can represent the smallest rectangle or sample box that includes the entire current block and / or the entire predicted block of the current block. Alternatively, the bounding box or bounding rectangle can be any sample rectangle or square that includes the entire current block and / or the entire predicted block of the current block.
[0558] i. For example, the condition is true if the size of the bounding box (denoted as S) is less than (or not greater than) the threshold (denoted as T), and false if S is not less than (or greater than) T.
[0559] ii. For example, if the size of the bounding box (denoted as S) is less than (or not greater than) a threshold (denoted as T), the condition is false, and if S is not less than (or greater than) T, the condition is true.
[0560] iii. For example, if f(S) is less than (or not greater than) T, the condition is true, and if S is not less than (or greater than) T, the condition is false, where f(S) is a function of S.
[0561] iv. For example, if f(S) is less than (or not greater than) T, the condition is false, and if S is not less than (or greater than) T, the condition is true, where f(S) is a function of S.
[0562] v. In one example, T depends on whether the block is predicted bidirectionally or unidirectionally.
[0563] vi. In one example, f(S) = (S <<p)> >(Lw+Lh), for example, p=7, Lw=log2(W) and Lh=log2(H).
[0564] (i) Alternatively, f(S) = (S*p) / (W*H), for example, p = 128.
[0565] (ii) Alternatively, f(S) = S*p, for example, P = 16*4, or P = 4*4, or P = 8*8.
[0566] vii. In one example, T equals T'.
[0567] viii. In one example, T equals k*T', for example, k = 2.
[0568] ix. In one example, when the block is predicted bidirectionally, T equals T', and when the block is predicted unidirectionally, T equals 2*T'.
[0569] x. In one example, T' = ((Pw + offsetW) * (Ph + offsetH) <<p)> >(Log2(Pw)+Log2(Ph)), for example, p=7.
[0570] (i) Alternatively, T = ((Pw + offsetW) * (Ph + offsetH) * p) / (Pw * Ph), for example, p = 128.
[0571] (ii) Alternatively, T = (Pw + offsetW) * (Ph + offsetH) * W * H.
[0572] (iii) For example, Pw = 16, Ph = 4, offsetW = offsetH = 7.
[0573] (iv) For example, Pw = 4, Ph = 16, offsetW = offsetH = 7.
[0574] (v) For example, Pw = 4, Ph = 4, offsetW = offsetH = 7.
[0575] (vi) For example, Pw = 8, Ph = 8, offsetW = offsetH = 7.
[0576] xi. The left boundary of the bounding rectangle represented by BL can be derived as
[0577] BL=min{x 0 +MV 0 x x 1 +MV 1 x , ..., x N1-1 +MV N1-1 x}+offsetBL,
[0578] Where x k (k = 0, 1, ..., N1-1) is the horizontal coordinate of the top-left corner of the sub-block (denoted as sub-block k), MV k x It is the horizontal component of the MV of sub-block k, and offsetBL is the offset, for example, offsetBL=0.
[0579] xii. The right boundary of the bounding rectangle represented by BR can be derived as
[0580] BR = max{x 0 +MV 0 x x 1 +MV 1 x, ..., x N2-1 +MV N2-1 x +offsetBR
[0581] Where x k (k = 0, 1, ..., N²-1) is the horizontal coordinate of the bottom right corner of the sub-block (denoted as sub-block k), MV k x It is the horizontal component of the MV of sub-block k, and offsetBR is the offset, for example, offsetBR=0.
[0582] xiii. The top boundary of the bounding rectangle represented by BT can be derived as
[0583] BT=min{y 0 +MV 0 y y 1 +MV 1 y , ..., y N3-1 +MV N3-1 y}+offsetBT,
[0584] Where y k (k = 0, 1, ..., N3-1) is the vertical coordinate of the top-left corner of the sub-block (denoted as sub-block k), MV k y It is the vertical component of the MV of sub-block k, and offsetBT is the offset, for example, offsetBT=0.
[0585] xiv. The bottom boundary of the bounding rectangle represented by BB can be derived as follows:
[0586] BB = max{y 0 +MV 0 y y 1 +MV 1 y , ..., y N4-1 +MV N4-1 y}+offsetBB,
[0587] Where x k (k = 0, 1, ..., N⁴⁻¹) is the vertical coordinate of the bottom right corner of the sub-block (denoted as sub-block k), MV k y It is the horizontal component of the MV of sub-block k, and offsetBB is the offset, for example, offsetBB=0.
[0588] xv. The size of the bounding box is derived as (BR-BL+IntW)*(BB-BT+IntH), for example, IntW=IntH=7.
[0589] xvi. Sub-blocks used to derive bounding boxes can be split from blocks in any style specified by interleaving prediction.
[0590] xvii. The N1 sub-blocks used to derive BL can include the four corner sub-blocks divided from the block using the original partitioning pattern of affine projection.
[0591] xviii. The N2 sub-blocks used to derive BR may include the four corner sub-blocks divided from the block using the original partitioning pattern of affine projection.
[0592] xix. The N3 sub-blocks used to derive BT can include the four corner sub-blocks divided from the block using the original partitioning pattern of affine mapping.
[0593] The N4 sub-blocks used to derive BB can include the four corner sub-blocks divided from the block using the affine primitive partitioning pattern.
[0594] xxi. The N1 sub-blocks used to derive BL may include the four corner sub-blocks divided from the block using a second partitioning pattern defined by interleaving prediction.
[0595] xxii. The N2 sub-blocks used to derive the BR may include four corner sub-blocks divided from the block using a second partitioning pattern defined by the interleaving prediction.
[0596] xxiii. The N3 sub-blocks used to derive BT may include four corner sub-blocks divided from the block using a second partitioning pattern defined by interleaving prediction.
[0597] xxiv. The N4 sub-blocks used to derive BB may include four corner sub-blocks divided from the block using a second partitioning pattern defined by interleaving prediction.
[0598] xxv. can further trim at least one of BL, BR, BT, and BB.
[0599] (i) BL and BR are cropped to the range of [startX, endX], where startX and endX can be defined as the X coordinates of the top left and top right samples within the current block's image / piece / strip / block.
[0600] a. In one example, startX is 0 and endX is (picW-1), where picW represents the image width in the luminance sample.
[0601] (ii) BT and BB are cropped to the range of [startY, endY], where Y can be defined as the X coordinates of the top left and top right samples within the current block's image / piece / strip / block.
[0602] a. In one example, startY is 0 and endY is (picH-1), where picH represents the image height in the luminance sample.
[0603] xxvi. Before the MV of a sub-block is used to derive BL and / or BR and / or BT and / or BB, the MV of the sub-block may be rounded or clipped to integer precision.
[0604] (i) For example, MV k x Recalculated as (MV) before being used to derive BL and / or BR and / or BT and / or BB. k x +offset)>>P.
[0605] (ii) For example, MV k y Recalculated as (MV) before being used to derive BL and / or BR and / or BT and / or BB. k y +offset)>>P.
[0606] (iii) For example, offset = 0, p = 4.
[0607] i. In one example, based on the MV of the sub-block x s and MV y s determines the conditions.
[0608] i. In one example, if the MV of a sub-block in the first partition style x Subtract MV x The condition is true if s is greater than (or not less than) the threshold (denoted as T).
[0609] ii. In one example, if MV x s minus the MV of a child block in the first partition style x The condition is true if the value is greater than (or not less than) a threshold (denoted as T).
[0610] iii. In one example, if the MV of a child block in the first partition style y Subtract MV y The condition is true if s is greater than (or not less than) the threshold (denoted as T).
[0611] iv. In one example, if MV ys minus the MV of a child block in the first partition style y The condition is true if the value is greater than (or not less than) a threshold (denoted as T).
[0612] 4. Whether and / or how motion compensation is applied to the sub-block at sub-block index (xSbIdx, ySbIdx) depends on the affine mode indicator, color components, and color format. In the following example, MotionModelIdc[xCb][yCb] equals 0 when the current block is not encoded in affine mode, and MotionModelIdc[xCb][yCb] is greater than 0 when the current block is encoded in affine mode. SubWidthC and SubHeightC are defined according to the color format. For example, SubWidthC and SubHeightC are defined in JVET-R2001 as follows:
[0613]
[0614]
[0615] a. Whether motion compensation is applied separately to the sub-block at sub-block index (xSbIdx, ySbIdx) can depend on the affine mode indicator, color components, and color format.
[0616] i. For example, when the color component is not luminance (in JVET-R2001, the color component ID is not equal to 0), if the current block is not encoded in affine mode (MotionModelIdc[xCb][yCb)).
[0617] If xSbIdx%SubWidthC and ySbIdx%SubHeightC are both equal to 0 in JVET-R2001, then motion compensation is applied individually to the subblock at subblock index (xSbIdx, ySbIdx).
[0618] b. The width and / or height (denoted as cW and / or cH) of the chroma sub-block for motion compensation at sub-block index (xSbIdx, ySbIdx) may depend on the affine mode indicator and color format.
[0619] i. For example, cW = MotionModelIdc[xCb][yCb]? sbWidth: sbWidth / SubWidthC, where sbWidth is the width of the luma sub-block. For example, sbWidth equals 4.
[0620] ii. For example, cH = MotionModelIdc[xCb][yCb]? sbHeight: sbHeight / SubHeightC, where sbHeight is the height of the luma sub-block. For example, sbHeight equals 4.
[0621] 5. Whether and / or how to apply weighted prediction to the sub-block at sub-block index (xSbIdx, ySbIdx) can depend on the affine mode indicator, color components, and color format. In the following example, MotionModelIdc[xCb][yCb] equals 0 when the current block is not encoded in affine mode, and MotionModelIdc[xCb][yCb] is greater than 0 when the current block is encoded in affine mode. SubWidthC and SubHeightC are defined according to the color format. For example, SubWidthC and SubHeightC are defined in JVET-R2001 as follows:
[0622]
[0623]
[0624] a. Whether to apply weighted prediction separately to the sub-block at sub-block index (xSbIdx, ySbIdx) can depend on the affine mode indicator, color components, and color format.
[0625] i. For example, when the color component is not luminance (the color component ID is not equal to 0 in JVET-R2001), if the current block is not encoded in affine mode (MotionModelIdc[xCb][yCb] is equal to 0 in JVET-R2001), or if both xSbIdx%SubWidthC and ySbIdx%SubHeightC are equal to 0, then a weighted prediction is applied separately to the sub-block at the sub-block index (xSbIdx, ySbIdx).
[0626] b. The width and / or height (denoted as cW and / or cH) of the chroma sub-blocks for weighted prediction at sub-block indices (xSbIdx, ySbIdx) may depend on the affine mode indicator and color format.
[0627] i. For example, cW =
[0628] MotionModelIdc[xCb][yCb]? sbWidth: sbWidth / SubWidthC, where sbWidth is the width of the luma sub-block. For example, sbWidth equals 4.
[0629] ii. For example, cH = MotionModelIdc[xCb][yCb]? sbHeight: sbHeight / SubHeightC, where sbHeight is the height of the luma sub-block. For example, sbHeight equals 4.
[0630] 6. The derivation of MotionModelIdc[x][y] can depend on or be conditional on inter_affine_flag[x0][y0] and / or sps_6param_affine_enabled_flag.
[0631] a. In one example, when inter_affine_flag[x0][y0] equals 0, MotionModelIdc[x][y] is set to equal 0.
[0632] b. In one example, when sps_6param_affine_enabled_flag equals 0, MotionModelIdc[x]y] is set to equal 0.
[0633] c. For example, MotionModelIdc[x][y]=inter_affine_flag[x0][y0]? (1+(sps_6param_affine_enabled_flag?cu_affine_type_flag[x0][y0]:0):0
[0634] 6. It is suggested that the application of interleaved prediction can depend on the relationship between the MV (denoted as SBMV1) of the sub-block of mode 1 (denoted as SB1) and the MV (denoted as SBMV0) of the sub-block of mode 0 (denoted as SB0). In the following discussion, it is assumed that the affine model is defined as
[0635]
[0636] The MV at (xPos, yPos) is derived.
[0637] a. Given SB1, SB0 can be selected from a sub-block of pattern 0 that overlaps with SB1.
[0638] i. The choice can depend on the affine parameters.
[0639] ii. For example, if condition A is satisfied, then the center of SB0 is above the center of SB1. If condition B is satisfied, then the center of SB0 is below the center of SB1.
[0640] (i) For example, condition A is SB1 in the last row.
[0641] (ii) For example, condition B is SB1 on the first row.
[0642] (iii) For example, condition B is (mv0_y>0 and mv1_y>0) or (mv0_y>0 and mv1_y<0 and |mv0_y|>|mv1_y|) or (mv0_y<0 and mv1_y>0 and |mv0_y|<|mv1_y|).
[0643] (iv) For example, condition A is satisfied when condition B is not satisfied.
[0644] in
[0645]
[0646] iii. For example, if condition C is satisfied, then the center of SB0 is to the left of the center of SB1. If condition D is satisfied, then the center of SB0 is to the right of the center of SB1.
[0647] (i) For example, condition C is that SB1 is on the rightmost row.
[0648] (ii) For example, condition D is that SB1 is in the leftmost row.
[0649] (iii) For example, condition D is (mv0_x>0 and mv1_x>0) or (mv0_x>0 and mv1_x<0 and |mv0_x|>|mv1_x|) or (mv0_x<0 and mv1_x>0 and |mv0_x|<|mv1_x|)
[0650] (iv) For example, condition C is satisfied when condition D is not satisfied.
[0651] in
[0652] If the center of SB0 is above the center of SB1
[0653]
[0654] ·otherwise
[0655]
[0656] b. If the center of SB0 is to the left of the center of SB1, then align_left equals 1; otherwise, it equals 0. If the center of SB0 is above the center of SB1, then align_top equals 1; otherwise, it equals 0. The selection of SB0 is then based on align_left and align_top, as follows: Figure 41 As shown.
[0657] c. If the relationship between the MVs in SB0 and SB1 satisfies condition E, then disable interleaved prediction.
[0658] i. Condition E may depend on align_left and / or align_top.
[0659] ii. If align_left equals 1, then the condition that satisfies condition E is (SBMV1x>>N)-(SBMV0x>>N)>HMargin1.
[0660] iii. If align_left equals 1, then the condition that satisfies condition E is (SBMV0x>>N)-(SBMV1x>>N)>HMargin2.
[0661] iv. If align_left equals 0, then the condition that satisfies condition E is (SBMV0x>>N)-(SBMV1x>>N)>HMargin3.
[0662] v. If align_left equals 0, then the condition that satisfies condition E is (SBMV1x>>N)-(SBMV0x>>N)>HMargin4.
[0663] vi. If align_top equals 1, then the condition that satisfies condition E is (SBMV1y>>N)-(SBMV0y>>N)>VMargin1.
[0664] vii. If align_top equals 1, then the condition that satisfies condition E is (SBMV0y>>N)-(SBMV1y>>N)>VMargin2.
[0665] viii. If align_top equals 0, then the condition that satisfies condition E is (SBMV0x>>N)-(SBMV1y>>N)>VMargin3.
[0666] ix. If align_top equals 0, then the condition that satisfies condition E is (SBMV1y>>N)-(SBMV0y>>N)>VMargin4.
[0667] xN represents the MV precision, for example, N=2 or 4.
[0668] xi.HMargin1~HMargin4 and VMargin1~VMargin4 are integers, which may depend on encoding and decoding information, such as block size or whether the block is bidirectional or unidirectional prediction.
[0669] (i) If the block is bidirectionally predicted, then HMargin1 = HMargion3 = 6, HMargin2 = HMargion4 = 8 + 6.
[0670] (ii) If the block is predicted in one direction, then HMargin1 = HMargion3 = 4, HMargin2 = HMargion4 = 8 + 4.
[0671] (iii) If the block is bidirectionally predicted, then VMargin1 = VMargion3 = 4, HMargin2 = HMargion4 = 8 + 4.
[0672] (iv) If the block is predicted in one direction, then HMargin1 = HMargion3 = 1, HMargin2 = HMargion4 = 8 + 1.
[0673] (v) If the block size is 8x8, then HMargin1 = HMargion3 = 6, HMargin2 = HMargion4 = 8 + 6.
[0674] (vi) If the block size is 4x4, then HMargin1 = HMargion3 = 4, HMargin2 = HMargion4 = 8 + 4.
[0675] (vii) If the block size is 8x8, then VMargin1 = VMargion3 = 4, HMargin2 = HMargion4 = 8 + 4.
[0676] (viii) If the block size is 4x4, then HMargin1 = HMargin3 = 1.
[0677] HMargin2 = HMargion4 = 8 + 1.
[0678] 5. Examples
[0679] In this patent application, the deletion of text is indicated by double brackets (e.g., [[]]), with the deleted text located between the double brackets; newly added text is indicated by bold italic text.
[0680] 5.1 Example 1: Limiting bandwidth by disabling interleaving affines
[0681] Figure 31 and Figure 32Two partitioning patterns in interleaved prediction are illustrated. Assume mvA0, mvB0, mvC0, and mvD0 represent the motion vectors of sub-blocks A, B, C, and D in partitioning pattern 0; and mvA1, mvB1, mvC1, and mvD1 represent the motion vectors of sub-blocks A, B, C, and D in partitioning pattern 1. Assume the width and height of each sub-block in partitioning pattern 0 are subwith and subheihgt, respectively. The Unified Prefetch Integer Sample Number (UPISN) is derived as follows: 1)
[0683] 2)
[0685] 3)
[0687]
[0688] For unidirectional predictions with interleaved predictions, if sampleNum is greater than 129536, the prediction sample for partition pattern 1 is set to be equal to the prediction sample for partition pattern 0.
[0689] For bidirectional forecasting using interleaved forecasting,
[0690] For the MV derivation of list 0, if sampleNum0 is greater than 64768, then set the predicted sample of partition style 1 of list 0 to be equal to the predicted sample of partition style 0.
[0691] For the MV derivation of sampleNum1 in list 1, if sampleNum1 is greater than 64768, then set the predicted sample of partition style 1 in list 1 to be equal to the predicted sample of partition style 0.
[0692] 5.2. Example 2: Limiting bandwidth by disabling affine mapping
[0693] sampNum, sampNum0, and sampNum1 are derived in the same manner as in Example 1.
[0694] The requirement is that in a consistent bitstream,
[0695] For unidirectional forecasts using interleaved forecasts, sampNum should be less than or equal to 129536.
[0696] For bidirectional forecasts using interleaved forecasts, sampNum0 and sampNum1 should be less than or equal to 64768.
[0697] 5.3. Example 3: Motion compensation of chromaticity components in affine mode
[0698] These changes are based on JVET-R2001-v10
[0699] 8.5.6 Inter-frame block decoding process
[0700] 8.5.6.1 Overview
[0701] …
[0702] For each codec subblock at subblock index (xSbIdx, ySbIdx), where xSbIdx = 0..numSbX-1 and ySbIdx = 0..numSbY-1, the following applies:
[0703] – Luminance position (xSb, ySb), specifies the top-left sample of the current codec sub-block relative to the top-left luminance sample of the current image, derived as follows:
[0704] (xSb, ySb) = (xCb+xSbIdx*sbWidth, yCb+ySbIdx*sbHeight) (923)
[0705] – For X values of 0 and 1 respectively, when predFlagLX[xSbIdx][ySbIdx] equals 1, the following applies:
[0706] – composed of an ordered two-dimensional array of brightness samples, refPicLX L The two ordered two-dimensional arrays refPicLX for the chromaticity samples Cb and refPicLX Cr The resulting reference image is derived by calling the procedure specified in Clause 8.5.6.2, with X and refIdxLX as input.
[0707] – The motion vector offset mvOffset is set to be equal to refMvLX[xSbIdx][xSbIdx]-mvLX[xSbIdx][ySbIdx].
[0708] – If cIdx equals 0, the following applies:
[0709] – array predSamplesLX L It is derived by calling the fractional sample interpolation procedure specified in Clause 8.5.6.3, where the luminance position (xSb, ySb), the codec sub-block width sbWidth, the codec sub-block height sbHeight in the luminance sample, the luminance motion vector offset mvOffset, the refined luminance motion vector refMvLX[xSbIdx][xSbIdx], and the reference array refPicLX LbdofFlag, dmvrFlag, hpelIfIdx, cIdx, RprConstraintsActive[X][refIdxLX], and RefPicScale[X][refIdxLX] are used as inputs.
[0710] – When cbProfFlagLX equals 1, invoke the prediction refinement process using optical flow specified in Clause 8.5.6.4, where the array predSamplesLX is composed of sbWidth, sbHeight, and (sbWidth+2)x(sbHeight+2). L The motion vector difference array diffMvLX[xIdx][yIdx] (where xIdx = 0..cbWidth / numSbX-1, yIdx = 0..cbHeight / numSbY-1) is used as input, and a refined (sbWidth)x(sbHeight) array predSamplesLX is used as input. L As output.
[0711]
[0712] Otherwise, if If cIdx equals 1, then the following applies:
[0713] – array predSamplesLX Cb It is derived by calling the fractional sample interpolation procedure specified in Clause 8.5.6.3, where the luminance position (xSb, ySb) and the width of the encoded / decoded sub-block are... sbWidth / SubWidthC, Encoder / Decoder Subblock Height sbHeight / SubHeightC, chroma motion vector offset mvOffset, refined chroma motion vector refMvLX[xSbIdx][ySbIdx], reference array refPicLX Cb bdofFlag, dmvrFlag, hpelIfIdx, cIdx, RprConstraintsActive[X][refIdxLX], and RefPicScale[X][refIdxLX] are used as inputs.
[0714] – Otherwise (cIdx equals 2), the following applies:
[0715] – array predSamplesLX Cr It is derived by calling the fractional sample interpolation procedure specified in Clause 8.5.6.3, where the luminance position (xSb, ySb) and the width of the encoded / decoded sub-block are... sbWidth / SubWidthC, Encoder / Decoder Subblock Height sbHeight / SubHeightC, chroma motion vector offset mvOffset, thin chroma motion vector refMvLX[xSbIdx][xSbIdx], reference array refPicLX Cr bdofFlag, dmvrFlag, hpelIfIdx, cIdx, RprConstraintsActive[X][refIdxLX], and RefPicScale[X][refIdxLX] are used as inputs.
[0716] …[Standard text continued]
[0717] 5.4. Example 4: Motion compensation of chromaticity components in affine mode
[0718] 8.5.6 Inter-frame block decoding process
[0719] 8.5.6.1 Overview
[0720] This procedure is invoked when decoding a codec unit encoded and decoded in inter-frame prediction mode.
[0721] The input to this process is:
[0722] – Luminance position (xCb, yCb), specifies the top-left sample of the current codec block relative to the top-left luminance sample of the current image.
[0723] – The variable cbWidth specifies the width of the current encoded block in the luma sample.
[0724] – The variable cbHeight specifies the height of the current codec block in the luminance sample.
[0725] – Variables numSbX and numSbY specify the number of luminance codec sub-blocks in the horizontal and vertical directions, respectively.
[0726] – Motion vectors mvL0[xSbIdx][ySbIdx] and mvL1[xSbIdx][ySbIdx], where xSbIdx = 0..numSbX-1, ySbIdx = 0..numSbY-1,
[0727] – Refine the motion vectors refMvL0[xSbIdx][ySbIdx] and refMvL1[xSbIdx][ySbIdx], where xSbIdx = 0..numSbX-1, ySbIdx = 0..numSbY-1,
[0728] –Refer to indices refIdxL0 and refIdxL1,
[0729] – The prediction list uses the flags predFlagL0[xSbIdx][ySbIdx] and predFlagL1[xSbIdx][ySbIdx], where xSbIdx = 0..numSbX-1, ySbIdx = 0..numSbY-1,
[0730] – Half-sample interpolation filter index hpelIfIdx,
[0731] – Bidirectional prediction weight index bcwIdx
[0732] – The minimum sum of absolute differences during motion vector refinement on the decoder side: dmvrSad[xSbIdx][ySbIdx], where xSbIdx = 0..numSbX-1, ySbIdx = 0..numSbY-1,
[0733] – Decoder-side motion vector refinement flag dmvrFlag
[0734] – The variable cIdx specifies the color component index of the current block.
[0735] –Prediction refinement utilizes the flags cbProfFlagL0 and cbProfFlagL1.
[0736] – Motion vector difference arrays diffMvL0[xIdx][yIdx] and diffMvL1[xIdx][yIdx], where xIdx = 0..cbWidth / numSbX-1, yIdx = 0..cbHeight / numSbY-1.
[0737] The output of this process is:
[0738] – An array of predicted samples, predSamples.
[0739] Let predSamplesL0 L ,predSamplesL1 L and predSamplesIntra L It is an array of (cbWidth) x (cbHeight) predicted brightness sample values, and predSamplesL0 is set to... Cb ,predSamplesL1 Cb 、predSamplesL0 Cr and predSamplesL1 Cr 、predSamplesIntra Cband predSamplesIntra Cr It is an array of (cbWidth / SubWidthC)x(cbHeight / SubHeightC) predicting the chromaticity sample values.
[0740] – The variable currPic specifies the current image, and the derivation of the variable bdofFlag is as follows:
[0741] – bdofFlag is set to TRUE if all of the following conditions are true.
[0742] –ph_bdof_disabled_flag equals 0.
[0743] Both predFlagL0[xSbIdx][ySbIdx] and predFlagL1[xSbIdx][ySbIdx] are equal to 1.
[0744] –DiffPicOrderCnt(currPic, RefPicList[0][refIdxL0]) is equal to DiffPicOrderCnt(RefPicList[1][refIdxL1], currPic).
[0745] –RefPicList[0][refIdxL0] is STRP, RefPicList[1][refIdxL1] is STRP.
[0746] –MotionModelIdc[xCb][yCb] equals 0.
[0747] –merge_subblock_flag[xCb][yCb] equals 0.
[0748] –sym_mvd_flag[xCb][yCb] equals 0.
[0749] –ciip_flag[xCb][yCb] equals 0.
[0750] –BcwIdx[xCb][yCb] equals 0.
[0751] Both `luma_weight_l0_flag[refIdxL0]` and `luma_weight_l1_flag[refIdxL1]` are equal to 0.
[0752] Both chroma_weight_l0_flag[refIdxL0] and chroma_weight_l1_flag[refIdxL1] are equal to 0.
[0753] –cbWidth is greater than or equal to 8.
[0754] –cbHeight is greater than or equal to 8.
[0755] –cbHeight*cbWidth is greater than or equal to 128.
[0756] –RprConstraintsActive[0][refIdxL0] equals 0, RprConstraintsActive[1][refIdxL1] equals 0.
[0757] –cIdx equals 0.
[0758] Otherwise, bdofFlag is set to equal FALSE.
[0759] – If numSbY equals 1 and numSbX equals 1, then the following applies:
[0760] – When bdofFlag equals TRUE, the variables numSbY and numSbX are modified as follows:
[0761] numSbX=(cbWidth>16)? (cbWidth>>4):1(919)
[0762] numSbY=(cbHeight>16)? (cbHeight>>4): 1(920)
[0763] – For X = 0..1, xSbIdx = 0..numSbX-1 and ySbIdx = 0..numSbY-1, the following applies:
[0764] –predFlagLX[xSbIdx][ySbIdx] is set to be equal to predFlagLX[0][0].
[0765] –refMvLX[xSbIdx][ySbIdx] is set to be equal to refMvLX[0][0].
[0766] –mvLX[xSbIdx][ySbIdx] is set to be equal to mvLX[0][0].
[0767] The width and height sbWidth and sbHeight of the current codec sub-block in the luminance sample are derived as follows:
[0768] sbWidth=cbWidth / numSbX (921)
[0769] sbHeight=cbHeight / numSbY (922)
[0770] For each codec subblock at subblock index (xSbIdx, ySbIdx), where xSbIdx = 0..numSbX-1 and ySbIdx = 0..numSbY-1, the following applies:
[0771] – Luminance position (xSb, ySb), specifies the top-left sample of the current codec sub-block relative to the top-left luminance sample of the current image, derived as follows:
[0772] (xSb, ySb) = (xCb+xSbIdx*sbWidth, yCb+ySbIdx*sbHeight) (923)
[0773] – For X values of 0 and 1 respectively, when predFlagLX[xSbIdx][ySbIdx] equals 1, the following applies:
[0774] – composed of an ordered two-dimensional array of brightness samples, refPicLX L The two ordered two-dimensional arrays refPicLX for the chromaticity samples Cb and refPicLX Cr The resulting reference image is derived by calling the procedure specified in Clause 8.5.6.2, with X and refIdxLX as input.
[0775] – The motion vector offset mvOffset is set to be equal to refMvLX[xSbIdx][xSbIdx]-mvLX[xSbIdx][ySbIdx].
[0776] – If cIdx equals 0, the following applies:
[0777] – array predSamplesLX L It is derived by calling the fractional sample interpolation procedure specified in Clause 8.5.6.3, where the luminance position (xSb, ySb), the codec sub-block width sbWidth, the codec sub-block height sbHeight in the luminance sample, the luminance motion vector offset mvOffset, the refined luminance motion vector refMvLX[xSbIdx][xSbIdx], and the reference array refPicLX L bdofFlag, dmvrFlag, hpelIfIdx, cIdx, RprConstraintsActive[X][refIdxLX], and RefPicScale[X][refIdxLX] are used as inputs.
[0778] – When cbProfFlagLX equals 1, invoke the prediction refinement process using optical flow specified in Clause 8.5.6.4, where the array predSamplesLX is composed of sbWidth, sbHeight, and (sbWidth+2)x(sbHeight+2). L The motion vector difference array diffMvLX[xIdx][yIdx] (where xIdx = 0..cbWidth / numSbX-1, yIdx = 0..cbHeight / numSbY-1) is used as input, and a refined (sbWidth)x(sbHeight) array predSamplesLX is used as input. L As output.
[0779]
[0780] Otherwise, if If cIdx equals 1, then the following applies:
[0781] – array predSamplesLX Cb It is derived by calling the fractional sample interpolation procedure specified in Clause 8.5.6.3, where the luminance position (xSb, ySb) and the width of the encoded / decoded sub-block are... sbWidth / SubWidthC, Encoding / Decoding Subblock Height sbHeight / SubHeightC, chroma motion vector offset mvOffset, refined chroma motion vector refMvLX[xSbIdx][ySbIdx], reference array refPicLX Cb The following parameters are used as inputs: bdofFlag, dmvrFlag, hpelIfIdx, cIdx, RprConstraintsActive[X][refIdxLX], and RefPicScale[X][refIdxLX]. Otherwise (if cIdx equals 2), the following applies:
[0782] – array predSamplesLX Cr It is derived by calling the fractional sample interpolation procedure specified in Clause 8.5.6.3, where the luminance position (xSb, ySb) and the width of the encoded / decoded sub-block are... sbWidth / SubWidthC, Encoder / Decoder Subblock Height sbHeight / SubHeightC, chroma motion vector offset mvOffset, thin chroma motion vector refMvLX[xSbIdx][xSbIdx], reference array refPicLX Cr bdofFlag, dmvrFlag, hpelIfIdx, cIdx, RprConstraintsActive[X][refIdxLX], and RefPicScale[X][refIdxLX] are used as inputs.
[0783] The variable sbBdofFlag is set to equal FALSE.
[0784] – When bdofFlag equals TRUE, the variable sbBdofFlag is further modified as follows:
[0785] – If dmvrFlag equals 1 and the variable dmvrSad[xSbIdx][ySbIdx] is less than (2*sbWidth*sbHeight), then the variable sbBdofFlag is set to FALSE.
[0786] Otherwise, the variable sbBdofFlag is set to TRUE.
[0787] – The derivation of the array predSamples for predicting samples is as follows:
[0788] – If cIdx equals 0, then the predicted samples predSamples[x] within the current luminance codec subblock are... L +xSb][y L +ySb], where x L =0..sbWidth-1 and y L =0..sbHeight-1, the derivation is as follows:
[0789] – If sbBdofFlag equals TRUE, then invoke the bidirectional optical flow sample prediction procedure specified in Clause 8.5.6.5, where nCbW is set to equal the luminance codec subblock width sbWidth, nCbH is set to equal the luminance codec subblock height sbHeight, and the sample array predSamplesL0 is... L and predSamplesL1 L The variables predFlagL0[xSbIdx][ySbIdx], predFlagL1[xSbIdx][ySbIdx], refIdxL0, and refIdxL1 are used as inputs, and predSamples[x L +xSb][yL +ySb] is used as the output.
[0790] - Otherwise (sbBdofFlag equals FALSE), invoke the weighted sample prediction procedure specified in Item 8.5.6.6, where the luma codec sub-block width sbWidth, the luma codec sub-block height sbHeight, and the sample array predSamplesL0 are used. L and predSamplesL1 L The variables predFlagL0[xSbIdx][ySbIdx], predFlagL1[xSbIdx][ySbIdx], refIdxL0, refIdxL1, bcwIdx, dmvrFlag, and cIdx are used as inputs, and the variable predSamples[x L +xSb][y L +ySb] is used as the output.
[0791]
[0792] Otherwise, if If cIdx equals 1, then the predicted samples within the current chroma component Cb encoding / decoding subblock are derived by calling the weighted sample prediction process specified in Clause 8.5.6.6, predSamples[x C +xSb / SubWidthC][y C +ySb / SubHeightC], where x C =0.. And y C =0.. Where nCbW is set to equal to sbWidth / SubWidthC and nCbH are set to equal to sbHeight / SubHeightC, sample array predSamplesL0 Cb and predSamplesL1 Cb The variables predFlagL0[xSbIdx][ySbIdx], predFlagL1[xSbIdx][ySbIdx], refIdxL0, refIdxL1, bcwIdx, dmvrFlag, and cIdx are used as inputs.
[0793] – Otherwise (cIdx equals 2), derive the predicted samples within the current chroma component Cr encoding / decoding subblock by invoking the weighted sample prediction procedure specified in Clause 8.5.6.6, predSamples[x C+xSb / SubWidthC][y C +ySb / SubHeightC], where x C =0.. And y C =0.. Where nCbW is set to equal to sbWidth / SubWidthC and nCbH are set to equal to sbHeight / SubHeightC, sample array predSamplesL0 Cr and predSamplesL1 Cr The variables predFlagL0[xSbIdx][ySbIdx], predFlagL1[xSbIdx][ySbIdx], refIdxL0, refIdxL1, bcwIdx, dmvrFlag, and cIdx are used as inputs.
[0794] – When cIdx equals 0, for x = 0..sbWidth-1 and y = 0..sbHeight-1, the following assignments are made:
[0795] MvL0[xSb+x][ySb+y]=mvL0[xSbIdx][ySbIdx] (924)
[0796] MvL1[xSb+x][ySb+y]=mvL1[xSbIdx][ySbIdx] (925)
[0797] MvDmvrL0[xSb+x][ySb+y]=refMvL0[xSbIdx][ySbIdx] (926)
[0798] MvDmvrL1[xSb+x][ySb+y]=refMvL1[xSbIdx][ySbIdx] (927)
[0799] RefIdxL0[xSb+x][ySb+y]=refIdxL0 (928)
[0800] RefIdxL1[xSb+x][ySb+y]=refIdxL1 (929)
[0801] PredFlagL0[xSb+x][ySb+y]=predFlagL0[xSbIdx][ySbIdx] (930)
[0802] PredFlagL1[xSb+x][ySb+y]=predFlagL1[xSbIdx][ySbIdx] (931)
[0803] HpelIfIdx[xSb+x][ySb+y]=hpelIfIdx (932)
[0804] BcwIdx[xSb+x][ySb+y]=bcwIdx (933)
[0805] When ciip_flag[xCb][yCb] equals 1, the array predSamples of the predicted samples is modified as follows:
[0806] – If cIdx equals 0, the following applies:
[0807] – Invoke the generic intra-sample prediction procedure specified in Clause 8.4.5.2.5, where the position (xTbCmp, yTbCmp) is set to equal to (xCb, yCb), the intra-prediction mode predModeIntra is set to equal to INTRA_PLANAR, the transform block width nTbW and height nTbH are set to equal to cbWidth and cbHeight, the codec block width nCbW and height nCbH are set to equal to cbWidth and cbHeight, and the variable cIdx is used as input, and the output is assigned to the (cbWidth)x(cbHeight) array predSamplesIntra. L .
[0808] – Invoke the weighted sample prediction procedure specified in Clause 8.5.6.7 for combination merging and intra-frame prediction, where the position (xTbCmp, yTbCmp) is set to equal to (xCb, yCb), the codec block width cbWidth, the codec block height cbHeight, and the sample arrays predSamplesInter and predSamplesIntra are set to equal to predSamples and predSamplesIntra, respectively. L The color component index cIdx is taken as input, and the output is assigned to the (cbWidth)x(cbHeight) array predSamples.
[0809] Otherwise, if cIdx equals 1 and cbWidth / SubWidthC is greater than or equal to 4, then the following applies:
[0810] – Invoke the generic intra-sample prediction procedure specified in Clause 8.4.5.2.5, where the position (xTbCmp, yTbCmp) is set to equal to (xCb / SubWidthC, yCb / SubHeightC), the intra-prediction mode predModeIntra is set to equal to INTRA_PLANAR, the transform block width nTbW and height nTbH are set to equal to cbWidth / SubWidthC and cbHeight / SubHeightC, the codec block width nCbW and height nCbH are set to equal to bWidth / SubWidthC and cbHeight / SubHeightC, and the variable cIdx is used as input, and the output is assigned to the array predSamplesIntra (cbWidth / SubWidthC) x (cbHeight / SubHeightC). Cb .
[0811] – Invoke the weighted sample prediction procedure specified in Clause 8.5.6.7 for combining Merge and intra-prediction, where the position (xTbCmp, yTbCmp) is set to equal to (xCb, yCb), the codec block width cbWidth / SubWidthC, the codec block height cbHeight / SubHeightC, and the sample arrays predSamplesInter and predSamplesIntra are each set to equal to predSamples. Cb and predSamplesIntra Cb The color component index cIdx is taken as input, and the output is assigned to the array predSamples (cbWidth / SubWidthC)x(cbHeight / SubHeightC).
[0812] Otherwise, if cIdx equals 2 and cbWidth / SubWidthC is greater than or equal to 4, then the following applies:
[0813] – Invoke the generic intra-sample prediction procedure specified in Clause 8.4.5.2.5, where the position (xTbCmp, yTbCmp) is set to equal to (xCb / SubWidthC, yCb / SubHeightC), the intra-prediction mode predModeIntra is set to equal to INTRA_PLANAR, the transform block width nTbW and height nTbH are set to equal to cbWidth / SubWidthC and cbHeight / SubHeightC, the codec block width nCbW and height nCbH are set to equal to bWidth / SubWidthC and cbHeight / SubHeightC, and the variable cIdx is used as input, and the output is assigned to the array (cbWidth / SubWidthC)x(cbHeight / SubHeightC)predSamplesIntra. Cr .
[0814] – Invoke the weighted sample prediction procedure specified in Clause 8.5.6.7 for combining Merge and intra-prediction, where the position (xTbCmp, yTbCmp) is set to equal to (xCb, yCb), the codec block width cbWidth / SubWidthC, the codec block height cbHeight / SubHeightC, and the sample arrays predSamplesInter and predSamplesIntra are each set to equal to predSamples. Cr and predSamplesIntra Cr The color component index cIdx is taken as input, and the output is assigned to the array predSamples (cbWidth / SubWidthC)x(cbHeight / SubHeightC).
[0815] 5.5. Example 5: Derivation of MotionModelIdc
[0816] `cu_affine_type_flag[x0][y0]` equal to 1 specifies that for the current codec unit, when decoding P or B stripes, motion compensation based on a 6-parameter affine model is used to generate the prediction samples for the current codec unit. `cu_affine_type_flag[x0][y0]` equal to 0 specifies that motion compensation based on a 4-parameter affine model is used to generate the prediction samples for the current codec unit.
[0817] MotionModelIdc[x][y] represents the motion model of the encoding / decoding unit, as shown in Table 15. The array indices x and y specify the position (x, y) of the brightness sample relative to the top-left brightness sample of the image.
[0818] For x = x0..x0 + cbWidth-1 and y = y0..y0 + cbHeight-1, the derivation of the variable MotionModelIdc[x][y] is as follows:
[0819] – If general_merge_flag[x0][y0] equals 1, then the following applies:
[0820] MotionModelIdc[x][y]=merge_subblock_flag[x0][y0] (165)
[0821] – Otherwise (general_merge_flag[x0][y0] equals 0), the following applies:
[0822] [[MotionModelIdc[x][y]=inter_affine_flag[x0][y0]+cu_affine_type_flag[x0][y0]]]
[0823]
[0824] Table 15 - Explanation of MotionModelIdc[x0][y0]
[0825] 0 Translational motion 1 Four-parameter affine motion 2 Six-parameter affine motion
[0826] 5.6. Example 5: 6-tap interpolation filtering for interleaving prediction
[0827] Text changes based on AVS-N2865
[0828] 9.9.2.3 Affine Brightness Sample Interpolation Process
[0829] Luminance interpolation filter coefficients of affine prediction template 1
[0830]
[0831]
[0832] 5.7. Example 6: Cache Bandwidth Constraints Based on Interleaving Prediction
[0833] Text changes based on AVS-N2865
[0834] 9.19 Derivation of the motion vector array of the affine motion unit sub-block
[0835] If there are 3 motion vectors in the affine control point motion vector group, the motion vector group is represented as mvsAffine(mv0, mv1, mv2); otherwise (if there are 2 motion vectors in the affine control point motion vector group), the motion vector group is represented as MVSAFFINEmvsAffine(mv0, mv1).
[0836] a) Calculate the variables dHorX, dVerX, dHorY, and dVerY:
[0837]
[0838] b) If mvsAffine has 3 motion vectors, then:
[0839]
[0840] c) Otherwise (mvsAffine has 2 motion vectors):
[0841]
[0842] (xE, yE) represents the position of the top-left sample of the current prediction unit's brightness prediction block within the current image's brightness sample matrix. The width and height of the current prediction unit are width and height, respectively, and the width and height of each sub-block are subwidth and subheight, respectively. See [link to documentation]. Figure 40 The sub-block containing the top-left sample of the current prediction unit's brightness prediction block is A, the sub-block containing the top-right sample is B, and the sub-block containing the bottom-left sample is C.
[0843] c) Initialize OutBandwidthConstrain to 0, and initialize AlignTop, AlignLeft, AlignLeftFirstRow and AlignLeftLastRow to 1.
[0844] b) If the prediction reference mode of the current prediction unit is 'Pred_List01' or AffineSubblockSizeFlag is equal to 1, then subwidth and subheight are both equal to 8, MarginVer is equal to 4, MarginHor is equal to 6, and (x, y) is the coordinate of the top-left corner of the sub-block. Calculate the motion vector of each luma sub-block:
[0845] d) For template 0:
[0846] 1) If the current sub-block is A, then xPos and yPos are both equal to 0;
[0847] 2) Otherwise, if the current sub-block is B, then xPos equals width and yPos equals 0;
[0848] 3) Otherwise, if the current sub-block is C and there are 3 motion vectors in mvsAffine, then xPos equals 0 and yPos equals height;
[0849] 4) Otherwise, xPos equals (x-xE)+4, and yPos equals (y-yE)+4;
[0850] 5) The motion vector mvE0 of the current sub-block:
[0851]
[0852] For template 1:
[0853] 1) If the current sub-block is A, then xPos and yPos are both equal to 0;
[0854] 2) Otherwise, if the current sub-block is B, then xPos equals width and yPos equals 0;
[0855] 3) Otherwise, if the current sub-block is C and there are 3 motion vectors in mvsAffine, then xPos equals 0 and yPos equals height;
[0856] 4) Otherwise,
[0857] a) If y is less than yE + subheight / 2, then xPos equals (x - xE) + 4, and yPos equals (y - yE) + 2;
[0858] b) Otherwise, if x is less than xE + subwidth / 2, then xPos equals (x - xE) + 2, and yPos equals (y - yE) + 4;
[0859] c) Otherwise, if x is greater than xE+width-subwidth / 2 and y is greater than yE+height-subheight / 2 or x is less than xE+subwidth / 2 and y is greater than yE+height-subheight / 2, then xPos equals (x-xE)+2 and yPos equals (y-yE)+2.
[0860] d) Otherwise, if y is greater than yE+height-subheight / 2, then xPos equals (x-xE)+4 and yPos equals (y-yE)+2;
[0861] e) Otherwise, if x is greater than xE + width - subwidth / 2, then xPos equals (x - xE) + 2, and yPos equals (y - yE) + 4;
[0862] f) Otherwise, xPos equals (x-xE)+4, and yPos equals (y-yE)+4;
[0863] 5) The motion vector mvE1 of the current sub-block:
[0864]
[0865] 1) If the sub-block is in the first row, AlignTop equals 0; if the sub-block is in the last row, AlignTop equals 1; otherwise, calculate mv0_y and mv1_y. If mv0_y > 0 and mv1_y > 0, or mv0_y > 0 and mv1_y < 0 and |mv0_y| > |mv1_y|, or mv0_y < 0 and mv1_y > 0 and |mv0_y| < |mv1_y|, then AlignTop equals 0; otherwise, AlignTop equals 1.
[0866]
[0867] 1) If the child block is in the first column, AlignLeft equals 0; if the child block is in the last column, AlignLeft equals 1; otherwise, calculate mv0_x and mv1x. If AlignTop equals 1:
[0868]
[0869] otherwise:
[0870]
[0871] If mv0_x>0 and mv1_x>0, or mv0_x>0 and mv1_x<0 and |mv0_x|>|mv1_x|, or mv0_x<0 and mv1_x>0 and |mv0_x|<|mv1_x|, then AlignLeft equals 0; otherwise, AlignLeft equals 1.
[0872] Based on AlignTop and AlignLeft, the motion vector of the corresponding sub-block in template 0 is obtained and recorded as follows: Figure 44 (mvx_pattern0, mvy_pattern0) in the example.
[0873] 1) If AlignLeft equals 1 and (mvE1_x>>4)-(mvx_pattern0>>4)>MarginHor or (mvx_pattern0>>4)-(mvE1_x>>4)>8+MarginHor, or AlignLeft equals 0 and (mvx_pattern0>>4)-(mvE1_x>>4)>MarginHor or (mvE1_x>>4)-(mvx_pattern0>>4)>8+MarginHor, then OutBandwidthConstrain equals 1; If AlignTop equals 1 and (mvE1_y>>4)–(mvy_pattern0>>4)>MarginVer or (mvy_pattern0>>4)-(mvE1_y>>4)>8+MarginVer, or AlignTop equals 0 and (mvy_pattern0>>4)-(mvE1_y>>4)>MarginVer or (mvE1_y>>4)-(Mvy_pattern0>>4)>8+MarginVer, then OutBandwidthConstrain equals 1.
[0874] a) If the prediction reference mode of the current prediction unit is 'Pred_List0' or 'Pred_List1' and AffineSubblockSizeFlag is equal to 0, then subwidth and subheight are both equal to 4, MarginVer is equal to 1, MarginHor is equal to 4, and (x, y) is the top-left corner position of the sub-block. Calculate the motion vector for each luma sub-block:
[0875] b) For template 0:
[0876] 1) If the current sub-block is A, then xPos and yPos are both equal to 0;
[0877] 2) Otherwise, if the current sub-block is B, then xPos equals width and yPos equals 0;
[0878] 1) Otherwise, if the current sub-block is C and there are 3 motion vectors in mvAffine, then xPos equals 0 and yPos equals height;
[0879] 2) Otherwise, xPos equals (x-xE)+2, and yPos equals (y-yE)+2;
[0880] 3) The motion vector mvE0 of the current sub-block:
[0881]
[0882] For template 1:
[0883] 1) If the current sub-block is A, then xPos and yPos are both equal to 0;
[0884] 2) Otherwise, if the current sub-block is B, then xPos equals width and yPos equals 0;
[0885] 3) Otherwise, if the current sub-block is C and there are 3 motion vectors in mvsAffine, then xPos equals 0 and yPos equals height;
[0886] 4) Otherwise,
[0887] a) If y is less than yE + subheight / 2, then xPos equals (x - xE) + 2, and yPos equals (y - yE) + 1;
[0888] b) Otherwise, if x is less than xE + subwidth / 2, then xPos equals (x - xE) + 1, and yPos equals (y - yE) + 2;
[0889] c) Otherwise, if x is greater than xE+width-subwidth / 2 and y is greater than yE+height-subheight / 2 or x is less than xE+subwidth / 2 and y is greater than yE+height-subheight / 2, then xPos equals (x-xE)+1 and yPos equals (y-yE)+1.
[0890] d) Otherwise, if y is greater than yE+height-subheight / 2, then xPos equals (x-xE)+2 and yPos equals (y-yE)+1;
[0891] e) Otherwise, if x is greater than xE + width - subwidth / 2, then xPos equals (x - xE) + 1, and yPos equals (y - yE) + 2;
[0892] f) Otherwise, xPos equals (x-xE)+2, and yPos equals (y-yE)+2;
[0893] 5) The motion vector mvE1 of the current sub-block:
[0894]
[0895] 1) If the sub-block is in the first row, AlignTop equals 0; if the sub-block is in the last row, AlignTop equals 1; otherwise, calculate mv0_y and mv1_y. If mv0_y > 0 and mv1_y > 0, or mv0_y > 0 and mv1_y < 0 and |mv0_y| > |mv1_y|, or mv0_y < 0 and mv1_y > 0 and |mv0_y| < |mv1_y|, then AlignTop equals 0; otherwise, AlignTop equals 1.
[0896]
[0897] 1) If the child block is in the first column, AlignLeft equals 0; if the child block is in the last column, AlignLeft equals 1; otherwise, calculate mv0_x and mv1_x. If AlignTop equals 1:
[0898]
[0899] otherwise:
[0900]
[0901] If mv0_x>0 and mv1_x>0, or mv0_x>0 and mv1_x<0 and |mv0_x|>|mv1_x|, or mv0_x<0 and mv1_x>0 and |mv0_x|<|mv1_x|, then AlignLeft equals 0; otherwise, AlignLeft equals 1.
[0902] Based on AlignTop and AlignLeft, the motion vector of the corresponding sub-block in template 0 is obtained and recorded as follows: Figure 45 (mvx_pattern0, mvy_pattern0) in the example.
[0903] If AlignLeft equals 1 and (mvE1_x>>4)-(mvx_pattern0>>4)>MarginHor or (mvx_pattern0>>4)-(mvE1_x>>4)>4+MarginHor, or AlignLeft equals 0 and (mvx_pattern0>>4)-(mvE1_x>>4)>MarginHor or (mvE1_x>>4)-(mvx_pattern0>>4)>4+MarginHor, then OutBandwidthConstrain equals 1. If AlignTop equals 1 and (mvE1_y>>4)–(mvy_pattern0>>4)>MarginVer or (mvy_pattern0>>4)-(mvE1_y>>4)>4+MarginVer, or AlignTop equals 0 and (mvy_pattern0>>4)-(mvE1_y>>4)>MarginVer or (mvE1_y>>4)-(mvy_pattern0>>4)>4+MarginVer, then OutBandwidthConstrain equals 1.
[0904] 1) If the prediction model of the current prediction unit is an interleaved affine motion model, then obtain the motion vectors of the four sub-blocks A0, B0, C0, D0 corresponding to template 0 and A1, B1, C1 corresponding to template 1 according to the above method, and the D1 motion vector of the four sub-blocks, and calculate the number of integer pixels sampled by the prediction unit, sampleNum:
[0905] 1) According to the method in d) or e), obtain the motion vectors mvA0, mvB0, mvC0, and mvD0 corresponding to the four sub-blocks A0, B0, C0, and D0 of template 0, respectively, as well as the height cornerWidth0 and width cornerHeight0 of the four sub-blocks.
[0906]
[0907] 1) Obtain the motion vectors mvA1, mvB1, mvC1 and mvD1 corresponding to the four sub-blocks A1, B1, C1 and D1 of template 1, respectively, according to the method in d) or e), as well as the height cornerWidth1 and width cornerHeight1 of the four sub-blocks;
[0908]
[0909] a)
[0910] 1) Number of integer pixels prefetched for the current prediction unit:
[0911]
[0912]
[0913] 9.20 Derivation of Interleaved Affine Prediction Weights
[0914] Weightrounding equals 2, weightshift equals 2. Remember that the width of the current block is width, and the height is height. The derivation of the interleaved affine prediction weights is as follows:
[0915] b) When subblockwidth equals 4 and subblockheight equals 4:
[0916] a) Iterate through all sub-blocks in template 0, and the value of weight0 in each 4x4 sub-block is as follows: Figure 46 As shown.
[0917] 1) Traverse all sub-blocks in template 1, marking the top-left corner coordinates of each sub-block unit as (x, y), and the width and height as sub_w and sub_h:
[0918] 1. All weight1 values in the sub-block are initialized to 1;
[0919] 2. weight1[x+sub_w / 2][y+sub_h / 2] equals 3;
[0920] 3. weight1[x+sub_w / 2-1][y+sub_h / 2] equals 3;
[0921] 4. weight1[x+sub_w / 2][y+sub_h / 2-1] equals 3;
[0922] 5. weight1[x+sub_w / 2-1][y+sub_h / 2-1] equals 3;
[0923] 6. If x < 2 and y < 2, x < 2 and y > height - 2, x > width - 2 and y < 2, or x > width - 2 and y > height - 2, then the weight0 of all positions within the sub-block is equal to 4, and the weight1 is equal to 0.
[0924] 2) Iterate through all positions in the current block for weight0 and weight1. If weight0 and weight1 are equal at the same position, then weight0 and weight1 at that position are both equal to 2;
[0925] b) When subblockwidth equals 8 and subblockheight equals 8:
[0926] Iterate through all sub-blocks in template 0, and the value of weight0 in each 8x8 sub-block is as follows: Figure 47 As shown.
[0927] 1) Traverse all sub-blocks in template 1, marking the top-left corner coordinates of each sub-block unit as (x, y), and the width and height as sub_w and sub_h:
[0928] 1. All weight1 values in the sub-block are initialized to 1;
[0929] 2. weight1[x+sub_w / 2][y+sub_h / 2] equals 3, weight1[x+sub_w / 2+1][y+sub_h / 2] equals 3, weight1[x+sub_w / 2][y+sub_h / 2+1] equals 3, weight1[x+sub_w / 2+1][y+sub_h / 2+1] equals 3;
[0930] 3. weight1[x+sub_w / 2-2][y+sub_h / 2] equals 3, weight1[x+sub_w / 2-1][y+sub_h / 2] equals 3, weight1[x+sub_w / 2-2][y+sub_h / 2+1] equals 3, weight1[x+sub_w / 2-1][y+sub_h / 2+1] equals 3;
[0931] 4. weight1[x+sub_w / 2][y+sub_h / 2-2] equals 3, weight1[x+sub_w / 2+1][y+sub_h / 2-2] equals 3, weight1[x+sub_w / 2][y+sub_h / 2-1] equals 3, weight1[x+sub_w / 2+1][y+sub_h / 2-1] equals 3;
[0932] 5. weight1[x+sub_w / 2-2][y+sub_h / 2-2] equals 3, weight1[x+sub_w / 2-1][y+sub_h / 2-2] equals 3, weight1[x+sub_w / 2-2][y+sub_h / 2-1] equals 3, weight1[x+sub_w / 2-1][y+sub_h / 2-1] equals 3;
[0933] 2) Iterate through all positions in the current block for weight0 and weight1. If weight0 and weight1 are equal at the same position, then weight0 and weight1 at that position are both equal to 2;
[0934] Figure 33 This is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values), or in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0935] System 1900 may include codec component 1904, which may implement the various codec or encoding methods described in this document. Codec component 1904 may reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of codec component 1904 may be stored or transmitted via communication through the connection represented by component 1906. Component 1908 may use the stored or communicated bitstream (or codec) representation of the video received at input 1902 to generate pixel values or displayable video to be sent to display interface 1910. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it should be understood that codec tools or operations are used at the encoder, and the decoder will perform a corresponding decoding tool or operation that reverses the result of the codec.
[0936] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be found in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0937] Figure 34This is a block diagram of a video processing apparatus 3600. Apparatus 3600 can be used to implement one or more methods described herein. Apparatus 3600 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. Processor 3602 can be configured to implement one or more methods described in this document. Memory 3604 can be used to store data and code for implementing the methods and techniques described herein. Video processing hardware 3606 can be used to implement some of the techniques described in this document in hardware circuitry.
[0938] Figure 36 This is a block diagram illustrating an example video codec system 100 that can utilize the techniques disclosed herein.
[0939] like Figure 36 As shown, the video encoding / decoding system 100 may include a source device 110 and a target device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The target device 120 can decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device.
[0940] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0941] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations of these sources. Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded images and associated data. An encoded image is an encoded representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to target device 120 via I / O interface 116 through network 130a. Encoded video data may also be stored on storage medium / server 130b for access by target device 120.
[0942] The target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.
[0943] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with target device 120, or may be external to target device 120 configured to interface with an external display device.
[0944] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other current and / or further standards.
[0945] Figure 37 It shows that it can be Figure 36 A block diagram of an example of a video encoder 200 in system 100, showing a video encoder 114.
[0946] The video encoder 200 can be configured to perform any or all of the technologies disclosed herein. Figure 37 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor may be configured to perform any or all of the techniques described in this disclosure.
[0947] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.
[0948] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.
[0949] Furthermore, some components, such as motion estimation unit 204 and motion compensation unit 205, can be highly integrated, but for illustrative purposes... Figure 37 The examples represent the examples respectively.
[0950] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0951] The mode selection unit 203 can, for example, select a coding / decoding mode (intra-frame or inter-frame) based on the error result, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 207 to generate residual block data, and provide it to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combined intra-frame and inter-frame prediction (CIIP) mode, where the prediction is based on the inter-frame prediction signal and the intra-frame prediction signal. The mode selection unit 203 can also select the resolution of the motion vector for the block (e.g., sub-pixel or integer pixel precision) in the case of inter-frame prediction.
[0952] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of the images from buffer 213 (excluding images associated with the current video block).
[0953] For example, the motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block depending on whether the current video block is in an I-strip, a P-strip, or a B-strip.
[0954] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Then, motion estimation unit 204 can generate a reference index indicating the reference image in list 0 or list 1 containing the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0955] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in the reference images in list 1. Then, motion estimation unit 204 can generate reference indices indicating the reference images in lists 0 and 1 containing the reference video blocks, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0956] In some examples, the motion estimation unit 204 can output a complete set of motion information for the decoder's decoding processing.
[0957] In some examples, motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, motion estimation unit 204 may signal the motion information of the current video block to another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0958] In one example, the motion estimation unit 204 may instruct the video decoder 300 in the syntax structure associated with the current video block to indicate that the current video block has the same motion information as another video block.
[0959] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) within the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0960] As described above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0961] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples of other video blocks in the same frame. The prediction data for the current video block may include the predicted video block and various syntax elements.
[0962] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) one or more predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.
[0963] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.
[0964] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0965] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0966] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to the corresponding sample of one or more predicted video blocks generated by prediction unit 202 to produce a reconstructed video block associated with the current block, which is stored in buffer 213.
[0967] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0968] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0969] Figure 38 It shows that it can be Figure 36 A block diagram of an example of video decoder 300 in system 100, showing video decoder 114.
[0970] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 37 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0971] exist Figure 38 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform operations related to the video encoder 200 ( Figure 37 The encoding sequence described is roughly the opposite of the decoding sequence.
[0972] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information from the entropy-encoded video data, including motion vectors, motion vector precision, reference image list index, and other motion information. For example, motion compensation unit 302 can determine such information by executing AMVP and Merge modes.
[0973] The motion compensation unit 302 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. Identifiers of the interpolation filters used at sub-pixel precision can be included in the syntax elements.
[0974] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during video block encoding to calculate the interpolation of sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and use the interpolation filter to generate the prediction block.
[0975] The motion compensation unit 302 may use some syntax information to determine the size of the blocks used to encode one or more frames and / or one or more stripes of the encoded video sequence, segmentation information describing how each macroblock of the picture of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.
[0976] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially neighboring blocks. Inverse quantization unit 303 performs inverse quantization (i.e., dequantization) on the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.
[0977] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0978] The following provides a list of preferred solutions for some embodiments.
[0979] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 1).
[0980] 1. A video processing method (e.g., Figure 35 The method described in the text (3500) includes: a conversion between video blocks and the codec representation of the video, determining (3502) the gradient of the prediction vector at the sub-block level of the video block using a rule specifying that the same gradient value is used for all samples in each sub-block, and performing (3504) the conversion based on the determination.
[0981] 2. According to the method of Solution 1, the rule specifies that all samples in each sub-block of the video block use the same motion displacement and the same gradient value.
[0982] 3. The method of any of the solutions 1-2, wherein the video block is divided into sub-blocks, wherein at least some of the sub-blocks have a size of 2x2, 2x1 or 4x1 pixels.
[0983] 4. The method of any of the solutions 1-3, wherein the same gradient value is determined by averaging the sample values of at least some samples of the sub-block.
[0984] 5. The method of any of the solutions 1-3, wherein the same gradient value is determined as the average of the gradients for each sample or the average of the gradients from a subset of samples within a sub-block.
[0985] 6. The same gradient is determined based on the method of any of the solutions 1-5, wherein the first sub-block has a different size than the second sub-block used to determine the motion vector of the video block.
[0986] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 2).
[0987] 7. A video processing method comprising: for a conversion between a current video block of a video and a bitstream representation of the video, determining, according to rules, predictive refinement using motion information employing optical flow techniques; and performing the conversion based on the determination, wherein the rules specify that predictive refinement is performed after performing bidirectional prediction to obtain motion information.
[0988] 8. According to the method of Solution 7, the prediction refinement is calculated by determining the time prediction signal using at least two decoded motion information values and refining the time prediction signal.
[0989] 9. According to the method of Solution 7, the prediction refinement is calculated by determining the first prediction sub-value from the first reference image list and the second prediction sub-value from the second prediction sub-list, as well as the offset.
[0990] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 3).
[0991] 10. A video processing method comprising: a conversion between a current video block and a bitstream representation of the video, determining that an encoding / decoding tool is enabled for the current video block because the current video block satisfies a condition; and performing the conversion based on the determination, wherein the condition relates to a control point motion vector of the current video block or the size of the current video block or a prediction mode of the current video block.
[0992] 11. The method of Solution 10, wherein the current video block is represented in the codec representation using affine prediction-based encoding and decoding.
[0993] 12. According to the method of any of the solutions 10-11, wherein determining further includes determining to disable the encoding / decoding tools for the current video block because the current video block does not meet the conditions.
[0994] 13. The method of any of the solutions 10-12, wherein the current video block is represented in the codec representation using interleaved prediction or affine prediction.
[0995] 14. The method of any of the solutions 10-13, wherein the condition is determined based on the bounding box of the current video block using one of the following rules: (a) the condition is considered satisfied only if the size of the bounding box is higher than a first threshold, or (b) the condition is considered satisfied only if the size of the bounding box is lower than a second threshold.
[0996] 15. According to the method of solution 14, the value of the first threshold or the second threshold depends on whether one-way prediction or two-way prediction is used to represent the current video block in the encoded and decoded representation.
[0997] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., items 5 and 6).
[0998] 16. A video processing method comprising: for a conversion between a video block and a encoded / decoded representation of the video, determining, according to a rule, whether and how to compute a prediction of a sub-block of the video block; and performing the conversion based on the determination, wherein the rule is based on one or more of whether an affine mode is enabled for the video block or the color component to which the video block belongs or the color format of the video.
[0999] 17. According to the method of solution 16, wherein the rule relates to the size of the sub-blocks to which weighted predictions are performed.
[1000] 18. The method according to any of the solutions 1-17, wherein performing the conversion includes encoding the video to generate a encoded / decoded representation.
[1001] 19. The method of any of the solutions 1-17, wherein performing the conversion includes parsing and decoding the encoded / decoded representation to generate a video.
[1002] 20. A video decoding apparatus, including a processor configured to implement the methods described in one or more of solutions 1 to 19.
[1003] 21. A video encoding apparatus, including a processor configured to implement the methods described in one or more of solutions 1 to 19.
[1004] 22. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the methods described in any of solutions 1 to 19.
[1005] 23. A method, apparatus or system described in this document.
[1006] Figure 48 This is a flowchart of example method 4800 for video processing. Operation 4802 includes a transformation between video blocks and the video bitstream, determining the gradient of the prediction vector at the sub-block level of the video block according to a rule, wherein the rule specifies the use of the same gradient value assigned to all samples within the sub-blocks of the video block. Operation 4804 includes performing the transformation based on the determined gradient.
[1007] In some embodiments of method 4800, the rule specifies that all samples in sub-blocks of a video block use the same motion displacement information and the same gradient values. In some embodiments of method 4800, the refinement operation performed at the sub-block level includes deriving a refinement of the sub-block (xs, ys) (Vx(xs, ys) × Gx(xs, ys) + Vy(xs, ys) × Gy(xs, ys)), where Gx(xs, ys) and Gy(xs, ys) are the gradients of the sub-block, and where Vx(xs, ys) and Vy(xs, ys) are the motion displacement vectors of the sub-block. In some embodiments of method 4800, the motion displacement vector V(x, y) of the video block is derived at the sub-block level. In some embodiments of method 4800, the video block is divided into sub-blocks, wherein at least some sub-blocks have a size of 2x2 pixels, 2x1 pixels, or 4x1 pixels. In some embodiments of method 4800, the same gradient value is determined by averaging the sample values of multiple samples of a sub-block or by averaging the sample values of some samples among multiple samples of a sub-block. In some embodiments of method 4800, the same gradient value is determined by averaging each gradient of multiple samples of a sub-block, or by averaging the gradients of some samples among multiple samples of a sub-block.
[1008] In some embodiments of method 4800, the rule specifies that the size of the sub-block is predefined, derived based on information included in the bitstream, or indicated in the bitstream. In some embodiments of method 4800, the rule specifies that the size of the sub-block is derived based on the size of the video block. In some embodiments of method 4800, sub-blocks of the same size are used to determine the same gradient value and to determine the motion vector or motion vector difference.
[1009] Figure 49 This is a flowchart of example method 4900 for video processing. Operation 4902 includes, for the conversion between the current video block and the video bitstream, determining the prediction refinement using optical flow technology after performing bidirectional prediction techniques to obtain motion information of the current video block. Operation 4904 includes performing the conversion based on the determination.
[1010] In some embodiments of method 4900, prediction refinement using optical flow techniques is employed by first determining a temporal prediction signal using at least two decoded motion information values and then refining the temporal prediction signal. In some embodiments of method 4900, the gradient of the current video block is derived from the temporal prediction signal. In some embodiments of method 4900, prediction refinement using optical flow techniques includes performing sub-block-based affine motion compensation to obtain prediction samples, refining the prediction samples by adding differences derived from the optical flow equation, and the bidirectional prediction technique includes using two lists of reference images to generate prediction units from a weighted average of the two sample blocks.
[1011] Figure 50 This is a flowchart of an example method 5000 for video processing. Operation 5002 includes, for a conversion between a current video block and a video bitstream, determining whether to apply encoding / decoding tools related to affine prediction to the current video block based on whether the current video block satisfies a condition. Operation 5004 includes performing the conversion based on the determination, wherein the condition relates to one or more control point motion vectors of the current video block, or a first size of the current video block, or a second size of a sub-block of the current video block, or one or more motion vectors derived for one or more sub-blocks of the current video block, or a prediction mode of the current video block.
[1012] In some embodiments of method 5000, an affine prediction-based codec is used to represent the current video block in the bitstream. In some embodiments of method 5000, the codec tool includes an affine prediction tool. In some embodiments of method 5000, the codec tool includes an interleaving prediction tool. In some embodiments of method 5000, in response to a condition not being met, it is determined that the codec tool is not applied to the current video block in the coherent bitstream. In some embodiments of method 5000, in response to a condition not being met, the codec tool is replaced with another codec tool applied to the current video block. In some embodiments of method 5000, in response to a condition not being met and in response to encoding / decoding the current video block using an affine Merge mode, the codec tool is replaced with another codec tool applied to the current video block. In some embodiments of method 5000, in response to a condition not being met, the codec tool is turned off or disabled for the current video block at the encoder and decoder.
[1013] In some embodiments of method 5000, in response to a condition not being met, the bitstream does not include syntax elements indicating whether the encoding / decoding tool is applied to the current video block. In some embodiments of method 5000, in response to the absence of syntax elements in the bitstream, it is inferred that the encoding / decoding tool is turned off or disabled. In some embodiments of method 5000, in response to a condition being met, the encoding / decoding tool is an affine prediction tool, and wherein in response to a condition not being met, the encoding / decoding tool is a prediction tool other than an affine prediction tool. In some embodiments of method 5000, one or more control point motion vectors are used as the motion vector for the current video block without applying an affine motion model to the current video block. In some embodiments of method 5000, in response to a condition being met, the encoding / decoding tool is an interleaving prediction tool, and wherein in response to a condition not being met, the encoding / decoding tool is a prediction tool other than an affine prediction tool. In some embodiments of method 5000, in response to a condition not being met, the prediction sample of the second partitioning style in the interleaving prediction tool is set to be equal to the prediction sample of the first partitioning style at the corresponding position.
[1014] In some embodiments of method 5000, a condition is determined based on the size of the bounding box of the current video block. In some embodiments of method 5000, the condition is satisfied in response to the size of the bounding box being less than or equal to a threshold, and the condition is not satisfied in response to the size of the bounding box being greater than the threshold. In some embodiments of method 5000, the condition is based on a first motion vector in the horizontal direction of a plurality of sub-blocks of the current video block and a second motion vector in the vertical direction of a plurality of sub-blocks of the current video block. In some embodiments of method 5000, the condition is true in response to the result of a third motion vector in the horizontal direction of a sub-block of the current video in a first partitioning pattern minus the first motion vector being greater than the threshold. In some embodiments of method 5000, the condition is true in response to the result of a first motion vector minus a third motion vector in the horizontal direction of a sub-block of the current video in a first partitioning pattern being greater than the threshold. In some embodiments of method 5000, the condition is true in response to the result of a fourth motion vector in the vertical direction of a sub-block of the current video in a first partitioning pattern minus the second motion vector being greater than the threshold. In some embodiments of method 5000, the condition is true in response to the second motion vector minus the fourth motion vector in the vertical direction of a sub-block of the current video in the first partitioning pattern being greater than a threshold.
[1015] Figure 51 This is a flowchart of an example method 5100 for video processing. Operation 5102 includes, for the conversion between the current video block and the video bitstream, determining whether and how to apply motion compensation to the sub-block at the sub-block index, based on the following rules: (1) whether to apply an affine mode to the current video block, wherein the bitstream includes an affine motion indicator indicating whether an affine mode is applied, (2) the color component to which the current video block belongs, and (3) the color format of the video. Operation 5104 includes performing the conversion based on the determination.
[1016] In some embodiments of method 5100, in response to the following, a rule specifies that motion compensation is applied individually to the sub-block at the sub-block index: (1) the color component is the luma component, or (2) the affine mode is not applied to the current video block, or (3) the remainder when the horizontal sub-block index is divided by the width of the current video block is equal to 0 and the second remainder when the vertical sub-block index is divided by the height of the current video block is equal to 0. In some embodiments of method 5100, the rule specifies that the first width or first height of the chroma sub-block used to perform motion compensation at the sub-block index of the chroma sub-block of the video depends on the affine mode indication and the color format of the video. In some embodiments of method 5100, in response to an affine mode indication indicating that the affine mode is applied to the current video block, the rule specifies that the first width of the chroma sub-block is the second width of the luma sub-block, and in response to an affine mode indication indicating that the affine mode is not applied to the current video block, the rule specifies that the first width of the chroma sub-block is the second width of the luma sub-block divided by the third width of the current video block. In some embodiments of method 5100, in response to an affine mode indication indicating that an affine mode is applied to the current video block, a rule specifies that the first height of the chroma sub-block is the second height of the luma sub-block, and in response to an affine mode indication indicating that an affine mode is not applied to the current video block, a rule specifies that the first height of the chroma sub-block is the second height of the luma sub-block divided by the third height of the current video block.
[1017] Figure 52 This is a flowchart of an example method 5200 for video processing. Operation 5202 includes, for the conversion between the current video block and the video bitstream, determining, according to rules, whether and how to use weighted prediction to compute predictions for sub-blocks of the current video block based on: (1) whether an affine mode is applied to the current video block, wherein the bitstream includes an affine motion indicator indicating whether an affine mode is applied, (2) the color component to which the current video block belongs, and (3) the color format of the video. Operation 5204 includes performing the conversion based on the determination.
[1018] In some embodiments of method 5200, in response to the following, a rule specifies that a weighted prediction is applied individually to the sub-block at the sub-block index: (1) the color component is not a luma component, or (2) an affine mode is not applied to the current video block, or (3) a second remainder of the horizontal sub-block index divided by the width of the current video block is equal to 0 and a second remainder of the vertical sub-block index divided by the height of the current video block is equal to 0. In some embodiments of method 5200, the rule specifies that a first width or a first height of the chroma sub-block used to perform weighted prediction at the sub-block index of the chroma sub-block of the video depends on the affine mode indication and the color format of the video. In some embodiments of method 5200, in response to an affine mode indication indicating that an affine mode is applied to the current video block, the rule specifies that a first width of the chroma sub-block is a second width of the luma sub-block, and in response to an affine mode indication indicating that an affine mode is not applied to the current video block, the rule specifies that a first width of the chroma sub-block is the second width of the luma sub-block divided by a third width of the current video block.
[1019] In some embodiments of method 5200, in response to an affine mode indication indicating that an affine mode is applied to the current video block, a rule specifies that the first height of the chroma sub-block is the second height of the luma sub-block, and in response to an affine mode indication indicating that an affine mode is not applied to the current video block, a rule specifies that the first height of the chroma sub-block is the second height of the luma sub-block divided by the third height of the current video block.
[1020] Figure 53 This is a flowchart of an example method 5300 for video processing. Operation 5302 includes, for the conversion between a current video block and a bitstream of video, determining the value of a first syntax element in the bitstream according to a rule, wherein the value of the first syntax element indicates whether an affine mode is applied to the current video block, and wherein the rule specifies that the value of the first syntax element is based on: (1) a second syntax element in the bitstream, which indicates whether motion compensation based on an affine model is used to generate a predicted sample of the current video block, or (2) a third syntax element in the bitstream, which indicates whether motion compensation based on a 6-parameter affine model is enabled for a codec layer video sequence (CLVS). Operation 5304 includes performing the conversion based on the determination.
[1021] In some embodiments of method 5300, in response to the second syntax element being equal to 0, the rule specifies that the value of the first syntax element is set to zero. In some embodiments of method 5300, in response to the third syntax element being equal to 0, the rule specifies that the value of the first syntax element is set to zero.
[1022] Figure 54This is a flowchart of example method 5400 for video processing. Operation 5402 includes, for the conversion between the current video block and the video bitstream, determining whether to enable the interleaving prediction tool for the current video block based on the relationship between the first motion vector of the first sub-block of the first style of the current video block and the second motion vector of the second sub-block of the second style of the current video block. Operation 5404 includes performing the conversion based on the determination.
[1023] In some embodiments of method 5400, a second sub-block is selected from a sub-block of a second style, wherein the sub-block overlaps with a first sub-block. In some embodiments of method 5400, the second sub-block is selected from a sub-block based on affine parameters. In some embodiments of method 5400, in response to a first condition, the second center of the second sub-block is located above the first center of the first sub-block, and in response to a second condition, the second center of the second sub-block is located below the first center of the first sub-block. In some embodiments of method 5400, the first condition is that the first sub-block is in the last line of the current video block. In some embodiments of method 5400, the second condition is that the first sub-block is in the first line of the current video block. In some embodiments of method 5400, the second condition is: (mv0_y>0 and mv1_y>0), or (mv0_y>0 and mv1_y<0 and |mv0_y|>|mv1_y|), or (mv0_y<0 and mv1_y>0 and |mv0_y|<|mv1_y|), where mv0_y is a second motion vector in the vertical direction, and where mv1_y is a first motion vector in the vertical direction. In some embodiments of method 5400, the first condition is satisfied in response to the failure to satisfy the second condition.
[1024] In some embodiments of method 5400, in response to satisfying a third condition, the second center of the second sub-block is located to the left of the first center of the first sub-block, and in response to satisfying a fourth condition, the second center of the second sub-block is located to the right of the first center of the first sub-block. In some embodiments of method 5400, the third condition is that the first sub-block is in the rightmost row of the current video block. In some embodiments of method 5400, the fourth condition is that the first sub-block is in the leftmost row of the current video block. In some embodiments of method 5400, in response to the relationship between the first motion vector and the second motion vector satisfying a fifth condition, the interleaving prediction tool is disabled. In some embodiments of method 5400, the fifth condition depends on a first syntax element and a second syntax element in the bitstream, the first syntax element indicating whether the center of the first sub-block is to the left of the center of the second sub-block, and the second syntax element indicating whether the center of the first sub-block is above the center of the second sub-block.
[1025] In some embodiments of methods 4800-5400(one or more), performing the conversion includes encoding video into a bitstream. In some embodiments of methods 4800-5400(one or more), performing the conversion includes generating a bitstream from video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium. In some embodiments of methods 4800-5400(one or more), performing the conversion includes decoding video from the bitstream. In some embodiments, a video decoding apparatus includes a processor configured to implement one or more methods described in embodiments 4800-5400(one or more). In some embodiments, a video encoding apparatus includes a processor configured to implement one or more methods described in embodiments 4800-5400(one or more). In some embodiments, a computer program product having computer instructions stored thereon, which, when executed by a processor, cause the processor to implement one or more methods described in embodiments 4800-5400(one or more). In some embodiments, a non-transitory computer-readable storage medium stores a bitstream generated according to one or more methods described in embodiments 4800-5400(one or more). In some embodiments, a non-transitory computer-readable storage medium stores instructions that enable a processor to execute one or more methods described in embodiments 4800-5400. In some embodiments, a bitstream generation method includes: generating a bitstream of video according to one or more methods described in embodiments 4800-5400, and storing the bitstream on a computer-readable program medium. In some embodiments, a method, apparatus, or bitstream generated according to the disclosed method or system described in this document.
[1026] Some embodiments of the disclosed technology involve making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement that tool or mode in the processing of video blocks, but may not necessarily modify the resulting bitstream based on the use of that tool or mode. That is, when a video processing tool or mode is enabled based on a decision or determination, the conversion from a video block to a bitstream representation of the video will use the video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that it has been modified based on the video processing tool or mode. That is, the conversion from a bitstream representation of the video to a video block will be performed using the video processing tool or mode enabled based on a decision or determination.
[1027] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, the encoder will not use the tool or mode in the conversion from video block to bitstream representation of video. In another example, when a video processing tool or mode is disabled, the decoder will process the bitstream knowing that the bitstream has not been modified based on the decision or determination that the disabled video processing tool or mode has been used.
[1028] The disclosures and other solutions, examples, embodiments, modules, and functional operations described in this document may be implemented in digital electronic circuits or in computer software, firmware, or hardware (including the structures disclosed in this document and their structural equivalents, or combinations thereof). The disclosed embodiments and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a storage device, a composition of substances influencing machine-readable propagation signals, or combinations thereof. The term "data processing apparatus" includes all means, devices, and machines for processing data, such as programmable processors, computers, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or combinations thereof. Propagation signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.
[1029] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), a single file dedicated to the program in question, or multiple coordination files (e.g., files storing one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected via a communication network.
[1030] The processes and logic flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. Processing and logic flows can also be executed by dedicated logic circuitry, and the devices can be implemented as dedicated logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).
[1031] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors in any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include, or be operatively coupled to, receiving data from or transferring data to one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or both. However, a computer does not need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[1032] Although this patent document contains numerous details, these details should not be construed as limiting the scope of any subject matter or potentially claimed content, but rather as descriptions of features of specific embodiments that may be specific to particular technologies. Certain features described in this patent document within the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although the foregoing features may be described as functioning in certain combinations, or even initially claimed to be so, in some cases, one or more features from the claimed combination may be removed from the combination, and the claimed combination may be directed to sub-combinations or variations thereof.
[1033] Similarly, although operations are described in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order or sequence shown, or requiring all illustrated operations to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[1034] Only some implementation methods and examples are described, and other implementation methods, enhancements and variations can be made based on the content described and explained in this patent document.
Claims
1. A method for processing video data, comprising: For the conversion between the current video block and the video bitstream, the decision to apply fractional sample interpolation to the color components of the sub-blocks of the current video block is based on whether certain conditions are met. The color component is a chromaticity component; and The conversion is performed based on the determination. The condition mentioned is one of the following: The current video block is not encoded or decoded in affine mode, or The remainders of xSbIdx divided by SubWidthC and ySbIdx divided by SubHeightC are both 0, where xSbIdx is the horizontal index of the sub-block, ySbIdx is the vertical index of the sub-block, and SubWidthC and SubHeightC represent the color format of the current image including the current video block.
2. The method according to claim 1, wherein, When the condition is met, the fractional sample interpolation process is enabled for the color components of the sub-block; when the condition is not met, the fractional sample interpolation process is disabled for the color components of the sub-block.
3. The method according to claim 1, wherein, When the fractional sample interpolation process is enabled, the size of the color component of the sub-block used in the fractional sample interpolation process is based on whether the current video block is encoded and decoded in affine mode.
4. The method according to claim 3, wherein, When the current video block is encoded and decoded in the affine mode, the dimensions of the color components of the sub-block used in the fractional sample interpolation process are sbWidth and sbHeight, where sbWidth and sbHeight specify the width of the sub-block in units of luminance samples and the height of the sub-block in units of luminance samples, respectively.
5. The method according to claim 3, wherein, When the current video block is not encoded or decoded in the affine mode, the dimensions of the color components of the sub-block used in the fractional sample interpolation process are sbWidth / SubWidthC and sbHeight / SubHeightC, where sbWidth and sbHeight specify the width and height of the sub-block in units of luminance samples, respectively.
6. The method according to any one of claims 1-5, wherein, Performing the conversion includes encoding the video into the bitstream.
7. The method according to any one of claims 1-5, wherein, Performing the conversion includes decoding the video from the bitstream.
8. The method according to claim 1, further comprising: For the first conversion between the first video block and the bitstream of the video, the gradient of the prediction vector at the sub-block level of the first video block is determined according to a first rule. The first rule specifies the use of the same gradient value assigned to all samples within the first sub-block of the first video block; and The first conversion is performed based on the determination.
9. The method according to claim 8, wherein, The first rule specifies that all samples in the first sub-block of the first video block use the same motion displacement information and the same gradient value.
10. The method according to claim 9, wherein, The refinement operation performed at the first sub-block level includes deriving a refinement of the first sub-block (xs, ys) (Vx(xs, ys) × Gx(xs, ys) + Vy(xs, ys) × Gy(xs, ys)), where Gx(xs, ys) and Gy(xs, ys) are the gradients of the first sub-block, and where Vx(xs, ys) and Vy(xs, ys) are the motion displacement vectors of the first sub-block.
11. The method according to claim 9, wherein, The motion displacement vector V(x, y) of the first video block is derived at the sub-block level.
12. The method according to claim 8, wherein, The first video block is divided into multiple sub-blocks, wherein at least some of the multiple sub-blocks have a size of 2x2 pixels, 2x1 pixels, or 4x1 pixels.
13. The method according to claim 8, wherein, The same gradient value is determined by averaging the sample values of multiple samples in the first sub-block or by averaging the sample values of some of the multiple samples in the first sub-block.
14. The method of claim 8, wherein the same gradient value is determined by averaging the gradients of each of the plurality of samples of the first sub-block, or by averaging the gradients of some of the plurality of samples of the first sub-block.
15. The method according to claim 8, wherein, The first rule specifies that the size of the first sub-block is predefined, derived based on information included in the bitstream, or indicated in the bitstream.
16. The method according to claim 15, wherein, The first rule specifies that the size of the first sub-block is derived based on the size of the first video block.
17. The method according to claim 8, wherein, Sub-blocks of the same size are used to determine the same gradient value and to determine the motion vector or motion vector difference.
18. The method according to claim 1, further comprising: For the second conversion between the second video block of the video and the bitstream of the video, after performing bidirectional prediction technology to obtain motion information of the second video block, it is determined to use prediction refinement using optical flow technology; as well as The second conversion is performed based on the determination.
19. The method according to claim 18, wherein, The prediction refinement using optical flow technology is achieved by first determining the time prediction signal using at least two decoded motion information values and then refining the time prediction signal.
20. The method according to claim 19, wherein, The gradient of the second video block is derived from the time prediction signal.
21. The method according to claim 19, The prediction refinement using optical flow technology includes performing sub-block-based affine motion compensation to obtain prediction samples, and refining the prediction samples by adding a difference derived from the optical flow equation. The bidirectional prediction technique described therein involves using two lists of reference images to generate prediction units from a weighted average of two sample blocks.
22. The method according to claim 1, further comprising: For the third conversion between the third video block and the bitstream of the video, it is determined whether to apply the encoding and decoding tools related to affine prediction to the third video block based on whether the third video block satisfies the third condition; as well as Based on the determination, the third transformation is performed. The third condition refers to one or more control point motion vectors of the third video block, or a first size of the third video block, or a second size of a sub-block of the third video block, or one or more motion vectors derived for one or more sub-blocks of the third video block, or a prediction mode of the third video block.
23. The method according to claim 22, wherein, The third video block is represented in the bitstream using an affine prediction-based codec.
24. The method according to claim 22, wherein, The encoding / decoding tools include affine prediction tools.
25. The method according to claim 22, wherein, The encoding / decoding tools include an interleaving prediction tool.
26. The method according to claim 22, wherein, In response to the failure to meet the third condition, it is determined that the encoding / decoding tool is not applied to the third video block in the consistent bitstream.
27. The method according to claim 26, wherein, In response to the failure to meet the third condition, the codec tool is replaced with another codec tool applied to the third video block.
28. The method according to claim 27, wherein, In response to the failure to meet the third condition and in response to encoding / decoding the third video block using an affine Merge mode, the encoding / decoding tool is replaced with another encoding / decoding tool applied to the third video block.
29. The method according to claim 22, wherein, In response to the failure to meet the third condition, the encoding / decoding tools are turned off or disabled for the third video block at both the encoder and decoder.
30. The method according to claim 29, wherein, In response to the failure to meet the third condition, the bitstream does not include a syntax element indicating whether the encoding / decoding tool is applied to the third video block.
31. The method according to claim 30, wherein, In response to the absence of the syntax element in the bitstream, it is inferred that the encoding / decoding tool is turned off or disabled.
32. The method according to claim 22, wherein, In response to the satisfaction of the third condition, the encoding / decoding tool is an affine prediction tool, and in response to the non-satisfaction of the third condition, the encoding / decoding tool is a prediction tool other than the affine prediction tool.
33. The method of claim 32, wherein the motion vectors of one or more control points are used as motion vectors for the third video block without applying an affine motion model to the third video block.
34. The method according to claim 22, wherein, In response to the satisfaction of the third condition, the encoding / decoding tool is an interleaving prediction tool, and in response to the non-satisfaction of the third condition, the encoding / decoding tool is a prediction tool other than an affine prediction tool.
35. The method according to claim 34, wherein, In response to the failure to meet the third condition, the prediction sample of the second partitioning pattern in the interleaving prediction tool is set to be equal to the prediction sample of the first partitioning pattern at the corresponding position.
36. The method according to claim 22, wherein, The third condition is determined based on the size of the bounding box of the third video block.
37. The method according to claim 36, wherein, The third condition is satisfied in response to the size of the bounding box being less than or equal to a threshold, and wherein the third condition is not satisfied in response to the size of the bounding box being greater than the threshold.
38. The method according to claim 22, wherein, The third condition is based on a first motion vector in the horizontal direction of a plurality of sub-blocks of the third video block and a second motion vector in the vertical direction of the plurality of sub-blocks of the third video block.
39. The method according to claim 38, wherein, The third condition is true when the result of subtracting the first motion vector from the third motion vector in the horizontal direction of a sub-block of the third video in the first segmentation pattern is greater than a threshold value.
40. The method of claim 38, wherein, The third condition is true when the result of subtracting the third motion vector in the horizontal direction of a sub-block of the third video in the first segmentation pattern from the first motion vector is greater than a threshold value.
41. The method according to claim 40, wherein, The third condition is true when the result of subtracting the second motion vector from the fourth motion vector in the vertical direction of a sub-block of the third video in the first segmentation pattern is greater than a threshold.
42. The method according to claim 38, wherein, The third condition is true when the result of subtracting the fourth motion vector in the vertical direction of a sub-block of the third video in the first segmentation pattern from the second motion vector is greater than a threshold.
43. The method according to claim 1, further comprising: For the fourth conversion between the fourth video block of the video and the bitstream of the video, according to the fourth rule, it is determined whether and how motion compensation is applied to the sub-block at the sub-block index: (1) Whether to apply an affine mode to the fourth video block, wherein the bitstream includes an affine motion indicator indicating whether the affine mode is applied. (2) The color component to which the fourth video block belongs, and (3) The color format of the video; and The fourth transformation is performed based on the determination.
44. The method according to claim 43, wherein, In response to the following, the fourth rule specifies that the motion compensation be applied individually to the sub-block at the sub-block index: (1) The color component is a luminance component, or (2) The affine mode is not applied to the fourth video block, or (3) The remainder of the horizontal sub-block index divided by the width of the fourth video block is equal to 0 and the second remainder of the vertical sub-block index divided by the height of the fourth video block is equal to 0.
45. The method of claim 43, wherein the fourth rule specifies that the first width or first height of the chroma sub-block used to perform the motion compensation at the sub-block index of the chroma sub-block of the video depends on the affine mode indication and the color format of the video.
46. The method according to claim 45, in, In response to the affine mode indication that the affine mode is applied to the fourth video block, the fourth rule specifies that the first width of the chroma sub-block is the second width of the luma sub-block, and In response to the affine mode indication indicating that the affine mode is not applied to the fourth video block, the fourth rule specifies that the first width of the chroma sub-block is the second width of the luma sub-block divided by the third width of the fourth video block.
47. The method according to claim 45, in, In response to the affine mode indication that the affine mode is applied to the fourth video block, the fourth rule specifies that the first height of the chroma sub-block is the second height of the luma sub-block, and In response to the affine mode indication indicating that the affine mode is not applied to the fourth video block, the fourth rule specifies that the first height of the chroma sub-block is the second height of the luma sub-block divided by the third height of the fourth video block.
48. The method according to claim 1, further comprising: For the fifth transition between the fifth video block and the bitstream of the video, according to the fifth rule, it is determined whether and how to use weighted prediction to calculate the prediction of the sub-blocks of the fifth video block: (1) Whether to apply an affine mode to the fifth video block, wherein the bitstream includes an affine motion indicator indicating whether the affine mode is applied. (2) The color component to which the fifth video block belongs, and (3) The color format of the video; and Based on the determination, the fifth transformation is performed.
49. The method according to claim 48, wherein, In response to the following, the fifth rule specifies that the weighted prediction be applied individually to the sub-block at the sub-block index: (1) The color component is not the luminance component, or (2) The affine mode is not applied to the fifth video block, or (3) The remainder of the horizontal sub-block index divided by the width of the fifth video block is equal to 0 and the second remainder of the vertical sub-block index divided by the height of the fifth video block is equal to 0.
50. The method of claim 48, wherein the fifth rule specifies that the first width or first height of the chroma sub-block used to perform the weighted prediction at the sub-block index of the chroma sub-block of the video depends on the affine mode indication and the color format of the video.
51. The method according to claim 50, in, In response to the affine mode indication that the affine mode is applied to the fifth video block, the fifth rule specifies that the first width of the chroma sub-block is the second width of the luma sub-block, and In response to the affine mode indication indicating that the affine mode is not applied to the fifth video block, the fifth rule specifies that the first width of the chroma sub-block is the second width of the luma sub-block divided by the third width of the fifth video block.
52. The method according to claim 50, in, In response to the affine mode indication that the affine mode is applied to the fifth video block, the fifth rule specifies that the first height of the chroma sub-block is the second height of the luma sub-block, and In response to the affine mode indication indicating that the affine mode is not applied to the fifth video block, the fifth rule specifies that the first height of the chroma sub-block is the second height of the luma sub-block divided by the third height of the fifth video block.
53. The method according to claim 1, further comprising: For the sixth transformation between the sixth video block and the bitstream of the video, the value of a first syntax element in the bitstream is determined according to the sixth rule, wherein the value of the first syntax element indicates whether an affine mode is applied to the sixth video block, and The sixth rule specifies that the value of the first syntax element is based on: (1) A second syntax element in the bitstream, the second syntax element indicating whether motion compensation based on an affine model is used to generate the predicted samples of the sixth video block, or (2) A third syntax element in the bitstream, the third syntax element indicating whether motion compensation based on a 6-parameter affine model is enabled for a codec layer video sequence (CLVS); and Based on the determination, the sixth transformation is performed.
54. The method according to claim 53, wherein, In response to the second syntax element being equal to 0, the sixth rule specifies that the value of the first syntax element is set to zero.
55. The method according to claim 53, wherein, In response to the third syntax element being equal to 0, the sixth rule specifies that the value of the first syntax element is set to zero.
56. The method according to claim 1, further comprising: For the seventh conversion between the seventh video block and the bitstream of the video, based on the relationship between the first motion vector of the first sub-block of the first style of the seventh video block and the second motion vector of the second sub-block of the second style of the seventh video block, it is determined whether to enable the interleaving prediction tool for the seventh video block; as well as Based on the determination, the seventh transformation is performed.
57. The method according to claim 56, wherein, Select the second sub-block from one of the sub-blocks of the second style, wherein the sub-block overlaps with the first sub-block.
58. The method according to claim 57, wherein, The second sub-block is selected from the first sub-block based on the affine parameters.
59. The method according to claim 57, In response to the fulfillment of the first condition, the second center of the second sub-block is located above the first center of the first sub-block, and in, In response to the fulfillment of the second condition, the second center of the second sub-block is located below the first center of the first sub-block.
60. The method according to claim 59, wherein, The first condition is that the first sub-block is in the last line of the seventh video block.
61. The method according to claim 59, wherein, The second condition is that the first sub-block is in the first row of the seventh video block.
62. The method of claim 59, wherein the second condition is: (mv0_y>0 and mv1_y>0), or (mv0_y>0 and mv1_y<0 and |mv0_y|>|mv1_y|), or (mv0_y<0 and mv1_y>0 and |mv0_y|<|mv1_y|), Where mv0_y is the second motion vector in the vertical direction, and Where mv1_y is the first motion vector in the vertical direction.
63. The method according to claim 59, wherein, In response to the failure to meet the second condition, the first condition is met.
64. The method according to claim 57, in, In response to the fulfillment of the third condition, the second center of the second sub-block is to the left of the first center of the first sub-block, and In response to the fulfillment of the fourth condition, the second center of the second sub-block is to the right of the first center of the first sub-block.
65. The method according to claim 64, wherein, The third condition is that the first sub-block is in the rightmost row of the seventh video block.
66. The method according to claim 64, wherein, The fourth condition is that the first sub-block is in the leftmost row of the seventh video block.
67. The method according to claim 56, wherein, In response to the relationship between the first motion vector and the second motion vector satisfying the fifth condition, the interleaving prediction tool is disabled.
68. The method according to claim 67, in, The fifth condition depends on the first and second syntax elements in the bitstream. The first syntax element indicates whether the center of the first sub-block is to the left of the center of the second sub-block. The second syntax element indicates whether the center of the first sub-block is above the center of the second sub-block.
69. The method according to claim 56, wherein, Performing the seventh conversion includes encoding the seventh video block into the bitstream.
70. The method according to claim 56, wherein, Performing the seventh conversion includes generating the bitstream from the seventh video block, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium.
71. The method according to claim 56, wherein, Performing the seventh conversion includes decoding the seventh video block from the bitstream.
72. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: For the conversion between the current video block and the video bitstream, the decision to apply fractional sample interpolation to the color components of the sub-blocks of the current video block is based on whether certain conditions are met. The color component is a chromaticity component; and The conversion is performed based on the determination. The condition mentioned is one of the following: The current video block is not encoded or decoded in affine mode, or The remainders of xSbIdx divided by SubWidthC and ySbIdx divided by SubHeightC are both 0, where xSbIdx is the horizontal index of the sub-block, ySbIdx is the vertical index of the sub-block, and SubWidthC and SubHeightC represent the color format of the current image including the current video block.
73. The apparatus according to claim 72, wherein, When the condition is met, the fractional sample interpolation process is enabled for the color components of the sub-block; when the condition is not met, the fractional sample interpolation process is disabled for the color components of the sub-block.
74. The apparatus according to claim 72, wherein, When the fractional sample interpolation process is enabled, the size of the color component of the sub-block used in the fractional sample interpolation process is based on whether the current video block is encoded and decoded in affine mode.
75. The apparatus according to claim 74, wherein, When the current video block is encoded and decoded in the affine mode, the dimensions of the color components of the sub-block used in the fractional sample interpolation process are sbWidth and sbHeight, where sbWidth and sbHeight specify the width of the sub-block in units of luminance samples and the height of the sub-block in units of luminance samples, respectively.
76. The apparatus according to claim 74, wherein, When the current video block is not encoded or decoded in the affine mode, the dimensions of the color components of the sub-block used in the fractional sample interpolation process are sbWidth / SubWidthC and sbHeight / SubHeightC, where sbWidth and sbHeight specify the width and height of the sub-block in units of luminance samples, respectively.
77. A non-transitory computer-readable storage medium for storing instructions, said instructions causing a processor to: For the conversion between the current video block and the video bitstream, the decision to apply fractional sample interpolation to the color components of the sub-blocks of the current video block is based on whether certain conditions are met. The color component is a chromaticity component; and The conversion is performed based on the determination. The condition mentioned is one of the following: The current video block is not encoded or decoded in affine mode, or The remainders of xSbIdx divided by SubWidthC and ySbIdx divided by SubHeightC are both 0, where xSbIdx is the horizontal index of the sub-block, ySbIdx is the vertical index of the sub-block, and SubWidthC and SubHeightC represent the color format of the current image including the current video block.
78. The non-transitory computer-readable storage medium according to claim 77, wherein, When the condition is met, the fractional sample interpolation process is enabled for the color components of the sub-block; when the condition is not met, the fractional sample interpolation process is disabled for the color components of the sub-block.
79. A non-transitory computer-readable recording medium having stored thereon computer-readable instructions and a bit stream, the computer-readable instructions, when executed by a processor, causing the processor to perform: For the current video block, the decision to apply fractional sample interpolation to the color components of the sub-blocks of the current video block is based on whether certain conditions are met. The color component is a chromaticity component; and The bit stream is generated based on the determination. The condition mentioned is one of the following: The current video block is not encoded or decoded in affine mode, or The remainders of xSbIdx divided by SubWidthC and ySbIdx divided by SubHeightC are both 0, where xSbIdx is the horizontal index of the sub-block, ySbIdx is the vertical index of the sub-block, and SubWidthC and SubHeightC represent the color format of the current image including the current video block.
80. A method for storing a video bitstream, comprising: For the current video block, determine whether to apply fractional sample interpolation to the color components of the sub-blocks of the current video block based on whether certain conditions are met; where the color components are chroma components; and Based on the determination to generate the bit stream, and to store the bit stream, The condition mentioned is one of the following: The current video block is not encoded or decoded in affine mode, or The remainders of xSbIdx divided by SubWidthC and ySbIdx divided by SubHeightC are both 0, where xSbIdx is the horizontal index of the sub-block, ySbIdx is the vertical index of the sub-block, and SubWidthC and SubHeightC represent the color format of the current image including the current video block.
81. A video decoding apparatus, comprising a processor configured to perform the method of any one of claims 8-68, 71.
82. A video encoding apparatus, comprising a processor configured to implement the method of any one of claims 8-70.
83. A non-transitory computer-readable storage medium having stored thereon a bit stream and computer-readable instructions, the computer-readable instructions, when executed by a processor, causing the processor to perform the method of any one of claims 8-69 to generate the bit stream.
84. A non-transitory computer-readable storage medium storing instructions that enable a processor to perform the method described in any one of claims 8-71.
85. A method for storing a bit stream, comprising: A video bitstream is generated according to the method described in any one of claims 8-69, and The bitstream is stored on a computer-readable storage medium.
Citation Information
Patent Citations
Affine mode in video encoding and decoding
CN110891175A
Affine motion vector prediction in video coding
US20190149838A1