Screen content encoding / decoding based on the interaction with motion information
By modifying the motion information of the intra-block copy mode, using block vector difference and weighting factors, combining the triangle segmentation mode and the combined intra-inter prediction mode, the encoding efficiency of screen content is optimized, and the intra-block copy codec mode in the prior art is solved.
Patent Information
- Application Number
- CN202080043774.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-16
- Filing Date
- 2020-06-15
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-06-15
AI Technical Summary
When the existing video encoding and decoding technology processes screen content, the motion information encoding and decoding efficiency of the intra-block copying and decoding mode is not high, and it is difficult to effectively utilize the repeated patterns in the screen content, resulting in low encoding efficiency.
By modifying the motion information of the intra-block copy mode, using block vector difference and weighting factors to optimize the intra-block copy codec tool, combining the triangle segmentation mode and the combined intra-inter prediction mode to improve encoding efficiency.
Improve the efficiency of screen content encoding and decoding, especially when processing screen content, reduces redundancy and improves encoding efficiency and decoding quality.
Smart Images

Figure CN113966612B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] Under the provisions of the applicable Patent Law and / or the Paris Convention, this application timely claims the priority and benefits of International Patent Application No. PCT / CN2019 / 091446, filed on June 16, 2019. The entire disclosure of International Patent Application No. PCT / CN2019 / 091446 is incorporated herein by reference as part of the disclosure of this application. Technical field
[0003] This patent document generally relates to video encoding and decoding technologies. Background art
[0004] Video coding and decoding standards have evolved mainly through the development of well - known ITU - T and ISO / IEC standards. ITU - T developed H.261 and H.263, ISO / IEC developed MPEG - 1 and MPEG - 4 Visual, and the two organizations jointly developed the H.262 / MPEG - 2 video, H.264 / MPEG - 4 Advanced Video Coding (AVC), and H.265 / High Efficiency Video Coding (HEVC) standards. Since H.262, video coding and decoding standards are based on a hybrid video coding and decoding structure, in which temporal prediction plus transform coding is adopted. To explore future video coding and decoding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and applied them to a reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Team (JVET) was created between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG), which is dedicated to researching the next - generation Versatile Video Coding (VVC) standard targeting a 50% bit - rate reduction compared to HEVC. Summary of the invention
[0005] In the case of using the disclosed video encoding, transcoding, or decoding technologies, embodiments of a video encoder or decoder can manipulate the virtual boundaries of coding and decoding tree blocks to provide higher compression efficiency and a simpler implementation of encoding or decoding tools.
[0006] In one exemplary aspect, a video processing method is disclosed. The method includes: during the conversion between a current video block of a video picture and a bitstream representation of the current video block, determining a block vector difference (BVD) representing the difference between a block vector corresponding to the current video block and its predictor; and performing the conversion between the current video block and the bitstream representation of the current video block using the block vector. Here, the block vector indicates a motion match for the current video block in the video picture, and a modified value of the BVD is encoded and decoded into the bitstream representation.
[0007] In another exemplary aspect, a video processing method is disclosed. The method includes: during the conversion between a current video block of a video picture and a bitstream representation of the current video block, determining to use an intra-block copy tool for the conversion; and performing the conversion using a modified block vector predictor corresponding to a modified value of a block vector difference (BVD) of the current video block.
[0008] In another exemplary aspect, a video processing method is disclosed. The method includes performing the conversion between a current video block of a video picture and a bitstream representation of the current video block using an intra-block copy encoding / decoding tool, where a modified block vector corresponding to a modified value of the block vector is used for the conversion and is included in the bitstream representation.
[0009] In another exemplary aspect, a video processing method is disclosed. The method includes: during the conversion between a current video block and a bitstream representation of the current video block, determining a weighting factor wt based on the conditions of the current video block; and performing the conversion using a combined intra-inter prediction encoding operation, where the weighting factor wt is used to weight the motion vector of the current video block.
[0010] In another exemplary aspect, a video processing method is disclosed. The method includes: determining to use a triangle partition mode (TPM) encoding / decoding tool for the conversion between a current video block and a bitstream representation of the current video block, where at least one operation parameter of the TPM encoding / decoding tool depends on the characteristics of the current video block, and the TPM encoding / decoding tool partitions the current video block into two separately encoded / decoded non-rectangular partitions; and performing the conversion by applying the TPM encoding / decoding tool that uses the said one operation parameter.
[0011] In another exemplary aspect, a video processing method is disclosed. The method includes: modifying at least one of the motion information associated with a block encoded / decoded in an intra-block copy (IBC) mode for the conversion between a block of a video and a bitstream representation of the block; and performing the conversion based on the modified motion information.
[0012] In another exemplary aspect, a method for video processing is disclosed. The method includes: for the conversion between a block of a video and the bitstream representation of the block, determining a weighting factor for an intra prediction signal in a combined intra-inter prediction (CIIP) mode based on a motion vector (MV) associated with the block; and performing the conversion based on the weighting factor.
[0013] In another exemplary aspect, a method for video processing is disclosed. The method includes: for the conversion between a block of a video and the bitstream representation of the block, determining weights used in a triangle prediction mode (TPM) based on motion information of one or more partitions of the block; and performing the conversion based on the weights.
[0014] In another exemplary aspect, a method for video processing is disclosed. The method includes: for the conversion between a block of a video and the bitstream representation of the block, determining whether to apply a hybrid process in a triangle prediction mode (TPM) based on transform information of the block; and performing the conversion based on the determination.
[0015] In yet another exemplary aspect, a video coding apparatus configured to perform the methods described above is disclosed.
[0016] In yet another exemplary aspect, a video decoder configured to perform the methods described above is disclosed.
[0017] In yet another exemplary aspect, a machine-readable medium is disclosed. The medium stores code that, when executed, causes a processor to implement one or more of the methods described above.
[0018] The above and other aspects and features of the disclosed technology are described in more detail in the drawings, the description, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 Examples of multi-type tree partitioning patterns are shown.
[0020] Figure 2 Examples of signaling of partition flags in a quadtree codec tree structure with nested multi-type trees are shown.
[0021] Figure 3 Examples of the fast structure of a quadtree with nested multi-type trees are shown.
[0022] Figure 4 Examples of a No TT (no TT) partition for a 128×128 codec block are shown.
[0023] Figure 5 Sixty-seven intra prediction modes are shown.
[0024] Figure 6 Shows exemplary positions of samples for the derivation of α and β.
[0025] Figure 7 Shows an example of four reference lines adjacent to a prediction block.
[0026] Figures 8A - 8B Shows the illustration of sub - partitions for 4x8 and 8x4 CUs and examples of sub - partitions for CUs other than 4x8, 8x4, and 4x4.
[0027] Figure 9 Is an illustration of the ALWIP of a 4×4 block.
[0028] Figure 10 Is an illustration of the ALWIP for an 8x8 block.
[0029] Figure 11 Is an illustration of the ALWIP for an 8×4 block.
[0030] Figure 12 Is an illustration of the ALWIP for a 16x16 block.
[0031] Figure 13 Is an example of the derivation process for merge candidate list construction.
[0032] Figure 14 Shows exemplary positions of spatial domain merge candidates.
[0033] Figure 15 Shows an example of candidate pairs considered for redundancy check of spatial domain merge candidates.
[0034] Figure 16 Shows an example of the positions of the second PUs for N×2N and 2N×N partitions.
[0035] Figure 17 Is an illustration of the motion vector scaling for temporal domain merge candidates.
[0036] Figure 18 Shows example candidate positions C0 and C1 of temporal domain merge candidates.
[0037] Figure 19 Shows an example of combined bidirectional prediction merge candidates.
[0038] Figure 20 Shows an example of the derivation process of motion vector prediction candidates.
[0039] Figure 21 Is an illustration of the motion vector scaling for spatial domain motion vector candidates.
[0040] Figures 22A - 22B Examples of the 4-parameter affine model and the 6-parameter affine model are shown respectively.
[0041] Figure 23 An example of the affine MVF for each sub-block is shown.
[0042] Figure 24 An example of the encoding / decoding process of history-based motion vector prediction is shown.
[0043] Figure 25 An example of the merge candidate construction process is shown.
[0044] Figure 26 An example of the inter-frame prediction mode based on triangle partitioning is shown.
[0045] Figure 27 An example of the ultimate motion vector representation (UMVE) search is shown.
[0046] Figure 28 An example of the UMVE search points is shown.
[0047] Figure 29 It is an illustration of the intra-block copy encoding / decoding mode.
[0048] Figure 30 Examples of the top and left adjacent blocks used in the CIIP weight derivation are shown.
[0049] Figure 31 The segmentation of the video block is shown.
[0050] Figure 32 It is a block diagram of an exemplary apparatus for video processing.
[0051] Figure 33 It is a flowchart of an exemplary method for video processing.
[0052] Figure 34 It is a flowchart of an exemplary method for video processing.
[0053] Figure 35 It is a flowchart of an exemplary method for video processing.
[0054] Figure 36 It is a flowchart of an exemplary method for video processing.
[0055] Figure 37 It is a flowchart of an exemplary method for video processing. Detailed implementation manners
[0056] Section headings are used in this document for ease of understanding, and are not intended to limit the embodiments disclosed in a section to that section only. Further, although some embodiments are described with reference to multi-functional video coding or other specific video codecs, the disclosed techniques are also applicable to other video codecs. Further, although some embodiments describe in detail the video encoding steps, it should be understood that the corresponding decoding steps for decoding the encoded data will be implemented by a decoder. Further, the term video processing encompasses video encoding or compression, video decoding or decompression, and video transcoding, in which a video is represented in one compression format and converted to another compression format or with a different compression bit rate.
[0057] 1. Brief Summary
[0058] This document relates to video coding and decoding techniques. Specifically, it relates to the generation of prediction blocks. It can be applied to existing video coding and decoding standards, such as HEVC, or standards under consideration (multi-functional video coding). It can also be applicable to future video coding and decoding standards or video codecs.
[0059] 2. Background
[0060] Video coding and decoding standards have mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards. Since H.262, video coding and decoding standards have been based on a hybrid video coding structure, in which temporal prediction and transform coding are employed. To explore future video coding and decoding techniques beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and applied them to a reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Team (JVET) was created between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG), which is dedicated to researching the VVC standard targeting a 50% bit rate reduction compared to HEVC.
[0061] 2.1. CTU Partitioning Using a Tree Structure
[0062] In HEVC, the CTU is partitioned into CUs by using a quadtree structure (represented as a codec tree) to adapt to various local characteristics. At the leaf CU level, it is determined whether to use inter-picture (temporal) prediction or intra-picture (spatial) prediction to codec the picture region. Depending on the partition type of the PU, each leaf CU can be further partitioned into one, two, or four PUs. In a PU, the same prediction process is applied, and the relevant information is transmitted to the decoder on a PU basis. After obtaining the residual block by applying the prediction process based on the PU partition type, the CU can be split into transform units (TUs) according to another quadtree structure similar to the codec tree of the leaf CU. An important feature of the HEVC structure is that it has multiple partition concepts, including CUs, PUs, and TUs.
[0063] In VVC, a quadtree with a nested multi-type tree using a binary and ternary partition segmentation structure replaces the concept of multiple segmentation unit types, that is, it eliminates the separation of the CU, PU, and TU concepts, unless the CU is too large for the maximum transform length, and supports greater flexibility in the CU split shape. In the codec tree structure, the CU can be square or rectangular. First, the codec tree unit (CTU) is segmented according to the quadtree (also known as the quaternary tree) structure. Then, the quadtree leaf nodes can be further segmented according to the multi-type tree structure. As Figure 1 shown, there are four partition types in the multi-type tree structure, vertical binary partition (SPLIT_BT_VER), horizontal binary partition (SPLIT_BT_HOR), vertical ternary partition (SPLIT_TT_VER), and horizontal ternary partition (SPLIT_TT_HOR). The multi-type tree leaf nodes are called codec units (CUs), and unless the CU is too large for the maximum transform length, this segmentation is used for prediction and transform processing without any further splitting. This means that in most cases, in the quadtree codec block structure with a nested multi-type tree, the CUs, PUs, and TUs have the same block size. An exception occurs when the maximum supported transform length is less than the width or height of the color component of the CU.
[0064] Figure 2A signaling mechanism for split partition information in a quadtree codec tree structure with nested multi-type trees is shown. A codec tree unit (CTU) is treated as the root of a quadtree and is first split according to the quadtree structure. Then, each quadtree leaf node is further split according to the multi-type tree structure (when it is large enough to allow). In the multi-type tree structure, a first flag (mtt_split_cu_flag) is signaled to indicate whether the node is further split; when the node is further split, a second flag (mtt_split_cu_vertical_flag) is signaled to indicate the split direction, and then a third flag is signaled to indicate whether the split is a binary split or a ternary split. Based on the values of mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the multi-type tree partition mode (MttSplitMode) of the CU is derived, as shown in Table 1.
[0065] Table 1 – MttSplitMode Derivation Based on Multi-Type Tree Syntax Elements
[0066] MttSplitMode mtt_split_cu_vertical_flag mtt_split_cu_binary_flag SPLIT_TT_HOR 0 0 SPLIT_BT_HOR 0 1 SPLIT_TT_VER 1 0 SPLIT_BT_VER 1 1
[0067] Figure 3 A CTU divided into multiple CUs by means of a quadtree and a nested multi-type tree codec block structure is shown, where the thick block edges represent quadtree splits and the remaining edges represent multi-type tree splits. The split of the quadtree with nested multi-type trees provides a content-adaptive codec tree structure composed of CUs. In the luma sample unit, the size of the CU can be as large as the CTU or as small as 4×4. For the case of 4:2:0 chroma format, the maximum chroma CB size is 64×64 and the minimum chroma CB size is 2×2.
[0068] In VVC, the maximum supported luma transform size is 64×64 and the maximum supported chroma transform size is 32×32. When the width or height of the CB is greater than the maximum transform width or height, the CB is automatically divided along the horizontal and / or vertical directions to comply with the transform size limit in that direction.
[0069] For the quadtree codec tree scheme with nested multi-type trees, the following parameters are defined and specified by SPS syntax elements.
[0070] – CTU size: the size of the root node of the quadtree
[0071] – MinQTSize: the minimum allowed quadtree leaf node size
[0072] – MaxBtSize: the maximum allowed binary tree root node size
[0073] – MaxTtSize: The maximum allowable size of the trinary tree root node
[0074] – MaxMttDepth: The maximum allowable hierarchical depth of the multi-type tree partitioning made from the quadtree leaf
[0075] – MinBtSize: The minimum allowable size of the binary tree leaf node
[0076] – MinTtSize: The minimum allowable size of the trinary tree leaf node
[0077] In an example of the quadtree encoding / decoding tree structure with a nested multi-type tree, the CTU size is set to 128×128 luma samples together with two corresponding 64×64 blocks of 4:2:0 chroma samples, MinQTSize is set to 16×16, MaxBtSize is set to 128×128, and MaxTtSize is set to 64×64, MinBtSize and MinTtSize (for both width and height) are set to 4×4, and MaxMttDepth is set to 4. The quadtree splitting is first applied to the CTU to generate quadtree leaf nodes. The size of the quadtree leaf nodes can have a size ranging from 16×16 (i.e., MinQTSize) to 128×128 (i.e., the CTU size). If the leaf QT node is 128×128, it will not be further partitioned according to the binary tree because its size exceeds MaxBtSize and MaxTtSize (i.e., 64×64). Otherwise, the leaf quadtree node can be further split according to the multi-type tree. Therefore, the quadtree leaf node is also the root node of the multi-type tree and has a multi-type tree depth (mttDepth) of 0. When the multi-type tree depth reaches MaxMttDepth (i.e., 4), further partitioning is not considered. When the width of the multi-type tree node is equal to MinBtSize and less than or equal to 2*MinTtSize, further horizontal partitioning is not considered. Similarly, when the height of the multi-type tree node is equal to MinBtSize and less than or equal to 2*MinTtSize, further vertical partitioning is not considered.
[0078] To allow for 64×64 luma block and 32×32 chroma pipeline designs in the VVC hardware decoder, TT partitioning is prohibited when the width or height of the luma encoding / decoding block is greater than 64, as Figure 4 shown. TT partitioning is prohibited when the width or height of the chroma encoding / decoding block is greater than 32.
[0079] In VTM5, the codec tree scheme supports the ability to have separate block tree structures for luminance and chrominance. Currently, for P slices and B slices, the luminance CTB and chrominance CTB in a CTU must share the same codec tree structure. However, for I slices, luminance and chrominance can have separate block tree structures. When applying the separate block tree structure, the luminance CTB is partitioned into CUs according to one codec tree structure, and the chrominance CTB is partitioned into chrominance CUs according to another codec tree structure. This means that a CU in an I slice can be composed of the coded blocks of the luminance component or the coded blocks of the two chrominance components, and a CU in a P slice or B slice is always composed of the coded blocks of all three color components, unless the video is monochrome.
[0080] 2.2. Intra Prediction in VVC
[0081] 2.2.1. 67 Intra Prediction Modes
[0082] To capture any edge directions present in natural videos, the number of directional intra modes in VTM4 was extended from 33 (as used in HEVC) to 65. The new directional modes not in HEVC are shown as red dashed arrows in Figure 5 and the planar and DC modes remain the same. These denser directional intra prediction modes apply to all block sizes and both luminance intra prediction and chrominance intra prediction.
[0083] 2.2.2. Position-Dependent Intra Prediction Combination (PDPC)
[0084] In VTM4, the intra prediction result of the planar mode is further modified by the position-dependent intra prediction combination (PDPC) method. PDPC is an intra prediction method that calls for a combination of unfiltered boundary reference samples and HEVC-style intra prediction using filtered boundary reference samples. PDPC is applied to the following intra modes without signaling: planar, DC, horizontal, vertical, lower left angular mode and its eight adjacent angular modes, and upper right angular mode and its eight adjacent angular modes.
[0085] The predicted sample pred(x,y) is predicted by a linear combination of the intra prediction mode (DC, planar, angular) and reference samples according to the following equation:
[0086] pred(x,y) = (wL × R -1,y + wT × R x,-1 – wTL × R -1,-1 + (64 – wL – wT + wTL) × pred(x,y) + 32) >> 6
[0087] where R x,-1 、R -1,ydenote reference samples that are respectively above and to the left of the current sample (x, y), and R -1,-1 denote the reference sample at the upper left corner of the current block.
[0088] If PDPC is applied to DC, planar, horizontal, and vertical intra modes, no additional boundary filters are required, as is the case for the HEVC DC mode boundary filter or the horizontal / vertical mode edge filter.
[0089] 2.2.3. Cross-Component Linear Model Prediction (CCLM)
[0090] To reduce cross-component redundancy, the cross-component linear model (CCLM) prediction mode is used in VTM4, for which chroma samples are predicted as follows based on the reconstructed luma samples of the same CU by using a linear model:
[0091] pred C (i,j) = α · rec L ′(i,j) + β
[0092] where pred C (i,j) represents the predicted chroma sample in the CU, and rec L (i,j) represents the downsampled reconstructed luma sample in the same CU. The linear model parameters α and β are derived from the relationship between the luma values and chroma values of two samples and their corresponding chroma samples, which are the luma samples with the minimum and maximum sample values within the set of the downsampled adjacent luma samples. Figure 6 An example of the positions of the left and above samples and the samples of the current block involved in the CCLM mode is shown.
[0093] This parameter calculation is performed as part of the decoding process, rather than only as an encoder search operation. Therefore, the α value and β value are not communicated to the decoder by syntax.
[0094] For chroma intra mode coding and decoding, a total of 8 intra modes are allowed for chroma intra mode coding and decoding. These modes include five traditional intra modes and three cross-component linear model modes (CCLM, LM_A, and LM_L). Chroma mode coding and decoding directly depends on the intra prediction mode of the corresponding luma block. Since a separate block splitting structure for the luma component and chroma component is enabled in the I slice, a chroma block can correspond to multiple luma blocks. Therefore, for the chroma DM mode, the intra prediction mode of the corresponding luma block covering the center position of the current chroma block is directly inherited.
[0095] 2.2.4. Multiple Reference Line (MRL) Intra Prediction
[0096] Multi-reference line (MRL) intra prediction uses more reference lines for intra prediction. In Figure 7 an example of 4 reference lines is illustrated, where the samples of segments A and F are not taken from the reconstructed neighboring samples, but are filled with the closest samples from segments B and E respectively. HEVC intra prediction uses the closest reference line (i.e., reference line 0). In MRL, 2 additional lines (reference line 1 and reference line 3) are used. The index of the selected reference line (mrl_idx) is signaled and used to generate the intra predictor. For reference line idx greater than 0, only the additional reference line modes are included in the MPM list, and only the mpm index is signaled without the remaining modes.
[0097] 2.2.5. Intra sub-division (ISP)
[0098] The intra sub-division (ISP) tool divides the luma intra prediction block vertically or horizontally into 2 or 4 sub-divisions according to the block size. For example, the minimum block size of ISP is 4x8 (or 8x4). If the block size is greater than 4x8 (or 8x4), the corresponding block is divided by 4 sub-divisions. Figures 8A - 8B Examples of two possibilities are shown. All sub-divisions satisfy the condition of having at least 16 samples.
[0099] For each sub-division, the reconstructed samples are obtained by adding the residual signal to the prediction signal. Here, the residual signal is generated through processes such as entropy decoding, inverse quantization, and inverse transformation. Therefore, the reconstructed sample values of each sub-division can be used to generate the prediction of the next sub-division, and the process for each sub-division is repeated. Additionally, the first sub-division to be processed is the sub-division containing the top-left sample of the CU, and then continues downward (for horizontal division) or to the right (for vertical division). Therefore, the reference samples used to generate the sub-division prediction signal are only located on the left and upper sides of the line. All sub-divisions share the same intra mode.
[0100] 2.2.6. Affine linear weighted intra prediction (ALWIP, also known as matrix-based intra prediction)
[0101] Affine linear weighted intra prediction (ALWIP, also known as matrix-based intra prediction (MIP)) is proposed.
[0102] Two tests are conducted. In test 1, ALWIP is designed with a storage limit of 8 Kbytes per sample and at most 4 multiplications. Test 2 is similar to test 1, but the design is further simplified in terms of storage requirements and model architecture.
[0103] · A single set consisting of a matrix and an offset vector for all block shapes.
[0104] · The number of modes is reduced to 19 for all block shapes
[0105] · The storage requirement is reduced to 5,760 10-bit values, i.e., 7.20 kilobytes.
[0106] · Perform linear interpolation of the predicted samples in a single step, thus replacing the iterative interpolation as in the first test.
[0107] 2.2.6.1. Test 1
[0108] To predict the samples of a rectangular block with width W and height H, affine linear weighted intra prediction (ALWIP) takes as input H reconstructed neighboring boundary samples of a line to the left of the block and W reconstructed neighboring boundary samples of a line above the block. If the reconstructed samples are not available, the reconstructed samples are generated as done in conventional intra prediction.
[0109] The generation of the prediction signal is based on the following three steps:
[0110] 1. Among the boundary samples, four samples in the case of W = H = 4 and eight samples in all other cases are extracted by taking the mean.
[0111] 2. Perform matrix-vector multiplication with the mean samples as input, and then add the offset. The result is a reduced prediction signal for the subsampled set of samples in the original block.
[0112] 3. The prediction signals at the remaining positions are generated from the prediction signal on the subsampled set by linear interpolation, which is single-step linear interpolation within each direction.
[0113] The matrices and offset vectors required to generate the prediction signal are taken from three matrix sets S0, S1, S2. Set S0 consists of 18 matrices each having 16 rows and 4 columns and 18 offset vectors each having a size of 16 and is used for blocks with a size of 4×4. Set S1 consists of 10 matrices each having 16 rows and 8 columns and 10 offset vectors each having a size of 16 and is used for blocks with sizes of 4×8, 8×4, and 8×8. Finally, set S2 consists of 6 matrices each having 64 rows and 8 columns and 6 offset vectors having a size of 64 and is used for all other block shapes, either these matrices and offset vectors or parts of these matrices and offset vectors.
[0114] The total number of multiplications required in the calculation of the matrix vector product is always less than or equal to 4×W×H. In other words, for the ALWIP mode, a maximum of four multiplications per sample point are required.
[0115] 2.2.6.2. Mean value of the boundary
[0116] In the first step, the input boundaries bdry top and bdry left are reduced to smaller boundaries and Here, and both consist of 2 sample points in the case of a 4×4 block and both consist of 4 sample points in all other cases.
[0117] In the case of a 4×4 block, for 0 ≤ i < 2,
[0118]
[0119] is defined and similarly
[0120] Otherwise, if the block width W is given as W = 4·2 k , then for 0 ≤ i < 4,
[0121]
[0122] is defined and similarly
[0123] The two reduced boundaries and are concatenated to the reduced boundary vector bdry red , so that this vector has a size of four for a block with a shape of 4×4 and a size of eight for blocks with all other shapes. If mode refers to the ALWIP mode, then this concatenation is defined as follows:
[0124]
[0125] Finally, for the interpolation of the downsampled prediction signal, on large blocks, a second version of the mean boundary is required. That is, if min(W,H) > 8 and W ≥ H, then write W = 8*2 l , and for 0 ≤ i < 8,
[0126]
[0127] If min(W,H) > 8 and H > W, then it is defined similarly
[0128] 2.2.6.3 Generation of the reduced prediction signal by matrix-vector multiplication
[0129] From the reduced input vector bdry red generate the reduced prediction signal pred red . The latter signal is a signal on a downsampling block having a width W red and a height H red . Here, W red and H red are defined as:
[0130]
[0131] The reduced prediction signal pred is calculated by computing the matrix-vector product and adding an offset red :
[0132] pred red = A · bdry red + b.
[0133] Here, A is a matrix having W red · H red rows and having 4 columns in the case of W = H = 4 and 8 columns in all other cases. b is a vector having a size of W red · H red .
[0134] The matrix A and the vector b are taken from one of the following sets S0, S1, S2. Define the following index idx = idx(W, H):
[0135]
[0136] In addition, let m be as follows:
[0137]
[0138] Thus, if idx ≤ 1 or idx = 2 and min(W, H)>4, then let and In the case of idx = 2 and min(W, H) = 4, let A be the matrix generated by omitting each row that corresponds to an odd x coordinate in the downsampling block when W = 4 or an odd y coordinate in the downsampling block when H = 4. .
[0139] Finally, replace the reduced prediction signal with its transpose in the following cases:
[0140] · W = H = 4 and mode ≥ 18
[0141] · max(W, H) = 8 and mode ≥ 10
[0142] · max(W, H) > 8 and mode ≥ 6
[0143] In the case where W = H = 4, pred red requires 4 multiplications for its calculation because in this case A has 4 columns and 16 rows. In all other cases, A has 8 columns and W red · H red rows, and it is immediately verified that in these cases 8 · W red · H red ≤ 4 · W · H multiplications are required, that is, in these cases as well, at most 4 multiplications per sample point are needed to calculate pred red .
[0144] 2.2.6.4. Illustration of the entire ALWIP process
[0145] In Figure 9 , Figure 10 , Figure 11 and Figure 12 the entire processes of mean value calculation, matrix - vector multiplication, and linear interpolation are illustrated for different shapes. Note that other shapes are treated as in one of the depicted cases.
[0146] 1. Given a 4×4 block, ALWIP takes two mean values along each axis of the boundary. The four resulting input sample points participate in matrix - vector multiplication. The matrix is taken from the set S0. After adding an offset, 16 final prediction sample points are obtained. Linear interpolation is not necessary for generating the prediction signal. Thus, a total of (4 · 16) / (4 · 4) = 4 multiplications are performed per sample point.
[0147] 2. Given an 8×8 block, ALWIP takes four mean values along each axis of the boundary. The eight resulting input sample points participate in matrix - vector multiplication. The matrix is taken from the set S1. It obtains 16 sample points at the odd positions of the prediction block. Thus, a total of (8 · 16) / (8 · 8) = 2 multiplications are performed per sample point. After adding an offset, these sample points are interpolated vertically using the reduced top boundary. Horizontal interpolation follows using the original left boundary.
[0148] 3. Given an 8×4 block, ALWIP takes four averages along the horizontal axis of the boundary and takes four original boundary values on the left boundary. The resulting eight input samples participate in a matrix-vector multiplication. The matrix is taken from set S1. It obtains 16 samples at the odd horizontal positions and each vertical position of the prediction block. Thus, each sample performs a total of (8·16) / (8·4) = 4 multiplications. After adding an offset, these samples are interpolated horizontally using the original left boundary.
[0149] The transposed case is treated accordingly.
[0150] 4. Given a 16×16 block, ALWIP takes four averages along each axis of the boundary. The resulting eight input samples participate in a matrix-vector multiplication. The matrix is taken from set S2. It obtains 64 samples at the odd positions of the prediction block. Thus, each sample performs a total of (8·64) / (16·16) = 2 multiplications. After adding an offset, these samples are interpolated vertically using the eight averages of the top boundary. Horizontal interpolation follows using the original left boundary. In this case, the interpolation process does not add any multiplications. Therefore, two multiplications per sample are required to calculate the ALWIP prediction.
[0151] For larger shapes, the process is basically the same, and it is easy to check that the number of multiplications per sample is less than four.
[0152] For a W×8 block (where W > 8), only horizontal interpolation is required because the samples are given at the odd horizontal positions and each vertical position.
[0153] Finally, for a W×4 block (where W > 8), let A k be the matrix generated by omitting each row corresponding to the odd entries along the horizontal axis of the downsampling block.
[0154] Thus, the output size is 32, and only horizontal interpolation still needs to be performed.
[0155] The transposed case is treated accordingly.
[0156] In the following discussion, the boundary samples used for multiplying with the matrix can be called "reduced boundary samples". The boundary samples used for interpolating the final prediction block by the downsampling block can be called "upsampling samples".
[0157] 2.2.6.5. Single-step linear interpolation
[0158] For a W×H block (where max(W,H) ≥ 8), the prediction signal is obtained by linear interpolation from the reduced prediction signal pred on W red ×H red red It is generated. According to the block size, linear interpolation is performed within the vertical direction, the horizontal direction, or both. If it is to be applied in both directions, then when W < H, it is first applied to the horizontal direction, otherwise it is first applied to the vertical direction.
[0159] Without loss of generality, consider a W×H block, where max(W, H) ≥ 8 and W ≥ H. Then one-dimensional linear interpolation is performed as follows. Without loss of generality, it is sufficient to describe the linear interpolation within the vertical direction. First, the reduced prediction signal is extended to the top by the boundary signal. Define the vertical upsampling factor U ver = H / H red , and write After that, the extended reduced prediction signal is defined by the following formula
[0160]
[0161] After that, from this extended reduced prediction signal, the vertically linearly interpolated prediction signal is generated by the following formula
[0162]
[0163] where, 0 ≤ x < W red , 0 ≤ y < H red and 0 ≤ k < U ver .
[0164] 2.2.6.6. Signaling of the proposed intra prediction mode
[0165] For each coding unit (CU) in the intra mode, a flag indicating whether the ALWIP mode will be applied to the corresponding prediction unit (PU) is sent in the bitstream. The signaling of the index of the latter is coordinated with the MRL. If the ALWIP mode will be applied, the index predmode of the ALWIP mode is signaled using an MPM list with 3 MPMs.
[0166] Here, the following intra modes of the upper and left PUs are used to perform the derivation of these MPMs. There are three fixed tables map_angular_to_alwip idx (idx ∈ {0, 1, 2}), which assign the ALWIP mode to each conventional intra prediction mode predmode Angular predmode
[0167] predmode ALWIp = map_angular_to_aiwip idx [predmode Angular .
[0168] For each PU with width W and height H, an index
[0169] idx(PU) = idx(W,H) ∈ {0,1,2}
[0170] is defined, which indicates from which of the three sets the ALWIP parameters will be taken, as in Section 2.2.6.3 above.
[0171] If the upper prediction unit PU above is available, belongs to the same CTU as the current PU, and is in the intra mode, if idc(PU) = idx(PU above ), and if ALWIP is applied to the PU in ALWIP mode above then let
[0172]
[0173] If the upper PU is available, belongs to the same CTU as the current PU, and is in the intra mode, and if the conventional intra prediction mode is applied to this upper PU then let
[0174]
[0175] In all other cases, let
[0176]
[0177] This means that this mode is not available. In the same way, without the restriction that the left PU must belong to the same CTU as the current PU, the mode
[0178] Finally, three fixed default lists list idx (idx ∈ {0,1,2}) are provided, each of which contains three distinct ALWIP modes. Three distinct MPMs are constructed from the default lists list idx(PU) and the modes and by substituting -1 with the default values and eliminating duplicates.
[0179] The left adjacent block and the upper adjacent block used in the construction of the ALWIP MPM list are A1 and B1.
[0180] 2.2.6.7. Adapted MPM List Derivation for Conventional Luminance and Chrominance Intra Prediction Modes
[0181] The proposed ALWIP mode is coordinated with the MPM-based encoding and decoding of the conventional intra prediction mode as follows. The luminance and chrominance MPM list derivation processes for the conventional intra prediction mode use a fixed table map_alwip_to_angular idx (idx ∈ {0, 1, 2}), so as to map the ALWIP mode predmode aLWIP on a given PU to one of the conventional intra prediction modes
[0182] predmode Angular = map_alwip_to_angular idx(PU) [predmode ALWIP .
[0183] For luminance MPM list derivation, whenever an adjacent luminance block using the ALWIP mode predmode ALWIP is encountered, this block is treated as if it uses the conventional intra prediction mode predmode Angular . For chrominance MPM list derivation, whenever the current luminance block uses the LWIP mode, the same mapping is adopted to convert the ALWIP mode to the conventional intra prediction mode.
[0184] 2.3. Inter Prediction in HEVC / H.265
[0185] For a coding unit (CU) for inter coding, it can be coded with the aid of a prediction unit (PU), two PUs according to the splitting mode. Each inter prediction PU has motion parameters for one or two reference picture lists. The motion parameters include a motion vector and a reference picture index. The use of one of the two reference picture lists can also be signaled using inter_pred_idc. The motion vector can be explicitly coded as Δ relative to a predictor.
[0186] When coding a CU in skip mode, a PU is associated with the CU, and there are no significant residual coefficients, no coded motion vector Δ or reference picture index. The merge mode is defined such that the motion parameters of the current PU are obtained from adjacent PUs including spatial candidates and temporal candidates. The merge mode can be applied to any inter prediction PU, not only for skip mode. An alternative to the merge mode is the explicit transmission of motion parameters, where, for each PU, the motion vector (more precisely, the motion vector difference (MVD) compared to the motion vector predictor), the corresponding reference picture index for each reference picture list, and the use of the reference picture list are explicitly signaled. Such a mode is referred to as advanced motion vector prediction (AMVP) in the present disclosure.
[0187] When the signaling indicates that one of these two reference picture lists will be used, the PU is generated from one sample block. This practice is called "unidirectional prediction". Unidirectional prediction is available for both P-bands and B-bands.
[0188] When the signaling indicates that both of these two reference picture lists will be used, the PU is generated from two sample blocks. This practice is called "bidirectional prediction". Bidirectional prediction is only available for B-bands.
[0189] Details of these inter-frame prediction modes specified in HEVC will be provided below. The description will start with the merge mode.
[0190] 2.3.1. Reference Picture Lists
[0191] In HEVC, the term inter-frame prediction is used to denote prediction derived from data elements (e.g., sample values or motion vectors) of reference pictures (rather than the current decoded picture). Similar to H.264 / AVC, a picture can be predicted from multiple reference pictures. The reference pictures used for inter-frame prediction are organized into one or more reference picture lists. The reference index identifies which reference picture in the list should be used to create the prediction signal.
[0192] A single reference picture list, i.e., List 0, is used for P-bands, and two reference picture lists, i.e., List 0 and List 1, are used for B-bands. It should be noted that the reference pictures included in List 0 / 1 can come from past pictures and future pictures with reference to the capture / display order.
[0193] 2.3.2. Merge Mode
[0194] 2.3.2.1. Derivation of merge mode candidates
[0195] When predicting the PU using the merge mode, an index pointing to an entry in the merge candidate list is parsed from the bitstream, and its motion information is retrieved. The construction of this list is specified in the HEVC standard and can be summarized according to the following sequence of steps:
[0196] · Step 1: Initial candidate derivation
[0197] ο Step 1.1: Spatial candidate derivation
[0198] ο Step 1.2: Redundancy check of spatial candidates
[0199] ο Step 1.3: Temporal candidate derivation
[0200] · Step 2: Additional candidate insertion
[0201] ο Step 2.1: Creation of Bidirectional Prediction Candidates
[0202] ο Step 2.2: Insertion of Zero Motion Candidates
[0203] These steps are also Figure 13 schematically illustrated. For spatial domain merge candidate derivation, up to four merge candidates are selected from among candidates located at five different positions. For temporal domain merge candidate derivation, up to one merge candidate is selected from among two candidates. Since a constant number of candidates are taken for each PU at the decoder, additional candidates are generated when the number of candidates obtained from step 1 does not reach the maximum number of merge candidates (MaxNumMergeCand) signaled in the slice header. Since the number of candidates is constant, truncated unary code binarization (TU) is used to encode the index of the best merge candidate. If the size of the CU is equal to 8, all PUs of the current CU share a single merge candidate list, which is equivalent to the merge candidate list of a 2N×2N prediction unit.
[0204] In the following, the operations associated with the foregoing steps will be described in detail.
[0205] 2.3.2.2. Spatial Domain Candidate Derivation
[0206] In the derivation of spatial domain merge candidates, up to four merge candidates are selected from among candidates located at the Figure 14 positions shown. The order of derivation is A1, B1, B0, A0, and B2. Position B2 is considered only if any of the PUs at positions A1, B1, B0, A0 are unavailable (e.g., because they belong to another strip or slice) or are intra-coded. After the candidate at position A1 is added, the addition of the remaining candidates is subject to a redundancy check, which ensures that candidates with the same motion information are excluded from the list, thereby improving coding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only pairs connected by Figure 15 the arrows are considered, and the candidate is added to the list only if the corresponding candidates for the redundancy check do not have the same motion information. Another source of duplicate motion information is the "second PU" associated with a partition different from 2N×2N. As an example, Figure 16 the second PUs for the N×2N and 2N×N cases are shown respectively. When the current PU is partitioned into N×2N, the candidate at position A1 is not considered for list construction. In fact, adding this candidate would cause two prediction units to have the same motion information, which is redundant for having only one PU within the coding unit. Similarly, when the current PU is partitioned into 2N×N, position B1 is not considered.
[0207] 2.3.2.3. Temporal candidate derivation
[0208] In this step, only one candidate is added to the list. In particular, in the derivation of this temporal merge candidate, the scaled motion vector is derived based on the collocated PUs of the picture with the smallest POC difference between the current picture and the pictures belonging to the given reference picture list. The reference picture list used for the derivation of the collocated PUs is signaled explicitly in the slice header. The scaled motion vector for the temporal merge candidate is obtained as shown by the dashed line in Figure 17 , which is scaled from the motion vector of the collocated PU using the POC distance (i.e., tb and td), where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the collocated picture and the collocated picture. The reference picture index of the temporal merge candidate is set to be equal to zero. The actual implementation of the scaling process is described in the HEVC specification. For B slices, two motion vectors are obtained (one for reference picture list 0 and the other for reference picture list 1), and they are combined to produce the bi-predictive merge candidate.
[0209] In the collocated PU (Y) of the reference frame, the position of the temporal candidate is selected between candidates C0 and C1, as shown in Figure 18 . If the PU at position C0 is not available, is intra-coded, or is outside the current coding tree unit (CTU, also known as LCU, largest coding unit) row, then position C1 is adopted. Otherwise, position C0 is used in the derivation of the temporal merge candidate.
[0210] 2.3.2.4. Additional candidate insertion
[0211] In addition to the spatial and temporal merge candidates, there are two additional types of merge candidates: the combined bi-predictive merge candidate and the zero merge candidate. The combined bi-predictive merge candidate is generated by using the spatial merge candidate and the temporal merge candidate. The combined bi-predictive merge candidate is only used for B slices. The combined bi-predictive candidate is generated by combining the motion parameters of the first reference picture list of the initial candidate with the motion parameters of the second reference picture list of the other. If the two tuples provide different motion hypotheses, then they form a new bi-predictive candidate. As an example, Figure 19Shows the situation when two candidates within the original list (left), which have mvL0 and refIdxL0 or mvL1 and refIdxL1, are used to create a combined bidirectional prediction merge candidate that is added to the final list (right). There are many rules regarding the combinations considered for generating these additional merge candidates.
[0212] Zero motion candidates are inserted to fill the remaining entries in the merge candidate list and thus hit the MaxNumMergeCand capacity. These candidates have zero spatial displacement and a reference picture index that starts at zero and increases each time a new zero motion candidate is added to the list. Finally, no redundancy checks are performed on these candidates.
[0213] 2.3.3.AMVP
[0214] AMVP utilizes the spatio-temporal correlation of motion vectors with adjacent PUs, which is used for the explicit transmission of motion parameters. For each reference picture list, the motion vector candidate list is constructed as follows: First, the availability of the left and upper temporally adjacent PU positions is checked, redundant candidates are removed, and a zero vector is added so that the candidate list has a constant length. Then, the encoder can select the best predictor from the candidate list and transmit the corresponding index indicating the selected candidate. Similar to merge index signaling, the index of the best motion vector candidate is encoded using a truncated unary code. The maximum value to be encoded in this case is 2 (refer to Figure 20 ). In the following sections, details of the derivation process of motion vector prediction candidates will be provided.
[0215] 2.3.3.1.Derivation of AMVP Candidates
[0216] Figure 20 Summarizes the derivation process of motion vector prediction candidates.
[0217] In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. For spatial motion vector candidate derivation, finally two motion vector candidates are derived based on the motion vectors of each PU at five different positions as shown in Figure 20 .
[0218] For temporal motion vector candidate derivation, one motion vector candidate is selected from two candidates derived from two different co-located positions. After creating the first list of spatio-temporal candidates, duplicate motion vector candidates in the list are removed. If the number of possible candidates is greater than two, then motion vector candidates are removed from the list whose reference picture index in the associated reference picture list is greater than 1. If the number of spatio-temporal motion vector candidates is less than two, then additional zero motion vector candidates are added to the list.
[0219] 2.3.3.2. Spatial motion vector candidates
[0220] In the derivation of spatial motion vector candidates, at most two candidates are considered among five possible candidates, which are derived from PUs located at positions as Figure 16 shown, and these positions are the same as those for motion merge. The derivation order for the left side of the current PU is defined as A0, A1, and scaled A0, scaled A1. The derivation order for the upper side of the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. For each side, there are thus four cases that can be used as motion vector candidates, where two cases do not require the use of spatial scaling, and two cases use spatial scaling. The summary of these four different cases is as follows:
[0221] · No spatial scaling
[0222] –(1) Same reference picture list and same reference picture index (same POC)
[0223] –(2) Different reference picture lists, but same reference picture (same POC)
[0224] · Spatial scaling
[0225] –(3) Same reference picture list, but different reference picture index (different POC)
[0226] –(4) Different reference picture lists and different reference pictures (different POC)
[0227] First, the no-spatial-scaling cases are checked, followed by spatial scaling. When there is a difference in the POC between the reference pictures of adjacent PUs and the reference picture of the current PU, spatial scaling is considered regardless of the reference picture list. If all PUs of the left-side candidate are unavailable or are intra-coded, then scaling for the upper-side motion vector is allowed, which helps in the parallel derivation of the left-side MV candidate and the upper-side MV candidate. Otherwise, spatial scaling for the upper-side motion vector is not allowed.
[0228] During the spatial scaling process, the motion vectors of adjacent PUs are scaled in a similar manner to temporal scaling, asFigure 21 As shown. The main difference is that the reference picture list and index of the current PU are given as input; the actual scaling process is the same as that of temporal scaling.
[0229] 2.3.3.3. Temporal Motion Vector Candidates
[0230] Except for the derivation of the reference picture index, all processes of the derivation of temporal merge candidates are the same as those of the derivation of spatial motion vector candidates (see Figure 20 ). The reference picture index is signaled to the decoder.
[0231] 2.4. Inter - frame Prediction Methods in VVC
[0232] There are several new codec tools for inter - frame prediction improvement, such as Adaptive Motion Vector Difference Resolution (AMVR) for signaling MVD, Merge with Motion Vector Difference (MMVD), Triangle Prediction Mode (TPM), Combined Intra - Inter Prediction (CIIP), Advanced TMVP (ATMVP, also known as SbTMVP), Affine Prediction Mode, Generalized Bi - directional Prediction (GBI), Decoder - side Motion Vector Refinement (DMVR), and Bi - directional Optical Flow (BIO, also known as BDOF).
[0233] There are three different merge list construction processes supported in VVC:
[0234] 1) Sub - block merge candidate list: It includes ATMVP and affine merge candidates. For both the affine mode and the ATMVP mode, a common merge list construction process is shared. Here, ATMVP and affine merge candidates can be added in sequence. The sub - block merge list size is signaled in the slice header, and the maximum value is 5.
[0235] 2) Regular merge list: For the remaining codec blocks, a common merge list construction process is shared. Here, spatial / temporal / HMVP paired - combination bi - directional prediction merge candidates and zero - motion candidates can be inserted in sequence. The regular merge list size is signaled in the slice header, and the maximum value is 6. MMVD, TPM, and CIIP rely on the regular merge list.
[0236] 3) IBC merge list: This list is completed in a similar way to the regular merge list.
[0237] Similarly, there are three AMVP lists supported in VVC:
[0238] 1) Affine AMVP candidate list
[0239] 2) Conventional AMVP candidate list
[0240] 3) IBC AMVP candidate list: The same construction process as the IBC merge list
[0241] 2.4.1. Coding and decoding block structure in VVC
[0242] In VVC, a quadtree / binary tree / trinary tree (QT / BT / TT) structure is adopted to divide the picture into square or rectangular blocks.
[0243] In addition to QT / BT / TT, a separate tree (also known as a dual coding tree) is adopted for I-frames in VVC. With the help of the separate tree, the coding and decoding block structure is signaled separately for the luminance component and the chrominance component.
[0244] In addition, except for the blocks coded and decoded using several specific coding and decoding methods (such as intra-subdivision prediction (where PU is equal to TU but smaller than CU) and sub-block transform for inter-coded blocks (where PU is equal to CU but TU is smaller than PU)), the CU is set to be equal to PU and TU.
[0245] 2.4.2. Affine prediction mode
[0246] In HEVC, only the translational motion model is applied for motion compensation prediction (MCP). However, in the real world, there are many types of motions, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, a simplified affine transform motion compensation prediction is applied using a 4-parameter affine model and a 6-parameter affine model. As shown in Figure 22, the affine motion field of the block is described by two control point motion vectors (CPMVs) of the 4-parameter affine model and three CPMVs of the 6-parameter affine model.
[0247] The motion vector field (MVF) of the block is described by the following equations respectively, where the 4-parameter affine model (where the 4 parameters are defined as variables a, b, e, and f) is in Equation (1), and the 6-parameter affine model (where the 6 parameters are defined as a, b, c, d, e, and f) is in Equation (2):
[0248]
[0249] where, (mv h 0, mv h 0) is the motion vector of the upper left control point, (mv h 1, mv h 1) is the motion vector of the upper right control point, (mv h 2, mv h2) is the motion vector of the lower left control point. All these three motion vectors are called control point motion vectors (CPMVs). (x, y) represents the coordinates of the representative point relative to the top - left sample within the current block. (mv h (x,y),mv v (x,y)) is the motion vector derived for the sample located at (x, y). The CP motion vector can be signaling - notified (such as in the affine AMVP mode) or can be derived on - the - fly (such as in the affine merge mode). w and h are the width and height of the current block. In practice, this division is implemented by right - shifting along with rounding operations. In VTM, the representative point is defined as the center position of the sub - block. For example, when the coordinates of the top - left corner of the sub - block relative to the top - left sample within the current block are (xs, ys), then the coordinates of the representative point are defined as (xs + 2, ys + 2). For each sub - block (e.g., 4x4 in VTM), this representative point is used to derive the motion vector for the entire sub - block.
[0250] To further simplify motion - compensated prediction, sub - block - based affine transform prediction is applied. To derive the motion vector for each M×N (in the current VVC, both M and N are set to 4) sub - block, the motion vector of the center sample of each sub - block is calculated according to Equation (1) and Equation (2) (as Figure 23 shown), and it is rounded to achieve 1 / 16 - fraction precision. Then, a motion - compensated interpolation filter for 1 / 16 pixels is applied to generate the prediction for each sub - block by means of the derived motion vector. The interpolation filter for 1 / 16 pixels is introduced by the affine mode.
[0251] After MCP, the high - precision motion vectors of each sub - block are rounded and saved with the same precision as the normal motion vectors.
[0252] 2.4.3. MERGE for the whole block
[0253] 2.4.3.1. Construction of the merge list for the translational conventional merge mode
[0254] 2.4.3.1.1. History - based motion vector prediction (HMVP)
[0255] Different from the merge list design, in VVC, the history - based motion vector prediction (HMVP) method is adopted.
[0256] In HMVP, the previously encoded / decoded motion information is stored. The motion information of the previously encoded block is defined as an HMVP candidate. Multiple HMVP candidates are stored in a table called the HMVP table, and this table is maintained on-the-fly during the encoding / decoding process. When starting to encode / decode a new slice / LCU row / strip, the HMVP table is cleared. Whenever there is an inter-coded block and non-sub-block non-TPM mode, the associated motion information is added as a new HMVP candidate to the last entry of the table. In Figure 24 the overall encoding / decoding process is illustrated.
[0257] 2.4.3.1.2. Conventional merge list construction process
[0258] The construction of the conventional merge list (for translational motion) can be summarized according to the following sequence of steps:
[0259] · Step 1: Derivation of spatial candidates
[0260] · Step 2: Insertion of HMVP candidates
[0261] · Step 3: Insertion of pairwise-averaged candidates
[0262] · Step 4: Default motion candidates
[0263] HMVP candidates can be used in both the AMVP candidate list construction process and the merge candidate list construction process. Figure 25 The modified merge list construction process is shown (highlighted in blue). When the merge candidate list is not full after the insertion of TMVP candidates, the HMVP candidates stored in the HMVP table can be used to fill the merge candidate list. Considering that a block often has a higher correlation with the nearest neighboring blocks in terms of motion information, the HMVP candidates in the table are inserted in descending order of the index. The last entry in the table is added to the list first, and the first entry is added last. Similarly, redundancy elimination is applied to the HMVP candidates. Once the total number of available merge candidates reaches the maximum number of merge candidates allowed to be signaled, the merge candidate list construction process is terminated.
[0264] Note that all spatial / temporal / HMVP candidates must be encoded / decoded in non-IBC mode. Otherwise, it is not allowed to add them to the conventional merge candidate list.
[0265] The HMVP table contains up to 5 conventional motion candidates, and each of them is unique.
[0266] 2.4.3.2. Triangular prediction mode (TPM)
[0267] In VTM4, triangular partitioning mode is supported for inter prediction. The triangular partitioning mode is applied only to CUs that are 8x8 or larger and are coded in merge mode rather than in MMVD or CIIP mode. For a CU that meets these conditions, a CU-level flag is signaled to indicate whether the triangular partitioning mode is applied.
[0268] When using this mode, the CU is evenly divided into two triangular partitions using a diagonal or anti-diagonal partition, as Figure 26 shown. Inter prediction is performed for each triangular partition within the CU using its own motion; only unidirectional prediction is allowed for each partition, i.e., each partition has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, like conventional bi-directional prediction, each CU requires only two motion-compensated predictions.
[0269] If the CU-level flag indicates that the current CU is coded using the triangular partitioning mode, then a flag indicating the direction of the triangular partition (diagonal or anti-diagonal) and two merge indices (one for each partition) are further signaled. After each type of triangular partition is predicted, hybrid processing with adaptive weights is used to adjust the sample values along the diagonal or anti-diagonal edges. This is the prediction signal for the entire CU, and the transform and quantization processes are applied to this entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the triangular partitioning mode is stored within 4x4 units.
[0270] The regular merge candidate list is reused for triangular partition merge prediction without additional motion vector pruning. For each merge candidate in the regular merge candidate list, only one of its L0 or L1 motion vectors is used for triangular prediction. In addition, the selection order of the L0 versus L1 motion vectors is based on the parity of their merge indices. With this scheme, the regular merge list can be used directly.
[0271] 2.4.3.3. MMVD
[0272] The ultimate motion vector representation (UMVE, also known as MMVD) will be introduced. UMVE is used for skip mode or merge mode in combination with the proposed motion vector representation method.
[0273] UMVE reuses the same merge candidates as those included in the regular merge candidate list in VVC. Among these merge candidates, a base candidate can be selected and further augmented by the proposed motion vector representation method.
[0274] UMVE provides a new method for representing the motion vector difference (MVD), where the MVD is represented by a starting point, a motion amplitude, and a motion direction.
[0275] The proposed technique uses the merge candidate list as it is. However, only candidates of the default merge type (MRG_TYPE_DEFAULT_N) are considered for the extension of UMVE.
[0276] The base candidate index defines the starting point. The base candidate index indicates the best candidate among the candidates in the list, as described below.
[0277] Table 2 Base candidate IDX
[0278] Base Candidate IDX 0 1 2 3 The Nth MVP The 1st MVP The 2nd MVP The 3rd MVP The 4th MVP
[0279] If the number of base candidates is equal to 1, then the base candidate IDX is not signaled.
[0280] The distance index is the motion amplitude information. The distance index indicates a predefined distance by the starting point information. Pre
[0281] The distance is defined as follows:
[0282] Table 3 Distance IDX
[0283]
[0284] The direction index represents the direction of the MVD relative to the starting point. The direction index can represent the four directions as shown below.
[0285] Table 4 Direction IDX
[0286] Direction IDX 00 01 10 11 x - axis + – N / A N / A y - axis N / A N / A + –
[0287] The UMVE flag is signaled immediately after the skip flag or the merge flag is sent. If the skip or merge flag is true, then the UMVE flag is parsed. If the UMVE flag is equal to 1, then the UMVE syntax is parsed. However, if it is not equal to 1, then the AFFINE flag is parsed. If the AFFINE flag is equal to 1, it is the AFFINE mode, but if it is not equal to 1, then the skip / merge index is parsed for the skip / merge mode of VTM.
[0288] No additional line buffer is required due to UMVE candidates. Because the software skip / merge candidates are directly used as base candidates. In the case of using the input UMVE index, the supplementation of the MV is determined just before motion compensation. There is no need to maintain a long line buffer for this.
[0289] Under current normal test conditions, the first merge candidate or the second merge candidate in the merge candidate list can be selected as the base candidate.
[0290] UMVE is also known as Merge with MV Difference (MMVD).
[0291] 2.4.3.4. Combined Intra-Inter Prediction (CIIP)
[0292] Multiple hypothesis prediction is proposed, where combined intra and inter prediction is a way to generate multiple hypotheses.
[0293] When applying multiple hypothesis prediction to improve the intra mode, multiple hypothesis prediction combines an intra prediction and a merge index prediction. In a merge CU, a flag is signaled for the merge mode, so that when the flag is true, the intra mode is selected from the intra candidate list. For the luminance component, the intra candidate list is derived from only one intra prediction mode (i.e., the planar mode). The weights applied to the predicted blocks from intra and inter prediction are determined by the coding / decoding modes (intra or non-intra) of two adjacent blocks (A1 and B1).
[0294] 2.4.4. MERGE for Sub-Block Based Techniques
[0295] It is proposed to place all sub-block related motion candidates in a separate merge list in addition to the regular merge list for non-sub-block merge candidates.
[0296] The separate merge list for placing sub-block related motion candidates is called the "sub-block merge candidate list".
[0297] In one example, the sub-block merge candidate list includes ATMVP candidates and affine merge candidates.
[0298] The sub-block merge candidate list is filled with candidates in the following order:
[0299] a. ATMVP candidates (may be available or not);
[0300] b. Affine merge list (including inherited affine candidates; and constructed affine candidates)
[0301] c. Zero-filled MV 4-parameter affine model
[0302] 2.2.4.1.1. ATMVP (also known as Sub-Block Temporal Motion Vector Predictor, SbTMVP)
[0303] The basic idea of ATMVP is to derive multiple sets of temporal motion vector predictors for a block. A set of motion information is assigned to each sub-block. When generating ATMVP merge candidates, motion compensation is performed at the 8×8 level instead of the entire block level.
[0304] 2.4.5. Conventional Inter Prediction Mode (AMVP)
[0305] 2.4.5.1. AMVP Motion Candidate List
[0306] Similar to the AMVP design in HEVC, up to 2 AMVP candidates can be derived. However, HMVP candidates can also be added after TMVP candidates. The HMVP candidates in the HMVP table are traversed in ascending order of index (i.e., starting from the index equal to 0, which is the earliest index). Up to 4 HMVP candidates can be checked to find out if their reference pictures are the same as the target reference picture (i.e., the same POC value).
[0307] 2.4.5.2. AMVR
[0308] In HEVC, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the motion vector of the PU and the predicted motion vector) is signaled in units of quarter luminance samples. In VVC, local adaptive motion vector resolution (AMVR) is introduced. In VVC, the MVD can be encoded and decoded in units of quarter luminance samples, integer luminance samples, and quarter luminance samples (i.e., 1 / 4 pixel, 1 pixel, 4 pixels). The MVD resolution is controlled at the coding unit (CU) level, and the MVD resolution flag is signaled in a conventional manner for each CU with at least one non-zero MVD component.
[0309] For a CU with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter luminance sample MV precision is used in the CU. When the first flag (equal to 1) indicates that quarter luminance sample MV precision is not used, another flag is signaled to indicate whether integer luminance sample MV precision or quarter luminance sample MV precision is adopted.
[0310] When the first MVD resolution flag of a CU is zero, or when this flag is not coded for the CU (meaning that all MVDs within the CU are zero), quarter luminance sample MV resolution is used for the CU. When the CU uses integer luminance sample MV precision or quarter luminance sample MV precision, the MVP in the AMVP candidate list of the CU is rounded to the corresponding precision.
[0311] 2.4.5.3. Symmetric Motion Vector Difference
[0312] Apply Symmetric Motion Vector Difference (SMVD) to the coding and decoding of motion information in bidirectional prediction.
[0313] First, as specified in N1001-v2, at the slice level, the following steps are used to derive variables RefIdxSymL0 and RefIdxSymL1 that respectively indicate the reference picture indices of list 0 / 1 used in the SMVD mode. When at least one of these two variables is equal to -1, the SMVD mode must be disabled.
[0314] 2.5. Multiple Transform Selection (MTS)
[0315] In addition to DCT-II already adopted in HEVC, the Multiple Transform Selection (MTS) scheme is also used for residual coding of both inter-coded and intra-coded blocks. It uses multiple selected transforms from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. The basis functions of the selected DST / DCT are shown.
[0316] Table 5 - Transform basis functions of DCT-II / III and DSTVII for N-point input
[0317]
[0318]
[0319] To maintain the orthogonality of the transform matrices, more accurate quantization of these transform matrices is made compared to the transform matrices in HEVC. To keep the intermediate values of the transformed coefficients within the 16-bit range, after horizontal transformation and after vertical transformation, all coefficients will have 10 bits.
[0320] To control the MTS scheme, separate enable flags are specified at the SPS level for intra and inter respectively. When MTS is enabled on the SPS, a CU-level flag is signaled to indicate whether MTS is applied. Here, MTS is only applied to the luminance. The MTS CU-level flag is signaled when the following conditions are met.
[0321] - Both the width and height are less than or equal to 32
[0322] - The CBF flag is equal to 1
[0323] If the MTS CU flag is equal to 0, then DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, then additionally two other flags are signaled to indicate the transform types in the horizontal and vertical directions respectively. The transform and signaling mapping table is shown in Table 6. When it comes to the accuracy of the transform matrix, an 8-bit main transform core is adopted. Therefore, all the transform cores used in HEVC are kept the same, which includes 4-point DCT-2 and DST-7 as well as 8-point, 16-point, and 32-point DCT-2. Moreover, other transform cores include 64-point DCT-2, 4-point DCT-8, and 8-point, 16-point, 32-point DST-7 and DCT-8 main transform cores.
[0324] Table 6 - Transform and Signaling Mapping Table
[0325]
[0326]
[0327] Similar to HEVC, the transform skip mode can be adopted to encode and decode the residual of a block. To avoid the redundancy of syntax encoding and decoding, the transform skip flag is not signaled when the MTS_CU_flag at the CU level is not equal to zero. The transform skip is enabled when the block width and height are equal to or less than 4.
[0328] 2.6. Intra Block Copy
[0329] Intra Block Copy (IBC), also known as current picture reference, has been adopted in the High Efficiency Video Coding Screen Content Coding Extension (HEVC-SCC) and the current Versatile Video Coding Test Model (VTM-4.0). IBC extends the concept of motion compensation from inter coding to intra coding. As shown in Figure 29 , when applying IBC, the current block is predicted by a reference block in the same picture. Before encoding or decoding the current block, the samples in the reference block must have been reconstructed. Although IBC is not that efficient for most camera-captured sequences, it shows significant coding and decoding gains for screen content. The reason is that there are a large number of repeating patterns in screen content pictures, such as icons and text characters. IBC can effectively eliminate the redundancy between these repeating patterns. In HEVC-SCC, if IBC selects the current picture as its reference picture, then the coding and decoding unit (CU) of inter coding can apply IBC. In this case, the motion vector (MV) is renamed as the block vector (BV), and the BV always has the accuracy of integer pixels. To be compatible with the main profile of HEVC, the current picture is marked as a "long-term" reference picture in the decoded picture buffer (DPB). It should be noted that, similarly, in multiple view / 3D video coding standards, the inter-view reference pictures are also marked as "long-term" reference pictures.
[0330] When BV is followed to find its reference block, prediction can be generated by copying the reference block. The residual can be obtained by subtracting the reference pixels from the initial signal. Then, transformation and quantization can be applied as in other coding and decoding modes.
[0331] However, when the reference block is outside the picture, or overlaps with the current block, or is outside the reconstructed area, or is outside the valid area restricted by certain constraints, some or all of the pixel values are not defined. Basically, there are two solutions to handle such problems. One solution is not to allow such situations, for example, in bitstream consistency. Another solution is to apply filling for those undefined pixel values. The following subsections will describe these solutions in detail.
[0332] 2.6.1. IBC in the VCC Test Model (VTM4.0)
[0333] In the current VCC test model, that is, in the VTM-4.0 design, the entire reference block should utilize the current coding tree unit (CTU) and not overlap with the current block. Therefore, there is no need to fill the reference block or the prediction block. The IBC flag is coded as the prediction mode of the current CU. Thus, there are three prediction modes for each CU, namely MODE_INTRA, MODE_INTER, and MODE_IBC.
[0334] 2.6.1.1. IBC Merge Mode
[0335] In the IBC Merge mode, an index pointing to an entry in the IBC merge candidate list is parsed from the bitstream. The construction of the IBC merge list can be summarized according to the following sequence of steps:
[0336] · Step 1: Derivation of spatial candidates
[0337] · Step 2: Insertion of HMVP candidates
[0338] · Step 3: Insertion of pairwise average candidates
[0339] In the derivation of spatial merge candidates, up to four merge candidates are selected from the candidates at the positions shown as A1, B1, B0, A0, and B2 in Figure 14 shown, and the order of derivation is A1, B1, B0, A0, and B2. Position B2 is considered only when any PU at positions A1, B1, B0, A0 is not available (e.g., because it belongs to another strip or slice) or is not coded in the IBC mode. After the candidate at position A1 is added, the insertion of the remaining candidates is subject to a redundancy check, which ensures the exclusion of candidates with the same motion information from the list, thereby improving the coding and decoding efficiency. 0、
[0340] After inserting the spatial candidate, if the IBC merge list size is still smaller than the maximum IBC merge list size, IBC candidates from the HMVP table can be inserted. A redundancy check is performed when inserting HMVP candidates. Finally, paired average candidates are inserted into the IBC merge list.
[0341] When the reference block of a merge candidate is outside the picture, or overlaps with the current block, or is outside the reconstruction region, or is outside the valid region restricted by certain constraints, the merge candidate is called an invalid merge candidate.
[0342] It should be noted that invalid merge candidates can be inserted into the IBC merge list.
[0343] 2.6.1.2. IBC AMVP Mode
[0344] In the IBC AMVP mode, the AMVP index pointing to an entry in the IBC AMVP list is parsed from the bitstream. The construction of the IBC AMVP list can be summarized according to the following step sequence:
[0345] · Step 1: Derivation of spatial candidates
[0346] ο Check A0, A1 until an available candidate is found.
[0347] ο Check B0, B1, B2 until an available candidate is found.
[0348] · Step 2: Insertion of HMVP candidates
[0349] · Step 3: Insertion of zero candidates
[0350] After inserting the spatial candidate, if the IBC AMVP list size is still smaller than the maximum IBC AMVP list size, IBC candidates from the HMVP table can be inserted.
[0351] Finally, zero candidates are inserted into the IBC AMVP list.
[0352] 2.6.1.3. Chrominance IBC Mode
[0353] In the current VVC, motion compensation in the chroma IBC mode is performed at the sub-block level. The chroma block is divided into several sub-blocks. Each sub-block determines whether the corresponding luma block has a block vector and, if so, determines its validity. There are encoder constraints in the current VTM, where if all sub-blocks in the current chroma CU have valid luma block vectors, the chroma IBC mode will be tested. For example, in a YUV 420 video, if the chroma block is N×M, the co-located luma region is 2N×2M. The sub-block size of the chroma block is 2×2. There are several steps to perform chroma mv derivation and then the block copy process.
[0354] 1) The chroma block is first divided into (N>>1)*(M>>1) sub-blocks.
[0355] 2) Each sub-block with the top-left sample coordinates (x,y) obtains the corresponding luma block covering the same top-left sample with coordinates (2x,2y).
[0356] 3) The encoder checks the block vector (bv) of the obtained luma block. If any of the following conditions is met, the bv is considered invalid.
[0357] a. The bv of the corresponding luma block does not exist.
[0358] b. The predicted block identified by the bv has not been reconstructed.
[0359] c. The predicted block identified by the bv partially or fully overlaps with the current block.
[0360] 4) The chroma motion vector of the sub-block is set to the motion vector of the corresponding luma sub-block.
[0361] When valid bvs are found for all sub-blocks, the IBC mode is allowed at the encoder.
[0362] The decoding process of the IBC block is listed below. The part related to chroma mv derivation in the IBC mode is highlighted as Grey .
[0363] 8.6.1 General decoding process of the coding unit for IBC prediction coding The inputs to this process are:
[0364] – The luma position (xCb,yCb), which specifies the top-left sample of the current coding block relative to the top-left luma sample of the current picture,
[0365] – The variable cbWidth, which specifies the width of the current coding block in luma samples,
[0366] – The variable cbHeight, which specifies the height of the current coding block in luma samples,
[0367] – The variable treeType specifies whether to use a unary tree or a binary tree. If a binary tree is used, it specifies whether the current tree corresponds to the luminance component or the chrominance component.
[0368] The output of this process is the modified reconstructed picture before loop filtering.
[0369] Call the derivation process for quantization parameters specified in Clause 8.7.1, which takes as inputs the luminance position (xCb, yCb), the width cbWidth of the current coded block in luminance samples, the height cbHeight of the current coded block in luminance samples, and the variable treeType.
[0370] The decoding process of a coded unit coded in the IBC prediction mode consists of the following ordered steps:
[0371] 1. Derive the motion vector components of the current coded unit as follows:
[0372] 1. If treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, the following applies:
[0373] – Call the derivation process for motion vector components specified in Clause 8.6.2.1, which takes as inputs the luminance coded block position (xCb, yCb), the luminance coded block width cbWidth, and the luminance coded block height cbHeight, and outputs the luminance motion vector mvL[0][0].
[0374] – When treeType is equal to SINGLE_TREE, call the derivation process for chrominance motion vectors in Clause 8.6.2.9, which takes the luminance motion vector mvL[0][0] as input and outputs the chrominance motion vector mvC[0][0].
[0375] – The number of luminance coded sub-blocks in the horizontal direction numSbX and the number in the vertical direction numSbY are both set to be equal to 1.
[0376] 1. Otherwise, if treeType is equal to DUAL_TREE_CHROMA, the following applies:
[0377] – Derive the number of luminance coded sub-blocks numSbX in the horizontal direction and the number numSbY in the vertical direction as follows:
[0378] numSbX = (cbWidth >> 2) (8 - 886)
[0379] numSbY = (cbHeight >> 2) (8 - 887)
[0380] – For xSbIdx = 0..numSbX - 1, ySbIdx = 0..numSbY - 1, derive the chrominance motion vector mvC [xSbIdx][ySbIdx] as follows:
[0381] – Derive the luma motion vector mvL[xSbIdx][ySbIdx] as follows:
[0382] – Derive the position (xCuY, yCuY) of the co - located luma coding / decoding unit as follows:
[0383] xCuY = xCb + xSbIdx * 4 (8 - 888)
[0384] yCuY = yCb + ySbIdx * 4 (8 - 889)
[0385] – If CuPredMode[xCuY][yCuY] is equal to MODE_INTRA, the following applies:
[0386] mvL[xSbIdx][ySbIdx][0] = 0 (8 - 890)
[0387] mvL[xSbIdx][ySbIdx][1] = 0 (8 - 891)
[0388] predFlagL0[xSbIdx][ySbIdx] = 0 (8 - 892)
[0389] predFlagL1[xSbIdx][ySbIdx] = 0 (8 - 893)
[0390] – Otherwise (CuPredMode[xCuY][yCuY] is equal to MODE_IBC), the following applies:
[0391] mvL[xSbIdx][ySbIdx][0] = MvL0[xCuY][yCuY][0] (8 - 894)
[0392] mvL[xSbIdx][ySbIdx][1] = MvL0[xCuY][yCuY][1] (8 - 895)
[0393] predFlagL0[xSbIdx][ySbIdx] = 1 (8 - 896)
[0394] predFlagL1[xSbIdx][ySbIdx] = 0 (8 - 897)
[0395] – Invoke the derivation process of the chrominance motion vector in Clause 8.6.2.9, which takes mvL[xSbIdx][ySbIdx] as the input and mvC[xSbIdx][ySbIdx] as the output.
[0396] – The requirement for bitstream consistency is that the chrominance motion vector mvC[xSbIdx][ySbIdx] shall comply with the following constraints:
[0397] – When invoking the derivation process for block availability as specified in Clause 6.4.X [Ed.(BB): Adjacent block availability check process to be determined], where the current chroma position (xCurr, yCurr) set to be equal to (xCb / SubWidthC, yCb / SubHeightC) and the adjacent chroma position (xCb / SubWidthC+(mvC[xSbIdx][ySbIdx][0]>>5), yCb / SubHeightC+(mvC[xSbIdx][ySbIdx][1]>>5)) are used as inputs, the output must be equal to true.
[0398] – When invoking the derivation process for block availability as specified in Clause 6.4.X [Ed.(BB): Adjacent block availability check process to be determined], where the current chroma position (xCurr, yCurr) set to be equal to (xCb / SubWidthC, yCb / SubHeightC) and the adjacent chroma position (xCb / SubWidthC+(mvC[xSbIdx][ySbIdx][0]>>5)+cbWidth / SubWidthC - 1, yCb / SubHeightC+(mvC[xSbIdx][ySbIdx][1]>>5)+cbHeight / SubHeightC - 1) are used as inputs, the output must be equal to true.
[0399] – One or both of the following conditions shall be true:
[0400] – (mvC[xSbIdx][ySbIdx][0]>>5)+xSbIdx*2 + 2 is less than or equal to 0.
[0401] – (mvC[xSbIdx][ySbIdx][1]>>5)+ySbIdx*2 + 2 is less than or equal to 0.
[0402] 2. Derive the predicted samples of the current coding unit as follows:
[0403] – If treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, then derive the predicted samples of the current coding unit as follows:
[0404] – Call the decoding process of the IBC block specified in Clause 8.6.3.1, which takes as input the luma coding block position (xCb, yCb), luma coding block width cbWidth and luma coding block height cbHeight, the number of luma coding sub-blocks in the horizontal direction numSbX and in the vertical direction numSbY, the luma motion vectors mvL[xSbIdx][ySbIdx] where xSbIdx = 0..numSbX-1 and ySbIdx = 0..numSbY-1, and a variable cIdx set to be equal to 0, and outputs an (cbWidth)x(cbHeight) array predSamples of predicted luma samples L as the IBC predicted samples (predSamples).
[0405] – Otherwise, if treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, derive the predicted samples of the current coding unit as follows:
[0406] – Call the decoding process of the IBC block specified in Clause 8.6.3.1, which takes as input the luma coding block position (xCb, yCb), luma coding block width cbWidth and luma coding block height cbHeight, the number of luma coding sub-blocks in the horizontal direction numSbX and in the vertical direction numSbY, the chroma motion vectors mvC[xSbIdx][ySbIdx] where xSbIdx = 0..numSbX-1 and ySbIdx = 0..numSbY-1, and a variable cIdx set to be equal to 1, and outputs an (cbWidth / 2)x(cbHeight / 2) array predSamples of predicted chroma samples for chrominance component Cb Cb as the IBC predicted samples (predSamples).
[0407] – Call the decoding process of the IBC block specified in Clause 8.6.3.1, which takes as input the luma coding block position (xCb, yCb), luma coding block width cbWidth and luma coding block height cbHeight, the number of luma coding sub-blocks in the horizontal direction numSbX and in the vertical direction numSbY, the chroma motion vectors mvC[xSbIdx][ySbIdx] where xSbIdx = 0..numSbX-1 and ySbIdx = 0..numSbY-1, and a variable cIdx set to be equal to 2, and outputs an (cbWidth / 2)x(cbHeight / 2) array predSamples of predicted chroma samples for chrominance component Cr CrThe ibc predicted samples (predSamples) are output.
[0408] 3. The variables NumSbX[xCb][yCb] and NumSbY[xCb][yCb] are set to be equal to numSbX and numSbY respectively.
[0409] 4. Derive the residual samples of the current coding / decoding unit as follows:
[0410] – When treeType is equal to SINGLE_TREE or treeType is equal to DUAL_TREE_LUMA, call the decoding process of the residual signal of the coded / decoded block coded in the inter prediction mode specified in Clause 8.5.8, which takes the position (xTb0, yTb0) set to be equal to the luma position (xCb, yCb), the width nTbW set to be equal to the luma coded / decoded block width cbWidth, the height nTbH set to be equal to the luma coded / decoded block height cbHeight, and the variable cIdx set to be equal to 0 as inputs, and takes the array resSamples L as output.
[0411] – When treeType is equal to SINGLE_TREE or treeType is equal to DUAL_TREE_CHROMA, call the decoding process of the residual signal of the coded / decoded block coded in the inter prediction mode specified in Clause 8.5.8, which takes the position (xTb0, yTb0) set to be equal to the chroma position (xCb / 2, yCb / 2), the width nTbW set to be equal to the chroma coded / decoded block width cbWidth / 2, the height nTbH set to be equal to the chroma coded / decoded block height cbHeight / 2, and the variable cIdx set to be equal to 1 as inputs, and takes the array resSamples Cb as output.
[0412] – When treeType is equal to SINGLE_TREE or treeType is equal to DUAL_TREE_CHROMA, call the decoding process of the residual signal of the coded / decoded block coded in the inter prediction mode specified in Clause 8.5.8, which takes the position (xTb0, yTb0) set to be equal to the chroma position (xCb / 2, yCb / 2), the width nTbW set to be equal to the chroma coded / decoded block width cbWidth / 2, the height nTbH set to be equal to the chroma coded / decoded block height cbHeight / 2, and the variable cIdx set to be equal to 2 as inputs, and takes the array resSamples Cr as output.
[0413] 5. Derive the reconstructed samples of the current coding unit as follows:
[0414] – When treeType is equal to SINGLE_TREE or treeType is equal to DUAL_TREE_LUMA, call the picture reconstruction process for color components specified in Clause 8.7.5, which takes the block position (xB, yB) set to be equal to (xCb, yCb), the block width bWidth set to be equal to cbWidth, the block height bHeight set to be equal to cbHeight, the variable cIdx set to be equal to 0, the (cbWidth) x (cbHeight) array predSamples set to be equal to predSamples L and the (cbWidth) x (cbHeight) array resSamples set to be equal to resSamples L as inputs, and the output is the modified reconstructed picture before loop filtering.
[0415] – When treeType is equal to SINGLE_TREE or treeType is equal to DUAL_TREE_CHROMA, call the picture reconstruction process for color components specified in Clause 8.7.5, which takes the block position (xB, yB) set to be equal to (xCb / 2, yCb / 2), the block width bWidth set to be equal to cbWidth / 2, the block height bHeight set to be equal to cbHeight / 2, the variable cIdx set to be equal to 1, the (cbWidth / 2) x (cbHeight / 2) array predSamples set to be equal to predSamples Cb and the (cbWidth / 2) x (cbHeight / 2) array resSamples set to be equal to resSamples Cb as inputs, and the output is the modified reconstructed picture before loop filtering.
[0416] – When treeType is equal to SINGLE_TREE or treeType is equal to DUAL_TREE_CHROMA, call the picture reconstruction process for color components specified in Clause 8.7.5, which takes the block position (xB, yB) set to be equal to (xCb / 2, yCb / 2), the block width bWidth set to be equal to cbWidth / 2, the block height bHeight set to be equal to cbHeight / 2, the variable cIdx set to be equal to 2, the (cbWidth / 2) x (cbHeight / 2) array predSamples set to be equal to predSamples CrThe (cbWidth / 2) x (cbHeight / 2) array predSamples and is set to be equal to resSamples Cr The (cbWidth / 2) x (cbHeight / 2) array resSamples as input, and the output is the modified reconstructed picture before loop filtering.
[0417] 2.6.2. Recent progress of IBC (in VTM5.0)
[0418] 2.6.2.1. Single BV list
[0419] The BV predictors for merge mode and AMVP mode in IBC will share a common predictor list, which consists of the following elements:
[0420] · 2 spatially adjacent positions (such as Figure 14 A1, B1 in)
[0421] · 5 HMVP entries
[0422] · Default zero vectors
[0423] The number of candidates in the list is controlled by a variable derived from the slice header. For the merge mode, up to the first 6 entries of this list will be used; for the AMVP mode, the first 2 entries of this list will be used. This list meets the shared merge list area requirements (the same list within the shared SMR).
[0424] In addition to the above BV predictor candidate list, pruning operations between simplified HMVP candidates and existing merge candidates (A1, B1) are also proposed. In the simplification, there will be at most 2 pruning operations because it only compares the first HMVP candidate with the spatial domain merge candidate.
[0425] 2.6.2.1.1. Decoding process
[0426] 8.6.2.2 Derivation process for IBC luma motion vector prediction
[0427] This process is only called when CuPredMode[xCb][yCb] is equal to MODE_IBC, where (xCb, yCb) specifies the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture.
[0428] The inputs to this process are:
[0429] – The luma position (xCb, yCb) of the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture,
[0430] – The variable cbWidth, which specifies the width of the current coding block in luminance samples,
[0431] – The variable cbHeight, which specifies the height of the current coding block in luminance samples.
[0432] The output of this process is:
[0433] – The luminance motion vector mvL with 1 / 16 fractional sample precision.
[0434] Derive the variables xSmr, ySmr, smrWidth, smrHeight, and smrNumHmvpIbcCand as follows:
[0435] xSmr = IsInSmr[xCb][yCb]? SmrX[xCb][yCb] : xCb (8 - 910)
[0436] ySmr = IsInSmr[xCb][yCb]? SmrY[xCb][yCb] : yCb (8 - 911)
[0437] smrWidth = IsInSmr[xCb][yCb]? SmrW[xCb][yCb] : cbWidth
[0438] (8 - 912)
[0439] smrHeight = IsInSmr[xCb][yCb]? SmrH[xCb][yCb] : cbHeight
[0440] (8 - 913)
[0441] smrNumHmvpIbcCand = IsInSmr[xCb][yCb]? NumHmvpSmrIbcCand : NumHmvpIbcCand
[0442] (8 - 914)
[0443] Derive the luminance motion vector mvL through the following ordered steps:
[0444] 1. Invoke the derivation process of the spatial motion vector candidates from adjacent codec units specified in Clause 8.6.2.3, with the luma codec block position (xCb, yCb) set to be equal to (xSmr, ySmr), the luma codec block width cbWidth and the luma codec block height cbHeight set to be equal to smrWidth and smrHeight as inputs, and the availability flags availableFlagA1, availableFlagB1 and the motion vectors mvA1 and mvB1 as outputs.
[0445] 2. Construct the motion vector candidate list mvCandList as follows:
[0446]
[0447] 3. The variable numCurrCand is set to be equal to the number of merging candidates in mvCandList.
[0448] 4. When numCurrCand is less than MaxNumMergeCand and smrNumHmvpIbcCand is greater than 0, invoke the derivation process of the motion vector candidates based on IBC history specified in 8.6.2.4, with mvCandList, isInSmr set to be equal to IsInSmr[xCb][yCb] and numCurrCand as inputs, and the modified mvCandList and numCurrCand as outputs.
[0449] 5. When numCurrCand is less than MaxNumMergeCand, the following applies until numCurrCand is equal to MaxNumMergeCand:
[0450] 1. Set mvCandList[numCurrCand][0] to be equal to 0.
[0451] 2. Set mvCandList[numCurrCand][1] to be equal to 0.
[0452] 3. Increment numCurrCand by 1.
[0453] 6. Derive the variable mvIdx as follows:
[0454] mvIdx = general_merge_flag[xCb][yCb]? merge_idx[xCb][yCb] : mvp_l0_flag[xCb][yCb] (8 - 916)
[0455] 7. Make the following assignments:
[0456] mvL[0] = mergeCandList[mvIdx][0] (8 - 917)
[0457] mvL[1] = mergeCandList[mvIdx][1] (8 - 918)
[0458] 2.6.2.2. Size Limitations of IBC
[0459] In the latest VVC and VTM5, it is proposed to explicitly use syntax constraints to disable the 128x128 IBC mode on top of the current bitstream constraints in the previous VTM and VVC versions, which makes the presence of the IBC flag depend on the CU size < 128x128.
[0460] 2.6.2.3. Shared Merge List for IBC
[0461] To reduce decoder complexity and support parallel encoding, it is proposed that all leaf codec units (CUs) of an ancestor node in the CU partition tree share the same merging candidate list, thus enabling parallel processing of small skip / merge codec CUs. The ancestor node is named the merge shared node. The shared merging candidate list is generated at the merge shared node, pretending the merge shared node is a leaf CU.
[0462] More specifically, the following can apply:
[0463] - If a block has luma samples no greater than 32 and is partitioned into 2 4×4 sub - blocks, a shared merge list is used between very small blocks (e.g., two adjacent 4×4 blocks).
[0464] - If the block has more than 32 luma samples, however, after partitioning, at least one sub - block is smaller than the threshold (32), then all sub - blocks of this partition share the same merge list (e.g., 16×4 or 4×16 ternary partition or 8×8 quaternary partition).
[0465] Such a restriction only applies to the IBC merge mode.
[0466] 2.6.3. Quantized Residual Differential Pulse Code Modulation (RDPCM)
[0467] VTM5 supports Quantized Residual Differential Pulse Code Modulation (RDPCM) for screen content coding and decoding.
[0468] When RDPCM is enabled, if the CU size is less than or equal to 32x32 luma samples and if the CU is intra-coded, then the transmission flag is at the CU level. This flag indicates whether conventional intra-coding or RDPCM is used. If RDPCM is used, then the RDPCM prediction direction flag is transmitted to indicate whether the prediction is horizontal or vertical. Thereafter, the block is predicted using the conventional horizontal or vertical intra-prediction process with unfiltered reference samples. The residual is quantized and the difference between each quantized residual and its predictor, where the predictor refers to the previously coded residual at the horizontal or vertical (depending on the RDPCM prediction direction) neighboring position, is coded.
[0469] For a block of size M (height) × N (width), let r i,j , 0 ≤ i ≤ M - 1, 0 ≤ j ≤ N - 1 be the prediction residual. Let Q(r i,j ), 0 ≤ i ≤ M - 1, 0 ≤ j ≤ N - 1 denote the quantized version of the residual r i,j . RDPCM is applied to the quantized residual values to obtain a modified M×N array where is predicted from its neighboring quantized residual values. For the vertical RDPCM prediction mode, for 0 ≤ j ≤ (N - 1), the following is derived
[0470]
[0471] For the horizontal RDPCM prediction mode, for 0 ≤ i ≤ (M - 1), the following is derived
[0472]
[0473] On the decoder side, the above process is reversed as follows to compute Q(r i,j ), 0 ≤ i ≤ M - 1, 0 ≤ j ≤ N - 1:
[0474]
[0475] The inverse quantized residual Q -1 (Q(r i,j )) is added to the intra-block prediction value to produce the reconstructed sample value.
[0476] The predicted quantized residual values are coded using the same residual coding process as in transform skip mode residual coding Sent to the decoder. For the MPM mode used for future intra-mode coding and decoding, since the intra-mode coding and decoding of the luminance frame is not performed for the RDPCM-coded / decoded CU, the first MPM intra-mode is associated with the current CU and used for the intra-mode coding and decoding of the chrominance blocks of the current CU and subsequent CUs. For deblocking, if both of the two blocks on both sides of the block boundary use RDPCM coding and decoding, then deblocking is not performed for this specific block boundary.
[0477] 2.7. Combined Inter and Intra Prediction (CIIP)
[0478] In VTM5, when coding and decoding a CU in merge mode, if the CU contains at least 64 luminance samples (that is, the CU width multiplied by the CU height is equal to or greater than 64), and if both the CU width and the CU height are less than 128 luminance samples, then an additional flag is signaled to indicate whether the combined inter / intra prediction (CIIP) mode is applied to the current CU. As its name indicates, CIIP prediction combines the inter-prediction signal and the intra-prediction signal. The inter-prediction signal P in the CIIP mode inter is derived using the same inter-prediction process applied in the regular merge mode; and the intra-prediction signal P intra is derived following the regular intra-prediction process using the planar mode. Then, the intra-prediction signal and the inter-prediction signal are combined using weighted averaging, where the weight values are calculated as follows based on the coding and decoding modes of the top and left neighboring blocks ( Figure 30 as shown):
[0479] – If the top neighbor is available and is intra-coded, then set isIntraTop to 1, otherwise set isIntraTop to 0;
[0480] – If the left neighbor is available and is intra-coded, then set isIntraLeft to 1, otherwise set isIntraLeft to 0;
[0481] – If (isIntraTop + isIntraLeft) equals 2, then set wt to 3;
[0482] – Otherwise, if (isIntraTop + isIntraLeft) equals 1, then set wt to 2;
[0483] – Otherwise, set wt to 1.
[0484] The CIIP prediction is formed as follows:
[0485]
[0486] 3. Technical problems solved by the technical solutions provided in this document
[0487] The current VVC design may have the following problems:
[0488] 1. When considering the motion information in some coding and decoding tools (such as TPM, RDPCM, CIIP), the coding and decoding efficiency of screen content coding and decoding can be improved.
[0489] 2. The motion / block vector coding and decoding in IBC mode may not be efficient enough.
[0490] 4. Enumeration of technologies and embodiments
[0491] The following enumeration should be regarded as an example to explain the general concept. These technologies should not be interpreted narrowly. In addition, these inventions can be combined in any way.
[0492] In the following, "block vector" represents a motion vector pointing to a reference block within the same picture. "IBC" represents a coding and decoding method that uses the information of samples (filtered or unfiltered) within the same picture. For example, if a block is coded and decoded using two reference pictures (one of which is the current picture and the other is not), then in such a case, it can also be classified as being coded and decoded in IBC mode.
[0493] In the following, we represent MVx and MVy as the horizontal and vertical components of the motion vector (MV); MVDx and MVDy as the horizontal and vertical components of the motion vector difference (MVD); BVx and BVy as the horizontal and vertical components of the block vector (BV); BVPx and BVPy as the horizontal and vertical components of the block vector predictor (BVP); BVDx and BVDy as the horizontal and vertical components of the block vector difference (BVD). CBx and CBy represent the position of the block relative to the upper left position of the picture; W and H represent the width and height of the block. (x, y) represents a sample relative to the upper left position of the block. Floor(t) is the largest integer not greater than t.
[0494] 1. It is proposed to modify the block vector difference (for example, the difference between the block vector and the block vector predictor) for IBC coded blocks, and code and decode the modified BVD.
[0495] a. In one example, BVDy can be modified to (BVDy + BH), where BH is a constant.
[0496] i. In one example, BH can be 128.
[0497] ii. In one example, BH can be -128.
[0498] iii. In one example, BH can be the height of the CTU.
[0499] iv. In one example, BH can be the negative height of the CTU.
[0500] v. In one example, BH can depend on the block dimension and / or the block position and / or the IBC reference region size.
[0501] b. In one example, BVDx can be modified to (BVDx + BW), where BW is a constant.
[0502] i. In one example, BW can be 128.
[0503] ii. In one example, BW can be -128.
[0504] iii. In one example, BW can be 128 * 128 / (CTU / CTB size).
[0505] iv. In one example, BW can be the negative of 128 * 128 / (CTU / CTB size).
[0506] v. In one example, BW can depend on the block dimension and / or the block position and / or the IBC reference region size.
[0507] c. In one example, such a modification can be performed for a set of BVD values BVDS1 and not performed for a set of BVD values BVDS2, and the syntax can be signaled for all other BVD values to indicate whether such a modification is applied.
[0508] d. Whether and / or how to modify the block vector difference can depend on the sign and / or magnitude of the difference.
[0509] i. In one example, when BVDy > BH / 2, BVDy can be modified to (BVDy - BH).
[0510] ii. In one example, when BVDy < -Bh / 2, BVDy can be modified to (BVDy + BH).
[0511] iii. In one example, when BVDx > BW / 2, BVDx can be modified to (BVDx - BW).
[0512] iv. In one example, when BVDx < -BH / 2, BVDx can be modified to (BVDx + BW).
[0513] 2. It is proposed to modify the block vector predictor for the IBC encoded / decoded block and decode the block using the modified BV predictor.
[0514] a. In one example, the modified BV predictor can be used together with the signaling notified BVD to derive the final BV of the block (e.g., in the AMVP mode).
[0515] b. In one example, the modified BV predictor can be directly used as the final BV of the block (e.g., in the merge mode).
[0516] c. The modification of the BV predictor can be done in a similar manner as in bullet item 1 by replacing BVD with BVP.
[0517] d. Whether and / or how to modify the BVP can depend on the sign and / or magnitude of the BVP.
[0518] i. In one example, when BVPy = BVPx = 0, (BVPx, BVPy) can be modified to (-64, -64).
[0519] 3. A block vector for IBC modified decoding is proposed, and the modified block vector is used to identify reference blocks (e.g., for sample copy).
[0520] a. In one example, when using BVy, BVy can be modified to (BVy + BH).
[0521] i. In one example, BH can be 128.
[0522] ii. In one example, BH can be -128.
[0523] iii. In one example, BH can be the height of the CTU / CTB.
[0524] iv. In one example, BH can be the negative height of the CTU / CTB.
[0525] v. In one example, BH can depend on the block dimension and / or block position and / or IBC reference region size.
[0526] b. In one example, when using BVx, BVx can be modified to (BVx + BW).
[0527] i. In one example, BW can be 128.
[0528] ii. In one example, BW can be -128.
[0529] iii. In one example, BW can be 128 * 128 / (CTU / CTB size).
[0530] iv. In one example, BW can be negative 128 * 128 / (CTU / CTB size).
[0531] v. In one example, BW can depend on the block dimension and / or block position and / or IBC reference region size.
[0532] c. In one example, the motion vector after the above modification must be valid.
[0533] i. In one example, a valid motion vector can correspond to a predicted block that does not overlap with the current block.
[0534] ii. In one example, a valid motion vector can correspond to a predicted block within the current picture.
[0535] iii. In one example, a valid motion vector can correspond to a predicted block in which all samples have been reconstructed.
[0536] d. In one example, the modified block vector can be stored and used for motion prediction or / and deblocking.
[0537] e. Whether and / or how to modify the block vector difference can depend on the magnitude of the component.
[0538] i. In one example, when (CBy + BVy) / (CTU / CTB height) < CBy / (CTU / CTB height), BVy can be modified to (BVy + (CTU / CTB height)).
[0539] ii. In one example, when (CBy + W + BVy) / (CTU / CTB height) > CBy / (CTU / CTB height), BVy can be modified to (BVy - (CTU / CTB height)).
[0540] iii. In one example, when (BVx, BVy) is invalid, if the block vector (BVx - BW, BVy) is valid, then BVx can be modified to (BVx + BW), where BVy can be the unmodified or modified vertical component.
[0541] iv. In one example, when (BVx, BVy) is invalid, if the block vector (BVx + BW, BVy) is valid, then BVx can be modified to (BVx + BW), where BVy can be the unmodified or modified vertical component.
[0542] 4. The modifications for different components can follow a certain order.
[0543] a. In one example, the modification of BVDx can be before the modification of BVDy.
[0544] b. In one example, BVy can be modified before modifying BVx.
[0545] 5. A motion vector predictor can be formed based on the modified BV / BVD.
[0546] a. In one example, HMVP can be based on the modified BV / BVD.
[0547] b. In one example, the merge candidate can be based on the modified BV / BVD.
[0548] c. In one example, AMVP can be based on the modified BV / BVD.
[0549] 6. The above constant values BW and BH can be predefined or can be signaled within the SPS / PPS / strip / group of pictures / slice / tile / CTU level.
[0550] 7. The above method can be disabled when the current block is on the picture / strip / group of pictures boundary.
[0551] a. In one example, when the top - left vertical position of the current block relative to the picture / strip / group of pictures is 0, BH can be set to 0.
[0552] b. In one example, when the top - left horizontal position of the current block relative to the picture / strip / group of pictures is 0, BW can be set to 0.
[0553] 8. The weighting factor (denoted by wt) applied to the intra - prediction signal in CIIP can depend on the motion vector.
[0554] a. In one example, when MVx is equal to 0 or MVy is equal to 0, wt for intra - prediction can be set to 1.
[0555] b. In one example, when the MV points to an integer pixel or has integer precision, wt for intra - prediction can be set to 1.
[0556] 9. The weights used in TPM can depend on the segmented motion information.
[0557] a. In one example, it can depend on whether one or more components (horizontal or / and vertical) of the MV associated with a segmentation are equal to 0.
[0558] b. In one example, when for the upper - right segmentation in TPM (i.e., Figure 31For the segmentation 1 shown in the left figure of , when MVx is equal to 0 or MVy is equal to 0, and W >= H, the weights of the samples that satisfy floor(x * H / W) <= y can be set to 1.
[0559] i. Alternatively, for other samples, the weights can be set to 0.
[0560] c. In one example, when for the upper - right segmentation in the TPM (i.e., Figure 31 For the segmentation 1 shown in the left figure of , when MVx is equal to 0 or MVy is equal to 0, and W < H, the weights of the samples that satisfy x <= floor(y * W / H) can be set to 1.
[0561] i. Alternatively, for other samples, the weights can be set to 0.
[0562] d. In one example, when for the lower - left segmentation in the TPM (i.e., Figure 31 For the segmentation 2 shown in the left figure of , when MVx is equal to 0 or MVy is equal to 0, and W >= H, the weights of the samples that satisfy floor(x * H / W) >= y can be set to 1.
[0563] i. Alternatively, for other samples, the weights can be set to 0.
[0564] e. In one example, when for the lower - left segmentation in the TPM (i.e., Figure 31 For the segmentation 2 shown in the left figure of , when MVx is equal to 0 or MVy is equal to 0, and W < H, the weights of the samples that satisfy x >= floor(y * W / H) can be set to 1.
[0565] i. Alternatively, for other samples, the weights can be set to 0.
[0566] f. In one example, when for the upper - left segmentation in the TPM (i.e., Figure 31 For the segmentation 1 shown in the right figure of , when MVx is equal to 0 or MVy is equal to 0, and W >= H, the weights of the samples that satisfy floor((W - 1 - x) * H / W) >= y can be set to 1.
[0567] i. Alternatively, for other samples, the weights can be set to 0.
[0568] g. In one example, when for the upper - left segmentation in the TPM (i.e., Figure 31 For the segmentation 1 shown in the right figure of , when MVx is equal to 0 or MVy is equal to 0, and W < H, the weights of the samples that satisfy (W - 1 - x) >= floor(y * W / H) can be set to 1.
[0569] i. Alternatively, for other samples, the weight can be set to 0.
[0570] h. In one example, when MVx is equal to 0 or MVy is equal to 0 for the lower-right split in the TPM (i.e., split 2 shown in the right figure of Figure 31 ), and W >= H, the weight of the samples that satisfy floor((W - 1 - x) * H / W) <= y can be set to 1.
[0571] i. Alternatively, for other samples, the weight can be set to 0.
[0572] i. In one example, when MVx is equal to 0 or MVy is equal to 0 for the lower-right split in the TPM (i.e., Figure 31 split 2 shown in the right figure of), and W < H, the weight of the samples that satisfy (W - 1 - x) <= floor(y * W / H) can be set to 1.
[0573] i. Alternatively, for other samples, the weight can be set to 0.
[0574] 10. This mixing process in the TPM can depend on the transform information.
[0575] a. In one example, when the transform skip mode is selected, this mixing process may not be applied.
[0576] b. In one example, when the block does not have non-zero residuals, this mixing process may not be applied.
[0577] c. In one example, when a transform is selected, this mixing process may be applied.
[0578] d. In one example, when this mixing process is not applied, only the motion information of split X can be used to generate the predicted samples of the split, where X can be 0 or 1.
[0579] 11. Whether to enable the above method can depend on
[0580] a. The characteristics of the encoded / decoded content
[0581] b. The content type (e.g., screen content, camera captured content)
[0582] c. Flags at the SPS / PPS / strip / group of pictures / picture / tile / CTU / other video unit level
[0583] d. The encoding / decoding information of a block
[0584] Figure 32It is a block diagram of a video processing device 3200. The device 3200 can be used to implement one or more of the methods described herein. The device 3200 can be embodied in a smart phone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The device 3200 can include one or more processors 3202, one or more memories 3204, and video processing hardware 3206. The (one or more) processors 3202 can be configured to implement one or more methods described in this document. The (one or more) memories 3204 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 3206 can be used to implement some of the techniques described in this document in hardware circuits.
[0585] In some embodiments, devices implemented on the hardware platforms described in connection Figure 32 can be used to implement these video coding and decoding methods.
[0586] Figure 33 It is a flowchart of an exemplary method 3300 for video processing. The method includes: during the conversion between a current video block of a video picture and the bitstream representation of the current video block, determining (3302) a block vector difference (BVD) representing the difference between the block vector corresponding to the current video block and its predictor, where the block vector indicates a motion match for the current video block in the video picture, and where the modified value of the BVD is encoded and decoded into the bitstream representation; and performing (3304) the conversion between the current video block and the bitstream representation of the current video block using the block vector. The predictor of the BVD can be calculated based on one or more previous block vectors used for the conversion of previous video blocks during the conversion.
[0587] The various solutions and embodiments described in this document will be further described with an enumeration of solutions.
[0588] The following solutions are examples of the techniques listed in item 1 of chapter 4.
[0589] 1. A video processing method, including: during the conversion between a current video block of a video picture and the bitstream representation of the current video block, determining a block vector difference (BVD) representing the difference between the block vector corresponding to the current video block and its predictor, where the block vector indicates a motion match for the current video block in the video picture, and where the modified value of the BVD is encoded and decoded into the bitstream representation; and performing the conversion between the current video block and the bitstream representation of the current video block using the block vector.
[0590] 2. The method according to solution 1, wherein the modified value of the BVD includes the result obtained by adding an offset BH to the y-component BVDy of the BVD.
[0591] 3. The method according to Solution 1, wherein the modified value of the BVD includes the result obtained by adding the offset BW to the x-component BVDx of the BVD.
[0592] 4. The method according to Solution 1, wherein the modified value of the BVD is obtained by checking whether the BVD comes from a set of BVD values.
[0593] 5. The method according to Solution 1, wherein the modified value of the BVD is obtained by using a modification technique depending on the value or magnitude of the BVD.
[0594] The following solutions are examples of the techniques listed in Item 2 of Section 4.
[0595] 1. A video processing method, comprising: determining to use an intra-block copy tool for the conversion during the conversion between a current video block of a video picture and a bitstream representation of the current video block; and performing the conversion using a modified block vector predictor corresponding to a modified value of a block vector difference (BVD) of the current video block.
[0596] 2. The method according to Solution 1, wherein the conversion further uses the value of the BVD corresponding to a field in the bitstream representation.
[0597] 3. The method according to any one of Solutions 1-2, wherein the conversion uses a block vector derived from the modified value and the BVD.
[0598] The following solutions are examples of the techniques listed in Item 3 of Section 4.
[0599] 1. A video processing method, comprising: performing a conversion between a current video block of a video picture and a bitstream representation of the current video block using an intra-block copy codec tool, wherein a modified block vector corresponding to a modified value of a block vector is used for the conversion and is to be included in the bitstream representation.
[0600] 2. The method according to Solution 1, wherein the conversion includes performing sample copying by another region referred to by the video picture through the modified block vector.
[0601] 3. The method according to any one of Solutions 1-2, wherein the modified value of the block vector is obtained by adding an offset BH to the y-component BVy of the block vector.
[0602] 4. The method according to any one of Solutions 1-2, wherein the modified value of the block vector is obtained by adding an offset BW to the x-component BVx of the block vector.
[0603] 5. The method according to any one of Solutions 1-4, wherein a further validity check is performed on the modified value.
[0604] The following solutions are examples of the techniques listed in Item 4 of Chapter 4.
[0605] 1. The method according to the above solutions, wherein the modification is performed sequentially.
[0606] 2. The method according to Solution 1, wherein the sequence includes first modifying the horizontal component of the block vector or the block vector difference, and then modifying its vertical component.
[0607] The following solutions are examples of the techniques listed in Items 5 and 6 of Chapter 4.
[0608] 1. The method according to any one of the above solutions, wherein a motion vector predictor is further formed using the modified block vector or the modified block vector difference.
[0609] 2. The method according to any one of the above solutions, wherein the modification uses an offset value signaled in the bitstream representation.
[0610] 3. The method according to Solution 2, wherein the offset value is signaled at the sequence parameter set level or the picture parameter set level or the slice level or the slice group level or the picture level or the tile level or the coding tree unit level.
[0611] The following solutions are examples of the techniques listed in Item 7 of Chapter 4.
[0612] 1. The method according to any one of the above solutions, wherein the modification is performed because the current video block is within the allowed region of the video picture.
[0613] 2. The method according to Solution 1, wherein the allowed region excludes the boundaries of the video picture or the video slice or the slice group.
[0614] The following solutions are examples of the techniques listed in Item 8 of Chapter 4.
[0615] 1. A video processing method, comprising: determining a weighting factor wt based on the conditions of the current video block during the conversion between the current video block and its bitstream representation; and performing the conversion using a combined intra-inter coding operation, wherein the weighting factor wt is used to weight the motion vector of the current video block.
[0616] 2. The method according to Solution 1, wherein the condition depends on whether the x or y component of the motion vector is zero.
[0617] 3. The method according to any one of Solutions 1-2, wherein the condition is based on the pixel resolution of the motion vector.
[0618] The following solutions are examples of the techniques listed in Items 9 and 10 of Chapter 4.
[0619] 1. A video processing method, comprising: determining to use a triangle partitioning mode (TPM) codec tool for conversion between a current video block and a bitstream representation of the current video block, wherein at least one operation parameter of the TPM codec tool depends on characteristics of the current video block, wherein the TPM codec tool partitions the current video block into two separately coded non-rectangular partitions; and performing the conversion by applying the TPM codec tool, the TPM codec tool using the said one operation parameter.
[0620] 2. The method according to Solution 1, wherein the operation parameter includes weights applied to the two non-rectangular partitions.
[0621] 3. The method according to any one of Solutions 1-2, wherein the characteristics of the current video correspond to whether one or both of the motion components of the current video block are zero.
[0622] 4. The method according to any one of Solutions 1-3, wherein the operation parameter includes the applicability of a mixing process used for the current video block during the conversion.
[0623] 5. The method according to Solution 4, wherein the mixing process is disabled when a transform skip mode is selected for the current video block.
[0624] The following solutions are examples of the techniques listed in Item 11 of Chapter 4.
[0625] 1. The method according to any one of the above solutions, wherein characteristics of the coded content of the current video block are used in the determination operation of the method.
[0626] 2. The method according to any one of the above solutions, wherein the content type of the current video block is used in the determination operation of the method.
[0627] Additional solutions include:
[0628] 1. A video encoding apparatus, comprising a processor configured to implement the method according to any one or more of the above solutions.
[0629] 2. A video decoding apparatus includes a processor configured to implement the method described in any one or more of the above solutions.
[0630] 3. A machine-readable medium having stored thereon code that implements one or more of the above methods by a processor.
[0631] Figure 34 is a flowchart of an exemplary method 3400 for video processing. The method includes: modifying (3402) at least one of the motion information associated with a block encoded or decoded in an intra block copy (IBC) mode for a conversion between the block of the video and a bitstream representation of the block; and performing (3404) the conversion based on the modified motion information.
[0632] In some examples, the motion information includes a block vector difference (BVD) representing a difference between a motion vector of the block and a motion vector predictor of the block, and the modified BVD is encoded or decoded into the bitstream representation.
[0633] In some examples, a vertical component BVDy of the BVD is modified to (BVDy + BH), where BH is a constant.
[0634] In some examples, BH is 128 or -128.
[0635] In some examples, BH is a height of a coding tree unit (CTU) of the block or a negative height of the CTU.
[0636] In some examples, BH depends on a block dimension and / or a block position and / or an IBC reference region size of the block.
[0637] In some examples, a horizontal component BVDx of the BVD is modified to (BVDy + BW), where BW is a constant.
[0638] In some examples, BW is 128 or -128.
[0639] In some examples, BW is 128 * 128 / S, where S is (a coding tree unit (CTU) size or a coding tree block (CTB) size), or BW is negative 128 * 128 / S (CTU size or CTB size).
[0640] In some examples, BW depends on a block dimension and / or a block position and / or an IBC reference region size of the block.
[0641] In some examples, the modification is performed for a set of BVD values BVDS1, and not performed for a set of BVD values BVDS2, and for all other BVD values the syntax is signaled to indicate whether the modification is applied.
[0642] In some examples, whether and / or how the BVP is modified depends on the sign and / or magnitude of the difference.
[0643] In some examples, when BVDy > BH / 2, BVDy is modified to (BVDy - BH).
[0644] In some examples, when BVDy < -Bh / 2, BVDy is modified to (BVDy + BH).
[0645] In some examples, when BVDx > BW / 2, BVDx is modified to (BVDx - BW).
[0646] In some examples, when BVDx < -BH / 2, BVDx is modified to (BVDx + BW).
[0647] In some examples, the motion information includes a block vector predictor (BVP) for the block, and the block is decoded using the modified BVP.
[0648] In some examples, in the advanced motion vector prediction (AMVP) mode, the modified BVP is used together with the signaled block vector difference (BVD) to derive the final block vector (BV) of the block, where BVD represents the difference between the BV of the block and the BVP of the block.
[0649] In some examples, in the merge mode, the modified BVP is directly used as the final BV of the block.
[0650] In some examples, the vertical component BVPy of the BVP is modified to (BVPy + BH), where BH is a constant.
[0651] In some examples, BH is 128 or -128.
[0652] In some examples, BH is the height of the coding tree unit (CTU) of the block or the negative of the CTU height.
[0653] In some examples, BH depends on the block dimension and / or block position and / or IBC reference region size of the block.
[0654] In some examples, the horizontal component BVPx of the BVP is modified to (BVPy + BW), where BW is a constant.
[0655] In some examples, BW is 128 or -128.
[0656] In some examples, BW is 128 * 128 / S, where S is the Coding Tree Unit (CTU) size or Coding Tree Block (CTB) size, or BW is negative 128 * 128 / S.
[0657] In some examples, BW depends on the block dimension and / or block position and / or IBC reference region size of the block.
[0658] In some examples, the modification is performed for a set of BVP values BVPS1, not performed for a set of BVP values BVPS2, and the syntax is signaled for all other BVP values to indicate whether the modification is applied.
[0659] In some examples, whether and / or how to modify BVP depends on the sign and / or magnitude of the BVP.
[0660] In some examples, when BVPy = BVPx = 0, (BVPx, BVPy) is modified to (-64, -64).
[0661] In some examples, the motion information includes a block vector (BV) for decoding in the IBC mode, and the reference block for sample copy is identified using the modified BV.
[0662] In some examples, when using the vertical component BVy of the BV, BVy is modified to (BVy + BH), where BH is a constant.
[0663] In some examples, BH is 128 or -128.
[0664] In some examples, BH is the height of the Coding Tree Unit (CTU) or Coding Tree Block (CTB) of the block or the negative height of the CTU or CTB.
[0665] In some examples, BH depends on the block dimension and / or block position and / or IBC reference region size of the block.
[0666] In some examples, when using the horizontal component BVx of the BV, BVx is modified to (BVy + BW), where BW is a constant.
[0667] In some examples, BW is 128 or -128.
[0668] In some examples, BW is 128 * 128 / (Coding Tree Unit (CTU) size or Coding Tree Block (CTB) size) or negative 128 * 128 / (CTU size or CTB size).
[0669] In some examples, the BW depends on the block dimension of the block and / or the block position and / or the IBC reference region size.
[0670] In some examples, the modified motion vector must be valid after the modification.
[0671] In some examples, the valid motion vector corresponds to a prediction block that does not overlap with the block.
[0672] In some examples, the valid motion vector corresponds to a prediction block within the current picture.
[0673] In some examples, the valid motion vector corresponds to a prediction block in which all samples have been reconstructed.
[0674] In some examples, the modified BV is stored and used for motion prediction or / and deblocking.
[0675] In some examples, whether and / or how to modify the BV depends on the magnitude of at least one component of the BV.
[0676] In some examples, when (CBy + BVy) / (CTU or CTB height) < CBy / (CTU or CTB height), BVy is modified to (BVy + (CTU or CTB height)), where CBx and CBy represent the position of the block relative to the upper left position of the picture.
[0677] In some examples, when (CBy + W + BVy) / (CTU / CTB height) > CBy / (CTU / CTB height), BVy is modified to (BVy - (CTU / CTB height)), where CBx and CBy represent the position of the block relative to the upper left position of the picture, and W and H represent the width and height of the block.
[0678] In some examples, when (BVx, BVy) is invalid, if the block vector (BVx - BW, BVy) is valid, then BVx is modified to (BVx + BW), where BVy is the unmodified or modified vertical component.
[0679] In some examples, when (BVx, BVy) is invalid, if the block vector (BVx + BW, BVy) is valid, then BVx is modified to (BVx + BW), where BVy is the unmodified or modified vertical component.
[0680] In some examples, the modifications for different components follow a predetermined order.
[0681] In some examples, the modification of BVDx is before the modification of BVDy.
[0682] In some examples, the modification of BVy is before the modification of BVx.
[0683] In some examples, the motion vector predictor associated with the block is formed based on a modified BV or BVD.
[0684] In some examples, history-based motion vector prediction (HMVP) is based on a modified BV or BVD.
[0685] In one example, a merge candidate is based on a modified BV or BVD.
[0686] In some examples, advanced motion vector prediction (AMVP) is based on a modified BV or BVD.
[0687] In some examples, the constant BW and BH are predefined or signaled at at least one of the SPS, PPS, slice, slice group, picture, tile, and CTU levels.
[0688] In some examples, the modification is disabled when the block is on at least one of a picture boundary, a slice boundary, and a slice group boundary.
[0689] In some examples, BH is set to 0 when the vertical position of the block relative to the top-leftmost position of at least one of the picture, slice, or slice group is 0.
[0690] In some examples, BW is set to 0 when the horizontal position of the block relative to the top-leftmost position of at least one of the picture, slice, or slice group is 0.
[0691] (7.b)
[0692] Figure 35 is a flowchart of an exemplary method 3500 of video processing. The method includes: determining (3502), for a block of a video and a conversion between the block and a bitstream representation of the block, a weighting factor for an intra prediction signal in a combined intra-inter prediction (CIIP) mode based on a motion vector (MV) associated with the block, and wherein a modified value of the BVD is encoded and decoded into the bitstream representation; and performing (3504) the conversion based on the weighting factor.
[0693] In some examples, the weighting factor for the intra prediction signal is set to 1 when the horizontal component of the MV is equal to 0 or the vertical component of the MV is equal to 0.
[0694] In some examples, the weighting factor for the intra prediction signal is set to 1 when the MV points to an integer pixel or has integer precision.
[0695] Figure 36It is a flowchart of an exemplary method 3600 for video processing. The method includes: for the conversion between a block of a video and the bitstream representation of the block, determining (3602) weights used in a triangular prediction mode (TPM) based on the motion information of one or more partitions of the block; and performing (3604) the conversion based on the weights.
[0696] In some examples, the block is partitioned into an upper-right partition and a lower-left partition by using an anti-diagonal partition or into an upper-left partition and a lower-right partition by using a diagonal partition, and the block has a width W and a height H.
[0697] In some examples, the weights are determined based on whether one or more components of a motion vector (MV) associated with a partition are equal to 0, where the components of the MV include a horizontal component MVx and a vertical component MVy.
[0698] In some examples, when for the upper-right partition in the TPM mode MVx is equal to 0 or MVy is equal to 0, and W >= H, for a first sample point satisfying floor(x * H / W) <= y, the weight is set to 1.
[0699] In some examples, for sample points other than the first sample point, the weight is set to 0.
[0700] In some examples, when for the upper-right partition in the TPM mode MVx is equal to 0 or MVy is equal to 0, and W < H, for a first sample point satisfying x <= floor(y * W / H), the weight is set to 1.
[0701] In some examples, for sample points other than the first sample point, the weight is set to 0.
[0702] In some examples, when for the lower-left partition in the TPM mode MVx is equal to 0 or MVy is equal to 0, and W >= H, for a first sample point satisfying floor(x * H / W) <= y, the weight is set to 1.
[0703] In some examples, for sample points other than the first sample point, the weight is set to 0.
[0704] In some examples, when for the lower-left partition in the TPM mode MVx is equal to 0 or MVy is equal to 0, and W < H, for a first sample point satisfying x <= floor(y * W / H), the weight is set to 1.
[0705] In some examples, for sample points other than the first sample point, the weight is set to 0.
[0706] In some examples, when MVx equals 0 or MVy equals 0 for the upper-left split in the TPM mode, and W >= H, for the first sample point that satisfies floor((W - 1 - x) * H / W) >= y, the weight is set to 1.
[0707] In some examples, for sample points other than the first sample point, the weight is set to 0.
[0708] In some examples, when MVx equals 0 or MVy equals 0 for the upper-left split in the TPM mode, and W < H, for the first sample point that satisfies (W - 1 - x) >= floor(y * W / H), the weight is set to 1.
[0709] In some examples, for sample points other than the first sample point, the weight is set to 0.
[0710] (9.g.i)
[0711] In some examples, when MVx equals 0 or MVy equals 0 for the lower-right split in the TPM mode, and W >= H, for the first sample point that satisfies floor((W - 1 - x) * H / W) <= y, the weight is set to 1.
[0712] In some examples, for sample points other than the first sample point, the weight is set to 0.
[0713] In some examples, when MVx equals 0 or MVy equals 0 for the lower-right split in the TPM mode, and W < H, for the first sample point that satisfies (W - 1 - x) <= floor(y * W / H), the weight is set to 1.
[0714] In some examples, for sample points other than the first sample point, the weight is set to 0.
[0715] Figure 37 is a flowchart of an exemplary method 3700 for video processing. The method includes: determining (3702) whether to apply a hybrid process in a triangular prediction mode (TPM) mode based on transform information of a block for the conversion between the block of the video and the bitstream representation of the block; and performing (3704) the conversion based on the determination.
[0716] In some examples, when the transform skip mode is selected, the hybrid process is not applied.
[0717] In some examples, when the block does not have non-zero residuals, the hybrid process is not applied.
[0718] In some examples, when a transform is selected, the hybrid process is applied.
[0719] In some examples, when the mixing process is not applied, the motion information of the split X of the block is used only to generate the predicted sample points of the split X, where X is 0 or 1.
[0720] In some examples, whether to enable the modification or the determination depends on at least one of the following:
[0721] a. The characteristics of the encoded / decoded content;
[0722] b. The content type including screen content and camera-captured content;
[0723] c. Flags at at least one level among SPS, PPS, slice group, slice, tile, CTU, other video unit levels; and
[0724] d. The encoding / decoding information of the block.
[0725] In some examples, the transformation generates a block of video from a bitstream representation.
[0726] In some examples, the transformation generates a bitstream representation from a block of video.
[0727] From the foregoing, it will be appreciated that specific embodiments of the techniques disclosed herein have been described for purposes of illustration, but various modifications may be made without departing from the scope of the invention. Accordingly, the techniques of this disclosure are subject to no other limitations than those imposed by the appended claims.
[0728] The embodiments of the subject matter and the functional operations described in this patent document may be implemented in various systems, digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or combinations of one or more of them. The embodiments of the subject matter described in this specification may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible and non-transitory computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a storage device, a composition of matter affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing unit" or "data processing apparatus" encompasses all apparatuses, devices, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer programs under consideration, e.g., code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0729] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one or more computers, which can be located at one site or distributed across multiple sites and interconnected by a communication network.
[0730] The processes and logical flows described in this specification can be performed by one or more programmable processors executing one or more computer programs, thereby performing functions by operating on input data and generating output. These processes and logical flows can also be performed by special-purpose logic circuitry, and the apparatus can also be implemented as special-purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[0731] For example, processors suitable for executing a computer program include general and special-purpose microprocessors, and any one or more processors of any kind of digital computer. In general, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Usually, a computer will also include one or more mass storage devices for storing data, e.g., magnetic disks, magneto-optical disks, or optical disks, or be operatively coupled to receive data from or transfer data to one or more mass storage devices, or both. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including (e.g.) semiconductor storage devices such as EPROMs, EEPROMs, and flash memory devices. The processor and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry.
[0732] The specification, together with the drawings, is intended to be exemplary only, where exemplary means an example. As used herein, the use of "or" is intended to include "and / or" unless the context clearly dictates otherwise.
[0733] Although this patent document contains many details, it should not be construed as limiting any invention or the scope of any claims, but rather as a description of specific features of particular embodiments of a specific invention. Certain features described in the context of individual embodiments of this patent document may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. In addition, although certain features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination may in some cases be excluded from that combination, and the claimed combination may cover a sub-combination or a variation of a sub-combination.
[0734] Similarly, although the operations are described in a particular order in the drawings, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order to obtain the desired result, or that all illustrated operations must be performed. In addition, the partitioning of various system components in the embodiments described in this patent document should not be construed as required in all embodiments.
[0735] Only a few embodiments and examples are described, and other embodiments, enhancements, and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A method for video processing, comprising: modifying at least one of the motion information associated with a block encoded and decoded in an Intra Block Copy (IBC) mode for the conversion between a block of a video and the bitstream of the block; and performing the conversion based on the modified motion information, wherein the motion information includes a Block Vector Difference (BVD) representing the difference between the motion vector of the block and the motion vector predictor of the block, and encoding and decoding the modified BVD into the bitstream, wherein the vertical component BVDy of the BVD is modified to (BVDy + BH), where BH is a constant.
2. The method according to claim 1, wherein, BH is 128 or -128.
3. The method according to claim 1, wherein, BH is the height of the Coding Tree Unit (CTU) of the block or the negative of the CTU height.
4. The method according to claim 1, wherein, BH depends on the block dimension and / or block position and / or IBC reference region size of the block.
5. The method according to claim 1, wherein, The horizontal component BVDx of the BVD is modified to (BVDy + BW), where BW is a constant.
6. The method according to claim 5, wherein, BW is 128 or -128.
7. The method according to claim 5, wherein BW is 128 * 128 divided by S or BW is negative 128 * 128 divided by S, where S is the Coding Tree Unit (CTU) size or Coding Tree Block (CTB) size.
8. The method according to claim 5, wherein BW depends on the block dimension and / or block position and / or IBC reference region size of the block.
9. The method according to any one of claims 1-8, wherein, Performing the modification for a certain set of BVD values BVDS1, not performing the modification for a certain set of BVD values BVDS2, and signaling a syntax for all other BVD values to indicate whether to apply the modification.
10. The method according to any one of claims 1-8, wherein Whether and / or how to modify the BVD depends on the sign and / or magnitude of the difference.
11. The method according to claim 10, wherein, When BVDy > BH / 2, modifying BVDy to (BVDy - BH).
12. The method according to claim 10, wherein, When BVDy < -Bh / 2, modifying BVDy to (BVDy + BH).
13. The method according to claim 10, wherein When BVDx > BW / 2, modifying BVDx to (BVDx - BW).
14. The method according to claim 10, wherein, When BVDx < -BH / 2, modifying BVDx to (BVDx + BW).
15. The method according to claim 1, wherein, The motion information further includes a Block Vector Predictor (BVP) for the block, and decoding the block using the modified BVP.
16. The method according to claim 15, wherein, In the Advanced Motion Vector Prediction (AMVP) mode, using the modified BVP together with the signaled Block Vector Difference (BVD) to derive the final Block Vector (BV) of the block, where BVD represents the difference between the BV of the block and the BVP of the block.
17. The method according to claim 15, wherein, In the merge mode, directly using the modified BVP as the final BV of the block.
18. The method according to any one of claims 15-17, wherein, Modifying the vertical component BVPy of the BVP to (BVPy + BH), where BH is a constant.
19. The method according to claim 18, wherein BH is 128 or -128.
20. The method according to claim 18, wherein BH is the height of the Coding Tree Unit (CTU) of the block or the negative of the CTU height.
21. The method according to claim 18, wherein BH depends on the block dimension and / or block position and / or IBC reference region size of the block.
22. The method according to any one of claims 15 - 17, wherein Modifying the horizontal component BVPx of the BVP to (BVPy + BW), where BW is a constant.
23. The method according to claim 22, wherein, BW is 128 or -128.
24. The method according to claim 22, wherein, BW is 128*128 / S or -128*128 / S, where S is the Coding Tree Unit (CTU) size or Coding Tree Block (CTB) size.
25. The method according to claim 22, wherein, BW depends on the block dimension and / or block position and / or Inter Block Copy (IBC) reference region size of the said block.
26. The method according to any one of claims 15 - 17, wherein The modification is performed for a certain set of BVP values BVPS1, not performed for a certain set of BVP values BVPS2, and for all other BVP values, the syntax is signaled to indicate whether the modification is applied.
27. The method according to any one of claims 15 - 17, wherein, Whether and / or how to modify BVP depends on the sign and / or magnitude of the said BVP.
28. The method according to claim 27, wherein When BVPy = BVPx = 0, (BVPx, BVPy) is modified to (-64, -64).
29. The method according to claim 1, wherein The motion information also includes a Block Vector (BV) for decoding in IBC mode, and the reference block for sample copy is identified using the modified BV.
30. The method according to claim 29, wherein, When using the vertical component BVy of the said BV, BVy is modified to (BVy + BH), where BH is a constant.
31. The method according to claim 30, wherein, BH is 128 or -128.
32. The method according to claim 30, wherein BH is the height of the Coding Tree Unit (CTU) of the said block, BH is the height of the Coding Tree Block (CTB) of the said block, BH is the negative CTU height of the said block, or BH is the negative CTB height of the said block.
33. The method according to claim 30, wherein, BH depends on the block dimension and / or block position and / or IBC reference region size of the said block.
34. The method according to claim 29, wherein, When using the horizontal component BVx of the said BV, BVx is modified to (BVy + BW), where BW is a constant.
35. The method according to claim 34, wherein, BW is 128 or -128.
36. The method according to claim 34, wherein, BW is 128*128 divided by the Coding Tree Unit (CTU) size, BW is 128*128 divided by the Coding Tree Block (CTB) size, BW is negative 128*128 divided by the CTU size, or BW is negative 128*128 divided by the CTB size.
37. The method according to claim 34, wherein, BW depends on the block dimension and / or block position and / or IBC reference region size of the said block.
38. The method according to any one of claims 29-37, wherein, The motion vector after modification must be valid.
39. The method according to claim 38, wherein, A valid motion vector corresponds to a prediction block that does not overlap with the said block.
40. The method according to claim 38, wherein, A valid motion vector corresponds to a prediction block within the current picture.
41. The method according to claim 38, wherein, A valid motion vector corresponds to a prediction block in which all samples have been reconstructed.
42. The method according to any one of claims 29 - 32, wherein, The modified BV is stored and used for motion prediction and / or deblocking.
43. The method according to any one of claims 29 - 37, wherein, Whether and / or how to modify BV depends on the magnitude of at least one component of the said BV.
44. The method according to claim 43, wherein, When (CBy + BVy) divided by the CTU height < CBy divided by the CTU height, BVy is modified to (BVy + CTU height), or when (CBy + BVy) divided by the CTB height < CBy divided by the CTB height, BVy is modified to (BVy + CTB height), where CBx and CBy represent the position of the block relative to the upper left position of the picture.
45. The method according to claim 43, wherein When (CBy + W + BVy) divided by the CTU height > CBy divided by the CTU height, modify BVy to (BVy - CTU height), or when (CBy + W + BVy) divided by the CTB height > CBy divided by the CTB height, modify BVy to (BVy - CTB height), where CBx and CBy represent the position of the block relative to the upper left position of the picture, and W and H represent the width and height of the block.
46. The method according to claim 43, wherein, When (BVx, BVy) is invalid, if the block vector (BVx - BW, BVy) is valid, then modify BVx to (BVx + BW), where BVy is the unmodified or modified vertical component.
47. The method according to claim 43, wherein When (BVx, BVy) is invalid, if the block vector (BVx + BW, BVy) is valid, then modify BVx to (BVx + BW), where BVy is the unmodified or modified vertical component.
48. The method according to claim 1, wherein The modifications for different components follow a predetermined order.
49. The method according to claim 48, wherein The modification of BVDx is before the modification of BVDy.
50. The method according to claim 48, wherein, The modification of BVy is before the modification of BVx.
51. The method according to claim 1, wherein, The motion vector predictor associated with the block is formed based on the modified BV or BVD.
52. The method according to claim 51, wherein, The history-based motion vector prediction HMVP is based on the modified BV or BVD.
53. The method according to claim 51, wherein, The merge candidate is based on the modified BV or BVD.
54. The method according to claim 51, wherein, The advanced motion vector prediction AMVP is based on the modified BV or BVD.
55. The method according to claim 1, wherein The constants BW and BH are predefined or signaled at least at one level among the SPS, PPS, slice, slice group, slice, tile, and CTU levels.
56. The method according to claim 1, wherein, When the block is on at least one of the picture boundary, slice boundary, and slice group boundary, disable the modification.
57. The method according to claim 56, wherein, When the vertical position of the block relative to the upper left position of at least one of the picture, slice, or slice group is 0, set BH to 0.
58. The method according to claim 56, wherein, When the horizontal position of the block relative to the upper left position of at least one of the picture, slice, or slice group is 0, set BW to 0.
59. The method according to claim 1, wherein The transformation generates the block from the bitstream.
60. The method according to claim 1, wherein, The transformation generates the bitstream from the block.
61. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions stored thereon, wherein, The instruction, when executed by the processor, causes the processor to implement the method according to any one of claims 1 to 60.
62. A computer program product stored on a non-transitory computer-readable medium, the computer program product including program code for implementing the method according to any one of claims 1 to 60.