Construction of Motion Candidate List Using Adjacent Block Information
The visual media processing method addresses inefficiencies in current video coding standards by employing non-rectangular partitioning and rules-based availability determination for adjacent blocks, resulting in improved coding efficiency and video quality.
Patent Information
- Application Number
- JP2023144057
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-09-07
- Filing Date
- 2023-09-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2040-06-04
AI Technical Summary
Current video coding standards face challenges in efficiently managing motion information and partitioning modes, particularly in handling non-rectangular partitions like triangular partitioning, which can lead to increased computational complexity and reduced coding efficiency.
The proposed solution involves a visual media processing method that utilizes non-rectangular partitioning, such as triangular partitioning, in video encoding and decoding. This method determines the availability of adjacent video blocks using specific rules and constructs a merge list based on these determinations, excluding motion information from unavailable blocks.
The approach enhances coding efficiency by optimizing motion information usage and partitioning modes, thereby improving the quality of decompressed video while reducing computational complexity.
Smart Images

Figure 0007687763000157 
Figure 0007687763000158 
Figure 0007687763000159
Abstract
Description
Background Art
[0001] Cross - reference to related applications This application is a divisional application of Japanese Patent Application No. 2021 - 571828, and the original application is based on International Patent Application No. PCT / CN2020 / 094310 filed on June 4, 2020. This application claims priority and the benefit thereof with respect to International Patent Application No. PCT / CN2019 / 089970 filed on June 4, 2019 and International Patent Application No. PCT / CN2019 / 104810 filed on September 7, 2019. All patent applications of the above - mentioned applications are hereby incorporated by reference in their entirety into this case.
[0002] Technical Field This specification is related to video and image encoding and decoding technologies.
[0003] Background Digital video occupies the largest bandwidth in the Internet and other digital communication networks. It is expected that as the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for the use of digital video will continue to increase.
Summary of the Invention
[0004] The disclosed technology may be used by embodiments of a video or image decoder or encoder that perform encoding or decoding of a video bitstream using non - rectangular partitioning such as triangular partitioning mode.
[0005] In one exemplary embodiment, a visual media processing method is disclosed. The method relates to the conversion between a first video block of visual media data and a bitstream representation of the visual media data, and includes a step of determining the availability of a second video block of the visual media data using rules, and a step of performing the conversion based on the determination. The rules are at least based on a coding mode used to code the first video block into a bitstream representation, and the rules define that by treating the second video block as unavailable, the motion information of the second video block is prohibited from being used in the construction of the merge list of the first block.
[0006] In one exemplary embodiment, a visual media processing method is disclosed. The method relates to the conversion between a first video block of visual media data and a bitstream representation of the visual media data, and includes a step of determining the availability of a second video block of the visual media data using rules, and a step of performing the conversion based on the determination. The rules define using an availability check process for the second video block at one or more positions of the visual media data.
[0007] In one exemplary embodiment, a visual media processing method is disclosed. The method relates to the conversion between a current video block of visual media data and a bitstream representation of the visual media data, and includes a step of determining two positions used to construct an intra block copy motion list for the current video block, and a step of performing the conversion based on the intra block copy motion list.
[0008] In one exemplary embodiment, a visual media processing method is disclosed. The method relates to the conversion between a current video block of visual media data and a coded representation of the visual media data, and includes determining the availability of adjacent blocks to derive one or more weights for the combined intra-inter prediction of the current video block based on rules, and performing the conversion based on the determination. The one or more weights include a first weight designated for the inter prediction of the current video block and a second weight designated for the intra prediction of the current video block, and the rules exclude using the comparison of the coding modes of the current video block and adjacent blocks.
[0009] In another exemplary embodiment, the above method may be executed by a video decoder device including a processor.
[0010] In another exemplary embodiment, the above method may be executed by a video encoder device including a processor.
[0011] In yet another exemplary embodiment, these methods may be embodied in the form of instructions executable by a processor and may be stored in a computer-readable program medium.
[0012] These and other aspects are further described in this specification.
Brief Description of the Drawings
[0013]
Figure 1
[0014]
Figure 2
[0015]
Figure 3
[0016]
Figure 4
[0017]
Figure 5
[0018]
Figure 6
[0019]
Figure 7
[0020]
Figure 8
[0021]
Figure 9
[0022]
Figure 10
[0023]
Figure 11
[0024]
Figure 12
[0025]
Figure 13
[0026]
Figure 14
[0027]
Figure 15
[0028]
Figure 16
[0029]
Figure 17
[0030]
Figure 18
[0031]
Figure 19
[0032]
Figure 20
[0033]
Figure 21
[0034]
Figure 22
[0035]
Figure 23
[0036]
Figure 24
[0037]
Figure 25
[0038]
Figure 26
[0039]
Figure 27
[0040]
Figure 28
[0041]
Figure 29
DETAILED DESCRIPTION OF THE INVENTION
[0042] This specification provides various techniques that can be used by a decoder of an image or video bitstream to improve the quality of a decompressed or decoded digital video or image. For the sake of brevity, the term "video" is used herein to include both a sequence of pictures (traditionally called video) and individual images. Further, it is also possible for a video encoder to implement these techniques during the encoding process to reconstruct decoded frames that are used for further encoding.
[0043] The headings of the sections are used in this specification for ease of understanding and do not limit the embodiments or techniques to the corresponding sections. Thus, embodiments in one section can be combined with embodiments in other sections.
[0044] 1. Overview This specification relates to video coding technology. Specifically, it relates to merge coding including triangular prediction mode. This may be applicable to existing video coding standards such as HEVC, or standards (general video coding) scheduled to be finalized. This may also be applicable to future video coding standards or video codecs.
[0045] 2. Initial Consideration Video coding standards have mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T created H.261 and H.263, ISO / IEC created MPEG-1 and MPEG-4 Visual, and the two organizations jointly created the H.262 / MPEG-2 video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC [1] standards. Since H.262, video coding standards have been based on a hybrid video coding structure, where temporal prediction and transform coding are used. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, many new methods have been adopted by JVET and incorporated into the reference software named Joint Exploration Model (JEM) [3][4]. In April 2018, the Joint Video Expert Team (JVET) of VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was launched and is working on the VVC standard aiming for a 50% bitrate reduction compared to HEVC.
[0046] The latest version of the VVC draft, namely Versatile Video Coding (Draft 5), is available at the following location: http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 14_Geneva / wg11 / JVET-N1001-v7.zip
[0047] The latest reference software of VVC, named VTM, is located at the following place: https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / tags / VTM-5.0
[0048] 2.1 Inter Prediction in HEVC / H.265 A coding unit (CU) to be inter-coded may be coded with one prediction unit (PU) or two PUs according to the partition mode. Each inter prediction PU has motion parameters regarding one or two reference picture lists. The motion parameters include a motion vector and a reference picture index. The use of one of the two reference picture lists may be signaled using inter_pred_idc. The motion vector may be explicitly coded as deltas with respect to the predictor.
[0049] When the CU is coded in skip mode, one PU is associated with the CU and there is no significant residual coefficient, coded motion vector delta, or reference picture index. The merge mode is defined, whereby the motion parameters for the current PU are obtained from neighboring PUs including spatial and temporal candidates. The merge mode can be applied not only to the skip mode but also to any inter-predicted PU. Instead of the merge mode, there is explicit transmission of motion parameters, in which case the motion vector (more precisely, the motion vector difference (MVD) when compared with the motion vector predictor), the corresponding reference picture index for each reference picture list, and the usage of the reference picture list are explicitly signaled for each PU. Such a mode is named advanced motion vector prediction (AMVP) in the present disclosure.
[0050] When signaling that one of the two reference picture lists should be used, the PU is generated from a single block of samples. This is called "uni-prediction". Uni-prediction is available for both P slices and B slices [2].
[0051] When signaling that both reference picture lists should be used, the PU is generated from two blocks of samples. This is called "bi-prediction". Bi-prediction is available only for B slices.
[0052] Details regarding the inter-prediction mode defined in HEVC are described below. The description will start from the merge mode.
[0053] 2.1.1. Reference Picture Lists In HEVC, the term inter prediction is used to mean a prediction derived from data elements (e.g., sample values, motion vectors) of reference pictures other than the current decoded picture. Similarly in H.264 / AVC, a picture can be predicted from multiple reference pictures. The reference pictures used for inter prediction are organized in one or more reference picture lists. The reference index identifies which of the reference pictures in the list should be used to create the prediction signal.
[0054] A single reference picture list, List0, is used for P slices, and two reference picture lists, List0 and List1, are used for B slices. The reference pictures included in List0 / 1 can be from past and future pictures in capture / display order.
[0055]
Number
[0056] These steps are also schematically shown in FIG. 1. For the derivation of spatial merge candidates, up to four merge candidates are selected from among candidates located at five different positions. For the derivation of temporal merge candidates, up to one merge candidate is selected from among two candidates. In the decoder, since a certain number of candidates are assumed for each PU, additional candidates are generated if the number of candidates obtained from step 1 does not reach the maximum number of merge candidates (MaxNumMergeCand) signaled in the slice header. Since the number of candidates is fixed, the index of the best merge candidate is encoded using truncated unary binarization (TU). When the size of the CU is equal to 8, all PUs of the current CU share a single merge candidate list, which coincides with the merge candidate list of the 2N×2N prediction unit.
[0057] The operations related to the above steps will be described in detail below.
[0058] 2.1.2.2 Spatial Candidate Derivation In the derivation of spatial merge candidates, up to four merge candidates are selected from among the candidates located at the positions shown in Figure 2. The order of derivation is A 1 , B 1 , B 0 , A 0 , B 2 . Position B 2 is considered only when any of the PUs in position A 1 , B 1 , B 0 , A 0 is not available (for example, because it belongs to another slice or tile) or is intra-coded. After the candidate in position A 1 is added, the addition of the remaining candidates is subject to redundancy checking, which ensures that candidates with the same motion information are excluded from the list, thereby improving the coding efficiency. To reduce the computational complexity, not all possible candidate pairs are considered in the above redundancy checking. Instead, only the pairs linked by the arrows in Figure 3 are considered, and the candidate is added to the list only if the corresponding candidates used for redundancy checking do not have the same motion information. Another source of duplicate motion information is the "second PU" associated with a partition different from 2Nx2N. As an example, Figure 4 shows the second PU for the cases of N×2N and 2N×N respectively. When the current PU is partitioned as N×2N, the candidate in position A 1 is not considered for list construction. In fact, by adding this candidate, two prediction units with the same motion information are derived, which is redundant, and the coding unit has only one PU. Similarly, when the current PU is partitioned as 2N×N, position B 1 is not considered
[0059] 2.1.2.3 Temporal Candidate Derivation In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal merge candidate, the scaled motion vector is derived based on the co-located PUs in the co-located picture. The scaled motion vector for the temporal merge candidate is obtained as shown by the dotted line in Figure 5, which is scaled from the motion vector of the co-located PU using the POC distance, tb, and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal merge candidate is set equal to zero. A practical implementation of the scaling process is described in the HEVC specification [1]. In a B slice, two motion vectors (one for reference picture list 0 and the other for reference picture list 1) are obtained and combined to create a bi-predicted merge candidate.
[0060]
Number
[0061] For the co-located PU (Y) belonging to the reference frame, the position for the temporal candidate is selected between candidate C 0 and candidate C 1 as shown in Figure 6. If the PU at position C 0 is not available, or is intra-coded, or is outside the row of the current coding tree unit (also known as CTU, largest coding unit (LCU)), position C 1 is used. Otherwise, position C 0 is used for the derivation of the temporal merge candidate. The related syntax elements are explained as follows: 7.3.6.1. General slice segment header syntax
[0062] [Number]
[0063] 2.1.2.5 Derivation of MV for TMVP Candidate More specifically, to derive a TMVP candidate, the following steps are performed: 1) Set the reference picture list X = 0, and the target reference picture is the reference picture with index 0 in list X (i.e., curr_ref). Start the derivation process for the motion vectors at the same position, and obtain the MV for list X pointing to curr_ref. 2) If the current slice is a B slice, set the reference picture list to X = 1, and the target reference picture is the reference picture with index equal to 0 in list X (i.e., curr_ref). Start the derivation process for the motion vectors at the same position, and obtain the MV for list X pointing to curr_ref.
[0064] The derivation process for the motion vectors at the same position is described in the following sub-section 2.1.2.5.1.
[0065] 2.1.2.5.1 Derivation Process for Motion Vectors at the Same Position For a block at the same position, it may be intra- or inter-coded with single-prediction or bi-prediction. If it is intra-coded, the TMVP candidate is set to be unavailable.
[0066] If it is single-prediction from list A, the motion vector of list A is scaled to the target reference picture list X.
[0067] If it is bi-prediction and the target reference picture list is X, the motion vector of list A is scaled to the target reference picture list X, and A is determined according to the following rules: - If the reference picture does not have a larger POC value compared to the current picture, A is set to X. - Otherwise, A is set equal to collocated_from_l0_flag.
[0068] The relevant working draft in JCTVC-W1005-v4 is described as follows:
[0069]
Number
[0070]
Number
[0071]
Number
[0072]
Number
[0073]
Number
[0074]
Number
[0075]
Number
[0076] 2.1.2.6 Additional Candidate Insertion In addition to spatial and temporal merge candidates, there are two additional types of merge candidates: combined bi-predictive merge candidates and zero merge candidates. Combined bi-predictive merge candidates are generated by utilizing spatial and temporal merge candidates. Combined bi-predictive merge candidates are only used for B slices. Combined bi-predictive candidates are generated by combining the first reference picture list motion parameters of the initial candidates with another second reference picture list motion parameter. If these two tuples result in different motion hypotheses, they will form a new bi-predictive candidate. As an example, FIG. 7 shows the case where two candidates in the original list (left side) having mvL0 and refIdxL0 or mvL1 and refIdxL1 are used to create a combined bi-predictive merge candidate that is added to the final list (right side). There are a number of rules regarding the combinations that are thought to generate these additional merge candidates defined in [1].
[0077] To fill the remaining entries of the merge candidate list, zero motion candidates are inserted, thus corresponding to the MaxNumMergeCand capacity. These candidates have a zero spatial displacement and a reference picture index, the latter starting from zero and incremented each time a new zero motion candidate is added to the list. Finally, no redundancy check is performed on these candidates.
[0078] 2.1.3 AMVP AMVP utilizes the spatio-temporal correlation of adjacent PUs of motion vectors, which is used for explicit transmission of motion parameters. For each reference picture list, first, the availability of the top-left temporally adjacent PU position is examined, redundant candidates are removed, and a motion vector candidate list is constructed by adding zero vectors to make the candidate list a fixed length. Then, the encoder can select the best predictor from the candidate list and transmit the corresponding index indicating the selected candidate. Similar to the merge index signaling, the index of the best motion vector candidate is encoded using truncated unary. The maximum value to be encoded in this case is 2 (see Figure 8). In the following section, details regarding the derivation process of motion vector prediction candidates are described.
[0079] 2.1.3.1 Derivation of AMVP Candidates Figure 8 summarizes the derivation process of motion vector prediction candidates.
[0080] In motion vector prediction, two types of motion vector candidates, spatial motion vector candidates and temporal motion vector candidates, are considered. For the derivation of spatial motion vector candidates, two motion vector candidates are derived based on the motion vectors of each PU located at five different positions as finally shown in Figure 2.
[0081] For the derivation of temporal motion vector candidates, one motion vector candidate is derived from among two candidates derived based on two different equivalent positions. After the first list of spatio-temporal candidates is created, duplicate motion vector candidates within the list are excluded. If the number of potential candidates is more than two, motion vector candidates whose reference picture index is greater than 1 within the relevant reference picture list are excluded from the list. If the number of spatio-temporal motion vector candidates is less than two, additional zero motion vector candidates are added to the list.
[0082] 2.1.3.2 Spatial Motion Vector Candidates In the derivation of spatial motion vector candidates, up to two candidates are considered from among five potential candidates, and the candidates are derived from PUs located at positions as shown in Figure 2, and those positions are the same as those for motion merging. The order of derivation for the left side of the current PU is A 0 , A 1 , scaled A 0 , scaled A 1 as defined. The order of derivation for the upper side of the current PU is B 0 , B 1 , B 2 , scaled B 0 , scaled B 1 , scaled B 2 as defined. Therefore, for each side, there are four cases that can be used as motion vector candidates, two cases do not require the use of spatial scaling, and two cases use spatial scaling. The four different cases are summarized as follows. · Without spatial scaling -(1) Same reference picture list, and same reference picture index (same POC) -(2) Different reference picture lists, but same reference picture (same POC) · With spatial scaling -(3) Same reference picture list, but different reference pictures (different POCs) -(4) Different reference picture lists, and different reference pictures (different POCs)
[0083] The cases without spatial scaling are inspected first, followed by spatial scaling. Spatial scaling is considered when the POC is different between the reference picture of the adjacent PU and the reference picture of the current PU, regardless of the reference picture list. If all PUs of the left candidate are not available or are intra-coded, scaling for the upper motion vector is allowed to assist in the parallel derivation of the left and upper MV candidates. Otherwise, spatial scaling is not allowed for the upper motion vector.
[0084] In the spatial scaling process, the motion vectors of adjacent PUs are scaled in the same way as in the case of temporal scaling, as shown in FIG. 9. The main difference is that the index of the current PU and the reference picture list are given as input, and the actual scaling process is the same as that by temporal scaling.
[0085] 2.1.3.3 Temporal motion vector candidates Except for the case of deriving the reference picture index, all processes for deriving temporal merge candidates are the same as those for deriving spatial motion vector candidates (see FIG. 6). The reference picture index is signaled to the decoder.
[0086] 2.2 Inter prediction methods in VVC There are several new coding tools for improving inter prediction, such as adaptive motion vector difference resolution (AMVR) for signaling MVD, merge with motion vector difference (MMVD), triangular prediction mode (TPM), combined intra-inter prediction (CIIP), advanced TMVP (also referred to as ATMVP or SbTMVP), affine prediction mode, generalized bi-prediction (GBI), decoder-side motion vector refinement (DMVR), bi-directional optical flow (also referred to as BIO or BDOF), etc.
[0087] There are three different merge list construction processes supported in VVC: 1) Sub-block merge candidate list: This includes ATMVP and affine merge candidates. One merge list construction process is shared by both the affine mode and the ATMVP mode. Here, the ATMVP and affine merge candidates may be added sequentially. The size of the sub-block merge list is signaled in the slice header, and the maximum value is 5. 2) Regular Merge List: For the blocks to be inter-coded, one merge list construction process is shared. Here, spatial / temporal merge candidates, HMVP, pairwise merge candidates, and zero motion candidates may be inserted sequentially. The regular merge list size is signaled in the slice header, and the maximum value is 6. MMVD, TPM, and CIIP depend on the regular merge list. 3) IBC Merge List: It is performed in the same way as the regular merge list.
[0088] Similarly, there are three AMVP lists supported in VVC: 1) Affine AMVP Candidate List 2) Regular AMVP Candidate List 3) IBC AMVP Candidate List: The same construction process as the IBC merge list due to the adoption of JVET-N0843
[0089] 2.2.1 Coding Block Structure in VVC In VVC, a quadtree / binary tree / trinary tree (QT / BT / TT) structure is adopted to divide the picture into square or rectangular blocks.
[0090] In addition to QT / BT / TT, for I-frames, VVC also adopts a separate tree (also referred to as a dual coding tree). In the separate tree, the coding block structure is signaled separately for the luma and chroma components.
[0091] Also, except for the blocks coded in combination with specific coding methods (such as intra-sub-partition prediction when the PU is equal to the TU but smaller than the CU, sub-block transform of the inter-coded block when the PU is equal to the CU but the TU is smaller than the PU, etc.), the CU is set equal to the PU and the TU.
[0092] 2.2.2 Affine Prediction Mode In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). In the real world, there are many types of motions such as zoom-in / out, rotation, viewpoint movement, and other irregular motions. In VVC, simplified affine transform motion compensation prediction is applied using a 4-parameter affine model and a 6-parameter affine model. As shown in Figure 10, the affine motion field of a block is described by two control point motion vectors (CPMVs) for the 4-parameter affine model and three CPMVs for the 6-parameter affine model.
[0093]
Number
[0094] To further simplify motion compensation prediction, sub-block-based affine transform prediction is applied. To derive the motion vector for each M×N sub-block (both M and N are set to 4 in the current VVC), the motion vector of the central sample of each sub-block is calculated according to Equations (1) and (2) as shown in Figure 11 and rounded to 1 / 16 fractional precision. Then, a 1 / 16-pel motion compensation interpolation filter is applied to generate the prediction for each sub-block with the derived motion vector. The 1 / 16-pel interpolation filter is introduced in the affine mode.
[0095] After MCP, the high-precision motion vector of each sub-block is rounded and saved with the same precision as a normal motion vector. 2.2.3 MERGE for the whole block 2.2.3.1 Construction of the merge list for the translational regular merge mode 2.2.3.1.1 History-based motion vector prediction (HMVP)
[0096] Unlike the merge list design, in VVC, the history-based motion vector prediction (HMVP) method is adopted.
[0097] In HMVP, the previously coded motion information is saved. The motion information of the previously coded block is defined as an HMVP candidate. Multiple HMVP candidates are saved in a table named the HMVP table, and this table is maintained on the fly during the encoding / decoding process. When starting to encode / decrypt a new tile / LCU row / slice, the HMVP table is emptied. If there are inter-coded blocks and non-sub-blocks and non-TPM modes, the relevant motion information is always added as a new HMVP candidate to the last entry of the table. The overall coding flow is shown in Figure 12.
[0098] 2.2.3.1.2 Regular Merge List Construction Process The construction of the regular merge list (for translational motion) can be summarized by the following step sequence: ● Step 1: Derivation of spatial candidates ● Step 2: Insertion of HMVP candidates ● Step 3: Insertion of pairwise average candidates ● Step 4: Default motion candidates
[0099] The HMVP candidates can be used in both the AMVP and the merge candidate list construction process. Figure 13 shows the modified merge candidate list construction process (highlighted with dotted boxes). After the insertion of the TMVP candidates, if the merge candidate list is not full, the HMVP candidates stored in the HMVP table can be used to fill up the merge candidate list. From the perspective of motion information, usually, considering that one block has a high correlation with its nearest neighboring block, the HMVP candidates in the table are inserted in descending order of index. The last entry in the table is added to the list first, and the first entry is added last. Similarly, redundancy removal is applied to the HMVP candidates. When the total number of available merge candidates reaches the maximum number of merge candidates allowed to be signaled, the construction process of the merge candidate list ends.
[0100] Note that all spatial / temporal / HMVP candidates are to be coded in non-IBC mode. Otherwise, adding to the regular merge candidate list is not allowed.
[0101] The HMVP table contains up to 5 normal motion candidates, each of which is unique.
[0102] 2.2.3.1.2.1 Pruning Process A candidate is only added to the list if the corresponding candidate used for redundancy check does not have the same motion information. Such a comparison process is called the pruning process.
[0103] The pruning process within the spatial candidates depends on the use of the TPM for the current block.
[0104] When the current block is coded without using the TPM mode (e.g., regular merge, MMVD, CIIP), the HEVC pruning process (i.e., 5-pruning) for the spatial merge candidates is used.
[0105] 2.2.4 Triangular Prediction Mode (TPM) In VVC, the triangular partitioning mode is supported for inter prediction. The triangular partitioning mode is only applied to CUs larger than 8×8 that are coded in merge mode rather than the MMVD or CIIP mode. For CUs that meet these conditions, a CU-level flag is signaled to indicate whether the triangular partitioning mode is applied or not.
[0106] When this mode is used, the CU is evenly divided into two triangular partitions using either a diagonal split or an anti-diagonal split, as shown in Figure 14. Each triangular partition in the CU is inter-predicted using its own motion: only uni-prediction is allowed for each partition, i.e., each partition has one motion vector and one reference index. Similar to conventional bi-prediction, uni-prediction motion constraints are applied to ensure that only two motion-compensated predictions are required for each CU.
[0107] If the CU-level flag indicates that the current CU is coded using the triangular partitioning mode, a flag indicating the direction of the triangular partition (diagonal or anti-diagonal) and two merge indices (one for each partition) are further signaled. After predicting each triangular partition, the sample values along the diagonal or anti-diagonal edges are adjusted using a mixing process with adaptive weights. This is the prediction signal for the entire CU, and the transform and quantization processes will be applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the triangular partitioning mode is stored in 4×4 units.
[0108] The regular merge candidate list is reused for triangular partition merge prediction without pruning the extra motion vectors. For each merge candidate in the regular merge candidate list, only one of its L0 or L1 motion vectors is used for triangular prediction. Furthermore, the order of selecting the L0 vs. L1 motion vectors is based on its merge index parity. Using this method, the regular merge list can be directly used.
[0109] 2.2.4.1 Merge Candidate List Construction for TPM Basically, as proposed in JVET-N0340, the regular merge list construction process is applied. However, some modifications are made.
[0110]
Number
[0111]
Number
[0112]
Number
[0113]
Number
[0114]
Number
[0115]
Number
[0116]
Number
[0117]
Number
[0118]
Number
[0119]
Number
[0120]
Number
[0121]
Number
[0122]
Number
[0123]
Number
[0124]
Number
[0125]
Number
[0126] [Number]
[0127] [Number]
[0128] [Number]
[0129] [Number]
[0130] [Number]
[0131] [Number]
[0132] [Number]
[0133] [Number]
[0134] [Number]
[0135] [Number]
[0136] 2.2.4.4.1 Decryption Process The decoding process as provided in JVET-N0340 is defined as follows:
[0137]
Number
[0138]
Number
[0139]
Number
[0140] The dual-prediction weight index bcwIdx is set equal to 0.
[0141]
Number
[0142]
Number
[0143]
Number
[0144]
Number
[0145]
Number
[0146]
Number
[0147]
Number
[0148]
Number
[0149] 2.2.5 MMVD In JVET-L0054, the final motion vector representation (also known as UMVE, MMVD) is presented. UMVE is used in either skip or merge mode using the proposed motion vector representation method.
[0150] UMVE re-uses the same merge candidates as those included in the regular merge candidate list of VVC. Among the merge candidates, it is possible to select the basic candidates, which are further extended by the proposed motion vector representation method.
[0151] UMVE provides a new motion vector difference (MVD) representation method, in which the starting point, the magnitude of the motion, and the direction of the motion are used to represent the MVD.
[0152] This proposed technique uses the merge candidate list as it is. However, only the candidates that are of the default merge type (MRG_TYPE_DEFAULT_N) are considered for the extension of UMVE.
[0153]
Number
[0154]
Number
[0155]
Number
[0156]
Number
[0157] The UMVE flag is signaled immediately after sending a skip flag or a merge flag. If the skip or merge flag is true, the UMVE flag is parsed. If the UMVE flag is equal to 1, the UMVE syntax is parsed. However, if it is not 1, the AFFINE flag is parsed. If the AFFINE flag is equal to 1, it is in the affine mode, otherwise, the skip / merge index is parsed with respect to the skip / merge mode of VTM.
[0158] No additional line buffer is required due to UMVE candidates. This is because the software skip / merge candidates are directly used as base candidates. Using the input UMVE index, the refinement of the MV is determined just before motion compensation. Therefore, there is no need to hold a long line buffer.
[0159] Under the current common test conditions, the first or second merge candidate in the merge candidate list can be selected as the base candidate.
[0160] UMVE is also called Merge with MV Differences (MMVD).
[0161] 2.2.6 Composite Intra-Inter Prediction (CIIP) In JVET-L0100, multiple hypothesis predictions are proposed, and composite intra & inter prediction is one of the methods to generate multiple hypotheses.
[0162] When multiple hypothesis predictions are applied to improve the intra mode, the multiple hypothesis predictions combine one intra prediction and one merge index prediction. In the merge CU, one flag is signaled for the merge mode to select the intra mode from the intra candidate list when the flag is true. For the luma component, the intra candidate list is derived from only one intra prediction mode, i.e., the planar mode. The weights applied to the predicted blocks from intra and inter predictions are determined by the coded modes (intra or non-intra) of two adjacent blocks (A1 and B1).
[0163] 2.2.7 Merge for sub-block-based technology In addition to the regular merge list for non-sub-block merge candidates, it is proposed that all sub-block related motion candidates be put into individual merge lists.
[0164] The sub-block related motion candidates are put into a separate merge list and named the "sub-block merge candidate list".
[0165] In one example, the sub-block merge candidate list includes ATMVP candidates and affine merge candidates.
[0166] The sub-block merge candidate list is filled with candidates in the following order: a. ATMVP candidates (which may or may not be available); b. Affine merge list (including inherited affine candidates; and constructed affine candidates) c. Padding as a zero MV 4-parameter affine model
[0167] 2.2.7.1.1 ATMVP (Sub-block Temporal Motion Vector Predictor, also known as SbTMVP) The basic idea of ATMVP is to derive multiple sets of temporal motion vector predictors for each block. Each sub-block is assigned a set of motion information. When an ATMVP merge candidate is generated, motion compensation is performed at the 8×8 level instead of the block level.
[0168] In the current design, ATMVP predicts the motion vectors of sub-CUs within a CU in two steps, which are described in the following two sub-sections 2.2.7.1.1.1 and 2.2.7.1.1.2 respectively.
[0169] [Number]
[0170] (The rounded MV is added to the center position of the current block and clipped within a certain range if necessary) The corresponding block is identified within the picture at the equivalent position signaled in the slice header using the initialized motion vector.
[0171] If the block is inter-coded, proceed to the second step. Otherwise, the ATMVP candidate is set as not available.
[0172] 2.2.7.1.1.2 Motion Derivation of Sub-CU The second step is to divide the current CU into sub-CUs and obtain the motion information of each sub-CU from the blocks corresponding to each sub-CU in the picture at the equivalent position.
[0173] When the corresponding block of the sub-CU is coded in inter mode, the motion information is used to derive the final motion information of the current sub-CU by calling the derivation process for the MV at the equivalent position, which is not different from the process of the conventional TMVP process. Basically, from the single-prediction or bi-prediction target list X, when the corresponding block is predicted, the motion vector is used; otherwise, it is predicted from list Y (Y = 1 - X) for single-prediction or bi-prediction, and if NoBackwardPredFlag is equal to 1, the MV in list Y is used. Otherwise, no motion candidate can be found.
[0174] When the block in the equivalent position picture identified by the initialized MV and the position of the current sub-CU is coded in intra or IBC, or when no motion candidate can be found as described above, the following is further applied:
[0175] The reference picture R col The motion vector used to obtain the motion field of is denoted as MV col To minimize the influence caused by MV scaling, the MV in the spatial candidate list used to derive MV col is selected in the following way: If the reference picture of the candidate MV is the equivalent position picture, this MV is selected and used as MV col without any scaling. Otherwise, the MV with the reference picture closest to the equivalent position picture is selected to derive MV col with scaling.
[0176] The decoding process related to the derivation process of the equivalent position motion vector in JVET-N1001 is described below together with the part related to ATMVP emphasized in bold underlined text:
[0177]
Number
[0178] [Number]
[0179] [Number]
[0180] [Number]
[0181] [Number] TIFF0007687763000065.tif186170
[0182] 2.2.8 Regular Inter-Mode (AMVP) 2.2.8.1 AMVP Motion Candidate List Similar to the AMVP design in HEVC, at most two AMVP candidates can be derived. However, the HMVP candidates may be added after the TMVP candidates. The HMVP candidates in the HMVP table are traversed in ascending order of the index (i.e., the index equal to 0, starting from the oldest). At most four HMVP candidates may be examined to find out whether their reference picture is the same as the target reference picture (i.e., the same POC value).
[0183] 2.2.8.2 AMVR In HEVC, when use_integer_mv_flag is equal to 0 in the slice header, the motion vector difference (MVD) (between the motion vector and the predicted motion vector of the PU) is signaled in units of 1 / 4 luma samples. In VVC, an Adaptive Motion Vector Resolution (AMVR) is introduced. In VVC, the MVD can be coded in units of 1 / 4 luma samples, integer luma samples, or 4 luma samples (i.e., 1 / 4-pel, 1-pel, 4-pel). The MVD resolution is controlled at the Coding Unit (CU) level, and the MVD resolution flag is conditionally signaled for each CU having at least one non-zero MVD component.
[0184] For a CU having at least one non-zero MVD component, a first flag is signaled to indicate whether 1 / 4 luma sample MV accuracy is used in the CU. If the first flag (equal to 1) indicates that 1 / 4 luma sample MV accuracy is not used, another flag is signaled to indicate whether integer luma sample MV accuracy or 4 luma sample MV accuracy is used.
[0185] If the first MVD resolution flag of a CU is zero or not coded for the CU (meaning that all MVDs within the CU are zero), 1 / 4 luma sample MV resolution is used for the CU. When a CU uses integer luma sample MV accuracy or 4 luma sample MV accuracy, the MVP of the AMVP candidate list for the CU is rounded to the corresponding accuracy.
[0186] 2.2.8.3 Symmetric Motion Vector Difference in JVET-N1001-v2 In JVET-N1001-v2, Symmetric Motion Vector Difference (SMVD) is applied to the coding of motion information in bi-prediction.
[0187] First, at the slice level, the variables RefIdxSymL0 and RefIdxSymL1 indicating the reference picture indices of list0 / list1 used in the SMVD mode are derived using the following steps as specified by N1001-v2. If at least one of the two variables is equal to -1, the SMVD mode shall be disabled.
[0188] 2.2.9 Refinement of Motion Information 2.2.9.1 Decoder-side Motion Vector Refinement (DMVR) In bi-prediction operation, for the prediction of one block region, two prediction blocks formed using the motion vectors (MVs) of list0 and list1 respectively are combined to form a single prediction signal. In the decoder-side motion vector refinement (DMVR) method, the two MVs of bi-prediction are further refined.
[0189] Regarding DMVR in VVC, as shown in Figure 19, MVD mirroring is assumed between list0 and list1, and bilateral matching is performed to refine the MV, that is, to find the best MVD among multiple MVD candidates. The MVs of the two reference picture lists are represented by MVL0 (L0X, L0Y), and MVL1 (L1X, L1Y). The MVD indicated by (MvdX, MvdY) for list0, which can minimize the cost function (such as SAD), is defined as the best MVD. For the SAD function, it is defined as the SAD between the reference block of list0 derived from the motion vector (L0X+MvdX, L0Y+MvdY) in the reference picture of list0 and the reference block of list1 derived from the motion vector (L1X-MvdX, L1Y-MvdY) in the reference picture of list1.
[0190] The motion vector refinement process may be repeated twice. In each iteration, as shown in FIG. 20, up to six MVDs (integer pel precision) can be examined in two steps. In the first step, MVDs (0, 0), (-1, 0), (1, 0), (0, -1), (0, 1) are examined. In the second step, one of MVDs (-1, -1), (-1, 1), (1, -1) or (1, 1) can be selected for further examination. Assume that the function Sad(x, y) returns the SAD value of MVD (x, y). The MVD indicated by (MvdX, MvdY) examined in the second step is determined as follows: MvdX = -1; MvdY = -1; If (Sad(1, 0) < Sad(-1, 0)) MvdX = 1; If (Sad(0, 1) < Sad(0, -1)) MvdY = 1;
[0191] In the first iteration, the starting point is the signaled MV. In the second iteration, the starting point is the signaled MV plus the best MVD selected in the first iteration. DMVR is applied only when one reference picture is the previous picture and the other reference picture is the subsequent picture, and the two reference pictures have the same picture order count distance from the current picture.
[0192]
Number
[0193]
Number
[0194]
Number
[0195] 2.3 Intra-block copy In High Efficiency Video Coding (HEVC) Screen Content Coding Extension (HEVC-SCC) and the current Versatile Video Coding (VVC) Test Model (VTM-4.0), Intra-block copy (IBC), also known as current picture reference, is adopted. IBC extends the concept of motion compensation from inter-frame coding to intra-frame coding. As shown in Fig. 21, when IBC is applied, the current block is predicted by a reference block within the same picture. The samples within the reference block must have been reconstructed already before the current block is encoded or decoded. IBC is not very efficient for sequences captured by most cameras, but it shows significant coding gain for screen content. The reason is that screen content pictures contain many repetitive patterns such as icons and text characters. IBC can effectively remove the redundancy between these repetitive patterns. In HEVC-SCC, an inter-coded coding unit (CU) can apply IBC when it selects the current picture as its reference picture. In this case, the motion vector (MV) is renamed as a block vector (BV), and the BV always has integer pixel accuracy. To maintain compatibility with the main profile of HEVC, the current picture is marked as a "long-term" reference picture in the decoded picture buffer (DPB). Similarly, in the multi-view / 3D video coding standard, the inter-view reference pictures should also be marked as "long-term" reference pictures.
[0196] After finding the reference block following the BV, prediction can be generated by copying the reference block. The residual can be obtained by subtracting the reference pixels from the original signal. Then, transformation and quantization can be applied in the same way as in other coding modes.
[0197] However, if the reference block is outside the picture, overlapping with the current block, outside the reconstructed area overlapping with the current block, or outside the valid area restricted by some constraints, all or some of the pixel values are not defined. To address this issue, there are basically two solutions. One is to not allow such situations, for example, in bitstream conformance. The other is to apply padding to the undefined pixel values. The following sub-sessions explain the solutions in detail.
[0198] 2.3.1 IBC in the VVC Test Model (VTM4.0) In the current VVC test model, i.e., the VTM-4.0 design, the entire reference block should be with the current coding tree unit (CTU) and not overlap with the current block. Therefore, there is no need to pad the reference or prediction block. The IBC flag is coded as the prediction mode of the current CU. Therefore, for each CU, there are three prediction modes: MODE_INTRA, MODE_INTER, and MODE_IBC
[0199]
Number
[0200] In the derivation of spatial merge candidates, among the candidates located at positions such as those shown in Figure 2 as A 1, B 1, B 0, A 0 and B 2 , a maximum of four merge candidates are selected. The derivation order is A 1, B 1, B 0, A 0 and B 2 . Position B 2 is the position A 1 , B 1 , B 0, A 0 It is only considered when none of the PUs are available (e.g., due to belonging to other slices or tiles), or when they are not coded in IBC mode. Position A 1 After a candidate at position A is added, the insertion of the remaining candidates is subject to a redundancy check, which ensures that candidates with the same motion information are excluded from the list to improve coding efficiency.
[0201] After the insertion of spatial candidates, if the IBC merge list size is still smaller than the maximum IBC merge list size, IBC candidates from the HMVP table may be inserted. When inserting HMVP candidates, a redundancy check is performed.
[0202] Finally, pairwise average candidates are inserted into the IBC merge list.
[0203] If the reference block identified by a merge candidate is outside the picture, overlapping with the current block, outside the reconstructed area, or outside the valid area restricted by some constraints, the merge candidate is called an invalid merge candidate.
[0204] Note that invalid merge candidates may be inserted into the IBC merge list.
[0205] 2.3.1.2 IBC AMVP Mode In IBC AMVP mode, an AMVP index indicating an entry in the IBC AMVP list is parsed from the bitstream. The construction of the IBC AMVP list can be summarized by the following sequence of steps: ● Step 1: Derivation of spatial candidates ○ Examine A 0 , A 1 until an available candidate is found. ○ Examine B 0 , B 1 , B2 Perform an inspection. ● Step 2: Insertion of HMVP candidates ● Step 3: Insertion of zero candidates
[0206] After the insertion of spatial candidates, if the IBC AMVP list size is still smaller than the IBC maximum AMVP list size, IBC candidates from the HMVP table may be inserted.
[0207] Finally, zero candidates are inserted into the IBC AMVP list.
[0208]
Number
[0209]
Number
[0210]
Number
[0211]
Number
[0212]
Number
[0213]
Number
[0214]
Number
[0215]
Number
[0216]
Number
[0217]
Number
[0218]
Number
[0219]
Number
[0220]
Number
[0221]
Number
[0222]
Number
[0223]
Number
[0224]
Number
[0225]
Number
[0226]
Number
[0227]
Number
[0228]
Number
[0229]
Number
[0230]
Number
[0231]
Number
[0232]
Number
[0233]
Number
[0234]
Number
[0235]
Number
[0236]
Number
[0237]
Number
[0238]
Number
[0239]
Number
[0240]
Number
[0241]
Number
[0242]
Number
[0243]
Number
[0244]
Number
[0245]
Number
[0246]
Number
[0247]
Number
[0248]
Number
[0249]
Number
[0250]
Number
[0251]
Number
[0252]
Number
[0253]
Number
[0254]
Number
[0255]
Number
[0256]
Number
[0257]
Number
[0258]
Number
[0259]
Number
[0260]
Number
[0261]
Number
[0262]
Number
[0263]
Number
[0264]
Number
[0265]
Number
[0266]
Number
[0267]
Number
[0268]
Number
[0269]
Number
[0270]
Number
[0271]
Number
[0272]
Number
[0273]
Number
[0274]
Number
[0275]
Number
[0276]
Number
[0277]
Number
[0278]
Number
[0279]
Number
[0280]
Number
[0281]
Number
[0282]
Number
[0283]
Number
[0284]
Number
[0285]
Number
[0286] FIG. 22 is a block diagram of a video processing apparatus 2200. The apparatus 2200 may be used to implement one or more of the methods described in this application. The apparatus 1500 may be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, or the like. The apparatus 2200 may include one or more processors 2202, one or more memories 2204, and video processing hardware 2206. The processor 2202 may be configured to implement one or more of the methods described in this document. The memories 2204 may be used to store data and code used to implement the methods and techniques described in this application. The video processing hardware 2206 may be used to implement some of the techniques described in this document in a hardware circuit. The video processing hardware 2206 may be partially or fully included within the processor 2202 in the form of dedicated hardware, or a graphics processing unit (GPU), or a dedicated signal processing block.
[0287] Some embodiments can be described using the following clauses.
[0288] Some exemplary embodiments of the techniques described in item 1 of section 4 include the following:
[0289] 1. A video processing method (e.g., method 2300 shown in FIG. 23), comprising: applying a pruning process to the merge list construction of a current video block partitioned using a triangular partition mode (TMP) in which the current video block is partitioned into at least two non-rectangular sub-blocks, the pruning process being the same as another pruning process for another video block partitioned using a non-TMP partition (step 2302); performing a conversion between the video block and a bitstream representation of the video block based on the merge list construction (step 2304); A method including the above.
[0290] 2. The method according to claim 1, wherein the pruning process includes using partial pruning on spatial merge candidates of the current video block.
[0291] 3. The method according to claim 1, wherein the pruning process includes applying full or partial pruning to the current video block based on a block size rule that defines using full or partial pruning based on the size of the current video block.
[0292] The method according to claim 1, wherein the pruning process includes using different orders of adjacent blocks during the merge list construction process.
[0293] Some exemplary embodiments of the technique described in item 2 of section 4 include the following:
[0294] 1. A video processing method, including: During the conversion between the current video block and the bitstream representation of the current video block, determining an alternative temporal motion vector predictor coding (ATMVP) mode for the conversion based on a list X of adjacent blocks of the current video block, where X is an integer and the value of X depends on the encoding conditions of the current video block; Performing the conversion based on the availability of the ATMVP mode; A method including the above.
[0295] 2. X indicates the position of the video picture at the equivalent position from which the temporal motion vector prediction used for the conversion between the current video block and the bitstream representation is performed, according to the method of claim 1.
[0296] X is determined by comparing the picture order count (POC) of all reference pictures in all reference lists for the current video block with the POC of the current video picture of the current video block, according to the method of claim 1.
[0297] If the above comparison indicates that the POC is less than or equal to the POC of the current picture, set X = 1, otherwise set X = 0, according to the method of claim 3.
[0298] The motion information stored in the history-based motion vector predictor table is used to initialize the motion vector in the ATMVP mode, according to the method of claim 1.
[0299] Some exemplary embodiments of the techniques described in item 3 of section 4 include the following:
[0300] 1. A video processing method, comprising: During the conversion between the current video block and the bitstream representation of the current video block, a step of determining that a sub-block-based coding technique is used for the conversion, wherein in the sub-block-based coding technique, the current video block is partitioned into at least two sub-blocks, and each sub-block is capable of deriving its own motion information; A step of performing the conversion by utilizing a merge list construction process for the current video block integrated using a block-based derivation process for the collocated motion vectors; And a method comprising.
[0301] 2. The merge list construction process and the derivation process include performing a single-prediction from list Y, and the motion vectors of list Y are scaled to the target reference picture list X, the method according to claim 1.
[0302] 3. The merge list construction process and the derivation process include performing a dual-prediction using the target reference picture list X, and the motion vectors of list Y are scaled to those of list X, and Y is determined according to a rule, the method according to claim 1.
[0303] Some exemplary embodiments of the technology described in item 4 of section 4 include the following:
[0304] 1. A video processing method, comprising: determining satisfied conditions and unsatisfied conditions based on the dimensions of the current video block of the video block and / or the availability of a merge sharing status in which merge candidates from different coding tools are shared; performing a conversion between the current video block and the bitstream representation of the current video block based on the conditions. A method comprising.
[0305] 2. The step of performing the conversion includes skipping the step of deriving a spatial merge candidate if the conditions are satisfied, the method according to claim 1.
[0306] 3. The step of performing the conversion includes skipping the step of deriving a history-based motion vector candidate if the conditions are satisfied, the method according to claim 1.
[0307] Based on the current video block being in a shared mode in the video picture, the conditions are determined to be satisfied, the method according to any one of claims 1-3.
[0308] Some exemplary embodiments of the technology described in item 5 of section 4 include the following:
[0309] 1. A video processing method, comprising: During conversion between a current video block and a bitstream representation of the current video block, determining that coding tools are disabled for the conversion, wherein the bitstream representation is configured to provide an indication that the maximum number of merge candidates for the coding tools is zero; Using the determination that the coding tools are disabled to perform the conversion; and A method comprising the above.
[0310] 2. The method according to claim 1, wherein the coding tools correspond to an intra block copy in which pixels of the current video block are coded from other pixels in the video region of the current video block.
[0311] 3. The method according to claim 1, wherein the coding tools are sub-block coding tools.
[0312] 4. The method according to claim 3, wherein the sub-block coding tools are affine coding tools or another motion vector predictor tool.
[0313] 5. The method according to any one of claims 1-4, wherein performing the conversion includes processing the bitstream by skipping syntax elements associated with the coding tools.
[0314] Some exemplary embodiments of the technology described in item 6 of section 4 include the following:
[0315] 1. A video processing method, comprising: A step of making a determination using rules during conversion between a current video block and a bitstream representation of the current video block, wherein the rule stipulates that a first syntax element in the bitstream representation conditionally exists based on a second syntax element indicating the maximum number of merge candidates used by a coding tool used during the conversion. A step of performing conversion between a current video block and a bitstream representation of the current video block based on the above determination. A method comprising the above.
[0316] 2. The method according to claim 1, wherein the first syntax element corresponds to a merge flag.
[0317] 3. The method according to claim 1, wherein the first syntax element corresponds to a skip flag.
[0318] The method according to any one of claims 1 - 3, wherein the coding tool is a sub - band coding tool, and the second syntax element corresponds to the maximum allowable merge candidates for the sub - band coding tool.
[0319] Regarding items 14 - 17 in the above section, the following clauses describe some technical solutions.
[0320] A video processing method, Regarding conversion between a coded representation of a first video block of a video and a second video block, a step of determining the availability of the second block using an availability check process during the conversion, wherein the availability check process checks at least a first position and a second position for the first video block. A step of performing conversion based on the result of the determination. A method comprising the above.
[0321] The above method, wherein the first position corresponds to the top - left position.
[0322] The above method, wherein the second position corresponds to the upper left position.
[0323] A video processing method, Regarding the conversion between the coded representation of a video block of a video and a second video block, a step of determining a list of intra-block copy motion candidates at a first position and a second position for the video block; A step of performing the conversion based on the result of the determination; A method including the above.
[0324] The above method, wherein the first position corresponds to the upper left position of the shared merge area of the video block.
[0325] The conversion includes generating a bitstream representation from the current video block, and is the method according to any one of the above clauses.
[0326] The conversion includes generating samples of the current video block from the bitstream representation, and is the method according to any one of the above clauses.
[0327] A video processing apparatus including a processor configured to execute the method according to any one or more of the above clauses.
[0328] A computer-readable medium storing code, which, when executed, causes a processor to execute the method according to any one or more of the above clauses.
[0329] FIG. 25 is a block diagram showing an exemplary video processing system 1900 in which various techniques disclosed in the present application can be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1922 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.
[0330] System 1900 may include a coding component 1904 capable of implementing various coding or encoding methods described in this document. The coding component 1904 can reduce the average bit rate of the video from the input 1902 to the output of the coding component 1904 to generate a coded representation of the video. Accordingly, coding techniques are sometimes referred to as video compression or video transcoding techniques. The output of the coding component 1904 may be stored or transmitted via a communication connection as represented by component 1906. The stored or communicated bitstream (or coded) representation of the video received at the input 1902 may be used by component 1908 to generate pixel values or viewable video to be transmitted to the display interface 1910. The process of generating a viewable video from the bitstream representation is sometimes referred to as video decompression. Further, certain video processing operations are referred to as "encoding" operations or tools, but it will be understood that encoding tools or operations are used in an encoder and corresponding decoding tools or operations that process the result of the encoding in reverse will be executed in a decoder.
[0331] Examples of a peripheral bus interface or a display interface may include a Universal Serial Bus (USB) or a High-Definition Multimedia Interface (HDMI (registered trademark)), a Displayport, etc. Examples of a storage interface may include a serial advanced technology attachment (SATA), a PCI, an IDE interface, etc. The techniques described in this document may be embodied in various electronic devices such as a mobile phone, a laptop, a smartphone, or other devices capable of performing digital data processing and / or video display.
[0332] FIG. 26 is a flowchart relating to an example of a visual media processing method. The steps of the flowchart are described in relation to Embodiment 13 in this specification. In step 2602, the process determines the availability of a second video block of the visual media data using rules with respect to the conversion between a first video block of the visual media data and the bitstream representation of the visual media data. In step 2604, the process performs the conversion based on the determination, and the rules are at least based on the coding mode used to code the first video block into the bitstream representation, and the rules define that by treating the second video block as unavailable, the motion information of the second video block is prohibited from being used in the merge list construction of the first block.
[0333] FIG. 27 is a flowchart relating to an example of a visual media processing method. The steps of the flowchart are described in relation to Embodiment 14 in this specification. In step 2702, the process determines the availability of a second video block of the visual media data using rules with respect to the conversion between a first video block of the visual media data and the bitstream representation of the visual media data. In step 2704, the process performs the conversion based on the determination, and the rules define using an availability check process for the second video block at one or more positions of the visual media data.
[0334] Figure 28 is a flowchart relating to an example of a visual media processing method. The steps of the flowchart are described in connection with Embodiment 15 herein. In step 2802, the process determines two positions used to construct an intra-block copy motion list for the current video block of the visual media data with respect to the conversion between the current video block of the visual media data and the bitstream representation of the visual media data. In step 2804, the process performs the conversion based on the intra-block copy motion list.
[0335] Figure 29 is a flowchart relating to an example of a visual media processing method. The steps of the flowchart are described in connection with Embodiment 16 herein. In step 2902, the process determines the availability of adjacent blocks to derive one or more weights for combined intra-inter prediction of the current video block based on rules with respect to the conversion between the current video block of the visual media data and the coded representation of the visual media data. In step 2904, the process performs the conversion based on the determination, and the one or more weights include a first weight designated for inter prediction of the current video block and a second weight designated for intra prediction of the current video block, and the rules exclude using a comparison of the coding modes of the current video block and adjacent blocks.
[0336] Some embodiments of this specification are exemplarily presented here in a numbered list format.
[0337] A1. A visual media processing method, comprising:
[0338] determining, using rules, the availability of a second video block of the visual media data with respect to the conversion between the first video block of the visual media data and the bitstream representation of the visual media data;
[0339] performing a conversion based on a determination; and including, the rule being at least based on a coding mode used to code a first video block into a bitstream representation, the rule defining that by treating a second video block as unavailable, the motion information of the second video block is prohibited from being used in constructing a merge list of the first block.
[0340] A2. The method according to clause A1, wherein the coding mode used to code the first video block corresponds to inter-coding, and the coding mode for coding the second video block of the visual media data corresponds to intra-block copy.
[0341] A3. The method according to clause A1, wherein the coding mode used to code the first video block corresponds to intra-block copy, and the coding mode for coding the second video block of the visual media data corresponds to inter-coding.
[0342] A4. Performing the inspection includes:
[0343] receiving, as an input parameter for determining the availability of the second video block, the coding mode used to code the first video block, the method according to clause A1.
[0344] B1. A visual media processing method,
[0345] for the conversion between a first video block of visual media data and a bitstream representation of the visual media data, determining the availability of a second video block of the visual media data using a rule;
[0346] performing a conversion based on a determination; and the rule specifies using an availability check process for a second video block at one or more positions of the visual media data.
[0347] B2. The one or more positions correspond to the upper left position of the first video block, and the rule further specifies that when the coding mode of the first video block is not the same as the coding mode of the second video block, the second video block is treated as unavailable for conversion. The method according to clause B1.
[0348] B3. The one or more positions correspond to the upper left position of a shared merge area, and the rule further specifies not comparing the coding mode of the first video block with the coding mode of the second video block. The method according to clause B1.
[0349] B4. The rule further specifies inspecting whether a video area covering the one or more positions is within the same slice and / or tile and / or brick and / or subpicture as the first video block. The method according to clause B3.
[0350] C1. A visual media processing method,
[0351] determining two positions used to construct an intra-block copy motion list for a current video block of visual media data with respect to the conversion between the current video block of visual media data and the bitstream representation of the visual media data;
[0352] performing a conversion based on the intra-block copy motion list; and the method includes.
[0353] C2. The method according to clause C1, wherein the two positions include a first position corresponding to the upper left position of the shared merge region of the current video block.
[0354] C3. The method according to clause C2, wherein the first position is used to determine the availability of adjacent video blocks of the current video block.
[0355] C4. The method according to clause C1, wherein the two positions include a second position corresponding to the upper left position of the current video block.
[0356] C5. The method according to clause C4, wherein the second position is used to determine the availability of adjacent video blocks of the current video block.
[0357] D1. A visual media processing method, comprising:
[0358] determining the availability of adjacent blocks to derive one or more weights for the combined intra-inter prediction of the current video block based on rules regarding the conversion between the current video block of visual media data and the coded representation of the visual media data;
[0359] performing the conversion based on the determination;
[0360] wherein the one or more weights include a first weight specified for the inter prediction of the current video block and a second weight specified for the intra prediction of the current video block;
[0361] and the rules exclude using a comparison of the coding modes of the current video block and adjacent blocks.
[0362] The method according to clause D1, where adjacent blocks are determined to be available when the coding mode of adjacent blocks is an intra mode, an intra block copy (IBC) mode, or a palette mode.
[0363] The method according to D1 or D2, where the first weight is different from the second weight.
[0364] The method according to any one of clauses A1 - D3, where the conversion includes generating a bitstream representation from the current video block.
[0365] The method according to any one of claims A1 - D3, where the conversion includes generating samples of the current video block from a bitstream representation.
[0366] A video processing apparatus including a processor configured to execute the method according to any one or more of clauses A1 - D3.
[0367] A video encoding apparatus including a processor configured to execute the method according to any one or more of clauses A1 - D3.
[0368] A video decoding apparatus including a processor configured to execute the method according to any one or more of clauses A1 - D3.
[0369] A computer - readable medium storing code, which, when executed, causes a processor to execute the method according to any one or more of clauses A1 - D3.
[0370] The method, system, or apparatus described herein.
[0371] As used herein, the terms "video processing" or "video media processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation or vice versa. The bitstream representation of a current video block may correspond to bits that are located at the same location or are spread to different locations within the bitstream, as defined by the syntax, for example. For example, a macroblock may be encoded from the perspective of the transformed coded error residual values and also using the bits of the headers and other fields within the bitstream. Further, during the conversion, the decoder may be able to analyze the bitstream with the knowledge that some fields may or may not be present based on the decisions as explained in the above solutions. Similarly, the encoder may be able to generate a coded representation by determining whether a particular syntax field is included or not and, accordingly, including or excluding the coded representation from the syntax field.
[0372] The disclosed and other solutions, examples, embodiments, modules, and functional operations described herein can be implemented in digital electronic circuitry, or in computer software, firmware, hardware, or in combinations of one or more of them, including the structures disclosed herein and their structural equivalents. The disclosed and other embodiments can be implemented as one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a mechanically propagated signal, or a combination of one or more of them. The term "data processing apparatus" includes all apparatus, devices, and machines for processing data, e.g., programmable processors, computers, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer programs of interest, e.g., processor firmware, protocol stack, database management system, operating system, or combinations of one or more of them. The propagated signal is an artificially generated signal, e.g., an electrical, optical, or electromagnetic signal generated by a machine that encodes information for transmission to the appropriate receiving device.
[0373] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, either as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored within a portion of a file that holds other programs or data (e.g., one or more scripts saved in a markup language document), within a single file dedicated to the program in question, or within a plurality of coordinated files (e.g., files that store one or more modules, sub-programs, or portions of code). A computer program can be deployed to be executed on one computer or on a plurality of computers, and the plurality of computers can be located at one site or distributed across a plurality of sites and interconnected by a communication network.
[0374] The processes and logic flows described herein can be executed by one or more programmable processors executing one or more computer programs, which can perform functions by acting on input data to generate output. The processes and logic flows can also be executed by special-purpose logic circuits, such as, for example, FPGAs (field-programmable gate arrays) or ASICs (application-specific integrated circuits), and can also be implemented as such devices.
[0375] Processors suitable for the execution of a computer program include, for example, microprocessors both general and special purpose, and any one or more processors of any kind of digital computer. In general, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical disks, or optical disks, or be operatively coupled to receive data from, transfer data to, or both, such devices. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks, including any form of nonvolatile memory, media, and memory devices. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0376] Although this specification contains many details, these should not be construed as limitations on any subject matter or as claims to what can be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular technology. The specific plurality of features described herein in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in a plurality of embodiments. Furthermore, although a feature may have been described above as acting in a particular combination or even initially claimed as such, one or more of the features in the claimed combination may, in some cases, be excised from the combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.
[0377] Similarly, in the figures, operations are shown in a particular order, but this should not be understood as requiring that such operations be performed in the particular order shown or in sequence in order to achieve the desired result, or that all of the illustrated operations be performed. Furthermore, the way in which the various system components are divided in the embodiments described in this patent document should not be understood as requiring such a division in all embodiments.
[0378] Only a few implementation examples and embodiments are described, and other implementations, extensions, and modifications can be made based on what is described and illustrated in this patent document.
[0379] (Appendix 1) A visual media processing method, Regarding the conversion between a first video block of visual media data and the bitstream representation of the visual media data, a step of determining the availability of a second video block of the visual media data using rules; a step of performing the conversion based on the determination and including, the rule being at least based on a coding mode used to code the first video block into the bitstream representation, the rule defining that by treating the second video block as unavailable, the motion information of the second video block is prohibited from being used in constructing the merge list of the first block, a method. (Appendix 2) The coding mode used to code the first video block corresponds to inter-coding, and the coding mode for coding the second video block of the visual media data corresponds to intra block copy, the method according to Appendix 1. (Appendix 3) The coding mode used to code the first video block corresponds to intra block copy, and the coding mode for coding the second video block of the visual media data corresponds to inter-coding, the method according to Appendix 1. (Appendix 4) Performing an inspection of receiving the coding mode used to code the first video block as an input parameter for determining the availability of the second video block, the method according to Appendix 1. (Appendix 5) A visual media processing method, regarding the conversion between a first video block of visual media data and the bitstream representation of the visual media data, a step of determining the availability of a second video block of the visual media data using a rule, a step of performing the conversion based on the determination and including, the rule defining that a availability inspection process is used for the second video block at one or more positions of the visual media data, a method. (Appendix 6) The above one or more positions correspond to the upper left position of the first video block, and the rule further stipulates that when the coding mode of the first video block is not the same as the coding mode of the second video block, the second video block is treated as unavailable for the conversion, the method described in Appendix 5. (Appendix 7) The above one or more positions correspond to the upper left position of the shared merge area, and the rule further stipulates that the coding mode of the first video block is not compared with the coding mode of the second video block, the method described in Appendix 5. (Appendix 8) The rule further stipulates that it inspects whether the video area covering the above one or more positions is within the same slice and / or tile and / or brick and / or sub-picture as the first video block, the method described in Appendix 7. (Appendix 9) A visual media processing method, Regarding the conversion between the current video block of visual media data and the bitstream representation of the visual media data, a step of determining two positions used to construct the intra-block copy motion list of the current video block; A step of performing the conversion based on the intra-block copy motion list A method including. (Appendix 10) The above two positions include a first position corresponding to the upper left position of the shared merge area of the current video block, the method described in Appendix 9. (Appendix 11) The first position is used to determine the availability of adjacent video blocks of the current video block, the method described in Appendix 10. (Appendix 12) The method according to appendix 9, wherein the two positions include a second position corresponding to the upper left position of the current video block. (Appendix 13) The method according to appendix 12, wherein the second position is used to determine the availability of an adjacent video block of the current video block. (Appendix 14) A visual media processing method, comprising: determining the availability of adjacent blocks to derive one or more weights for combined intra-inter prediction of the current video block based on rules for conversion between a current video block of visual media data and a coded representation of the visual media data; executing the conversion based on the determination; wherein the one or more weights include a first weight designated for inter prediction of the current video block and a second weight designated for intra prediction of the current video block; a method, wherein the rules exclude using a comparison of coding modes of the current video block and the adjacent blocks. (Appendix 15) The method according to appendix 14, wherein the adjacent block is determined to be available when the coding mode of the adjacent block is an intra mode, an intra block copy (IBC) mode, or a palette mode. (Appendix 16) The method according to appendix 14 or 15, wherein the first weight is different from the second weight. (Appendix 17) The method according to any one of appendices 1-16, wherein the conversion includes generating the bitstream representation from the current video block. (Appendix 18) The method according to any one of appendices 1-16, wherein the conversion includes generating samples of the current video block from the bitstream representation. (Appended Note 19) A video processing apparatus including a processor configured to execute the method according to any one of Appended Notes 1-16. (Appended Note 20) A video encoding apparatus including a processor configured to execute the method according to any one of Appended Notes 1-16. (Appended Note 21) A video decoding apparatus including a processor configured to execute the method according to any one of Appended Notes 1-16. (Appended Note 22) A computer-readable medium storing code, which, when executed, causes a processor to execute the method according to any one of Appended Notes 1-16. (Appended Note 23) The method, system or apparatus described herein.
Claims
Claim 1 A method for processing video data, Regarding the conversion between a first block of video and the bitstream of said video, performing a first determination about the availability of adjacent blocks for deriving one or more weights for predicting said first block by an availability check process based on rules, wherein said first block is coded in a first prediction mode, and in said first prediction mode, the prediction of said first block is generated based at least on inter prediction and intra prediction, a step; Executing said conversion based on said first determination; Including, said one or more weights include a first weight designated for the inter prediction of said first block and a second weight designated for the intra prediction of said first block, Said rules exclude using a comparison between the coding mode of said first block and the coding mode of said adjacent blocks, the value of the input parameter of said availability check process is set to false, and said value of the input parameter being false indicates that said availability check process does not depend on the coding mode of said adjacent blocks, a method. Claim 2 The method according to claim 1, wherein said adjacent blocks are allowed to be determined as available even when the coding mode of said adjacent blocks is an intra mode, an intra block copy mode, or a palette mode, In said intra block copy mode, the prediction samples are derived from a block of sample values of the decoded video region that is the same as determined by a block vector, and When said adjacent blocks are available, said one or more weights are designated depending on the coding mode of said adjacent blocks, a method. Claim 3 The method according to claim 1 or 2, further, Regarding the conversion between the second block of the video and the bitstream of the video, a step of making a second determination about the availability of the third block of the video based on the coding mode of the second block and the coding mode of the third block by the availability check process, wherein the third block is determined to be unavailable in response to the coding mode of the second block being different from the coding mode of the third block, the step and Based on the second determination, a step of performing the conversion between the second block and the bitstream Including, the motion information of the third block is prohibited from being used in the list construction of the second block in response to the third block being determined to be unavailable, the method.
4. In the method according to claim 3, the coding mode of the second block corresponds to an inter prediction mode, the coding mode of the third block corresponds to an intra block copy mode, and in the intra block copy mode, the prediction sample is derived from a block of sample values of the decoded video region that is the same as that determined by the block vector, or The coding mode of the second block corresponds to an intra block copy mode, the coding mode of the third block corresponds to an inter prediction mode, and in the intra block copy mode, the prediction sample is derived from a block of sample values of the decoded video region that is the same as that determined by the block vector, the method.
5. The method according to any one of claims 1 to 4, further Regarding the conversion between the fourth block of the video and the bitstream of the video, the fourth block is coded in a geometric partitioning mode which is a second prediction mode, the step and A step of constructing a candidate list for the fourth block, the constructing step including a step of checking the availability of spatial candidates in a specific adjacent block B2 based on the number of available candidates in the candidate list, the specific adjacent block B2 being at the upper left corner of the fourth block, the step and Determining the first motion information for the first geometric partition of the fourth block and the second motion information for the second geometric partition of the fourth block based on the candidate list; Applying a weighting process to generate a final prediction for the samples of the fourth block based on weighted addition of predicted samples derived based on the first motion information and the second motion information; Executing the transformation based on the applying step; A method comprising.
6. In the method according to claim 5, when the number of available candidates in the candidate list is 4, the spatial candidate in the specific adjacent block B2 is unavailable. The spatial candidate corresponding to the adjacent block coded in the intra-block copy mode is excluded from the candidate list for the fourth block. In the intra-block copy mode, the predicted samples of the adjacent block are derived from the block of sample values of the decoded video region that are the same as those determined by the block vector. A method.
7. In the method according to claim 5 or 6, the candidate list is constructed based on a pruning process, and the pruning process includes comparing motion information between at least two spatial candidates to avoid the same spatial candidate. The step of constructing the candidate list for the fourth block based on the pruning process includes checking the availability of the spatial candidate in the specific adjacent block A1. The specific adjacent block A1 is adjacent to the lower left corner of the fourth block. The checking step includes: when the candidate list includes the spatial candidate in the specific adjacent block B1, and the specific adjacent block B1 is adjacent to the upper right corner of the fourth block, comparing the motion information between the spatial candidates in the specific adjacent block A1 and the specific adjacent block B1; and based on the result of the comparison indicating that the motion information of the spatial candidate in A1 matches that in B1, determining that the spatial candidate in the specific adjacent block A1 is unavailable. The step of constructing the candidate list of the fourth block based on the pruning process includes the step of examining the availability of spatial candidates in a specific adjacent block B0, where the specific adjacent block B0 is at the upper right corner of the fourth block, and the examining step includes: when the candidate list includes spatial candidates in a specific adjacent block B1, and the specific adjacent block B1 is adjacent to the upper right corner of the fourth block, comparing the motion information between the spatial candidates in the specific adjacent block B0 and the specific adjacent block B1; and based on the result of the comparison indicating that the motion information of the spatial candidate in B0 matches that in B1, determining that the spatial candidate in the specific adjacent block B0 is unavailable; or The step of constructing the candidate list of the fourth block based on the pruning process includes the step of examining the availability of spatial candidates in a specific adjacent block A0, where the specific adjacent block A0 is at the lower left corner of the fourth block, and the examining step includes: when the candidate list includes spatial candidates in a specific adjacent block A1, and the specific adjacent block A1 is adjacent to the lower left corner of the fourth block, comparing the motion information between the spatial candidates in the specific adjacent block A0 and the specific adjacent block A1; and based on the result of the comparison indicating that the motion information of the spatial candidate in A0 matches that in A1, determining that the spatial candidate in the specific adjacent block A0 is unavailable, a method. Claim 8 In the method according to claim 7, the step of examining the availability of spatial candidates in the specific adjacent block B2 further includes: when the candidate list includes spatial candidates in specific adjacent blocks A1 and B1, the specific adjacent block B1 is adjacent to the upper right corner of the fourth block, and the specific adjacent block A1 is adjacent to the lower left corner of the fourth block, comparing the motion information between the spatial candidates in the specific adjacent block B2 and the specific adjacent blocks A1 and B1; Based on the result of the comparison indicating that the motion information of the spatial candidate in B2 matches that in A1 or B1, determining that the spatial candidate in the specific adjacent block B2 is unavailable; A method comprising. **Claim 9** A method according to any one of claims 1 to 8, further comprising: Regarding the conversion between the fifth block of the video and the bitstream of the video, performing a fourth determination that the fifth block is coded in a third prediction mode, wherein in the third prediction mode, the prediction samples are derived from a block of sample values of the same decoded video region determined by a block vector; Determining a motion candidate list construction process for the motion candidate list of the fifth block based on whether conditions regarding the width and height of the fifth block are satisfied, wherein the conditions are satisfied when the product of the width and height of the fifth block is less than or equal to a threshold; Executing the conversion based on the motion candidate list; A method, wherein when the conditions are satisfied, the derivation of the spatial merge candidates in the motion candidate list construction process is skipped. **Claim 10** In the method according to claim 9, the motion candidate list construction process comprises: Adding at least one spatial candidate to the motion candidate list; Adding at least one history-based motion vector predictor candidate to the motion candidate list, or Adding at least one zero candidate to the motion candidate list; Including at least one of the above, and the step of adding at least one spatial candidate to the motion candidate list comprises: Checking the availability of the spatial candidate in a specific adjacent block A1, wherein the specific adjacent block A1 is adjacent to the lower left corner of the fourth block; Adding the spatial candidate in the specific adjacent block A1 to the motion candidate list in response to the specific adjacent block A1 being available; Checking the availability of the spatial candidate in a specific adjacent block B1, wherein the specific adjacent block B1 is adjacent to the upper right corner of the fifth block; A method, including, in response to the specific adjacent block B1 being available, performing a first redundancy check, and ensuring that a spatial candidate in the specific adjacent block B1 having the same motion information as the spatial candidate in the specific adjacent block A1 is excluded from the motion candidate list.
11. In the method according to claim 10, the motion candidate list construction process includes: after adding the at least one spatial candidate, when the size of the motion candidate list is smaller than the maximum allowable list size of the third prediction mode, adding at least one history-based motion vector predictor candidate to the motion candidate list, performing a second redundancy check to ensure that candidates having the same motion information are excluded from the motion candidate list, which is applied when adding the at least one history-based motion vector predictor candidate, or the motion candidate list construction process includes: in response to the size of the motion candidate list being smaller than the maximum allowable list size of the third prediction mode, adding the at least one zero candidate to the motion candidate list.
12. A method according to any one of claims 1 to 9, further comprising performing a fifth determination regarding the conversion between the seventh block of the video and the bitstream of the video, the seventh block being coded in a fourth prediction mode, wherein in the fourth prediction mode, the signaled motion vector predictor is refined based on an offset explicitly signaled in the bitstream, determining a second motion candidate list construction process for the second motion candidate list of the seventh block based on whether conditions regarding the width and height of the seventh block are satisfied, the conditions regarding the width and height of the seventh block being satisfied when the product of the width and height of the seventh block is less than or equal to a second threshold, performing the conversion based on the second motion candidate list and including.
13. In the method according to any one of claims 1 to 12, the conversion includes encoding the video into the bitstream.
14. The method according to any one of claims 1 to 12, wherein the conversion includes decoding the video from the bitstream.
15. An apparatus for processing video data, comprising a non-transitory memory having instructions and a processor, the instructions, when executed by the processor, cause the processor to: Regarding the conversion between a first block of video and the bitstream of the video, perform a first determination regarding the availability of adjacent blocks for deriving one or more weights for predicting the first block by an availability check process based on rules, wherein the first block is coded in a first prediction mode, and in the first prediction mode, the prediction of the first block is generated based at least on inter prediction and intra prediction; Execute the conversion based on the first determination; wherein the one or more weights include a first weight designated for inter prediction of the first block and a second weight designated for intra prediction of the first block; the rules exclude using a comparison between the coding mode of the first block and the coding mode of the adjacent blocks, the value of the input parameter of the availability check process is set to false, and the false value of the input parameter indicates that the availability check process does not depend on the coding mode of the adjacent blocks.
16. A non-transitory computer-readable storage medium storing instructions, the instructions cause a processor to: Regarding the conversion between a first block of video and the bitstream of the video, perform a first determination regarding the availability of adjacent blocks for deriving one or more weights for predicting the first block by an availability check process based on rules, wherein the first block is coded in a first prediction mode, and in the first prediction mode, the prediction of the first block is generated based at least on inter prediction and intra prediction; Execute the conversion based on the first determination; Causing execution, the one or more weights include a first weight designated for inter prediction of the first block and a second weight designated for intra prediction of the first block. The rule excludes using a comparison between the coding mode of the first block and the coding mode of the adjacent block. The value of the input parameter of the availability check process is set to false, and the false value of the input parameter indicates that the availability check process does not depend on the coding mode of the adjacent block. A storage medium.
17. A method for storing a video bitstream, comprising: Performing a first determination on the availability of adjacent blocks for deriving one or more weights for predicting a first block of a video by an availability check process based on a rule, wherein the first block is coded in a first prediction mode, and in the first prediction mode, the prediction of the first block is generated based at least on inter prediction and intra prediction; Generating the video bitstream based on the first determination; Storing the bitstream in a non-transitory computer-readable storage medium. The one or more weights include a first weight designated for inter prediction of the first block and a second weight designated for intra prediction of the first block. The rule excludes using a comparison between the coding mode of the first block and the coding mode of the adjacent block. The value of the input parameter of the availability check process is set to false, and the false value of the input parameter indicates that the availability check process does not depend on the coding mode of the adjacent block. A method.
Citation Information
Patent Citations
JPP7346599B