Adjustment Method for Inter-Prediction Based on Sub-Blocks
Sub-block-based temporal motion vector prediction and inter-prediction techniques enhance video encoding efficiency, addressing bandwidth challenges in digital video communication by optimizing motion vector prediction and encoding processes.
Patent Information
- Application Number
- JP2023217004
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-09-22
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2039-11-22
AI Technical Summary
Digital video occupies the largest bandwidth in Internet and digital communication networks, and the increasing number of connected devices exacerbates this bandwidth demand, necessitating more efficient video encoding methods.
The use of sub-block-based temporal motion vector prediction (SbTMVP) and inter-prediction methods in video encoding, including techniques like affine motion compensation and triangular prediction, to enhance encoding efficiency and reduce bandwidth requirements.
These methods improve the quality of decompressed video and reduce bandwidth demands by optimizing motion vector prediction and encoding processes, making them suitable for existing and future video encoding standards.
Smart Images

Figure 0007698030000027 
Figure 0007698030000028 
Figure 0007698030000029
Abstract
Description
Technical Field
[0001] (Cross - reference to Related Applications) This This application claims , a divisional application based on Japanese Patent Application No. 2023-117792, which is a divisional application based on Japanese Patent Application No. 2021-526695, which is based on International Patent Application No. PCT / CN2019 / 120301 filed on November 22, 2019, and this international patent application the priority and benefit of International Patent Application No. PCT / CN2018 / 116889 filed on November 22, 2018, International Patent Application No. PCT / CN2018 / 125420 filed on December 29, 2018, International Patent Application No. PCT / CN2019 / 100396 filed on August 13, 2019, and International Patent Application No. PCT / CN2019 / 107159 filed on September 22, 2019 is the main claims one. All the above patents The entire disclosure of the above applications is incorporated by reference as part of the disclosure of this specification
[0002] This specification relates to image and video encoding and decoding
Background Art
[0003] Despite the progress of video compression, digital video still occupies the largest bandwidth usage in the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for the use of digital video is predicted to continue to increase
Summary of the Invention
[0004] Devices, systems, and methods related to digital video encoding, including an inter - prediction method based on sub - blocks, are described. The described method can be applied to both existing video encoding standards (e.g., High - Efficiency Video Coding (HEVC) and / or Versatile Video Coding (VVC)) and future video encoding standards or video codecs
[0005] In one representative aspect, the disclosed technology may be used to provide a method for video processing. This method is based on the maximum number of candidates (ML) in the merge candidate list based on sub-blocks and / or the sub-block-based temporal motion vector prediction (SbTMVP) candidate for the conversion between the current block of the video and the bitstream representation of the video. It includes determining whether to add them to the merge candidate list based on sub-blocks according to whether the temporal motion vector prediction (TMVP) is valid during conversion or whether the current picture reference (CPR) coding mode is used for conversion, and performing the conversion based on this determination. For the conversion between the current block of the video and the bitstream representation of the video, the maximum number of candidates (ML) in the merge candidate list based on sub-blocks and / or the sub-block-based temporal motion vector prediction (SbTMVP) candidate is determined based on whether the temporal motion vector prediction (TMVP) is valid during conversion or whether the current picture reference (CPR) coding mode is used for conversion. Whether to add it to the merge candidate list based on sub-blocks is determined, and the conversion is performed based on this determination. This includes determining whether to add to the merge candidate list based on sub-blocks according to whether the temporal motion vector prediction (TMVP) is valid during conversion or whether the current picture reference (CPR) coding mode is used for conversion, and performing the conversion based on this determination. This includes determining whether to add to the merge candidate list based on sub-blocks according to whether the temporal motion vector prediction (TMVP) is valid during conversion or whether the current picture reference (CPR) coding mode is used for conversion, and performing the conversion based on this determination. This includes determining whether to add to the merge candidate list based on sub-blocks according to whether the temporal motion vector prediction (TMVP) is valid during conversion or whether the current picture reference (CPR) coding mode is used for conversion, and performing the conversion based on this determination.
[0006] In another representative aspect, the disclosed technology may be used to provide a method for video processing. This method is for the conversion between the current block of the video and the bitstream representation of the video. Based on whether the temporal motion vector prediction (TMVP), the sub-block-based temporal motion vector prediction (SbTMVP) tool, and the affine coding mode are valid for use during conversion, the maximum number (ML) of candidates in the merge candidate list based on sub-blocks is determined, and the conversion is performed based on this determination. This includes determining the maximum number (ML) of candidates in the merge candidate list based on sub-blocks according to whether the temporal motion vector prediction (TMVP), the sub-block-based temporal motion vector prediction (SbTMVP) tool, and the affine coding mode are valid for use during conversion, and performing the conversion based on this determination. This includes determining the maximum number (ML) of candidates in the merge candidate list based on sub-blocks according to whether the temporal motion vector prediction (TMVP), the sub-block-based temporal motion vector prediction (SbTMVP) tool, and the affine coding mode are valid for use during conversion, and performing the conversion based on this determination.
[0007] In another representative aspect, the disclosed technology may be used to provide a method for video processing. This method is for the conversion between the current block of the first video segment of the video and the bitstream representation of the video. For the conversion between the current block of the first video segment of the video and the bitstream representation of the video, the sub-block-based motion vector prediction (SbTMVP) mode is based on whether the temporal motion vector prediction (TMVP) mode is valid for the first video segment. For the conversion between the current block of the first video segment of the video and the bitstream representation of the video, the sub-block-based motion vector prediction (SbTMVP) mode is based on whether the temporal motion vector prediction (TMVP) mode is valid for the first video segment. Determine that it is disabled for conversion because it is disabled for conversion by the bell and, based on that determination, perform a conversion, wherein the bitstream representation is Sb whether to include the display of the TMVP mode and / or the position of the display of the SbTMVP mode relative to the display of the TM VP mode in the merge candidate list, and perform a conversion based on the format that defines the position of the display of the SbTMVP mode relative to the display of the TMVP mode in the merge candidate list including.
[0008] In another representative aspect, the disclosed technology may be used to provide a method for video processing This method includes performing a conversion between the current block of the video encoded using the sub-block based temporal motion vector prediction (SbTMVP) tool or the temporal motion vector prediction (TMVP) tool and the bitstream representation of this video, and using a mask to selectively mask the coordinates of the current block or the corresponding position of this current block based on the compression of the motion vectors associated with the SbTMVP tool or the TMVP tool, and applying this mask includes performing a bitwise AND operation between the value of this coordinate and the value of this mask including performing a conversion between the current block of the video and the bitstream representation of this video and applying this mask includes performing a bitwise AND operation between the value of this coordinate and the value of this mask selectively mask the coordinates of the current block or the corresponding position of this current block based on the compression of the motion vectors associated with the SbTMVP tool or the TMVP tool selectively mask the coordinates of the current block or the corresponding position of this current block based on the compression of the motion vectors associated with the SbTMVP tool or the TMVP tool including performing a bitwise AND operation between the value of this coordinate and the value of this mask
[0009] In another representative aspect, the disclosed technology may be used to provide a method for video processing This method includes determining a valid corresponding region of the current block for applying the sub-block based motion vector prediction (SbTMVP ) tool based on one or more characteristics of the current block of the video segment of the video, and based on this determination, performing a conversion between this current block and the bitstream representation of this video ) tool based on one or more characteristics of the current block of the video segment of the video, and based on this determination, performing a conversion between this current block and the bitstream representation of this video including performing a conversion between this current block and the bitstream representation of this video including.
[0010] In another representative aspect, the disclosed technology may be used to provide a method for video processing. This method uses a sub-block based temporal motion vector prediction (SbTMVP) tool to determine a default motion vector for the current block of the video being encoded, and based on this determination, perform a conversion between the current block and the bitstream representation of this video, including a block including the corresponding position in the collocated picture associated with the center position of the current block, where the default motion vector is determined when a motion vector cannot be obtained from the block. For the current block of the video being encoded, determine a default motion vector using a sub-block based temporal motion vector prediction (SbTMVP) tool. Based on this determination, perform a conversion between the current block and the bitstream representation of this video. The default motion vector is determined when a motion vector cannot be obtained from the block including the corresponding position in the collocated picture associated with the center position of the current block. In another representative aspect, the disclosed technology may be used to provide a method for video processing. This method, for the current block of a video segment of a video, infers that the current picture of the current block is a reference picture with an index set to M in the reference picture list X, where M and X are integers and X = 0 or X = 1, and for the video segment, the sub-block based temporal motion vector prediction (SbTMVP) tool or the temporal motion vector prediction (TMVP) tool is disabled, and based on this inference, perform a conversion between the current block and the bitstream representation of the video.
[0011] In another representative aspect, the disclosed technology may be used to provide a method for video processing. This method, for the current block of a video, determines that the current picture of the current block is a reference picture with an index set to M in the reference picture list X, where M and X are integers and X = 0 or X = 1. When the current picture of the current block is a reference picture with an index set to M in the reference picture list X, where M and X are integers and X = 0 or X = 1, infer that the sub-block based temporal motion vector prediction (SbTMVP) tool or the temporal motion vector prediction (TMVP) tool is disabled for the video segment. Based on this inference, perform a conversion between the current block and the bitstream representation of the video. In another representative aspect, the disclosed technology may be used to provide a method for video processing. This method infers that for the current block of a video, if the current picture of the current block is a reference picture with an index set to M in the reference picture list X, where M and X are integers and X = 0 or X = 1, the sub-block based temporal motion vector prediction (SbTMVP) tool or the temporal motion vector prediction (TMVP) tool is disabled for the video segment. Based on this inference, perform a conversion between the current block and the bitstream representation of the video. In another representative aspect, the disclosed technology may be used to provide a method for video processing.
[0012] In another representative aspect, the disclosed technology may be used to provide a method for video processing. This method, for the current block of a video, determines that the current picture of the current block is a reference picture with an index set to M in the reference picture list X, where M and X are integers and X = 0 or X = 1. If the current picture of the current block is a reference picture with an index set to M in the reference picture list X, where M and X are integers and X = 0 or X = 1. When M and X are integers, determine that the application of the sub-block based temporal motion vector prediction (SbTM VP) tool is enabled, and based on this determination, perform a conversion between the current block and the bitstream representation of the video.
[0013] In another representative aspect, the disclosed technology may be used to provide a method for video processing. This method includes performing a conversion between the current block of the video and the bitstream representation of the video, where the current block is encoded using an encoding tool based on sub-blocks, and performing this conversion includes encoding a sub-block merge index in a unified manner using a plurality of bins (N) when the sub-block based temporal motion vector prediction (Sb TMVP) tool is enabled or disabled. TMVP) tool is enabled or disabled. TMVP) tool is enabled or disabled. block merge index in a unified manner using a plurality of bins (N).
[0014] In another representative aspect, the disclosed technology may be used to provide a method for video processing. This method includes determining a motion vector used to locate the position of a corresponding block in a picture different from the current picture including the current block for the current block of the video encoded using the sub-block based temporal motion vector prediction (SbTMVP) tool, and based on this determination, performing a conversion between the current block and the bitstream representation of the video. block and the bitstream representation of the video. block and the bitstream representation of the video. block and the bitstream representation of the video.
[0015] In another representative aspect, the disclosed technology may be used to provide a method for video processing. This method includes, for the conversion between the current block of the video and the bitstream representation of the video, based on whether affine prediction is enabled for the conversion of the current block, whether affine prediction is enabled for the conversion of the current block, to determine whether to insert a zero-motion affine merge candidate into the sub-block merge candidate list and to perform a conversion based on this determination.
[0016] In another representative aspect, the disclosed technology may be used to provide a method for video processing This method uses the current block of the video and the sub-block merge candidate list for conversion between the video and the bitstream representation of the video. When the sub-block merge candidate list is not satisfied, it includes inserting a zero-motion non-affine padding candidate into the sub-block merge candidate list and performing a conversion after the insertion.
[0017] In another representative aspect, the disclosed technology may be used to provide a method for video processing This method uses rules for determining that the motion vector is derived from one or more motion vectors of a block that includes the corresponding position in a collocated picture for conversion between the current block of the video and the bitstream representation of the video, and includes determining the motion vector and performing this conversion based on the motion vector.
[0018] In yet another exemplary aspect, a video encoder device is disclosed. This video encoder device includes a processing device configured to implement the methods described herein .
[0019] In yet another exemplary aspect, a video decoder device is disclosed. This video decoder device includes a processing device configured to implement the methods described herein.
[0020] In yet another aspect, a computer-readable medium storing code is disclosed. When this code is executed by a processing device, it causes the processing device to implement the methods described herein.
[0021] These and other aspects are described herein.
Brief Description of the Drawings
[0022]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13A
Figure 13B
Figure 14
Figure 15
Figure 16A
Figure 16B
Figure 17
Figure 18A
Figure 18B
Figure 19
Figure 20
Figure 21A
Figure 21B
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29A
Figure 29B
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Figure 40
Figure 41
Figure 42
Figure 43
Figure 44
[0023] The present specification provides a method for improving the quality of decompressed or decoded digital video or images. In addition, the present invention provides various techniques that can be used by a decoder of a video bitstream. The video encoder reconstructs the decoded frames to be used for further encoding. Alternatively, these techniques may be implemented during the encoding process.
[0024] Although section headings are used herein for ease of understanding, the embodiments and techniques Thus, embodiments from one chapter may be used interchangeably with other chapters. This can be combined with the examples from chapter .
[0025] 1. Overview
[0026] This patent specification relates to video coding technology. Specifically, the present invention relates to The present invention relates to motion vector coding that is compatible with existing video coding standards such as HEVC or The present invention can be applied to standards to be finalized (e.g., general purpose video coding). It is also applicable to future video coding standards or video codecs.
[0027] 2. Introduction
[0028] Video coding standards have mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T has developed H.261 and H.263, ISO / IEC has developed MPEG-1 and MPEG-4 Visual, and both organizations have jointly created H.262 / MPEG-2 Video and H.26 4 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. The video coding standard, H.262, is based on a hybrid video coding structure that utilizes temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, in 2015, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET). Since then, many new methods have been adopted by JVET and incorporated into a reference software called the Joint Exploration Model (JEM). In April 2018, a Joint Video Exploration Team (JVET) was formed between VCEG (Q6 / 16) and ISO / IEC JTC_1 SC29 / WG11 (MPEG), and they aimed to work on the VVC standard that reduces the bitrate by 50% compared to HEVC. In 2015, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET). Since then, many new methods have been adopted by JVET and incorporated into a reference software called the Joint Exploration Model (JEM). In April 2018, a Joint Video Exploration Team (JVET) was formed between VCEG (Q6 / 16) and ISO / IEC JTC_1 SC29 / WG11 (MPEG), and they aimed to work on the VVC standard that reduces the bitrate by 50% compared to HEVC. JTC_1 SC29 / WG11 (MPEG) to work on the VVC standard with the goal of reducing the bitrate by 50% compared to HEVC. The latest version of the VVC draft, namely, General Video Coding (Draft 3), can be referred to as follows. The latest version of the VVC draft, namely, General Video Coding (Draft 3), can be referred to as follows.
[0029] The latest version of the VVC draft, i.e., General Video Coding (Draft 3), can be referred to as follows. It can be referred to as follows.
[0030] http: / / phenix.it-sudparis.eu / jvet / doc_e nd_user / documents / 12_Macao / wg11 / JVET-L10 01-v2.zip The latest reference software for VVC, called VTM, is See below.
[0031] https: / / vcgit.hhi.fraunhofer.de / jvet / VV CSoftware_VTM / tags / VTM-3.0rC1
[0032] 2.1 Inter Prediction in HEVC / H.265
[0033] Each inter-predicted PU has motion parameters for one or two reference picture lists. The motion parameters include a motion vector and a reference picture index. The use of one of the two reference picture lists is inter _pred_idc. The motion vector for the predictor may be signaled using It may be explicitly coded as a delta.
[0034] If one CU is coded in skip mode, one PU is associated with this CU. There are no significant residual coefficients, no coded motion vector deltas, and no reference picture indexes. Specify the merge mode, which merges the motion parameters for the current PU into the spatial The merge mode is obtained from the neighboring PUs including the skip mode and the neighboring PUs including the temporal candidates. The merge mode can be applied to any inter-predicted PU, not just for the An alternative is to explicitly transmit the motion parameters, per PU, in each reference picture list. and a motion vector, which is a reference picture index corresponding to the use of a reference picture list. The motion vector difference compared to the motion vector predictor (MVD) is Surely signal notification is performed. Such a mode is referred to as Advanced Motion Vector Prediction (AMVP) in this disclosure. It is called.
[0035] When signal notification indicates using one of the two reference picture lists, a PU is generated from a block of one sample. This is called "single prediction". Single prediction is available for both P slices and B slices.
[0036] When signal notification indicates using both reference picture lists, a PU is generated from a block of two samples. This is called "dual prediction". Dual prediction is only available for B slices and is possible.
[0037] Hereinafter, the inter prediction mode defined in HEVC will be described in detail. First, the merge mode will be described.
[0038] 2.1.1 Reference Picture Lists
[0039] In HEVC, the term inter prediction is used to indicate a prediction derived from data elements (e.g., sample values or motion vectors) of reference pictures other than the currently decoded picture. Similar to H.264 / AVC, one picture can be predicted from multiple reference pictures. The reference pictures used for inter prediction are grouped into one or more reference picture lists. The reference index identifies which reference picture in the list is used to generate the prediction signal. One reference picture list, List0, is used for P slices, and two reference picture lists, List0 and List1, are used for B slices. Note that List0 / 1 contains
[0040] One reference picture list, List0, is used for P slices, and two reference picture lists, List0 and List1, are used for B slices. Note that List0 / 1 contains The reference pictures to be fetched may be from past and future pictures, regardless of the shooting / display order. It may be present.
[0041] 2.1.2 Merge Mode
[0042] 2.1.2.1 Derivation of Merge Mode Candidates
[0043] When predicting the PU using the merge mode, parse the index pointing to the entry in the merge candidate list from the bitstream, and use this to search for motion information. The construction of this list is defined in the HEVC standard and can be summarized based on the following sequence of steps. It can be summarized based on this.
[0044] · Step 1: Initial Candidate Derivation o Step 1.1: Spatial Candidate Derivation o Step 1.2: Redundancy Check of Spatial Candidates o Step 1.3: Temporal Candidate Derivation · Step 2: Additional Candidate Insertion o Step 2.1: Creation of Bi-Prediction Candidates o Step 2.2: Insertion of Motion Zero Candidates
[0045] These steps are also schematically shown in Figure 1. For spatial merge candidate derivation, select up to 4 merge candidates from among the candidates at 5 different positions. For temporal merge candidate derivation, select up to 1 merge candidate from among 2 candidates. On the decoder side, assuming a fixed number of candidates for each PU, if the number of candidates obtained in Step 1 does not reach the maximum number of merge candidates (MaxNumMergeCand ) signaled in the slice header, generate additional candidates. Since the number of candidates is fixed, the shortened unary ) is used. Encode the index of the best merge candidate using binary quantization (TU). When the size of the CU is equal to 8, all PUs of the current CU share the same one merge candidate list as the merge candidate list of the 2N×2N prediction unit.
[0046] The operations associated with the above steps are described in detail below.
[0047] 2.1.2.2 Spatial Candidate Derivation
[0048] In the derivation of spatial merge candidates, select up to four merge candidates from the candidates at the positions shown in Figure 2. The derivation order is A1, B1, B0, A0, B2. If any of the PUs at positions A1, B1, B0, A0 are not available (e.g., because they belong to another slice or tile), or if they are intra-coded, then position B2 is considered only. After adding the candidate at position A1, adding the remaining candidates is subject to a redundancy check, which can ensure that candidates with the same motion information are excluded from the list, improving the coding efficiency. To reduce the computational complexity, not all possible candidate pairs are considered in the above redundancy check. Instead, only the pairs connected by the arrows in Figure 3 are considered, and the candidate is added to the list only if the corresponding candidates used in the redundancy check do not have the same motion information. Another source of duplicate motion information is the "second PU" associated with a different partition than 2N×2N. As an example, Figures 4A and 4B show the second PU for the cases of N×2N and 2N×N respectively. When the current PU is partitioned into N×2N, the candidate at position A1 is not considered for list construction. In fact, adding this candidate would As a result, the dual prediction units have the same motion information, and it is redundant for one coding unit to have only one PU. Similarly, when the current PU is divided into 2N×N, position B1 is not considered.
[0049] 2.1.2.3 Temporal candidate derivation
[0050] In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal merge candidate, based on the same-position PU belonging to the picture with the minimum POC difference from the current picture in the given reference picture list, a scaled motion vector is derived. In the slice header, the reference picture list used for the derivation of the same-position PU is clearly signaled. As shown by the dotted line in FIG. 5, the scaled motion vector of the temporal merge candidate is obtained. This is scaled from the motion vector of the same-position PU by using the POC distances tb and td. tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the same-position PU and the same-position picture. The reference picture index of the temporal merge candidate is set equal to zero. The actual implementation of this scaling process is described in the HEVC specification. In the case of a B slice, two motion vectors, i.e., one for reference picture list 0 and the other for reference picture list 1, are obtained, and the dual prediction merge candidate is formed by combining these.
[0051]
[0052] That is, one is for reference picture list 0 and the other is for reference picture list 1, and the dual prediction merge candidate is formed by combining these.
[0051]
[0052] FIG. 5 is a diagram showing the scaling of the motion vector of the temporal merge candidate.
[0052] At the same position PU(Y) belonging to the reference frame, as shown in FIG. 6, select the position of the temporal candidate between candidate C0 and candidate C1. If the PU at position C0 is not available, is intra-coded, or is outside the current coding tree unit (CTU, also known as LCU, maximum coding unit) row, position C1 is used. Otherwise, position C0 is used for the derivation of the temporal merge candidate. FIG. 6 shows an example of candidate positions C0 and C1 of the temporal merge candidate.
[0053]
[0054] 2.1.2.4 Additional candidate insertion
[0055] In addition to the spatio-temporal merge candidate, there are two additional types of merge candidates, namely, combined bi-prediction merge candidate and zero merge candidate. By using the spatio-temporal merge candidate, a combined bi-prediction merge candidate is generated. The combined bi-prediction merge candidate is used only for B slices. By combining the first reference picture list motion parameter of the first candidate and the second reference picture list motion parameter of another candidate, a combined bi-prediction candidate is generated. If these two tuples provide different motion hypotheses, these tuples form a new bi-prediction candidate. As an example, FIG. 7 shows the generation of a combined bi-prediction merge candidate added to the final list (right side) using two candidates having mvL0, refIdx L0 or mvL1, refIdxL1 in the original list (left side). There are various rules regarding the combinations considered for generating these additional merge candidates. Insert motion zero candidates and fill the remaining entries in the merge candidate list to
[0056] , it hits the MaxNumMergeCand capacity. These candidates have a spatial displacement of zero, start from zero, and have a reference picture index that increases each time a new zero motion candidate is added to the list. chain index. Specifically, until the merge list is full, the following steps are performed in sequence.
[0057] Specifically, until the merge list is full, perform the following steps in order. 1. For P slices, set the variable numRef to either the number of reference pictures associated with list 0 or, for B slices, the minimum number of reference pictures in the two lists. Set it. 2. Add non-repetitive motion zero candidates. If the variable i is from 0 to numRef - 1, set the MV to (0, 0) and add the default motion candidate with the reference picture index set to i to list 0 (for P slices) and to both lists (for B slices). 3. Add the repetitive motion zero candidate with the MV set to (0, 0), the reference picture index of list 0 set to 0 (for P slices), and the reference picture indices of both lists set to 0 (for B slices).
[0058] Finally, no redundancy check is performed on these candidates.
[0059] 2.1.3 Advanced Motion Vector Prediction (AMVP)
[0060] AMVP utilizes the spatio-temporal correlation between PUs in the vicinity of the motion vector for the explicit transmission of motion parameters. In each reference picture list, check the availability of the left and upper temporally neighboring PU positions, remove redundant candidates, and add zero vectors to make the length of the candidate list constant, thereby constructing the motion vector candidate list. Next, wherein the encoder can select the best predictor from the candidate list and transmit the corresponding index indicating the selected candidate. Similar to the signaling of the merge index, the index of the best motion vector candidate is encoded using a shortened unary term. The maximum value of the object to be encoded in this case is 2 (see FIG. 8). In the following section, the details of the derivation process of the motion vector prediction candidate will be described.
[0061] 2.1.3.1 Derivation of AMVP Candidates
[0062] FIG. 8 summarizes the derivation process of the motion vector prediction candidate.
[0063] In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. To derive the spatial motion vector candidates, ultimately two motion vector candidates are derived based on the motion vectors of each PU at five different positions as shown in FIG. 2.
[0064] To derive the temporal motion vector candidate, one motion vector candidate is selected from two candidates derived based on two different co-located positions. After creating the first spatio-temporal candidate list, duplicate motion vector candidates in the list are removed. If the number of candidates is more than two, motion vector candidates with a reference picture index greater than 1 in the associated reference picture list are removed from the list. If the number of spatio-temporal motion vector candidates is less than two, additional zero motion vector candidates are added to the list.
[0065] 2.1.3.2 Spatial Motion Vector Candidates
[0066] In the derivation of spatial motion vector candidates, among the five candidates derived from the PU at the position as shown in FIG. 2, the positions of the maximum two candidates to be considered are the same as the positions of motion merge. The derivation order for the left side of the current PU is defined as A0, A1, scaled A0, scaled A1. The derivation order for the upper side of the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. Therefore, for each side, there are four cases where it can be used as a motion vector candidate, that is, two cases where spatial scaling is not necessary and two cases where spatial scaling is used. Summarizing the four different cases, it is as follows. Among the five candidates derived from the PU at the position as shown in FIG. 2 in the derivation of spatial motion vector candidates, the positions of the maximum two candidates to be considered are the same as the positions of motion merge. The derivation order for the left side of the current PU is defined as A0, A1, scaled A0, scaled A1. The derivation order for the upper side of the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. Therefore, for each side, there are four cases where it can be used as a motion vector candidate, that is, two cases where spatial scaling is not necessary and two cases where spatial scaling is used. Summarizing the four different cases, it is as follows. · Without spatial scaling - (1) The same reference picture list and the same reference picture index (the same POC) - (2) Different reference picture lists but the same reference picture (the same POC)
[0067] · With spatial scaling - (3) The same reference picture list but different reference pictures (different POCs) - (4) Different reference picture lists and different reference pictures (different POCs) First, check the case of non-spatial scaling, and then perform spatial scaling. Regardless of the reference picture list, if the POC is different between the reference picture of the neighboring PU and the reference picture of the current PU, consider spatial scaling. If all the PUs of the left side candidates are not available or are intra-coded, scale the upper motion vector. · With spatial scaling - (3) The same reference picture list but different reference pictures (different POCs) - (4) Different reference picture lists and different reference pictures (different POCs)
[0068] First, check the case of non-spatial scaling, and then perform spatial scaling. Regardless of the reference picture list, if the POC is different between the reference picture of the neighboring PU and the reference picture of the current PU, consider spatial scaling. If all the PUs of the left side candidates are not available or are intra-coded, scale the upper motion vector. Regardless of the reference picture list, if the POC is different between the reference picture of the neighboring PU and the reference picture of the current PU, consider spatial scaling. If all the PUs of the left side candidates are not available or are intra-coded, scale the upper motion vector. If all the PUs of the left side candidates are not available or are intra-coded, scale the upper motion vector. The ring is useful for the parallel derivation of the left and upper MV candidates. Otherwise, spatial scaling is not permitted for the upper motion vector.
[0069] In the spatial scaling process, as shown in FIG. 9, similar to the temporal scaling, the motion vectors of neighboring PUs are scaled. The main difference is that the reference picture list and index of the current PU are given as inputs, and the actual scaling process is the same as the temporal scaling.
[0070] 2.1.3.3 Temporal Motion Vector Candidates
[0071] All the processes for deriving the temporal merge candidates, except for deriving the reference picture index, are the same as the processes for deriving the spatial motion vector candidates (see FIG. 6). The reference picture index is signaled to the decoder.
[0072] 2.2 Sub - CU - Based Motion Vector Prediction Method in JEM
[0073] In JEM with QTBT, each CU can have at most one set of motion parameters for each prediction direction. In the encoder, by splitting the large CU into sub - CUs and deriving the motion information of all the sub - CUs of the large CU, two sub - CU - level motion vector prediction methods are considered. By the alternative temporal motion vector prediction (ATMVP) method, each CU can extract a set of multiple motion information from multiple blocks smaller than the current CU in the arranged reference pictures. In the spatio - temporal motion vector prediction (STMVP) method, the temporal motion vector predictor and the spatial neighboring motion vector Use a vector to recursively derive the motion vector of the sub-CU.
[0074] To maintain a more accurate motion field for sub-CU motion prediction, the motion compression of the reference frame is currently disabled.
[0075] Figure 10 shows an example of ATMVP motion prediction for a CU.
[0076] 2.2.1 Alternative Temporal Motion Vector Prediction
[0077] In alternative temporal motion vector prediction (ATMVP), the motion vector temporal motion vector prediction (TMVP) method is modified by extracting a set of multiple motion information (including motion vectors and reference indices) from blocks smaller than the current CU. In some implementations, the sub-CU is an N×N square block (N is set to 4 by default).
[0078] ATMVP predicts the motion vectors of the sub-CUs within a CU in two steps. In the first step, the corresponding block in the reference picture is specified with a temporal vector. This reference picture is called the motion source picture. In the second step, the current CU is divided into sub-CUs, and the motion vectors and reference indices of each sub-CU are obtained from the blocks corresponding to
[0079] each sub-CU. In the first step, the reference picture and the corresponding block are determined by the motion information of the spatially neighboring blocks of the current CU. To avoid repetitive scanning of neighboring blocks, the first Set the available motion vectors of 1 and their associated reference indices to the temporal vector and the index of the motion source picture. In this way, in ATMVP, compared with TMVP, the corresponding block can be identified more accurately, and the corresponding blo ck (which may be called an arrayed block) is always at the lower right or center position relative to the current CU.
[0080] In the second step, by adding the temporal vector to the coordinates of the current CU, the correspon ding block of the sub-CU is identified by the temporal vector in the motion source picture. For each sub-CU, the motion information of the corresponding block (the smallest motion grid covering the cent er sample) is used to derive the motion information of the sub-CU. After identifying the motion info rmation of the corresponding N×N block, similar to TMVP in HEVC, it is converted into the motion vect or and reference index of the current sub-CU, and motion scaling and other procedures are applied. F or example, the decoder checks whether the low-delay condition (i.e., the POC of all reference pictu res of the current picture is smaller than the POC of the current picture) is satisfied, and in som e cases, x the motion vector MV corresponding to the reference picture list X is used to y predict the motion vector MV of each sub-CU (where X is equal to 0 or 1, and Y is equal to 1 - X).
[0081] 2.2.2 Spatio-Temporal Motion Vector Prediction (STMVP)
[0082] In this method, the motion vectors of the sub-CUs are recursively is derived. FIG. 11 illustrates this concept. Consider an 8×8 CU including four 4×4 sub-CUs, A, B, C, and D. The 4×4 blocks in the neighborhood of the current frame are labeled a, b, c , d.
[0083] Derivation of the motion of sub-CU A begins by identifying its two spatial neighborhoods. The first neighborhood is the N×N block above sub-CU A (block c). If this block c is not available or is intra-coded, other N× N blocks above sub-CU A are checked (starting from block c and going from left to right). The second neighborhood is the block to the left of sub-CU A (block b). If block b is not available or is intra-coded, other blocks to the left of sub-CU A are checked (starting from block b and going from top to bottom). The motion information obtained from the blocks in each list of neighborhoods is scaled to the first reference frame of the given list. Next, a temporal motion vector predictor (TMVP) for sub-block A is derived following the same procedure as defined in HEVC . Motion information of the block at the same position in position D is retrieved and scaled accordingly . Finally, after searching and scaling the motion information, all available motion vectors (up to 3) for each reference list are separately averaged. This averaged motion vector is taken as the motion vector of the current sub-CU.
[0084] 2.2.3 Sub-CU Motion Prediction Mode Signal Notification
[0085] The sub-CU mode is made valid as an additional merge candidate, and to signal the mode, an additional Plus syntax elements are not required. As represented by the ATMVP mode and the STMVP mode, , add two additional merge candidates to the merge candidate list of each CU. When the sequence parameter set indicates that ATMVP and STMVP are valid, up to seven merge candidates are used. The encoding logic for the additional merge candidates is the same as that for the merge candidates in HM, i.e., for each CU in a P or B slice, more than two R-D checks are required for the two additional merge candidates.
[0086] In JEM, all binary values of the merge index are context-coded by CABAC. On the other hand, in HEVC, only the first binary value is context-coded, and the remaining binary values are context-bypass-coded.
[0087] 2.3 Inter prediction methods in VVC
[0088] There are several new coding tools for improving inter prediction, such as adaptive motion vector difference resolution (AMVR) for signaling MVD, affine prediction mode, triangle prediction mode (TPM), ATMVP, generalized bi-prediction (GBI), and bi-directional optical flow (BIO).
[0089] 2.3.1 Adaptive motion vector difference resolution
[0090] In HEVC, when use_integer_mv_flag is 0 in the slice header, the difference of the motion vector (MVD) (the difference between the motion vector and the predicted motion vector of the PU) is signaled in units of 1 / 4 luminance samples. In VVC, local adaptation The motion vector resolution (LAMVR) is introduced. In VVC, the MVD can be encoded in units of 1 / 4 luminance samples, integer luminance samples, or four luminance samples (i.e., 1 / 4 pixel, 1 pixel, 4 pixels). The MVD resolution is controlled at the coding unit (CU) level, and the MVD resolution flag is conditionally signaled for each CU having at least one non-zero MVD module. For a CU having at least one non-zero MVD module, a first flag is signaled to indicate whether 1 / 4 luminance sample MV accuracy is used in the CU. If the first flag (equal to 1) indicates that 1 / 4 luminance sample MV accuracy is not used, another flag is signaled to indicate whether integer luminance sample MV accuracy or 4-luminance sample MV accuracy is used.
[0091] If the first MVD resolution flag of the CU is zero, or if no coding is performed for the CU (i.e., all MVDs in the CU are zero), 1 / 4 luminance sample MV resolution is used for the CU. If the CU uses integer luminance sample MV accuracy or 4-luminance sample MV accuracy, the MVP in the AMVP candidate list of the CU is rounded to the corresponding accuracy. In the encoder, the CU-level RD check is used to determine which MVD resolution to use for the CU. That is, the CU-level RD check is performed three times for each MVD resolution. To increase the encoder speed, the following coding method is applied in JEM. In the encoder, the CU-level RD check is used to determine which MVD resolution to use for the CU. That is, the CU-level
[0092] RD check is performed three times for each MVD resolution. To increase the encoder speed, the following coding method is applied in JEM. If the first MVD resolution flag of the CU is zero, or if no coding is performed for the CU (i.e., all MVDs in the CU are zero), 1 / 4 luminance sample MV resolution is used for the CU. If the CU uses integer luminance sample MV accuracy or 4-luminance sample MV accuracy, the MVP in the AMVP candidate list of the CU is rounded to the corresponding accuracy. When the CU uses integer luminance sample MV accuracy or 4-luminance sample MV accuracy, the MVP in the AMVP candidate list of the CU is rounded to the corresponding accuracy. In the encoder, the CU-level RD check is used to determine which MVD resolution to use for the CU. That is, the CU-level
[0093] RD check is performed three times for each MVD resolution. To increase the encoder speed, the following coding method is applied in JEM. That is, the CU-level RD check is performed three times for each MVD resolution. To increase the encoder speed, the following coding method is applied in JEM. To increase the encoder speed, the following coding method is applied in JEM. The following coding method is applied in JEM.
[0094] ● During the RD check of a CU with the normal 1 / 4 luminance sample MVD resolution, the motion information (integer luminance sample accuracy) of the current CU is stored. During the RD check of the same CU with integer luminance samples and 4 luminance sample MVD resolution, the stored motion information (after rounding) is used as the starting point for further small-range motion vector refinement, so that the time-consuming motion estimation process does not repeat three times.
[0095] ● Conditionally call the RD check of a CU with 4 luminance sample MVD resolution. In the case of a CU, if the RD cost of the integer luminance sample MVD resolution is much larger than that of the 1 / 4 luminance sample MVD resolution, the RD check of the 4 luminance sample MVD resolution for the CU is skipped.
[0096] The encoding process is shown in FIG. 12. First, test the 1 / 4 pixel MV, calculate the RD cost, and represent it as RDCost0. Next, test the integer MV and represent the RD cost as RDCost1. If RDCost1 < th * RDCost0 (where th is a positive value), test the 4 pixel MV; otherwise, skip the 4 pixel MV. Basically, when checking the integer or 4 pixel MV, the motion information and RD cost, etc. for the 1 / 4 pixel MV are already known, and this can be reused to speed up the encoding process of the integer or 4 pixel MV.
[0097] 2.3.2 Triangular Prediction Mode
[0098] The concept of the triangular prediction mode (TPM) is to introduce a new triangular partition for motion compensation prediction. As shown in FIGS. 13A and 13B, the CU is divided diagonally or anti-diagonally. It is divided into two triangular prediction units in the direction. Each triangular prediction unit in the CU uses 1 unique single prediction motion vector and reference frame index derived from one single prediction candidate list for inter prediction. After predicting the triangular prediction unit, adaptive weighting processing is performed on the diagonal edges. Then, conversion and quantization processing are performed on the entire CU. Note that this mode is only applicable to the merge mode (note that the skip mode is treated as a special merge mode). is treated as a special merge mode). is treated as a special merge mode).
[0099] Figures 13A and 13B are explanatory diagrams for dividing the CU into two triangular prediction units (two division patterns). Figure 13A: 135-degree division type (division from the upper left corner to the lower right corner), Figure 13B: 45-degree division pattern. 13B: 45-degree division pattern.
[0100] 2.3.2.1 Single Prediction Candidate List of TPM
[0101] The single prediction candidate list called the TPM motion candidate list consists of five single prediction motion vector candidates. As shown in Figure 14, this is derived from seven neighboring blocks including five spatially neighboring blocks (1 to 5) and two temporally co-located blocks (6 to 7). The motion vectors of the seven neighboring blocks are collected in the order of the single prediction motion vector, the L0 motion vector of the dual prediction motion vector, the L1 motion vector of the dual prediction motion vector, and the average motion vector of the L0 and L1 motion vectors of the dual prediction motion vector, and put into the single prediction candidate list. If the number of candidates is less than 5, the motion vector zero is added to the list. The motion candidates added to this TPM list are called TPM candidates, and the motion information derived from spatial / temporal blocks is called normal motion candidates. list. If the number of candidates is less than 5, the motion vector zero is added to the list. The motion candidates added to this TPM list are called TPM candidates, and the motion information derived from spatial / temporal blocks is called normal motion candidates. The motion candidates added to this TPM list are called TPM candidates, and the motion information derived from spatial / temporal blocks is called normal motion candidates.
[0102] Specifically, the following steps are included.
[0103] 1) When adding regular motion candidates from spatially neighboring blocks, the full pruning operation by Obtain normal motion candidates from A1, B1, B0, A0, B2, Col, and Col2 (corresponding to blocks 1-7 in Figure 14).
[0104] 2) Set the variable numCurrMergeCand = 0.
[0105] 3) For each normal motion candidate derived from A1, B1, B0, A0, B2, Col, and Col2, if it has not been pruned and numCurrMergeCand is less than 5, and the normal motion candidate is a single prediction (from either List0 or List1), then it is directly added to the merge list as a TPM candidate with numCurrMergeCand incremented by 1. Such a TPM candidate is named "Candidate originally predicted as single". full pruning Apply.
[0106] 4) For each motion candidate derived from A1, B1, B0, A0, B2, Col, and Col2, if it has not been pruned and numCurrMergeCand is less than 5, and the normal motion candidate is a dual prediction, then the motion information from List0 is added to the TPM merge list as a new TPM candidate (i.e., modified from a single prediction in List0), and 1 is added to numCurrMergeCand. Such a TPM candidate is referred to as "Shortened List0 prediction candidate". full pruning Apply.
[0107] 5) For each motion candidate derived from A1, B1, B0, A0, B2, Col, and Col2, if it has not been pruned and numCurrMergeCand is less than 5, and if the normal motion candidate is a dual prediction, add the motion information from List 1 to the TPM merge list (i.e., modified from single prediction in List 1), and increment numCurrMergeCand by 1. Such a TPM candidate is called a "shortened List1 - prediction candidate". full pruning Apply it.
[0108] 6) For each motion candidate derived from A1, B1, B0, A0, B2, Col, and Col2, if it has not been pruned and the normal motion candidate is a dual prediction, then numCurrMergeCand is less than 5. - If the List0 - reference picture slice QP is smaller than the List1 - reference picture slice QP, first scale the motion information of List1 to the List0 - reference picture, and add the average of the two MVs (one from the original List0 and the other from the scaled - down MV from List1) to the TPM merge list. Such a candidate is called the average single prediction from the List0 motion candidate, and increment numCurrMergeCand by 1. - Otherwise, first scale the motion information of List0 to the List1 - reference picture, and add the average of the two MVs (one from the original List1 and the other from the scaled - down MV from List0) to the TPM merge list. Such a TPM candidate is called the average single prediction from the List1 motion candidate, and increment numCurrMergeCand by 1. full pruning Apply it.
[0109] 7) If numCurrMergeCand is less than 5, add zero motion vector candidates to it.
[0110] When inserting a candidate into the list, if it has to be compared with all the previously added candidates to check if it is the same as one of them, such a process is called loop pruning. of them, this kind of processing is called loop pruning. It is called loop pruning.
[0111] 2.3.2.2 Adaptive Weighting Process
[0112] After predicting each triangular prediction unit, perform an adaptive weighting process on the diagonal edge between two triangular prediction units to derive the final prediction of the entire CU. Define two groups of weight coefficients as follows. Define two groups of weight coefficients as follows. Define them. · The first group of weight coefficients uses {7 / 8, 6 / 8, 4 / 8, 2 / 8, 1 / 8} and {7 / 8 , 4 / 8, 1 / 8} for luma and chroma samples respectively. · The second group of weight coefficients uses {7 / 8, 6 / 8, 5 / 8, 4 / 8, 3 / 8, 2 / 8, 1 / 8 } and {6 / 8, 4 / 8, 2 / 8} for luma and chroma samples respectively.
[0113] Select a group of weight coefficients based on the comparison of the motion vectors of two triangular prediction units. The second group of weight coefficients is used when the reference pictures of the two triangular prediction units are different, or when the difference between their motion vectors is greater than 16 pixels. Otherwise, use the first group of weight coefficients.
[0114] 2.3.2.3 Signaling of Triangular Prediction Mode (TPM)
[0115] One bit flag indicating whether TPM is used or not may be signaled first. Then, two split patterns (as shown in FIGS. 13A and 13B), and the selected merge index for each of the two splits are further signaled.
[0116] 2.3.2.3.1 Signaling of TPM flag
[0117] Let the width and height of one luminance block be represented by W and H respectively. When W*H is less than 64, the triangular prediction mode is disabled.
[0118] When encoding one block in affine mode, the triangular prediction mode is also disabled. .
[0119] When one block is encoded in merge mode, one bit flag is signaled to indicate whether the triangular prediction mode is enabled or disabled for this block. This can be done.
[0120] This flag is encoded in three contexts based on the following formula. Ctx index = ((left block L available && L is coded with TPM?)? 1 : 0) + ((Above block A available && A is coded with TPM?)? 1 : 0);
[0121] FIG. 15 shows an example of neighboring blocks (A and L) used for context selection in TPM flag encoding. (A and L).
[0122] 2.3.2.3.2 Display of two split patterns (shown in FIG. 13), and Signal notification of the merge index selected for each
[0123] Note that the split pattern and the two merge indexes of the split are encoded with each other. In existing implementations, it was restricted that the two splits could not use the same reference index. Therefore, there are two (split patterns) * N (maximum number of merge candidates) * (N - 1 ) possibilities, and N is set to 5. One display is encoded, and the mapping between the split pattern, the two merge indexes, and the encoded instruction is derived from the array defined below.
[0124] const uint8_t g_TriangleCombination[TRIA NGLE_MAX_NUM_CANDS][3]={ {0,1,0},{1,0,1},{1,0,2},{0,0,1},{0,2,0}, {1,0,3},{1,0,4},{1,1,0},{0,3,0},{0,4,0}, {0,0,2},{0,1,2},{1,1,2},{0,0,4},{0,0,3}, {0,1,3},{0,1,4},{1,1,4},{1,1,3},{1,2,1}, {1,2,0},{0,2,1},{0,4,3},{1,3,0},{1,3,2}, {1,3,4},{1,4,0},{1,3,1},{1,2,3},{1,4,1}, {0,4,1},{0,2,3},{1,4,2},{0,3,2},{1,4,3}, {0,3,1},{0,2,4},{1,2,4},{0,4,2},{0,3,4}} ;
[0125] Split pattern (45 degrees or 135 degrees) = g_TriangleCombinatio n [signaled indication] [0]; Merge index of candidate A = g_TriangleCom bination [signaled indication] [1]; Merge index of candidate B = g_TriangleCom bination [signaled indication] [2];
[0126] When two motion candidates A and B are derived, the (P U1, PU2) motion information of the two splits can be set from either A or B, and whether PU1 uses the motion information of merge candidate A or B depends on the predicted directions of the two motion candidates. Table 1 shows the relationship between the two derived motion candidates A and B having two splits.
[0127]
Table 1
[0128] 2.3.2.3.3 (indicated by merge_triangle_idx) entropy - coding
[0129] merge_triangle_idx is within the range of [0, 39] (including each). The K_th order Exponential Golomb (EG) code is used for the binarization of merge_triangle_idx (K is set to 1 ).
[0130] K-th order EG
[0131] (at the expense of using more bits to encode a smaller number) fewer To encode larger numbers with bits, this can be generalized using a non - negative integer parameter k. To encode a non - negative integer x with an exp - Golomb code of degree k, do the following: 1. Encode [x / 2 using the order - 0 exp - Golomb code described above. Next, 1. Encode [x / 2 k using the aforementioned order - 0 exp - Golomb code. Next, 2. Encode x mod 2 2. Encode x mod 2 k in binary.
[0132]
Table 2
[0133] 2.3.3 Affine Motion Compensation Prediction
[0134] In HEVC, only the translational motion model is applied for motion - compensated prediction (MCP). In the real world, however, there are various types of motion, such as zoom - in / zoom - out, rotation, perspective motion, and other irregular motions. In VVC, a simple affine - transform motion - compensated prediction is applied using a 4 - parameter affine model and a 6 - parameter affine model. As shown in FIGS. 16A - 16B, the affine motion field of a block is represented by two control - point motion vectors (CPMV) in the case of the 4 - parameter affine model (FIG. 16A) and by three CPMVs in the case of the 6 - parameter affine model (FIG. 16B). In HEVC, only the translational motion model is applied for motion - compensated prediction (MCP). In the real world, however, there are various types of motion, such as zoom - in / zoom - out, rotation, perspective motion, and other irregular motions. In VVC, a simple affine - transform motion - compensated prediction is applied using a 4 - parameter affine model and a 6 - parameter affine model. As shown in FIGS. 16A - 16B, the affine motion field of a block is represented by two control - point motion vectors (CPMV) in the case of the 4 - parameter affine model (FIG. 16A) and by three CPMVs in the case of the 6 - parameter affine model (FIG. 16B). In HEVC, only the translational motion model is applied for motion - compensated prediction (MCP). In the real world, however, there are various types of motion, such as zoom - in / zoom - out, rotation, perspective motion, and other irregular motions. In VVC, a simple affine - transform motion - compensated prediction is applied using a 4 - parameter affine model and a 6 - parameter affine model. As shown in FIGS. 16A - 16B, the affine motion field of a block is represented by two control - point motion vectors (CPMV) in the case of the 4 - parameter affine model (FIG. 16A) and by three CPMVs in the case of the 6 - parameter affine model (FIG. 16B). In HEVC, only the translational motion model is applied for motion - compensated prediction (MCP). In the real world, however, there are various types of motion, such as zoom - in / zoom - out, rotation, perspective motion, and other irregular motions. In VVC, a simple affine - transform motion - compensated prediction is applied using a 4 - parameter affine model and a 6 - parameter affine model. As shown in FIGS. 16A - 16B, the affine motion field of a block is represented by two control - point motion vectors (CPMV) in the case of the 4 - parameter affine model (FIG. 16A) and by three CPMVs in the case of the 6 - parameter affine model (FIG. 16B). In HEVC, only the translational motion model is applied for motion - compensated prediction (MCP). In the real world, however, there are various types of motion, such as zoom - in / zoom - out, rotation, perspective motion, and other irregular motions. In VVC, a simple affine - transform motion - compensated prediction is applied using a 4 - parameter affine model and a 6 - parameter affine model. As shown in FIGS. 16A - 16B, the affine motion field of a block is represented by two control - point motion vectors (CPMV) in the case of the 4 - parameter affine model (FIG. 16A) and by three CPMVs in the case of the 6 - parameter affine model (FIG. 16B). In HEVC, only the translational motion model is applied for motion - compensated prediction (MCP). In the real world, however, there are various types of motion, such as zoom - in / zoom - out, rotation, perspective motion, and other irregular motions. In VVC, a simple affine - transform motion - compensated prediction is applied using a 4 - parameter affine model and a 6 - parameter affine model. As shown in FIGS. 16A - 16B, the affine motion field of a block is represented by two control - point motion vectors (CPMV) in the case of the 4 - parameter affine model (FIG. 16A) and by three CPMVs in the case of the 6 - parameter affine model (FIG. 16B). In HEVC, only the translational motion model is applied for motion - compensated prediction (MCP). In the real world, however, there are various types of motion, such as zoom - in / zoom - out, rotation, perspective motion, and other irregular motions. In VVC, a simple affine - transform motion - compensated prediction is applied using a 4 - parameter affine model and a 6 - parameter affine model. As shown in FIGS. 16A - 16B, the affine motion field of a block is represented by two control - point motion vectors (CPMV) in the case of the 4 - parameter affine model (FIG. 16A) and by three CPMVs in the case of the 6 - parameter affine model (FIG. 16B). In HEVC, only the translational motion model is applied for motion - compensated prediction (MCP). In the real world, however, there are various types of motion, such as zoom - in / zoom - out, rotation, perspective motion, and other irregular motions. In VVC, a simple affine - transform motion - compensated prediction is applied using a 4 - parameter affine model and a 6 - parameter affine model. As shown in FIGS. 16A - 16B, the affine motion field of a block is represented by two control - point motion vectors (CPMV) in the case of the 4 - parameter affine model (FIG. 16A) and by three CPMVs in the case of the 6 - parameter affine model (FIG. 16B).
[0135] The motion - vector field (MVF) of a block is given by the 4 - parameter affine model in Equation (1) (where the 4 parameters are defined as variables a, b, e, f) and the 6 - parameter affine model in Equation (2 The motion - vector field (MVF) of a block is given by the 4 - parameter affine model in Equation (1) (where the 4 parameters are defined as variables a, b, e, f) and the 6 - parameter affine model in Equation (2 ) (where the 4 parameters are defined as variables a, b, c, d, They are each represented by the following expressions using the definitions of e and f).
[0136]
Number
[0137]
Number
[0138] Here, (mv h 0, mv h 0) is the motion vector of the control point at the upper left corner, and (mv h 1, mv h 1) is the motion vector of the control point at the upper right corner, and (mv h 2, mv h 2) is the motion vector of the control point at the lower left corner. All three motion vectors are called control point motion vectors (CPMV). (x, y) represents the coordinates of the representative point for the upper left sample within the current block, and (mv tor (CPMV), and (x, y) represents the coordinates of the representative point for the upper left sample within the current block, and (mv (x, y), mv h (x, y)) is the motion vector derived for the sample located at (x, y). The CP motion vector may be signaled (such as in the affine AMVP mode) or derived on-the-fly (such as in the affine ma v (x, y)) is the motion vector derived for the sample located at (x, y). The CP motion vector may be signaled (such as in the affine AMVP mode) or derived on-the-fly (such as in the affine ma trix mode). w and h are the width and height of the current block. In fact, this division is implemented by a right shift with a rounding operation. In VTM, the representative point is (such as in the affine AMVP mode) or derived on-the-fly (such as in the affine ma trix mode). w and h are the width and height of the current block. In fact, this division is implemented by a right shift with a rounding operation. In VTM, the representative point is set to the center position of the sub-block. For example, if the coordinates of the upper left sample at the upper left corner of the sub-block in the current block are (xs, ys), the coordinates of the representative point are set to (xs + 2, y s + 2). For each sub-block (i.e., 4×4 in VTM), the representative point is set to (xs + 2, ys + 2). For each sub-block (i.e., 4×4 in VTM), the representative point Using this, the motion vector of the entire sub-block is derived.
[0139] To further simplify motion compensation prediction, affine transform prediction based on sub-blocks is applicable. For each M×N (in the current VVC, both M and N are set to 4) sub-block, as shown in FIG. 17, to derive the motion vector, the motion vector of the central sample of each sub-block is calculated according to equations (1) and (2) and rounded to a fractional precision of 1 / 16. Next, a motion compensation interpolation filter of 1 / 16 pixel is applied, and prediction of each sub-block is generated using the derived motion vector. The 1 / 16 pixel interpolation filter is introduced in the affine mode.
[0140] After MCP, the high-precision motion vector of each sub-block is rounded and stored with the same precision as the normal motion vector.
[0141] 2.3.3.1 Signal Notification of Affine Prediction
[0142] Similar to the translational motion model, there are also two modes for signal notification of side information by affine prediction. They are the AFFINE_INTER mode and the AFFINE_MERGE mode.
[0143] 2.3.3.2. AF_INTER Mode
[0144] For a CU where both the width and height are greater than 8, the AF_INTER mode can be applied. To indicate whether the AF_INTER mode is used, an affine flag at the CU level is signaled in the bitstream.
[0145] In this embodiment, for each reference picture list (list 0 or list 1), three types of affine motion predictors are used to construct an affine AMVP candidate list in the following order. Each candidate includes the estimated CPMV of the current block. The difference between the best CPMV found on the encoder side and the MV (mv in FIG. 20 0、 mv 1、 mv2, etc.) and the estimated CPMV are signaled. Further, the index of the affine AMVP candidate from which the estimated CPMV is derived is signaled.
[0146] 1) Inherited affine motion predictor
[0147] The checking order is similar to the checking order of the spatial MVP in HEVC AMVP list construction. First, among {A1, A0}, the left inherited affine motion predictor is derived from the first block having the same reference picture as the currently affine-encoded block. Next, the above-mentioned inherited affine motion predictor is derived from the first block among {B1, B0, B2} having the same reference picture as the currently encoded block. FIG. 19 shows five blocks A1, A0, B1, B 0, B2. When it is found that neighboring blocks are encoded in affine mode, the CPMV of the current block is predicted using the CPMV of the coding unit including this neighboring block. For example, if A1 is encoded in non-affine mode and A0 is encoded in 4-parameter affine mode, the left inherited affine MV predictor is derived from A0. In this case, for the upper left CPMV in FIG. 21B, it is MV0 For the upper right CPMV,
[0148] When it is found that neighboring blocks are encoded in affine mode, the CPMV of the current block is predicted using the CPMV of the coding unit including this neighboring block. For example, if A1 is encoded in non-affine mode and A0 is encoded in 4-parameter affine mode, the left inherited affine MV predictor is derived from A0. In this case, for the upper left CPMV in FIG. 21B, it is MV0 For the upper right CPMV, For the upper left CPMV in FIG. 21B, it is MV0 N For the upper right CPMV, Then, using the CPMV of the CU including A0 shown in N CPMV, for the upper left (coordinates (x0, y0)), upper right (coordinates (x1, y1)) and lower right position of the current block (coordinates (x2, y2)), MV0 , MV1 C , MV2 C is represented by, and the estimated CPMV of the current block C is derived.
[0149] 2) Constructed affine motion predictor
[0150] As shown in FIG. 20, the constructed affine motion predictor consists of control point motion vectors (CPMVs) having the same reference picture, derived from neighboring inter-coded blocks. When the current affine motion model is a 4-parameter affine, the number of CPMVs is 2, otherwise, when the current affine motion model is a 6-parameter affine, the number of CPMVs is 3. The upper left CPMV m-v0- is inter-coded and is derived by the MV of the first block of group {A, B, C} having the same reference picture as the current block. The upper right CPMV m-v1- is inter-coded and is derived by the MV of the first block of group {D, E} having the same reference picture as the current block. The lower left CPMV m-v2- is inter-coded and is derived by the MV of the first block of group {F, G} having the same reference picture as the current block. - When the current affine motion model is a 4-parameter affine, for the constructed affine motion predictor, when both m-v0- and m-v1- are established, i.e., m- is derived.
[0151] - When the current affine motion model is a 4-parameter affine, the constructed affine motion predictor, when both m-v0- and m-v1- are established, i.e., m- It is inserted into the candidate list only in the case of v0 ̄ and m ̄v1 ̄. It is used as the estimated CPMV at the upper left (coordinates (x0, y0)) and upper right (coordinates (x1, y1)) of the current block. .
[0152] - When the current affine motion model is a 6-parameter affine, if m ̄v0 ̄, m ̄v1 ̄, and m ̄v2 ̄ are all established, that is, m ̄v0 ̄, m ̄v 1 ̄, and m ̄v2 ̄ are all used as the estimated CPMV at the upper left (coordinates (x0, y0)) , upper right (coordinates (x1, y1)), and lower right (coordinates (x2, y2)) of the position of the current block only when used, the constructed affine motion predictor is inserted into the candidate list.
[0153] When inserting the constructed affine motion predictor into the candidate list, the pruning process is not applied .
[0154] 3) Normal AMVP motion predictor
[0155] The following is applied until the number of affine motion predictors reaches the maximum. 1) If available, set all CPMVs equal to m ̄v2 ̄ to derive an affine motion predictor. 2) If available, set all CPMVs equal to m ̄v1 ̄ to derive an affine motion predictor. 3) If available, set all CPMVs to m ̄v0 ̄ to derive an affine motion prediction predictor. 4) If available, by setting all CPMVs equal to HEVCTMVP , derive an affine motion predictor. 5) By setting all CPMVs to zero MV, derive an affine motion predictor Do.
[0156] Note that m ̄v i  ̄ has already been derived by the constructed affine motion predictor.
[0157] Figure 18A shows an example of a 4-parameter affine model. Figure 18B shows an example of a 6-parameter aff ine model.
[0158] Figure 19 shows an example of the MVP of AF_INTER of the inherited affine candidates.
[0159] Figure 20 shows an example of the MVP of AF_INTER of the constructed affine candidates.
[0160] In the AF_INTER mode, when the 4 / 6-parameter affine mode is used , 2 / 3 control points are required. Therefore, as shown in Figure 18, for these control points it is necessary to encode 2 / 3 MVDs. In the existing implementation, it is proposed to derive the MV as follows, that is, mvd1 and mvd2 are predicted from mvd0.
[0161] mv0 = m ̄v0 ̄ + mvd0 mv1 = m ̄v1 ̄ + mvd1 + mvd0 mv2 = m ̄v2 ̄ + mvd2 + mvd0
[0162] Here, m ̄v i  ̄, mvd i , mv1 are the predicted motion vectors, motion vector differences, and motion vectors of the upper left pixel (i = 0), upper right pixel (i = 1), and lower left pixel (i = 2) respectively, as shown in Figure 18B. Note that the addition of two motion vectors (for example, mvA( xA,yA) and mvB(xB,yB)) sums two modules separately That is, newMV = mvA + mvB, and the two modules of newMV are Set the rules to (xA+xB) and (yA+yB), respectively.
[0163] 2.3.3.3 AF_MERGE Mode
[0164] When applying a CU in AF_MERGE mode, the CU is a valid neighboring reconstruction block. Then, we obtain the first block coded in affine mode. Then, we select the candidate block The selection order is from left, top, top right, bottom left to top left, as shown in the order of A, B, C, D, E in Fig. 21A. For example, if the adjacent bottom left block is coded in affine mode as shown by A0 in Fig. 21B, the control point (CP) motion vector mv0 of the top left corner, top right corner, and bottom left corner of the neighboring CU / PU containing block A is N , mv1 N and mv2 N Then, take out mv0 N , mv1 N and mv2 N Based on the above, the top-left / top-right / bottom-left motion vector mv0 in the current CU / PU is calculated. C , mv1 C and mv2 C (used only for 6-parameter affine model). Note that in VTM-2.0, the sub-block located at the upper left corner (e.g., 4x4 block in VTM) stores MV0, and the sub-block located at the upper right corner stores mv1 if the current block is affine coded. If the current block is coded with a 6-parameter affine model, the sub-block located at the lower left corner stores mv2, otherwise (with a 4-parameter affine model), the LB stores mv2'. The other sub-blocks store the MV used for MC.
[0165] Current CU mv0 C , mv1 C , mv2 CAfter deriving the CPMV of, we generate the MVF of the current CU according to the simplified affine motion model equations (1) and (2). To identify whether the current CU is coded in AF_MERGE mode, we signal an affine flag in the bitstream if there is at least one neighboring block coded in affine mode.
[0166] In the existing implementation, the affine merge candidate list is constructed using the following steps: It will be built.
[0167] 1) Insert inherited affine candidates
[0168] The inherited affine candidate is the affine motion model of its valid neighboring affine coded blocks. This means deriving the candidate from the affine motion model of the neighboring blocks. Derive up to two inherited affine candidates and insert them into the candidate list. In the case of the predictor above, the scan order is {A0,A1}, and in the case of the predictor above, the scan order is {B0 ,B1,B2}.
[0169] 2) Insert the constructed affine candidates
[0170] The number of candidates in the affine merge candidate list is less than MaxNumAffineCand If so (e.g., 5), insert the constructed affine candidate into the candidate list. The affine candidates are constructed by combining the motion information of the neighborhood of each control point. This means:
[0171] a) First, the motion of the control points is estimated from the identified spatial and temporal neighborhoods shown in Figure 22. CPk (k=1,2,3,4) represents the kth control point. A0,A 1. A2, B0, B1, B2, B3 are the spatial positions for predicting CPk (k = 1, 2, 3), and T is the temporal position for predicting CP4. The coordinates of CP1, CP2, CP3, CP4 are (0, 0), (W, 0), (H , 0), (W, H) respectively, where W and H are the width and height of the current block.
[0172] The movement information of each control point is obtained according to the following priority. - For CP1, the check priority is B2 -> B3 -> A2. If available, B2 is used. Otherwise, if B2 is available, B3 is used. If both B2 and B3 are unavailable, A2 is used. If all three candidates are unavailable, the movement information of CP1 cannot be obtained. - For CP2, the check priority is B1 -> B0. - For CP3, the check priority is A1 -> A0. - For CP4, T is used.
[0173] b) Next, these combinations of control points are used to construct the affine merge candidates. I. To construct the 6-parameter affine candidates, the movement information of three control points is required. The three control points can be selected from one of the following four combinations ({CP1, CP2, CP4}, {C P1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4}) . The combinations of {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4} are converted into 6-parameter movement models represented by the upper-left, upper-right, and lower-left control points. II. To construct the 4-parameter affine candidates, the motion information of two control points is necessary. The two control points may be selected from one of two combinations ({CP1, CP2}, {CP1, CP3 }). These two combinations are converted into a 4-parameter motion model represented by the upper left and upper right control points. III. Insert the constructed combinations of affine candidates into the candidate list in the following order. {CP1, CP2, CP3}, {CP1, CP2, CP4}, {CP1, CP3, CP4 }, {CP2, CP3, CP4}, {CP1, CP2}, {CP1, CP3} i. For each combination, check the reference index of list X for each CP. If they are all the same, this combination has a valid CP MV for list X. If this combination does not have a valid C PMV for both list 0 and list 1, this combination is marked as invalid. Otherwise it is valid and the CPMV is put into the sub-block merge list.
[0174] 3) Padding of zero motion vectors
[0175] If the number of candidates in the affine merge candidate list is less than 5, insert zero motion vectors with a reference index of zero into the candidate list until the list is full. Specifically, for the sub-block merge candidate list, perform 4-parameter merge
[0176] candidates where the MV is set to (0, 0) and the prediction direction is set to single prediction from list 0 (in the case of P slices), and dual prediction (in the case of B slices).
[0177] 2.3.4 Reference to the current picture
[0178] Intra Block Copy (also known as IBC, Current Picture Reference (CPR) or intra-picture block compensation) is used in the HEVC Screen Content Coding Extension (SCC). The tool was adopted for text- and graphics-rich content. Screen capture is similar in that the repeated patterns occurring frequently within the same picture. It is very efficient for coding content video. Using the above reconstructed block as a predictor effectively reduces the prediction error, thus This can improve coding efficiency. Fig. 23 shows an example of intra-block compensation.
[0179] Similar to the design of the CRP in HEVC SCC and VVC, the use of IBC mode It is signaled at both the sequence level and the picture level. When IBC mode is enabled in a frame (SPS), it is enabled at the picture level. If IBC mode is enabled at the picture level, the current reconstruction The constructed picture is treated as a reference picture. Therefore, to signal the use of IBC mode, Therefore, on top of the existing VVC inter mode, no syntax changes at the block level are required. do not have.
[0180] Main Features: - It is treated as a normal inter mode. Therefore, even in IBC mode, Merge and skip modes are available. IBC or HEVC In We integrate the construction of a merge candidate list that contains merge candidates from neighboring positions encoded in the filter mode. Do the following. Based on the selected merge index, the current block in merge or skip mode is merged into the neighborhood encoded in IBC mode or encoded in normal inter mode with a different picture as the reference picture. The current block is merged into the neighborhood encoded in IBC mode or encoded in normal inter mode with a different picture as the reference picture. - The block vector prediction and encoding method for IBC mode reuses the methods used for motion vector prediction and encoding in the HEVC inter mode (AMVP and MVD encoding). - The motion vectors for IBC mode, also called block vectors, are encoded with integer pixel precision, but 1 / 4 pixel precision is required in the interpolation and deblocking stages, so they are stored in memory with 1 / 16 pixel precision after decoding. When used for motion vector prediction in IBC mode, the stored vector predictors are right-shifted by 4. - Search range: Limited within the current CTU. - If affine mode / triangular mode / GBI / weighted prediction is enabled, CPR is not permitted.
[0181] 2.3.5 Merge List Design in VVC
[0182] There are three different merge list construction processes supported in VVC.
[0183] 1) Sub-block merge candidate list: Includes ATMVP and affine merge candidates. One merge list construction process is common to both affine mode and ATMVP mode. Note that ATMVP and affine merge candidates may be added in order. The size of the sub-block merge list is signaled in the slice header, and its maximum value is 5.
[0184] 2) Single prediction TPM merge list: In the case of the triangular prediction mode, one merge list construction process is shared for two partitions, and the two partitions may select their own merge candidate indexes. When constructing this merge list, check the spatially neighboring blocks and the two temporal blocks of this block. The motion information derived from the spatial neighborhood and temporal blocks is called a normal motion candidate in the inventors' IDF. Using these normal motion candidates, a plurality of TPM candidates are further derived. Note that this conversion is performed at the block level as a whole, and different motion vectors may be used in the two partitions to generate their own predicted blocks. The size of the single prediction TPM merge list is fixed to 5.
[0185]
[0186] 3) Normal merge list: For the remaining coded blocks, one merge list construction process is shared. Spatially / temporally / HMVP, pairwise composite dual prediction merge candidates, and motion zero candidates may be inserted in order. The size of the normal merge list is signaled in the slice header, and its maximum value is 6.
[0187]
[0188] 2.3.5.1 Sub-block merge candidate list
[0187] In addition to the normal merge list of non-sub-block merge candidates, it is recommended to put all sub-block related motion candidates into a separate merge list.
[0188] Put the sub-block related motion candidates into a separate merge list, which is called the "sub-block merge candidate list".
[0189] In one example, the sub-block merge candidate list includes affine merge candidates, ATMVP candidates, and / or STMVP candidates based on sub-blocks.
[0190] 2.3.5.1.1 Another ATMVP Embodiment
[0191] In this contribution, the ATMVP merge candidates in the normal merge list are moved to the first position of the affine merge list. All merge candidates in the new list (i.e., , the merge candidate list based on sub-blocks) are based on the sub-block coding tool.
[0192] 2.3.5.1.2 ATMVP in VTM-3.0
[0193] In VTM-3.0, in addition to the normal merge candidate list, a special merge candidate list called the sub-block merge candidate list (also known as the affine merge candidate list) is added. The sub-block merge candidate list satisfies candidates in the following order. b. ATMVP candidates (which may or may not be available) c. Inherited affine candidates d. Constructed affine candidates e. Padding as a zero MV4 parameter affine model
[0194] The maximum number of candidates (denoted as ML) in the sub-block merge candidate list is derived as follows. Derivation. 1) The ATMVP usage flag (e.g., the flag may be named "sps_sbTMVp_enable d_flag") is on (equal to 1), but the affine usage flag (e.g., the flag is named "sps_affine_enabled_flag" If the name can be assigned (equal to 0), then next, ML is set to 1. It becomes. 2) If the ATMVP usage flag is off (equal to 0) and the affinity usage flag is off (equal to 0), then ML is set equal to 0. In this case, the sub-block merge candidate list is not used. 3) Otherwise (the affinity usage flag is on (equal to 1), the ATMVP usage flag is on or off), ML is signaled from the encoder to the decoder. Valid ML is 0 <= ML <= 5. When constructing the sub-block merge candidate list, first check the ATMVP candidates.
[0195] If any one of the following conditions is true, skip the ATMVP candidate and do not put it in the sub-block merge candidate list. 1) The ATMVP usage flag is OFF. 2) The TMVP usage flag (for example, when signaled at the slice level, the flag is named " slice_temporal_mvp_enabled_flag") is off. 3) The reference picture having the reference index 0 in the reference list 0 is the same as the current picture (CPR). 3) The reference picture having the reference index 0 in the reference list 0 is the same as the current picture (CPR).
[0196] The ATMVP in VTM-3.0 is much simpler than the ATMVP in JEM. When generating ATMVP merge candidates, the following processing is applied. a. As shown in Figure 22, check the neighboring blocks A1, B1, B0, A0 to find the first block, shown as block X, that is inter-coded but not CPR-coded. Find. b. Initialize TMV=(0,0). If there is one MV (denoted as MV’) in block X ), then (if signaled in the slice header) refer to the collocated reference picture and set TMV equal to MV’. c. Let the center point of the current block be (x0,y0). Next, place the corresponding position of (x0,y0) in the collocated picture at M=(x0+MV’x,y0+MV’y) . Find the block Z that contains M. i. If Z is intra-coded, ATMVP cannot be used. ii. If Z is inter-coded, the two lists of MVZ_0 and MVZ_1 of block Z are scaled as MVdefault0, MVdefault1 and stored in (Reflist 0 index0) and (Reflist1 index0). d. Assume that the center point of each 8×8 sub-block is (x0S,y0S). Next, locate the corresponding position of (x0S,y0S) in the collocated picture at MS=(x0S+MV’x,y0S+MV’y). Find the block ZS that contains MS. i. If ZS is intra-coded, MVdefault0 and MVdefault1 are assigned to the sub-block. ii. If ZS is inter-coded, the two lists of MVZS_0 and MVZS_1 of block ZS are scaled and assigned to the sub-block in (Reflist0 index0) and (Reflist1 index0).
[0197] MV clipping and masking in ATMVP:
[0198] In the collocated picture, the field that defines the corresponding position such as M or MS is clipped so as to be within a predetermined area. The size of the CTU is S×S, S = 128 in VTM-3.0. If the upper left position of the collocated CTU is (xC TU, yCTU), then the position at (xN, yN) of the corresponding position M or MS is clipped to the valid area xCTU <= xN < xCTU + S + 4; yCTU <= yN < yCTU + S.
[0199] In addition to clipping, (xN, yN) is also masked as xN = xN & MASK, yN = yN & M ASK, where MASK is an integer equal to ~(2 N - 1), N = 3, and the lowest 3 bits are set to 0. Thus, xN and yN must be multiples of 8. (“~” represents the bitwise complement operator.)
[0200] Figure 24 shows an example of the valid corresponding area in the collocated picture.
[0201] 2.3.5.1.3 Syntax Design in the Slice Header
[0202]
Table 3
[0203] 2.3.5.2 Normal Merge List
[0204] Unlike the design of the merge list, in VVC, the history-based motion vector prediction (HMV P) method is adopted.
[0205] The above-mentioned encoded motion information is stored in HMVP. The above-mentioned encoded block The motion information of CU is defined as an HMVP candidate. Multiple HMVP candidates are stored in a table called the HMVP table, and this table is maintained on-the-fly during the encoding / decoding process. When starting the encoding / decoding of a new slice, the HMVP table becomes empty. Whenever there is an inter-coded block, the relevant motion information is added as a new HMVP candidate to the last entry of the table. The overall encoding flow is shown in Figure 25. The HMVP candidates can be used in both the AMVP and the merge candidate list construction processes. Figure 26 shows the modified merge candidate list construction process (highlighted in blue). After inserting the TMV candidate, if the merge candidate list is not full, the HMVP candidates stored in the HMVP table can be used to fill in the merge candidate list. Considering that a block usually has a high correlation with the blocks in the closest neighborhood from the perspective of motion information, the HMVP candidates in the table are inserted in descending order of index. First, the last entry of the table is added to the list, and finally, the first entry is added. Similarly, redundancy removal is applied to the HMVP candidates. When the total number of available merge candidates reaches the maximum number of merge candidates that can be signaled, the merge candidate list construction process ends. When starting the encoding / decoding of a new slice, the HMVP table becomes empty. Whenever there is an inter-coded block, the relevant motion information is added as a new HMVP candidate to the last entry of the table. The overall encoding flow is shown in Figure 25. Whenever there is an inter-coded block, the relevant motion information is added as a new HMVP candidate to the last entry of the table. The overall encoding flow is shown in Figure 25.
[0206] The HMVP candidates can be used in both the AMVP and the merge candidate list construction processes. Figure 26 shows the modified merge candidate list construction process (highlighted in blue). After inserting the TMV candidate, if the merge candidate list is not full, the HMVP candidates stored in the HMVP table can be used to fill in the merge candidate list. Considering that a block usually has a high correlation with the blocks in the closest neighborhood from the perspective of motion information, the HMVP candidates in the table are inserted in descending order of index. First, the last entry of the table is added to the list, and finally, the first entry is added. Similarly, redundancy removal is applied to the HMVP candidates. When the total number of available merge candidates reaches the maximum number of merge candidates that can be signaled, the merge candidate list construction process ends. Considering that a block usually has a high correlation with the blocks in the closest neighborhood from the perspective of motion information, the HMVP candidates in the table are inserted in descending order of index. First, the last entry of the table is added to the list, and finally, the first entry is added. Similarly, redundancy removal is applied to the HMVP candidates. When the total number of available merge candidates reaches the maximum number of merge candidates that can be signaled, the merge candidate list construction process ends. When the total number of available merge candidates reaches the maximum number of merge candidates that can be signaled, the merge candidate list construction process ends.
[0207] 2.4 Rounding of MV
[0208] In VVC, when the MV is right-shifted, it is required to round the MV towards 0. In a formulated form, when the MV (MVx, MVy) is right-shifted by N bits, the resulting MV' (MVx', MVy') is derived as follows. In a formulated form, when the MV (MVx, MVy) is right-shifted by N bits, the resulting MV' (MVx', MVy') is derived as follows. The resulting MV' (MVx', MVy') is derived as follows.
[0209] MVx’=(MVx+((1<<N)>>1)-(MV_x>=0?1:0))>>N ;
[0210] MVy’=(MVy+((1<<N)>>1)-(MVy>=0?1:0))>>N;
[0211] 2.5 Embodiments of reference picture resampling (RPR)
[0212] ARC, also known as reference picture resampling (RPR), is incorporated into existing and future video standards and specifications.
[0213] In some embodiments of RPR, when the collocated picture has a different resolution from the current picture, TMVP is disabled. Also, when the resolution of the reference picture is different from the current picture, BDOF and DMVR are disabled.
[0214] When the resolution of the reference picture is different from the resolution of the current picture, for handling normal MC, the interpolation section is defined as follows.
[0215] 8.5.6.3 Fractional sample interpolation processing
[0216] 8.5.6.3.1 Overview
[0217] The input to this process is as follows.
[0218] - The luminance position (xSb, ySb) that defines the top - left sample of the current coded sub - block with respect to the top - left luminance sample of the current picture - A variable sbWidth that defines the width of the current coded sub - block, - A variable sbHeight that defines the height of the current coded sub - block, - Motion vector offset mvOffset, - Fine-tuned motion vector refMvLX, - Selected reference picture sample array refPicLX, - 1 / 2 sample interpolation filter index hpelIfIdx, - Bidirectional optical flow flag bdofFlag, - Variable cIdx that defines the color component index of the current block.
[0219] The output of this process is as follows. - (sbWidth + brdExtSize) × (sbHeight + brdExtSize) array predSamplesLX of predicted sample values.
[0220] The prediction block boundary expansion size brdExtSize is derived as follows. brdExtSize = (bdofFlag || (inter_affine_flag [xSb][ySb] && sps_affine_prof_enabled_fl ag))? 2 : 0 (8-752)
[0221] The variable fRefWidth is set equal to PicOutput WidthL of the reference picture in the luma samples.
[0222] The variable fRefHeight is set equal to PicOutpu tHeightL of the reference picture in the luma samples.
[0223] The motion vector mvLX is set equal to (refMvLX - mvOffset) . - When cIdx is equal to 0, the following applies. - The scaling factor and its fixed-point representation are defined as follows. hori_scale_fp=((fRefWidth<<14)+(PicOut putWidthL>>1)) / PicOutputWidthL (8-753) vert_scale_fp=((fRefHeight<<14)+(PicOu tputHeightL>>1)) / PicOutputHeightL (8-754 ) - Let (xIntL, yIntL) be the luminance position given in full sample units, and (xFracL, yFracL) be the offset given in 1 / 16 sample units. These variables are used only in this section to define the fractional sample positions in the reference sample array refPicLX. - Set the upper left coordinate of the bounding block (xSbInt L , ySbInt L ) for reference sample padding to be equal to (xSb+(mvLX[0]>>4), ySb+(mvLX[1]>>4)). >4)). - For each luminance sample position (x L = 0..sbWidth-1+brdExtSize, y L = 0..sbHeight- 1+brdExtSize) in the predicted luminance sample array predSamplesLX, the corresponding predicted luminance sample value predSampl esLX[x L [y L is derived as follows. - Let (refxSb L , refySb L ) and (refx L , refy L ) be the luminance positions pointed to by the motion vector (refMvLX[0], refMvLX 1]) given in 1 / 16 sample units. The variables refxSb L , refxL , refySb L , re fy L is derived as follows. refxSb L = ((xSb << 4) + refMvLX[0]) * hori_sca le_fp (8 - 755) refx L = ((Sign(refxSb) * ((Abs(refxSb) + 128 ) >> 8) + x L * ((hori_scale_fp + 8) >> 4)) + 32) >> 6 (8 - 756) refySb L = ((ySb << 4) + refMvLX[1]) * vert_sca le_fp (8 - 757) refyL = ((Sign(refySb) * ((Abs(refySb) + 128 ) >> 8) + yL * ((vert_scale_fp + 8) >> 4)) + 32) >> 6 (8 - 758) - Variable xInt L , yInt L , xFrac L , and yFrac L are derived as follows . xInt L = refx L >> 4 (8 - 759) yInt L = refy L >> 4 (8 - 760) xFrac L = refx L & 15 (8 - 761) yFrac L = refy L & 15 (8 - 762) - Whether bdofFlag is equal to TRUE (sps_affine_prof_en abled_flag is equal to TRUE, inter_affine_flag[xS b][ySb] is equal to TRUE), if one or more of the following conditions are true, the predicted luminance sample value predSamplesLX[x L [y L is, as defined in clause 8.5.6.3.3, using as input (xInt + (xFrac L >> 3) - 1), yInt L + (yFrac L >> 3) - 1) and refPicLX, to call the luminance integer sample L extraction process to derive. 1. x is equal to 0. L 2. x L is equal to sbWidth + 1. 3. y L is equal to 0. 4. y L is equal to sbHeight + 1. - Otherwise, as defined in clause 8.5.6.3.2, (xIntL - (b rdExtSize > 0? 1 : 0), yIntL - (brdExtSize > 0? 1 : 0 )), (xFracL, yFracL), (xSbInt L , ySbInt L ), ref PicLX, hpelIfIdx, sbWidth, sbHeight, and (xSb ), ySb) as input, to call the luminance sample 8-tap interpolation filtering process to derive the predicted luminance sample value predSamplesLX[xL][yL] - Otherwise (cIdx is not equal to 0), the following applies. - Let (xIntC, yIntC) be the chroma position given in full sample units and (xFracC, yFracC) be the offset given in 1 / 32 sample units These variables are used in this section only to represent common fractions in the reference sample sequence refPicLX. Used to define the location of the sample. - Bounding block for reference sample padding (xSbIntC, ySbIn The top left coordinate of tC is ((xSb / SubWidthC)+(mvLX[0]>>5), It is set equal to (ySb / SubHeightC)+(mvLX[1]>>5)). - At each chroma sample position in the predicted chroma sample array predSamplesLX ( xC=0..sbWidth-1, yC=0..sbHeight-1) The predicted chroma sample values predSamplesLX[xC][yC] are as follows: Derive as follows. - (refxSb C ,refySb C ) and (refx C ,refy C ) to 1 The motion vectors (mvLX[0], mvLX[1]) are given in units of 32 samples. The chroma position is the variable refxSb C , refySb C , refx C , refy C teeth , which is derived as follows. refxSb C =((xSb / SubWidthC<<5)+mvLX[0])* hori_scale_fp (8-763) refx C =((Sign(refxSb C )*((Abs(refxSb C )+ 256)>>9)+xC*((hori_scale_fp+8)>>4))+16)> >5 (8-764) refySb C= ((ySb / SubHeightC << 5) + mvLX[1]) *vert_scale_fp (8 - 765) refy C = ((Sign(refySb C ) * ((Abs(refySb C ) + 256) >> 9) + yC * ((vert_scale_fp + 8) >> 4)) + 16) > > 5 (8 - 766) - Variable xInt C 、yInt C 、xFrac C 、yFrac C are derived as follows. Derivation. xInt C = refx C >> 5 (8 - 767) yInt C = refy C >> 5 (8 - 768) xFrac C = refy C & 31 (8 - 769) yFrac C = refy C & 31 (8 - 770) - The predicted sample value predSamplesLX[xC][yC] is derived by calling the process specified in 8.5.6.3.4 with (xIntC, yIntC), (xFracC, yFracC), (xSbIntC, ySbIntC) , sbWidth, sbHeight, and refPicLX as inputs. Derivation by calling the process specified in 8.5.6.3.4.
[0224] 8.5.6.3.2 Luminance Sample Interpolation Filtering Process
[0225] The inputs to this process are as follows. - The luminance position in the full sample unit (xInt L , yInt L ), - The luminance position in the fractional sample unit (xFracL , yFrac L ), - Border for padding of reference samples with respect to the top - left luminance sample of the reference picture Full - sample unit (xSbInt L , ySb Int L ) defining the luminance position in, - Luminance reference sample array refPicLX L , - 1 / 2 - sample interpolation filter index hpelIfIdx, - Variable sbWidth defining the width of the current sub - block, - Variable sbHeight defining the height of the current sub - block, - Luminance position (xSb, ySb) defining the top - left sample of the current sub - block with respect to the top - left luminance sample of the current picture ,
[0226] The output of this process is the predicted luminance sample value predSampleLX L .
[0227] Variables shift1, shift2, and shift3 are derived as follows. - Set variable shift1 equal to Min(4, BitDepth Y - 8), set variable s hift2 equal to 6, and set variable shift3 equal to Max(2, 14 - BitDept h Y ). - Set variable picW equal to pic_width_in_luma_samples and set variable picH equal to pic_height_in_luma_samples .
[0228] Luminance for each 1 / 16 - fractional sample position p equal to L xFrac L or yFrac Interpolation filter coefficient f L [p] is derived as follows. - When MotionModelIdc[xSb][ySb] is greater than 0 and sbWidt h and sbHeight are both equal to 4, the luminance interpolation filter coefficient f L [p] is specified in Table 8-12. - Otherwise, based on hpelIfIdx, the luminance interpolation filter coefficient f L [p] is specified in Table 8-11.
[0229] When i = 0..7, the luminance i , yInt i position in the full sample unit (xInt is derived as follows. - When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies. xInt i = Clip3(SubPicLeftBoundaryPos, SubPi cRightBoundaryPos, xInt L + i - 3) (8-771) yInt i = Clip3(SubPicTopBoundaryPos, SubPic BotBoundaryPos, yInt L + i - 3) (8-772) - Otherwise (when subpic_treated_as_pic_flag[Sub PicIdx] is equal to 0), the following applies. xInt i = Clip3(0, picW - 1, sps_ref_wraparound _enabled_flag? ClipH((sps_ref_wraparound_offset_minus (1 + 1) * MinCbSizeY, picW, xInt L + i - 3): (8 - 773) xInt L + i - 3) yInt i = Clip3(0, picH - 1, yInt L + i - 3) (8 - 77 4)
[0230] When i = 0..7, the luminance position in the full - sample unit is further modified as follows. xInt i = Clip3(xSbInt L - 3, xSbInt L + sbWidth + 4 , xInt i ) (8 - 775) yInt i = Clip3(ySbInt L - 3, ySbInt L + sbHeight + 4, yInt i ) (8 - 776)
[0231] The predicted luminance sample value predSampleLX L is derived as follows. - When both xFrac L and yFrac L are equal to 0, the value of predSampleL X L is derived as follows. predSampleLX L = refPicLX L [xInt3][yInt3] << s hift3 (8 - 777) - Otherwise, when xFrac L is not equal to 0 and yFrac L is equal to 0, then predSampleLX L is derived as follows. predSampleLXL =(Σ 7 i=0 f L [xFrac L [i]*ref PicLX L [xInt i [yInt3])>>shift1 (8 - 778) - Otherwise, if xFrac L is equal to 0 and yFrac L is not equal to 0, the value of p redSampleLX L is derived as follows. predSampleLX L =(Σ 7 i=0 f L [yFrac L [i]*ref PicLX L [xInt3][yInt i )>>shift1 (8 - 779) - Otherwise, if xFrac L is not equal to 0 and yFrac L is not equal to 0 , the value of predSampleLX L is derived as follows. - The sample array temp[n] for n = 0..7 is derived as follows. temp[n]=(Σ 7 i=0 f L [xFrac L [i]*refPicLX L [xI nt i [yInt n )>>shift1 (8 - 780) - The predicted luminance sample value predSampleLX L is derived as follows. predSampleLX L =(Σ 7 i=0 f L [yFrac L [i]*temp[i )>>shift2 (8-781)
[0232]
Table 4
[0233]
Table 5
[0234] 8.5.6.3.3 Luminance Integer Sample Extraction Processing
[0235] The input to this process is as follows. - Luminance position in the full sample unit (xInt L , yInt L ), - Luminance reference sample array refPicLX L ,
[0236] The output of this process is the predicted luminance sample value predSampleLX L . This variable shift is set equal to Max(2, 14 - BitDepth Y ). . The variable picW is set equal to pic_width_in_luma_samples , and the variable picH is set equal to pic_height_in_luma_samples .
[0237] The luminance position in the full sample unit (xInt, yInt) is derived as follows . xInt = Clip3(0, picW - 1, sps_ref_wraparound _enabled_flag? ClipH((sps_ref_wraparound_offset_min us1 + 1)*MinCbSizeY, picW, xInt L ): xInt L )(8 - 7 82) yInt = Clip3(0, picH - 1, yInt L ) (8 - 783)
[0238] The predicted luminance sample value predSampleLX L is derived as follows. predSampleLX L = refPicLX L [xInt][yInt] << s hift3 (8 - 784)
[0239] 8.5.6.3.4 Chroma sample interpolation processing
[0240] The input to this process is as follows. - The chroma position in the full - sample unit (xInt C , yInt C ), - The chroma position in 1 / 32 fractional - sample units (xFrac C , yFrac C ), - The boundary for reference sample padding for the top - left chroma sample of the reference picture defining the top - left sample of the block, the chroma position in the full - sample unit (xSbIntC, ySb IntC), - The variable sbWidth defining the width of the current sub - block, - The variable sbHeight defining the height of the current sub - block, - The chroma reference sample array refPicLX C .
[0241] The output of this process is the predicted chroma sample value predSampleLX C .
[0242] The variables shift1, shift2, and shift3 are derived as follows. - Set the variable shift1 equal to Min(4, BitDepth C -8), set the variable sh ift2 equal to 6, and set the variable shift3 equal to Max(2, 14 - BitDepth C ). - The variable picW C is set equal to pic_width_in_luma_samples / SubW idthC, and the variable picH C is set equal to pic_height_in_luma _samples / SubHeightC.
[0243] Table 8 - 13 shows the chroma interpolation filter coefficients f C or yFrac C equal to each 1 / 32 fractional sample position p of f C [p].
[0244] The variable xOffset is set equal to (sps_ref_wraparound_offset_m inus1 + 1) * MinCbSizeY) / SubWidthC.
[0245] For i = 0..3, the chroma i position at the full - sample unit (xInt i ), yInt is derived as follows. - When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies. xInt i = Clip3(SubPicLeftBoundaryPos / SubWi dthC, SubPicRightBoundaryPos / SubWidthC, xI nt L+i) (8 - 785) yInt i = Clip3(SubPicTopBoundaryPos / SubHei ghtC, SubPicBotBoundaryPos / SubHeightC, yIn t L +i) (8 - 786) - Otherwise (subpic_treated_as_pic_flag[Sub PicIdx] equals 0), the following applies. xInt i = Clip3(0, picW C - 1, sps_ref_wraparoun d_enabled_flag? ClipH(xOffset, picW C , xInt C +i - 1): (8 - 787) xInt C +i - 1) yInt i = Clip3(0, picH C - 1, yInt C +i - 1) (8 - 78 8)
[0246] The chroma position in the full - sample unit (xInt i , yInt i ) for i = 0.. 3 is further modified as follows. xInt i = Clip3(xSbIntC - 1, xSbIntC + sbWidth + 2 , xInt i ) (8 - 789) yInt i = Clip3(ySbIntC - 1, ySbIntC + sbHeight + 2, yInt i ) (8 - 790)
[0247] The predicted chroma sample value predSampleLX Cis derived as follows. - xFrac C and yFrac C are both equal to 0, predSampleL X C The value of is derived as follows. predSampleLX C = refPicLX C [xInt1][yInt1] << shift3(8 - 791) - Otherwise, xFrac C is not equal to 0 and yFrac C is equal to 0, p redSampleLX C The value of is derived as follows. predSampleLX C =(Σ 3 i=0 f C [xFrac C [i]*ref PicLX C [xInt i [yInt1]) >> shift1(8 - 792) - Otherwise, xFrac C is equal to 0 and yFrac C is not equal to 0, p redSampleLX C The value of is derived as follows. predSampleLXC=(Σ 3 i=0 f C [yFrac C [i]*ref PicLX C [xInt1][yInt i ) >> shift1(8 - 793) - Otherwise, xFrac C is not equal to 0 and yFrac C is not equal to 0 predSampleLX C The value of is derived as follows. - The sample array temp[n] for n = 0..3 is derived as follows. temp[n]=(Σ 3 i=0 f C [xFrac C [i]*refPicLX C [xInt i [yInt n )>>shift1 (8 - 794) - The predicted chroma sample value predSampleLX C is derived as follows. predSampleLX C =(f C [yFrac C [0]*temp[0] +f C [yFrac C [1]*temp[1]+f C [yFrac C [2]*tem p[2]+ (8 - 795) f C [yFrac C [3]*temp[3])>>shift2
[0248]
Table 6
Table 7
[0249] 2.6 Embodiment Using Sub - pictures
[0250] In the current syntax design of sub - pictures in existing implementations, the position and dimensions of the sub - pictures are derived as follows.
[0251]
Table 8
[0252] When the subpics_present_flag is 1, it indicates that there are subpicture parameters in the current SPS RBSP syntax. When the subpics_present_flag is 0, it indicates that there are no subpicture parameters in the current SPS RBSP syntax. Note 2 - If the bitstream is the result of sub-bitstream extraction processing and only contains a subset of subpictures of the input bitstream to the sub-bitstream extraction processing, it may be necessary to set the value of subpics_present_flag to 1 in the RBSP of the SPS. max_subpics_minus1 plus 1 defines the maximum number of subpictures that can exist in the CVS. max_subpics_minus1 must be in the range from 0 to 254. The value 255 is reserved for future use by ITU-T|ISO / IEC. subpic_grid_col_width_minus1 plus 1 defines the width of each element of the subpicture identifier grid in units of 4 samples. The length of the syntax element is Ceil(Log2(pic_width_max_in_luma_samples / 4)) bits.
[0253] The variable NumSubPicGridCols is derived as follows. NumSubPicGridCols = (pic_width_max_in_luma_samples + subpic_grid_col_width_minus1 * 4 + 3) /
[0254]
[0255] (subpic_grid_col_width_minus1*4+4) (7- 5)
[0256] subpic_grid_row_height_minus1 plus1 specifies the height of each element of the sub-picture identifier grid in 4-sample units. The length of the syntax element is Ceil(Log2(pic_height_max_in_luma_samples / 4)) bits.
[0257] The variable NumSubPicGridRows is derived as follows. NumSubPicGridRows=(pic_height_max_in_lum a_samples+subpic_grid_row_height_minus1* 4+3) / (subpic_grid_row_height_minus1*4+4) (7 -6)
[0258] subpic_grid_idx[i][j] specifies the sub-picture index of the grid position (i, j). The length of the syntax element is Ceil(Log2(max_subp ics_minus1+1)) bits.
[0259] The variables SubPicTop[subpic_grid_idx[i][j]], SubP icLeft[subpic_grid_idx[i][j]], SubPicWidt h[subpic_grid_idx[i][j]], SubPicHeight[su bpic_grid_idx[i][j]], and NumSubPics are derived as follows
[0260] NumSubPics=0 for(i = 0; i < NumSubPicGridRows; i++){ for(j = 0; j < NumSubPicGridCols; j++){ if(i == 0) SubPicTop[subpic_grid_idx[i][j]] = 0 else if(subpic_grid_idx[i][j] != subpic_ grid_idx[i - 1][j]){ SubPicTop[subpic_grid_idx[i][j]] = i SubPicHeight[subpic_grid_idx[i - 1][j]] = i - SubPicTop[subpic_grid_idx[i - 1][j]] } if(j == 0) SubPicLeft[subpic_grid_idx[i][j]] = 0 ( 7 - 7) else if(subpic_grid_idx[i][j] != subpic_ grid_idx[i][j - 1]){ SubPicLeft[subpic_grid_idx[i][j]] = j SubPicWidth[subpic_grid_idx[i][j]] = j - SubPicLeft[subpic_grid_idx[i][j - 1]] } if(i == NumSubPicGridRows - 1) SubPicHeight[subpic_grid_idx[i][j]] = i - SubPicTop[subpic_grid_idx[i - 1][j]] + 1 if(j == NumSubPicGridRows - 1) SubPicWidth[subpic_grid_idx[i][j]] = j-SubPicLeft[subpic_grid_idx[i][j-1]]+1 if(subpic_grid_idx[i][j]>NumSubPics) NumSubPics=subpic_grid_idx[i][j] } }
[0261] If subpic_treated_as_pic_flag[i] is 1, The ith subpicture of each coded picture is filtered by the loop filtering operation. , which specifies that the picture is treated as a subpic_treated picture in the decoding process. If _as_pic_flag[i] is 0, the i-th coded picture in the CVS is Subpictures are not treated as pictures in the decoding process, except for in-loop filtering operations. If not present, subpic_treated_as_p The value of ic_flag[i] is inferred to be equal to 0.
[0262] 2.7 Combined Inter-Intra Prediction (CIIP)
[0263] Inter-intra combined prediction is a special merge candidate. We adopt Inter-Intra Prediction (IIP) for VVC. This means that W<= Can only be enabled for WxH blocks of 64 and H<=64.
[0264] 3. Shortcomings of existing implementations
[0265] In the current VVC design, ATMVP has the following problems: 1) There is an inconsistency between whether to apply ATMVP at the slice level and the CU level. 2) In the slice header, ATMVP may be valid whether or not TMVP is invalid. On the other hand, signal the ATMVP flag before the TMVP flag. 3) Masking is always performed regardless of whether the MV is compressed. 4) The corresponding valid area may be too large. 5) Deriving TMV is very complex. 6) Even if ATMVP is not available, a better default MV is desirable. 7) The MV scaling method in ATMVP may not be efficient. 8) ATMVP should consider CPR cases. 9) Even if affine prediction is disabled, the default zero-affine merge candidates may be included in the list. 10) The current picture is treated as a long-term reference picture and other pictures are treated as short-term reference pictures. For both ATMVP candidates and TMVP candidates, the motion information from the temporal blocks in the collocated picture is scaled to the reference picture with a fixed reference index (i.e., 0 for each reference picture list in the current design). However, when the CPR mode is enabled, the current picture is also treated as a reference picture and the current picture may be added to the reference picture list 0 (RefPicList0) with an index equal to 0. a. In TMVP, when the temporal block is encoded in the CPR mode and the reference picture in RefPicList0 is a short reference picture, the TMVP candidate is set to be unavailable. b. When the reference picture in RefPicList0 with index 0 is the current picture Yes, and the current picture is an Intra Random Access Point (IRAP) picture In this case, the ATMVP candidates are set to unavailable. c. For the ATMVP sub-blocks within one block, when deriving the motion information of the sub-blocks from one temporal block if this temporal block is coded in CPR mode a default ATMVP candidate (derived from one temporal block specified by the starting TMV and the center position of the current block is used to fill the motion information of this sub-block. ) 11) The MV is right-shifted to integer precision but does not follow the rounding rules in VVC. 12) In ATMVP, the MV (MVx, MVy) used to define the positions of the corresponding blocks in different pictures (e.g., when TMV is 0 ) is used as it is to point to the collocated picture. This is based on the assumption that all pictures have the same resolution. However, when RPR is enabled, different picture resolutions may be used. Similar issues also exist regarding identifying the corresponding blocks in the collocated picture for deriving the sub-block motion information. 13) If the width or height of one block is greater than 32 and the size of the maximum transform block is 32 for the CIIP coded block an intra prediction signal is generated in the CU size while the inter prediction signal is generated in the TU size (recursively splitting the current block into multiple 32× 32 blocks). Deriving the intra prediction signal using the CU results in lower efficiency.
[0266] There are some problems with the current design. First, the RefPicList0 If the reference picture of is the current picture and the current picture is not an IRAP picture The ATMVP procedure is still called, but any temporal motion vectors are Since it is not possible to scale the picture, the ATMVP procedure is not available. We were unable to find a suitable ATMVP candidate.
[0267] 4. Examples of embodiments and techniques
[0268] The following list of technologies and embodiments is to be considered as examples to illustrate the general concept. These techniques should not be interpreted narrowly. may be combined in any manner in the encoder or decoder embodiment. do.
[0269] 1. Whether TMVP is permitted and / or CPR is used depends on: To determine / parse the maximum number of candidates in a subblock merging candidate list, and and / or be considered to determine whether the ATMVP candidate should be added to the candidate list. Let ML be the maximum number in the subblock merging candidate list. a) In one example, the ATMVP selects the top candidate in the sub-block merge candidate list. In determining or parsing a large number, the ATMVP use flag is off (equal to 0). or is presumed not applicable if TMVP is disabled. In one example, the ATMVP use flag is on (equal to 1) and the TMVP If is disabled, the ATMVP candidate is selected from the subblock merge candidate list or the AT It will not be added to the MVP candidate list. ii. In one example, if the ATMVP usage flag is on (equal to 1), the TMV P is disabled, and the affinity usage flag is off (equal to 0), then ML is set equal to 0, which means that sub-block merge is not applicable. iii. In one example, if the ATMVP usage flag is on (equal to 1), the TM VP is enabled, and the affinity usage flag is off (equal to 0), then ML is set equal to 1 appropriately. b) In one example, the ATMVP is presumed not to be applicable when determining or parsing the maximum number of candidates in the sub-block merge candidate list if the ATMVP usage flag is off (equal to 0) or the collocated reference picture of the current picture is the current picture itself. i. In one example, if the ATMVP usage flag is on (equal to 1) and the collocated reference picture of the current picture is the current picture itself, the AT MVP candidate will not be added to the sub-block merge candidate list or the ATMVP candidate list. ii. In one example, if the ATMVP usage flag is on (equal to 1), the collocated reference picture of the current picture is the current picture itself, and the affinity usage flag is off (equal to 0), then ML is set equal to 0, which means that sub- block merge is not applicable. iii. In one example, if the ATMVP usage flag is on (equal to 1) and the collocated reference picture of the current picture is not the current picture itself, and the aff inity usage flag is off (equal to 0), then ML is set equal to 0, which means that sub- block merge is not applicable. iii. In one example, if the ATMVP usage flag is on (equal to 1) and the collocated reference picture of the current picture is not the current picture itself, and the aff inity usage flag is off (equal to 0), then ML is set equal to 0, which means that sub- block merge is not applicable. iii. In one example, if the ATMVP usage flag is on (equal to 1) and the collocated reference picture of the current picture is not the current picture itself, and the aff If the in-use flag is off (equal to 0), ML is set equal to 1. c) In one example, if the ATMVP use flag is off (equal to 0), or the reference picture having reference picture index 0 in reference list 0 is the current picture itself, ATMVP is presumed not to be applicable when determining or parsing the maximum number of candidates in the sub-block merge candidate list. i. In one example, if the ATMVP use flag is on (equal to 1), and the collocated reference picture having reference picture index 0 in reference list 0 is the current picture itself, the ATMVP candidates are not added to the sub-block merge candidate list or the ATMVP candidate list. ii. In one example, if the ATMVP use flag is on (equal to 1), and the reference picture having reference picture index 0 in reference list 0 is the current picture itself, and the affine use flag is off (equal to 0), ML is set equal to 0, which means that sub-block merge is not applicable. iii. In one example, if the ATMVP use flag is on (equal to 1), and the reference picture having reference picture index 0 in reference list 0 is not the current picture itself, and the affine use flag is off (equal to 0), ML is set equal to 1. d) In one example, if the ATMVP use flag is off (equal to 0), or the reference picture having reference picture index 0 in reference list 1 is the When determining or parsing numbers, it is presumed to be inapplicable. i. In one example, the ATMVP usage flag is on (equal to 1), and the collocated reference picture with reference picture index 0 in reference list 1 is the current picture itself, and the ATMVP candidate is not added to the sub-block merge candidate list or the ATMVP candidate list. ii. In one example, the ATMVP usage flag is on (equal to 1), and the reference picture with reference picture index 0 in reference list 1 is the current picture itself, and when the affine usage flag is off (equal to 0), ML is set to be equal to 0, which means that sub-block merge is not applicable. iii. In one example, the ATMVP usage flag is on (equal to 1), and the reference picture with reference picture index 0 in reference list 1 is not the current picture itself, and when the affine usage flag is off (equal to 0), ML is set to be equal to 1.
[0270] 2. When TMVP is disabled at the slice / tile / picture level, ATMVP is implicitly disabled, and the ATMVP flag is not signaled. a) In one example, the ATMVP flag is signaled after the TMVP flag in the slice header / tile header / PPS. b) In one example, the ATMVP or / and TMVP flag may not be signaled in the slice header / tile header / PPS, and is only signaled in the SPS header.
[0271] 3. Whether to mask the corresponding position in ATMVP and how to mask it depends on whether the MV is compressed and how it is compressed. Let (xN,yN) be the corresponding position calculated using the coordinates of the current block / sub-block and the starting motion vector (e.g., TMV) in the collocated picture. a) In one example, if there is no need to compress the MV (e.g., when sps_disable_motioncompression signaled in SPS is 1), (xN ,yN) is not masked. Otherwise, (since the MV needs to be compressed) (xN,y N) is masked as xN=xN&MASK, yN=yN&MASK. Here, MASK is equal to ~(2 -1), and M is an integer such as 3 or 4. M b) The MV compression method for the MV storage result in each 2 K ×2 K block shares the same motion information and defines the mask in ATMVP processing as ~(2 -1). K may be equal to M, for example, it is assumed that M = K + 1. M c) The MASK used for ATMVP and TMV may be the same or different.
[0272] 4. In one example, the MV compression method may be flexible. a) In one example, the MV compression method can be selected between non-compression, 8×8 compression (M = 3 in Bullet3.a), or 16×16 compression (M = 4 in Bullet3.a). b) In one example, the MV compression method is signaled by VPS / SPS / PPS / slice header / temporal layer. It may be signaled in the I group header. c) In one example, the MV compression method may be set differently in different standard profiles / levels / layers.
[0273] 5. The effective corresponding region in ATMVP may be adaptable. a) For example, the effective corresponding region may depend on the width and height of the current block. b) For example, the effective corresponding region may depend on the MV compression method. i. In one example, when the MV compression method is not used, the effective corresponding region is smaller, and when the MV compression
[0274] method is used, the effective corresponding region is larger. 6. The effective corresponding region in ATMVP may be based on a basic region having a size M×N smaller than the CTU region. For example, in VTM-3.0, the size of the CTU may be 128×128, and the size of the basic region may be 64×64. Let the width and height of the current block be W and H. 7 is shown as an example. i. For example, assuming that the upper left position of the placed basic region is (xBR, yBR), the corresponding position at (xN, yN) is clipped to the valid region xBR <= xN < xBR + M + 4; yBR <= yN <
[0275] Figure 27 shows an exemplary embodiment of the proposed valid region when the current block is within the basic region (BR). Exemplary embodiments are shown.
[0276] Figure 28 shows an exemplary embodiment of the valid region when the current block is not within the basic region. shown. b) In one example, when W > M and H > N, it means that the current block is not within one basic region, and the current block is divided into multiple parts. Each part has an individual valid corresponding region within the ATMVP. For position A in the current block, its corresponding position B in the located block should be within the valid corresponding region of the part where position A is located. i. For example, divide the current block into non - overlapping basic regions. The valid region corresponding to one basic region is the extension in its collocation basic region and the collocated picture. Figure 28 shows an example. 1. For example, assume that position A of the current block is in one basic region R. Let the collocation basic region of R in the collocated picture be CR. The corresponding position of A in the collocated block is position B, and the upper - left position of CR is ( xCR, yCR). Next, (xN, yN) of position B is clipped to xCR <= xN < xCR + M + 4; yCR <= yN < yCR + N.
[0277] 7. The motion vectors for defining the positions of corresponding blocks in different pictures used in the ATMVP can be derived as follows (for example, TMV in 2.3 .5.1.2). . a) In one example, the TMV is always set equal to a default MV such as (0, 0). It is obtained. i. In one example, the default MV is signaled in the VPS / SPS / PPS / slice header / tile group header / CTU / CU. b) In one example, the TMV is set to one MV stored in the HMVP table in the following way MV. i. If the HMVP list is empty, the TMV is set equal to the default MV, for example (0,0 ). ii. Otherwise (if the HMVP list is not empty), 1. The TMV may be set equal to the first element stored in the HMVP table Yes. 2. Alternatively, the TMV may be set equal to the last element stored in the HMVP table Set. 3. Alternatively, the TMV may be set equal to a specific MV stored in the HMVP table Set. a. In one example, the specific MV refers to reference list 0. b. In one example, the specific MV refers to reference list 1. c. In one example, the specific MV refers to a specific reference picture in reference list 0 , for example, the reference picture having index 0. d. In one example, the specific MV refers to a specific reference picture in reference list 1 , for example, the reference picture having index 0. e. In one example, the specific MV refers to the collocated picture . 4. Alternatively, if a specific MV stored in the HMVP table (for example, as described in bulle t3.) is not found, the TMV may be set equal to the default MV. a. In one example, only the first element stored in the HMVP table is searched, Find a specific MV. b. In one example, only the last element stored in the HMVP table is retrieved, to find a specific MV. c. In one example, some or all of the elements stored in the HMVP table are examined to find a specific MV. 5. Alternatively, further, the TMV obtained from the HMVP cannot reference the current picture itself. 6. Alternatively, further, the TMV obtained from the HMVP table may be scaled according to the collocated picture if not referenced. c) In one example, the TMV is set to one MV of one specific neighboring block. Other neighboring blocks are not included. i. The specific neighboring blocks may be blocks A0, A1, B0, B1, B2 in FIG. 22. ii. The TMV may be set equal to the default MV in the following cases. 1. There is no specific neighboring block. 2. The specific neighboring blocks are not inter-coded. iii. The TMV may be set equal to the specific MV stored in the specific neighboring block. 1. In one example, the specific MV refers to reference list 0. 2. In one example, the specific MV refers to reference list 1. 3. In one example, the specific MV refers to a specific reference picture in reference list 0, for example, the reference picture having index 0. 4. In one example, the specific MV refers to a specific reference picture in reference list 1, for example, the reference picture having index 0. 5. In one example, the specific MV refers to the collocated picture. 6. If a specific MV stored in a block in a specific neighborhood cannot be found, the TMV may be set equal to the default MV. iv. The TMV obtained from a block in a specific neighborhood may be scaled according to the picture it references if it does not reference the picture that is cached. v. The TMV obtained from a block in a specific neighborhood cannot reference the current picture itself.
[0278] 8. As disclosed in 2.3.5.1.2, the MVs default0 and MVdefault1 used in ATMVP may be derived as follows. a) In one example, MVdefault0 and MVdefault1 are set equal to (0,0 ). b) In one example, MVdefaultX (X = 0 or 1) is derived from HMVP. i. If the HMVP list is empty, MVdefaultX is set equal to a predefined default MV such as (0,0). 1. The predefined default MV may be signaled in the VPS / SPS / PPS / slice header / tile group header / CTU / CU. ii. Otherwise (if the HMVP list is not empty), 1. MVdefaultX may be set equal to the first element stored in the HMVP table. 2. MVdefaultX may be set equal to the last element stored in the HMVP table. 3. MVdefaultX may be set equal to only a specific MV stored in the HMVP table. a. In one example, the specific MV references reference list X. b. In one example, a particular MV refers to a particular reference picture in reference list X , for example, the reference picture having index 0. 4. If a particular MV stored in the HMVP table is not found, MVdef aultX may be set equal to a predefined default MV. a. In one example, only the first element stored in the HMVP table is searched . b. In one example, only the last element stored in the HMVP table is searched . c. In one example, some or all of the elements stored in the HMVP table are searched . 5. MVdefaultX obtained from the HMVP table may be scaled to match a collocated picture if not referenced . 6. MVdefaultX obtained from HMVP cannot reference the current picture itself . c) In one example, MVdefaultX (X = 0 or 1) is derived from neighboring blocks . i. The neighboring blocks may include blocks A0, A1, B0, B1, B2 in FIG. 22 . 1. For example, MVdefaultX is derived using only one of these blocks . 2. Alternatively, MVdefaul tX is derived using some or all of these blocks. a. These blocks are checked in order until a valid MVdefaultX is found . 3. A valid MVdefaultX is found from one or more selected neighboring blocks If not available, it is set equal to a predefined default MV such as (0,0). It is set to this value. a. The predefined default MV may be signaled in the VPS / SPS / PPS / slice header / tile group header / CTU / CU. ii. In the following cases, no valid MVdefaultX can be found from the blocks in a specific neighborhood. 1. There are no blocks in a specific neighborhood. It cannot be found. 1. There are no blocks in a specific neighborhood. 2. The blocks in a specific neighborhood are not inter-coded. iii. MVdefaultX may be set equal only to a specific MV stored in a block in a specific neighborhood. It may be set equal only to this value. 1. In one example, the specific MV refers to reference list X. 2. In one example, the specific MV refers to a specific reference picture in reference list X, for example, the reference picture with index 0. It may be mentioned that the reference picture has an index of 0. iv. The MVdefaultX obtained from a block in a specific neighborhood may be scaled to a specific reference picture, for example, the reference picture with index 0 in reference list X. It may be scaled to this value. It may be scaled to this value. v. The MVdefaultX obtained from a block in a specific neighborhood cannot refer to the current picture itself. It cannot refer to the current picture itself.
[0279] 9. For either a sub-block or a non-sub-block ATMVP candidate, if one temporal block for one sub-block / entire block in the collocated picture is coded in CPR mode, instead, one default motion candidate may be utilized. a) In one example, the default motion candidate is associated with the center position of the current block. It may be associated with the center position of the current block. It may be used. a) In one example, the default motion candidate is associated with the center position of the current block. It may be defined as a determined motion candidate (e.g., as disclosed in 2.3.5.1.2). MVdefault0 and / or MVdefault1 used in ATMVP t1). b) In one example, the default motion candidate may be defined as the (0,0) motion vector and a reference picture index equal to 0 for both reference picture lists, if available.
[0280] 10. Note that the default motion information in the ATMVP process (e.g., MVdefault0, MVde fault1 used in ATMVP as disclosed in 2.3.5.1.2) may be derived based on the location of the position used in the sub-block motion information derivation process. In this proposed method, since the default motion information is directly assigned to the sub-block, there is no need to further derive the motion information. a) In one example, instead of using the center position of the current block, the center position of the sub-block (e.g., the central sub-block) in the current block may be utilized. b) Examples of existing and proposed implementations are shown in FIGS. 29A and 29B, respectively.
[0281] 11. ATMVP candidates are always to be made available in the following manner. a) Assuming the center point of the current block is (x0, y0), the corresponding position of (x0, y0) in the collocated picture is M = (x0 + MV’x, y0 + MV’y). Find the block Z containing M. If Z is intra-coded, derive MVdefault0, MVdefault1 by any method proposed in item 6. b) Alternatively, block Z is not arranged to obtain motion information and is proposed in item 8 Some of the methods presented are applied directly to obtain MVdefault0 and MVdefault1 for use. c) Alternatively, the default motion candidates used in the ATMVP process are always available If, based on the current design, it is set to unavailable (e.g., a temporal block is intra-coded), other motion vectors may be used instead of the default motion candidates for use. i. In one example, the solution of International Application PCT / CN20 18 / 124639, incorporated herein by reference, may be applied. d) Alternatively, furthermore, whether the ATMVP candidates are always available depends on other high-level syntax information. i. In one example, the ATMVP candidates may be set to always be available only if the ATMVP enable flag in the slice / tile / picture header or other video units is presumed to be true. ii. In one example, the above method may be applicable only if the ATMVP enable flag in the slice header / picture header or other video units is set to true, and the current picture is not an IRAP picture and the current picture is not inserted into RefPicList0 with a reference index equal to 0
[0282] e) The ATMVP candidates are assigned a fixed index or a fixed group of indexes. If the ATMVP candidates are not always available, the fixed index / group index may be inferred to other types of motion candidates (e.g., affine candidates).
[0282] 12. Whether to include the motion zero affine merge candidate in the sub-block merge candidate list should depend on whether the affine prediction is enabled or not. a) For example, when the affine usage flag is off (when sps_affine_enabl ed_flag is equal to 0), the motion zero affine merge candidate cannot be included in the sub-block merge candidate list. b) Alternatively, instead, add the default motion vector candidates that are non-affine candidates.
[0283] 13. It is also possible to include non-affine padding candidates in the sub-block merge candidate list. a) If the sub-block merge candidate list is not full, motion zero non-affine padding candidates may be added. b) When selecting such padding candidates, it is necessary to set the affine_flag of the current block to 0. c) Alternatively, if the sub-block merge candidate list is not full and the affine usage flag is off, zero motion non-affine padding candidates are included in the sub-block merge candidate list.
[0284] 14. Assume that MV0 and MV1 represent the MVs of reference list 0 and reference list 1 of the block including the corresponding positions (for example, MV0 and MV1 may be MVZ_0 and MVZ_1, or MVZS_0 and MVZS_1 described in section 2.3.5.1.2). MV0' and MV1' represent the MVs in reference list 0 and reference list 1 to be derived for the current block or sub-block. In that case, MV0' and MV1' are It should be derived by collaring. a) When MV0 and the collocated picture are in reference list 1. b) When MV1 and the collocated picture are in reference list 0.
[0285] 15. In the reference picture list X (PicRefListX, for example, X = 0), When the current picture is treated as a reference picture with an index set to M (for example, 0), the ATMVP and / or TMVP permission / forbidden flag may be presumed to be false for a slice / tile or other types of video units. Here, M is the object reference picture index that scales the motion information of the temporal block in the ATMVP / TMVP process for PicRefListX. It may be equal to the object reference picture index that scales the motion information of the temporal block in the ATMVP / TMVP process for PicRefListX. It may be equal to the object reference picture index that scales the motion information of the temporal block in the ATMVP / TMVP process for PicRefListX. a) Alternatively, the above method is applicable only when the current picture is an intra-random access point (IRAP) picture. IRAP) picture. b) In one example, when the current picture in PicRefListX is treated as a reference picture with an index set to M (for example, 0), and / or when the reference picture with an index set to N (for example, 0) in PicRefListY is treated as such, the ATMVP and / or TMVP permission / forbidden flag may be presumed to be false. The variables M and N represent the object reference picture indices used in the TMVP or ATMVP process. such, the ATMVP and / or TMVP permission / forbidden flag may be presumed to be false. The variables M and N represent the object reference picture indices used in the TMVP or ATMVP process. The variables M and N represent the object reference picture indices used in the TMVP or ATMVP process. The variables M and N represent the object reference picture indices used in the TMVP or ATMVP process. c) In the case of the ATMVP process, the verification bitstream is restricted to follow the rule that the collocated picture from which the motion information of the current block is derived is not the current picture. the collocated picture from which the motion information of the current block is derived is not the current picture. the collocated picture from which the motion information of the current block is derived is not the current picture. d) Alternatively, if the above conditions are true, the ATMVP or TMVP process is not called. It is not called.
[0286] 16. If the reference picture with index M (e.g., 0) in the reference picture list X (PicRefListX, e.g., X = 0) for the current block is the current picture, the ATMVP can still be made effective for this block. If the reference picture with index M (e.g., 0) in the reference picture list X (PicRefListX, e.g., X = 0) for the current block is the current picture, the ATMVP can still be made effective for this block. is the current picture, the ATMVP can still be made effective for this block. is still valid. a) In one example, the motion information of all sub - blocks points to the current picture. b) In one example, when obtaining the motion information of sub - blocks from a temporal block, the temporal block is encoded with at least one reference picture that points to the current picture of the temporal block. In one example, when obtaining the motion information of sub - blocks from a temporal block, the temporal block is encoded with at least one reference picture that points to the current picture of the temporal block. is encoded. c) In one example, when obtaining the motion information of sub - blocks from a temporal block, no scaling operation is applied. is not applied.
[0287] 17. Regardless of the use of ATMVP, unify the encoding method of the sub - block merge index. Unify it. a) In one example, for the first L bins, they are context - encoded. For the remaining bins, they are bypass - encoded. In one example, L is set to 1. b) Alternatively, for all bins, they are context - encoded.
[0288] 18. In ATMVP, the MV (MVx, MVy) (e.g., TMV = 0) used to find the corresponding block in different pictures may be right - shifted to integer precision (denoted as MVx’, MVy’) by a rounding method similar to the MV scaling process. In ATMVP, the MV (MVx, MVy) (e.g., TMV = 0) used to find the corresponding block in different pictures may be right - shifted to integer precision (denoted as MVx’, MVy’) by a rounding method similar to the MV scaling process. by a rounding method similar to the MV scaling process to integer precision (denoted as MVx’, MVy’). and shifted right. a) Alternatively, the MV (e.g., TMV = 0) used to find corresponding blocks in different pictures in ATMVP may be right-shifted to integer precision in the same rounding method as the MV averaging process. b) Alternatively, the MV (e.g., TMV = 0) used to find corresponding blocks in different pictures in ATMVP may be right-shifted to integer precision in the same rounding method as the adaptive MV resolution (AMVR) process.
[0289] 19. In ATMVP, the MV (MVx, MVy) (e.g., TMV = 0) used to find corresponding blocks in different pictures may be right-shifted to integer precision ((denoted as (MVx’, MVy’)) by rounding in the direction approaching 0. a) For example, MVx’ = (MVx + ((1 << N) >> 1) - (MVx >= 0? 1 : 0)) >> N; N is an integer representing the resolution of MV, e.g., N = 4. i. For example, MVx’ = (MVx + (MVx >= 0? 7 : 8)) >> 4. b) For example, MVy’ = (MVy + ((1 << N) >> 1) - (MVy >= 0? 1 : 0)) >> N; N is an integer representing the resolution of MV, e.g., N = 4. i. For example, MVy’ = (MVy + (MVy >= 0? 7 : 8)) >> 4.
[0290] 20. In one example, the MV (MVx, , MVy) in bullet18 and bullet19 is used to define the position of the corresponding block by using the center position of the sub-block and the shifted MV, or the upper left position of the current block and the shifted MV, in order to derive the default motion information used in ATMVP. a) In one example, MV (MVx, MVy) is used to define the position of the corresponding block by using, for example, the center position of the sub-block and the shifted MV to derive the motion information of the sub-block in the current block during ATMVP processing. In one example, MV (MVx, MVy) is used to define the position of the corresponding block by using, for example, the center position of the sub-block and the shifted MV to derive the motion information of the sub-block in the current block during ATMVP processing. In one example, MV (MVx, MVy) is used to define the position of the corresponding block by using, for example, the center position of the sub-block and the shifted MV to derive the motion information of the sub-block in the current block during ATMVP processing. In one example, MV (MVx, MVy) is used to define the position of the corresponding block by using, for example, the center position of the sub-block and the shifted MV to derive the motion information of the sub-block in the current block during ATMVP processing.
[0291] 21. The methods proposed in bullet 18, 19, 20 may also be applied to other coding tools that require defining the position of the reference block in a different picture or the current picture with a motion vector. 21. The methods proposed in bullet 18, 19, 20 may also be applied to other coding tools that require defining the position of the reference block in a different picture or the current picture with a motion vector. 21. The methods proposed in bullet 18, 19, 20 may also be applied to other coding tools that require defining the position of the reference block in a different picture or the current picture with a motion vector.
[0292] 22. The MV (MVx, MVy) (e.g., TMV at 0) used to find the corresponding blocks in different pictures in ATMVP may point to or be scaled for the collocated picture. 22. The MV (MVx, MVy) (e.g., TMV at 0) used to find the corresponding blocks in different pictures in ATMVP may point to or be scaled for the collocated picture. 22. The MV (MVx, MVy) (e.g., TMV at 0) used to find the corresponding blocks in different pictures in ATMVP may point to or be scaled for the collocated picture. a) In one example, if the width and / or height of the collocated picture (or the conforming window therein) is different from the width and / or height of the current picture (or the conforming window therein), the MV may be scaled. a) In one example, if the width and / or height of the collocated picture (or the conforming window therein) is different from the width and / or height of the current picture (or the conforming window therein), the MV may be scaled. a) In one example, if the width and / or height of the collocated picture (or the conforming window therein) is different from the width and / or height of the current picture (or the conforming window therein), the MV may be scaled. a) In one example, if the width and / or height of the collocated picture (or the conforming window therein) is different from the width and / or height of the current picture (or the conforming window therein), the MV may be scaled. b) Let the width and height of the collocated picture (of the conforming window) be denoted as W1 and H1 respectively. Let the width and height of the current picture (of the conforming window) be W2 and H2 respectively. And the MV (MVx, MVy) may be scaled as MVx’ = MVx * W1 / W2 and MVy’ = MVy * H1 / H2. b) Let the width and height of the collocated picture (of the conforming window) be denoted as W1 and H1 respectively. Let the width and height of the current picture (of the conforming window) be W2 and H2 respectively. And the MV (MVx, MVy) may be scaled as MVx’ = MVx * W1 / W2 and MVy’ = MVy * H1 / H2. b) Let the width and height of the collocated picture (of the conforming window) be denoted as W1 and H1 respectively. Let the width and height of the current picture (of the conforming window) be W2 and H2 respectively. And the MV (MVx, MVy) may be scaled as MVx’ = MVx * W1 / W2 and MVy’ = MVy * H1 / H2. b) Let the width and height of the collocated picture (of the conforming window) be denoted as W1 and H1 respectively. Let the width and height of the current picture (of the conforming window) be W2 and H2 respectively. And the MV (MVx, MVy) may be scaled as MVx’ = MVx * W1 / W2 and MVy’ = MVy * H1 / H2. b) Let the width and height of the collocated picture (of the conforming window) be denoted as W1 and H1 respectively. Let the width and height of the current picture (of the conforming window) be W2 and H2 respectively. And the MV (MVx, MVy) may be scaled as MVx’ = MVx * W1 / W2 and MVy’ = MVy * H1 / H2.
[0293] 23. The current block used to derive the motion information in ATMVP processing The center point (e.g., the position (x0, y0) in 2.3.5.1.2) may be further corrected by scaling and / or adding an offset. a) In one example, if the width and / or height of the collocated picture (or the conformity window therein) is different from the width and / or height of the current picture (or the conformity window therein), the center point may be further corrected. . b) Let the upper left position of the conformity window in the collocated picture be X1 and Y1. Let the upper left position of the conformity window defined in the current picture be X2 and Y2. Let the width and height of the collocated picture (of the conformity window) be shown as W1 and H1 respectively. Let the width and height of the current picture (of the conformity window) be W2 and H2 respectively. In that case, (x0, y0) may be corrected as x0’ = (x0 - X2) * W1 / W2 + X1, y0’ = (y0 - Y2) * H1 / H2 + Y1. i. Alternatively, x0’ = x0 * W1 / W2, y0’ = y0 * H1 / H2.
[0294] 24. The corresponding position (e.g., the position M in 2.3.5.1.2) used to derive the motion information in the ATMVP process may be further corrected by scaling and / or adding an offset. a) In one example, if the width and / or height of the collocated picture (or the conformity window therein) is different from the width and / or height of the current picture (or the conformity window therein), the corresponding position may be further corrected. b) Let the upper left position of the conformity window in the collocated picture be X1 and Y1. Let the upper left position of the conformity window defined in the current picture be X2 and Y2. Let the width and height of the collocated picture (of the conformity window) be shown as W1 and H1 respectively. Let the width and height of the current picture (of the conformity window) be W2 and H2 respectively. In that case, M(x, y) may be corrected as x’ = (x - X2) * W1 / W2 + X1 and y’ = (y - Y2) * H1 / H2 + Y1. i. Alternatively, x’ = x * W1 / W2, y’ = y * H1 / H2.
[0295] Sub-picture related
[0296] 25. In one example, when the positions (i, j) and (i, j - 1) belong to different sub - pictures, the width of the sub - picture S ending at the (j - 1) - th column may be set equal to j - the left - most column of the sub - picture S. a) Emphasize the following embodiments based on existing implementation examples. NumSubPics = 0 for(i = 0; i < NumSubPicGridRows; i++){ for(j = 0; j < NumSubPicGridCols; j++){ if(i == 0) SubPicTop[subpic_grid_idx[i][j]] = 0 else if(subpic_grid_idx[i][j]!= subpic_ grid_idx[i - 1][j]){ SubPicTop[subpic_grid_idx[i][j]] = i SubPicHeight[subpic_grid_idx[i - 1][j]] = i - SubPicTop[subpic_grid_idx[i - 1][j]] } if (j == 0) SubPicLeft[subpic_grid_idx[i][j]] = 0 ( 7 - 7) else if (subpic_grid_idx[i][j] != subpic_ grid_idx[i][j - 1]) { SubPicLeft[subpic_grid_idx[i][j]] = j SubPicWidth[subpic_grid_idx[i][j - 1]] = j - SubPicLeft[subpic_grid_idx[i][j - 1]] } if (i == NumSubPicGridRows - 1) SubPicHeight[subpic_grid_idx[i][j]] = i - SubPicTop[subpic_grid_idx[i - 1][j]] + 1 if (j == NumSubPicGridRows - 1) SubPicWidth[subpic_grid_idx[i][j]] = j - SubPicLeft[subpic_grid_idx[i][j - 1]] + 1 if (subpic_grid_idx[i][j] > NumSubPics) NumSubPics = subpic_grid_idx[i][j] } }
[0297] 26. In one example, the height of the sub - picture S ending at the (NumSubPicGridRows - 1) - th row may be set equal to (NumSubPicGridRows - 1) - the top row of the sub - picture S + 1 + 1. a) Emphasize the following embodiments based on existing implementation examples. NumSubPics = 0 for (i = 0; i < NumSubPicGridRows; i++) { for (j = 0; j < NumSubPicGridCols; j++) { if (i == 0) SubPicTop[subpic_grid_idx[i][j]] = 0 else if (subpic_grid_idx[i][j]!= subpic_ grid_idx[i - 1][j]) { SubPicTop[subpic_grid_idx[i][j]] = i SubPicHeight[subpic_grid_idx[i - 1][j]] = i - SubPicTop[subpic_grid_idx[i - 1][j]] } if (j == 0) SubPicLeft[subpic_grid_idx[i][j]] = 0 ( 7 - 7) else if (subpic_grid_idx[i][j]!= subpic_ grid_idx[i][j - 1]) { SubPicLeft[subpic_grid_idx[i][j]] = j SubPicWidth[subpic_grid_idx[i][j]] = j - SubPicLeft[subpic_grid_idx[i][j - 1]] } if (i == NumSubPicGridRows - 1) SubPicHeight[subpic_grid_idx[i][j]] = i - SubPicTop[subpic_grid_idx[i][j]] + 1 if (j == NumSubPicGridRows - 1) SubPicWidth[subpic_grid_idx[i][j]] = j - SubPicLeft[subpic_grid_idx[i][j - 1]] + 1 if (subpic_grid_idx[i][j] > NumSubPics) NumSubPics = subpic_grid_idx[i][j]
[0298] 27. In one example, ending with the (NumSubPicGridColumns - 1)th column the width of sub - picture S may be set equal to (NumSubPicGridColumns - 1) - the left - most column of sub - picture S, and then adding 1. a) Embodiments based on existing implementation examples are emphasized below. NumSubPics = 0 for (i = 0; i < NumSubPicGridRows; i++) { for (j = 0; j < NumSubPicGridCols; j++) { if (i == 0) SubPicTop[subpic_grid_idx[i][j]] = 0 else if (subpic_grid_idx[i][j]!= subpic_ grid_idx[i - 1][j]) { SubPicTop[subpic_grid_idx[i][j]] = i SubPicHeight[subpic_grid_idx[i - 1][j]] = i - SubPicTop[subpic_grid_idx[i - 1][j]] } if (j == 0) SubPicLeft[subpic_grid_idx[i][j]] = 0 ( 7 - 7) else if (subpic_grid_idx[i][j]!= subpic_ grid_idx[i][j - 1]) { SubPicLeft[subpic_grid_idx[i][j]] = j SubPicWidth[subpic_grid_idx[i][j]] = j - SubPicLeft[subpic_grid_idx[i][j - 1]] } if (i == NumSubPicGridRows - 1) SubPicHeight[subpic_grid_idx[i][j]] = i - SubPicTop[subpic_grid_idx[i - 1][j]] + 1 if (j == NumSubPicGridColumns - 1) SubPicWidth[subpic_grid_idx[i][j]] = j - SubPicLeft[subpic_grid_idx[i][j]] + 1 if (subpic_grid_idx[i][j] > NumSubPics) NumSubPics = subpic_grid_idx[i][j]
[0299] 28. The sub - picture grid must be an integer multiple of the CTU size. a) Emphasize the embodiments based on the existing implementation examples as follows. subpic_grid_col_width_minus1 plus1 defines the width of each element of the sub - picture identifier grid in units of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) bits. / CtbSizeY)) bits. The variable NumSubPicGridCols is derived as follows. NumSubPicGridCols = (pic_width_max_in_luma _samples + subpic_grid_col_width_minus1 * Ct (bSizeY + CtbSizeY - 1) / (subpic_grid_col_width_minus1 * CtbSizeY + CtbSizeY) (7 - 5) subpic_grid_row_height_m inus1 plus1 specifies the height of each element of the subpicture identifier grid in 4 - sample units. The length of the syntax element is Ceil(Log2(pic_height_max_ in_luma_samples / CtbSizeY)) bits. in_luma_samples / CtbSizeY)) bits. The variable NumSubPicGridRows is derived as follows. NumSubPicGridRows = (pic_height_max_in_lum a_samples + subpic_grid_row_height_minus1 * CtbSizeY + CtbSizeY - 1) / (subpic_grid_row_height_minus1 * CtbSize Y + CtbSizeY) (7 - 6)
[0300] 29. Conformance constraints are added to ensure that subpictures do not overlap with each other and that all subpictures cover the entire picture. a) Embodiment examples based on existing implementations are emphasized below. a) Embodiment examples based on existing implementations are emphasized below. When both of the following conditions are met, subpic_grid_idx[i][j] must be equal to idx. i >= SubPicTop[idx] and i < SubPicTop[idx] + Su bPicHeight[idx]. j >= SubPicLeft[idx] and j < SubPicLeft[idx] + SubPicWidth[idx]. If both of the following conditions are not satisfied, subpic_grid_idx[i][j] shall be different from idx. i >= SubPicTop[idx] and i < SubPicTop[idx] + Su bPicHeight[idx]. j >= SubPicLeft[idx] and j < SubPicLeft[idx] + SubPicWidth[idx].
[0301] RPR related
[0302] 30. The syntax element (such as a flag) indicated as RPR_flag is signaled to indicate whether RPR can be used in the video unit (such as a sequence). RPR_f lag may be signaled in the SPS, VPS, or DPS. a) In one example, when it is signaled that RPR is not used (for example, when the RPR_ flag is 0), all the widths / heights signaled in the PPS shall be the same as the maximum width / maximum height signaled in the SPS. b) In one example, when it is signaled that RPR is not used (for example, when the RPR_ flag is 0), all the widths / heights of the PPS are not signaled and are presumed to be the maximum width / maximum height signaled in the SPS. c) In one example, when it is signaled that RPR is not used (for example, when the RP R_flag is 0), the compliance window information is not used in the decoding process. Otherwise (when it is signaled to use RPR), the compliance window information may be used in the decoding process.
[0303] 31. Used to derive the predicted block of the current block in motion compensation processing The interpolation filter may be selected based on whether the resolution of the reference picture is different from that of the current picture or whether the width and / or height of the reference picture is larger than the resolution of the current picture than that of the current picture and may be selected accordingly a. In one example, if condition A is satisfied and condition A depends on the dimensions of the current picture and / or the reference picture, an interpolation filter with fewer taps may be applied i. In one example, condition A is that the resolution of the reference picture is different from that of the current picture ii. In one example, condition A is that the width and / or height of the reference picture is larger than that of the current picture iii. In one example, condition A is W1 > a * W2 and / or H1 > b * H2 where (W1, H1) represents the width and height of the reference picture, and (W2, H2) represents the width and height of the current picture, and a and b are two factors, for example, a = b = 1 .5 iv. In one example, condition A may depend on whether dual prediction is used 1) Condition A is satisfied only when dual prediction is used for the current block v. In one example, condition A may depend on M and N, where M and N represent the width and height of the current block 1) For example, condition A is satisfied only when M * N <= T, where T is an integer such as 6 2) For example, condition A is satisfied only when M <= T1 or N <= T2, where T1 and T2 are integers, for example, T1 = T2 = 4 3) For example, condition A is satisfied only when M <= T1 and N <= T2, where 44. For example, condition A is satisfied only when M <= T1 and N <= T2, where 4) For example, condition A is satisfied only when M <= T1 and N <= T2, where T1 and T2 are integers, for example, T1 = T2 = 4 3) For example, condition A is satisfied only when M <= T1 and N <= T2, where Here, T1 and T2 are integers. For example, T1 = T2 = 4. 4) For example, condition A is satisfied when M * N = T, or M = T1 or N = T2, where T, T1, and T2 are integers. For example, T = 64, T1 = T2 = 4. 5) In one example, the smaller condition in the above sub - bullet may be replaced by the larger one. vi. In one example, a 1 - tap filter is applied. That is, non - filtered integer pixels are output as interpolation results. vii. In one example, when the resolution of the reference picture is different from that of the current picture, a bilinear filter is applied. viii. In one example, when the resolution of the reference picture is different from that of the current picture, or when the width and / or height of the reference picture is larger than the resolution of the current picture , a 4 - tap filter or a 6 - tap filter is applied. 1) The 6 - tap filter may be used for affine motion compensation. 2) The 4 - tap filter may be used for chroma sample interpolation. b. Whether to apply the method disclosed in bullet31 and / or how to apply it may depend on the color component. i. For example, these methods are applied only to the luminance component. c. Whether to apply the method disclosed in bullet31 and / or how to apply it may depend on the interpolation filtering direction. i. For example, this method is applied only to horizontal filtering. ii. For example, this method is applied only to vertical filtering.
[0304] CIIP related
[0305] 32. The intra prediction signal used in CIIP processing may be performed at the TU level instead of the CU level (e.g., using the reference samples outside the TU instead of the CU). a) In one example, if either the width or height of the CU is larger than the maximum transform block size, the CU may be divided into multiple TUs. For example, the reference samples outside the TU may be used to generate intra / inter prediction for each TU. b) In one example, when the maximum transform size K is smaller than 64 (e.g., K = 32), the intra prediction used in CIIP is performed in a recursive manner as in a normal intracoded block. c) For example, by dividing the KM×KN CIIP coding block where M and N are integers into MN of K×K blocks, intra prediction is performed for each K×K block. Subsequently, the intra prediction of the encoded / decoded K×K block may depend on the reconstructed samples of the encoded / decoded K×K block.
[0306] 5. Additional exemplary embodiments
[0307] 5.1 Embodiment #1: Examples of syntax design in SPS / PPS / slice header / tile group header
[0308] The changes compared with the VTM3.0.1rC1 standard software are emphasized in large bold font as follows.
[0309]
Table 9
[0310] 5.2 Embodiment #2: Example of Syntax Design in SPS / PPS / Slice Header / Tile Group Header Example of Syntax Design
[0311] 7.3.2.1 Sequence Parameter Set RBSP Syntax
[0312] [Table 10]
[0313] When sps_sbtmvp_enabled_flag is equal to 1, a sub-block based temporal motion vector predictor may be used, and it is specified that pictures including all slices where slice_type in CVS is not equal to I can be decoded. sps_sbtmvp_enabled_flag equal to 0 specifies that the sub-block based temporal motion vector predictor is not used in CVS. If sps_sbtmvp_enabled_flag does not exist, it is assumed to be equal to 0. Based on the sub-block, it is possible to use a temporal motion vector predictor, and it is specified that pictures including all slices where slice_type in CVS is not equal to I can be decoded. Equal to, all slices where slice_type in CVS is not equal to I, it is specified that pictures can be decoded. Equal to 0 sps_sbtmvp_enabled_flag equal to 0 specifies that the sub-block based temporal motion Vector predictor is not used in CVS. sps_sbtmvp_en abled_flag does not exist, it is assumed to be equal to 0.
[0314] five_minus_max_num_subblock_merge_cand Specifies the maximum number of merge motion vector prediction (MVP) candidates based on sub-blocks supported in slices subtracted from 5. If five_minus_max_num_su bblock_merge_cand does not exist, it is assumed to be equal to 5 - sps_sbtmvp_e nabled_flag. The maximum number of merge MVP candidates based on sub-blocks, MaxNumSubblockMergeCand, is derived as follows. The maximum number of merge MVP candidates based on sub-blocks, MaxNumSubblockMergeCand, is derived as follows.
[0315] MaxNumSubblockMergeCand = 5 - five_minus_ma x_num_subblock_merge_cand (7-45)
[0316] The value of MaxNumSubblockMergeCand is within the range of 0 to 5.
[0317] 8.3.4.2 Derivation Process of Motion Vectors and Reference Indexes in Subblock Merge Mode
[0318] The input to this process is as follows.
[0319] ..[No change in the current VVC specification draft]
[0320] The output of this process is as follows.
[0321] ...[No change in the current VVC specification draft]
[0322] The variables numSbX, numSbY, and the subblock merge candidate list subblo ckMergeCandList are derived by the following sequential steps.
[0323] When sps_sbtmvp_enabled_flag is equal to 1 and (the current picture is IR AP and the index 0 of reference picture list 0 is not the current picture) is false the following applies.
[0324] To merge candidates from neighboring coding units defined in Clause 8.3.2.3 the derivation process is called with the position (xCb, yCb) of the luma coding block, the luma coding block width cbWidth, the luma coding block height cbHeight, and the luma coding block width as input, with X being 0 or 1, and the output is the availability flag availableFlag A0, availableFlagA1, availableFlagB0, avail ableFlagB1 and availableFlagB2, reference index ref IdxLXA0, refIdxLXA1, refIdxLXB0, refIdxLXB1 and refIdxLXB2, and prediction list usage flag predFlagLXA0, predFlagLXA1, predFlagLXB0, predFlagLXB1 and predFlagLXB2, and motion vectors mvLXA0, mvLXA1, mvL XB0, mvLXB1 and mvLXB2.
[0325] Derivation process of sub-block-based temporal merge candidates defined in item 8.3.4.3 is called with xSbIdx = 0..numSbX - 1, ySbIdx = 0..numSbY - 1 and X being 0 or 1, luminance position (xCb, yCb), luminance coded block width cbW idth, luminance coded block height cbHeight, availability flags, availabl eFlagA0, availableFlagA1, availableFlagB0, availableFlagB1, reference indexes refIdxLXA0, refId xLXA1, refIdxLXB0, efIdxLXB1, prediction list usage flags pre dFlagLXA0, predFlagLXA1, predFlagLXB0, pred FlagLXB1, motion vectors mvLXA0, mvLXA1, mvLXB0, mvLX B1 as inputs, and the output is the availability flag availableFlagSbC ol, luminance coded sub-blocks in the horizontal direction numSbX and vertical direction numSbY The number of CUs, reference index refIdxLXSbCol, and luminance motion vector mvLXSb Col[xSbIdx][ySbIdx] and prediction list usage flag predFlag It is LXSbCol[xSbIdx][ySbIdx].
[0326] When sps_affine_enabled_flag is equal to 1, the sample location (xNbA0, yNbA0), (xNbA1, yNbA1), (xNbA2, yNbA2 ), (xNbB0, yNbB0), (xNbB1, yNbB1), (xNbB2, yNb B2), (xNbB3, yNbB3), and variables numSbX and numSbY are derived as follows as follows.
[0327] [There is no change to the current VVC specification draft]
[0328] 5.3 Example of rounding of MV in Embodiment #3 The syntax change is based on the existing implementation form.
[0329] 8.5.5.3 Derivation process of temporal merge candidates based on sub-blocks … - The position (xColSb, y ColSb) of the collocated sub-block within ColPic is derived as follows.
[0330]
Chemical formula
[0331] 8.5.5.4 Derivation process of motion data based on temporal merge base for sub-blocks … The position (xColCb, yColCb ) of the collocated block within ColPic is derived as follows.
[0332] [Chemistry]
[0333] 5.3 Embodiment #3: Example of Rounding of MV The syntax change is based on the existing implementation form.
[0334] 8.5.5.3 Derivation Process of Temporal Merge Candidates Based on Sub-Blocks …
[0335] [Chemistry]
[0336] 8.5.5.4 Derivation Process of Motion Data Based on Temporal Merge Base for Sub-Blocks …
[0337] [Chemistry]
[0338] 5.4 Embodiment #4: Second Example of Rounding of MV
[0339] 8.5.5.3 Derivation Process of Temporal Merge Candidates Based on Sub-Blocks
[0340] The input to this process is as follows. - The upper left sample of the current luminance coding block for the upper left luminance sample of the current picture, the luminance position (xCb, yCb) of the sample, - The variable cbWidth that defines the width of the current coding block in the luminance sample, - The variable cbHeight that defines the height of the current coding block in the luminance sample . - The availability flag availableFlagA1 of the neighboring coding unit, - The reference index refIdxLXA1 of the neighboring encoding unit, where X is 0 or 1, - The prediction list usage flag predFlagLXA1 of the neighboring encoding unit, where X is 0 or 1, - The motion vector at 1 / 16 fractional sample accuracy mvLXA1 of the neighboring encoding unit where X is 0 or 1.
[0341] The output of this process is as follows. - The availability flag availableFlagSbCol, - The number of luminance encoding sub-blocks in the horizontal direction numSbX and the vertical direction numSbY of, - The reference indices refIdxL0SbCol and refIdxL1SbCol, - The luminance motion vectors at 1 / 16 fractional sample accuracy mvL0SbCol[xSbIdx][ySbIdx] and mvL1SbCol[xSbIdx][ySbIdx], where xSbIdx = 0..numSbX-1, ySbIdx = 0..numSbY-1, - The prediction list usage flags predFlagL0SbCol[xSbIdx][ySbI dx] and predFlagL1SbCol[xSbIdx][ySbIdx], where xSbIdx = 0..numSbX-1, ySbIdx = 0..numSbY-1.
[0342] The availability flag availableFlagSbColl is derived as follows. - If one or more of the following conditions are true, 0 is set for availableFlagSbCol is set. - slice_temporal_mvp_enabled_flag is equal to 0. - sps_sbtmvp_enabled_flag is equal to 0. - The cbWidth is less than 8. - The cbHeight is less than 8. - Otherwise, the following ordered steps apply.
[0343] 1. The position of the top - left sample of the luminance coding tree block containing the current coding block ( xCtb, yCtb) and the position of the center sample at the bottom - right of the current luminance coding block (x Ctr, yCtr) are derived as follows. xCtb=(xCb >> CtuLog2Size)< <ctulog2size ( 8-542) yctb="(yCb">(CtuLog2Size) << CtuLog2Size ( 8 - 543) xCtr = xCb+(cbWidth / 2) (8 - 544) yCtr = yCb+(cbHeight / 2) (8 - 545)
[0344] 2. The luminance positions (xColCtrCb, yColCtrCb) are set equal to the top - left luminance sample of the collocated picture specified by ColPic, at the position given by (xCtr, yCtr) within ColPic, in the same - position luminance - coded block containing that position, at its
[0345] top - left sample. 3. The derivation process of the time - merge - based motion data based on the sub - blocks specified in Clause 8.5.5.4 is called with X as 0 and 1, (xCtb, yCtb), the location (xCol CtrCb, yColCtrCb), the availability flag Availability flag A1, the prediction - list utilization flag predFlagLXA1, and the reference index refIdxLXA1, and the motion vector mvLXA1 as inputs, and the output is the motion vector ctrMVLX
[0346] - When both ctrPredFlagL0 and ctrPredFlagL1 are equal to 0 availableFlagSbCol is set equal to 0. - Otherwise, availableFlagSbCol is set equal to 1.
[0347] When availableFlagSbCol is equal to 1, the following applies. - Variables numSbX, numSbY, sbWidth, sbHeight, refId xLXSbCol is derived as follows. numSbX = cbWidth >> 3 (8 - 546) numSbY = cbHeight >> 3 (8 - 547) sbWidth = cbWidth / numSbX (8 - 548) sbHeight = cbHeight / numSbY (8 - 549) refIdxLXSbCol = 0 (8 - 550) - xSbIdx = 0..numSbX - 1 and ySbIdx = 0...numSbY In the case of -1, the motion vector mvLXSbCol[xSbIdx][ySbIdx] and the prediction list usage flag predFlags LXSbCol[xSbIdx] are as follows derived. - The luminance position (xSb, ySb) that defines the top - left sample of the current coded sub - block with respect to the top - left luminance sample of the current picture is derived as follows. The luminance position (xSb, ySb) that defines the top - left sample of the current coded sub - block with respect to the top - left luminance sample of the current picture is derived as follows. xSb = xCb + xSbIdx * sbWidth + sbWidth / 2 (8 - 55 1) ySb = yCb + ySbIdx * sbHeight + sbHeight / 2 (8 - 552) - The position (xColSb, y ColSb) of the collocated sub - block within ColPic is derived as follows.
[0348]
Chemical
[0349] - The variable currCb defines the luminance coding block containing the current coded sub-block within the current picture. - The variable colCb defines the luminance coding block containing the modified position given by ((xColSb >> 3) << 3, (yColSb >> 3) << 3) within ColPic. - The luminance position (xColCb, yColCb) is set equal to the top-left luminance sample of the same-position luminance coding block specified by colCb, relative to the top-left luminance sample of the picture located at the location specified by ColPic. - The same-position motion vector derivation process defined in clause 8.5.2.12 is called with inputs equal to cur rCb, colCb, (xColCb, yColCb), 0 and refI set equal to 1, and sbFlag set equal to 1. The output is assigned to the motion vectors of sub-block mvL0SbCol[xSbIdx][ySbIdx] and available FlagL0SbCol. - The same-position motion vector derivation process defined in clause 8.5.2.12 is called with inputs equal to cur rCb, colCb, (xColCb, yColCb), 0 and refI set equal to 1, and sbFlag set equal to 1. The output is assigned to the motion vectors of sub-block MVL1SbCol[xSbIdx][ySbIdx] and availableF lagL1SbCol. - If both availableFlagL0SbCol and availableFlagL 1SbCol are equal to 0, and X is 0 and 1, the following applies. mvLXSbCol[xSbIdx][ySbIdx] = ctrMvLX(8 - 5 56) predFlagLXSbCol[xSbIdx][ySbIdx] = ctrPre dFlagLX(8 - 557)
[0350] 8.5.5.4 Derivation Process of Temporal Merge - Based Motion Data Based on Sub - Blocks
[0351] The inputs to this process are as follows. - The position (xCtb, yCtb) of the top - left sample of the luminance coding tree block containing the current coding block, - The position (xColCtrCb, yColCtrCb) of the top - left sample of the luminance coding block located at the same place and containing the bottom - right center sample. - The availability flag availableFlagA1 of the neighboring coding unit, - The reference index refIdxLXA1 of the neighboring coding unit, - The prediction list utilization flag predFlagLXA1 of the neighboring coding unit, - The motion vector at the 1 / 16 - fraction sample accuracy mvLXA1 of the neighboring coding unit.
[0352] The outputs of this process are as follows. - Motion vectors ctrMvL0 and ctrMvL1, - Prediction list utilization flags ctrPredFlagL0, ctrPredFlagL1, - Temporal motion vector tempMv.
[0353] The variable tempMv is set as follows. tempMv[0] = 0(8 - 558) tempMv[1] = 0(8 - 559)
[0354] The variable currPic defines the current picture.
[0355] If availableFlagA1 is equal to TRUE, the following applies. - If all of the following conditions are true, tempMv is set equal to mvL0A1. . - predFlagL0A1 is equal to 1, - DiffPicOrderCnt(ColPic,RefPicList[0] refIdxL0A1]) is equal to 0, - Otherwise, if all of the following conditions are true, tempMv is set equal to mvL1A1 . - The slice type is the same as B, - predFlagL1A1 is equal to 1, - DiffPicOrderCnt(ColPic,RefPicList[1] refIdxL1A1]) is equal to 0.
[0356]
Chemical formula
[0357] The array colPredMode is set equal to the prediction mode array CuPredMode[0] of the collocated picture specified by ColPic.
[0358] The motion vectors ctrMvL0, ctrMvL1, and the prediction list usage flags ctrPredFlagL0, ctrPredFlagL1 are derived as follows. . - If colPredMode[xColCb][yColCb] is equal to MODE_INTER , the following applies. - The variable currCb is within the current picture and contains (xCtrCb,yCtrCb) Define a luminance quantization block. - The variable colCb defines a luminance quantization block that includes the modified position given by ((xColCb>>3)<< 3,(yColCb>>3)<<3) within ColPic. Define. - The luminance position (xColCb,yColCb) is set equal to the top-left luminance sample of the picture located at the location specified by ColPic for the same-position luminance quantization block specified by colCb. The top-left sample of the same-position luminance quantization block specified by colCb. - The derivation process of the same-position motion vector specified in clause 8.5.2.12 is called by setting cur rCb, colCb, (xColCb,yColCb), 0 equal to refI dxL0, and sbFlag set equal to 1 as inputs, and assigning the outputs to ctrMvL0 and ctrPredFlagL0. - The derivation process of the same-position motion vector specified in clause 8.5.2.12 is called by setting cur rCb, colCb, (xColCb,yColCb), 0 equal to refI dxL1, and sbFlag set equal to 1 as inputs, and assigning the outputs to ctrMvL1 and ctrPredFlagL1. - Otherwise, the following applies. ctrPredFlagL0 = 0 (8-563) ctrPredFlagL1 = 0 (8-564)
[0359] 5.5 Embodiment #5: Third Example of Rounding of MV
[0360] 8.5.5.3 Derivation Process of Temporal Merge Candidates Based on Sub-Blocks
[0361] The inputs to this process are as follows. - The luminance position (xCb, yCb) of the top-left sample of the current luminance coding block with respect to the top-left luminance sample of the current picture, - The variable cbWidth that defines the width of the current coding block in the luminance sample, - The variable cbHeight that defines the height of the current coding block in the luminance sample . - The availability flag availableFlagA1 of the neighboring coding unit, - The reference index refIdxLXA1 of the neighboring coding unit, where X is 0 or 1, - The prediction list usage flag predFlagLXA1 of the neighboring coding unit, where X is 0 or 1, - The motion vector at the 1 / 16 fractional sample accuracy mvLXA1 of the neighboring coding unit, where X is 0 or 1.
[0362] The output of this process is as follows. - The availability flag availableFlagSbCol, - The number of luminance coding sub-blocks in the horizontal direction numSbX and the vertical direction numSbY, - The reference indices refIdxL0SbCol and refIdxL1SbCol, - The luminance motion vectors at the 1 / 16 fractional sample accuracy mvL0SbCol[xSbIdx][ySbIdx] and mvL1SbCol[xSbIdx][ySbIdx], where xSbIdx = 0..numSbX-1, ySbIdx = 0..numSbY-1, - The prediction list usage flags predFlagL0SbCol[xSbIdx][ySbIdx] and predFlagL1SbCol[xSbIdx][ySbIdx], where xSbIdx = 0..numSbX-1, ySbIdx = 0..numSbY-1.
[0363] The availability flag availableFlagSbColl is derived as follows. - If one or more of the following conditions are true, 0 is set to availableFlagSbCol. - slice_temporal_mvp_enabled_flag is equal to 0. - sps_sbtmvp_enabled_flag is equal to 0. - cbWidth is less than 8. - cbHeight is less than 8. - Otherwise, the following ordered steps apply.
[0364] 5. The position of the top-left sample of the luma coding tree block containing the current coding block ( (xCtb, yCtb)) and the position of the center sample at the bottom-right of the current luma coding block (x Ctr, yCtr) are derived as follows. xCtb = (xCb >> CtuLog2Size) < <ctulog2size ( 8-542) yctb="(yCb">>CtuLog2Size) << CtuLog2Size ( 8 - 543) xCtr = xCb+(cbWidth / 2) (8 - 544) yCtr = yCb+(cbHeight / 2) (8 - 545)
[0365] 6. The luminance position (xColCtrCb, yColCtrCb) is determined by ColPic for the top - left luminance sample of the collocated picture specified by ColPic, and is set equal to the top - left sample of the same - position luminance - coded block containing the position given by (xCtr, yCtr) within ColPic.
[0366] 7. The derivation process of the temporal merge - based motion data based on the sub - blocks specified in 8.5.5.4 is called with X being 0 and 1, (xCtb, yCtb), location (xCol CtrCb, yColCtrCb), availability flag Availability flag Availabilit y flag A1, prediction - list usage flag predFlagLXA1, and reference index refIdxLXA1, and motion vector mvLXA1 as inputs, and the output is, with X being 0 and 1, motion vector ctrMVLX, and the prediction - list usage flag ctrPredFlagLX of the collocated block, and temporal motion vector te mpMv.
[0367] 8. The variable availableFlagSbCol is derived as follows. - If both ctrPredFlagL0 and ctrPredFlagL1 are equal to 0, availableFlagSbCol is set equal to 0. - Otherwise, availableFlagSbCol is set equal to 1. .
[0368] When availableFlagSbCol is equal to 1, the following applies. - Variables numSbX, numSbY, sbWidth, sbHeight, refId xLXSbCol is derived as follows. numSbX = cbWidth >> 3 (8 - 546) numSbY = cbHeight >> 3 (8 - 547) sbWidth = cbWidth / numSbX (8 - 548) sbHeight = cbHeight / numSbY (8 - 549) refIdxLXSbCol = 0 (8 - 550) - xSbIdx = 0..numSbX - 1 and ySbIdx = 0...numSbY In the case of -1, the motion vector mvLXSbCol[xSbIdx][ySbIdx] and the prediction list usage flag predFlags LXSbCol[xSbIdx] are as follows derived. - The luminance position (xSb, ySb) that defines the top - left sample of the current coded sub - block with respect to the top - left luminance sample of the current picture is derived as follows. xSb = xCb + xSbIdx * sbWidth + sbWidth / 2 (8 - 55 1) ySb = yCb + ySbIdx * sbHeight + sbHeight / 2 (8 - 552) - The position (xColSb, y ColSb) of the collocated sub - block inside ColPic is derived as follows.
[0369]
Chemical
[0370] - The variable currCb defines the luminance coding block containing the current coded sub-block within the current picture. - The variable colCb defines the luminance coding block containing the modified position given by ((xColSb>>3)<< 3,(yColSb>>3)<<3) within ColPic. - The luminance position (xColCb, yColCb) is set equal to the top-left luminance sample of the same-position luminance coding block specified by colCb with respect to the top-left luminance sample of the picture located at the location specified by ColPic. - The derivation process of the same-position motion vector defined in clause 8.5.2.12 is called with cur rCb, colCb, (xColCb, yColCb), 0 set equal to refI dxL0 and sbFlag set equal to 1 as inputs, and the output is assigned to the motion vectors of sub-block mvL0SbCol[xSbIdx][ySbIdx] and available FlagL0SbCol. - The derivation process of the same-position motion vector defined in clause 8.5.2.12 is called with cur rCb, colCb, (xColCb, yColCb), 0 set equal to refI dxL1 and sbFlag set equal to 1 as inputs, and the output is assigned to the motion vectors of sub-block MVL1SbCol[xSbIdx][ySbIdx] and availableF lagL1SbCol. - If both availableFlagL0SbCol and availableFlagL 1SbCol are equal to 0 and X is 0 and 1, the following applies. mvLXSbCol[xSbIdx][ySbIdx]=ctrMvLX(8 - 556) predFlagLXSbCol[xSbIdx][ySbIdx]=ctrPre dFlagLX(8 - 557)
[0371] 8.5.5.4 Derivation Process of Temporal Merge - Based Motion Data Based on Sub - Blocks
[0372] The inputs to this process are as follows. - The position (xCtb, yCtb) of the top - left sample of the luminance coding tree block containing the current coding block, - The position (xColCtrCb, yColCtrCb) of the top - left sample of the luminance coding block located at the same place and containing the bottom - right center sample. - The availability flag availableFlagA1 of the neighboring coding unit, - The reference index refIdxLXA1 of the neighboring coding unit, - The prediction list usage flag predFlagLXA1 of the neighboring coding unit, - The motion vector at the 1 / 16 - fraction sample accuracy mvLXA1 of the neighboring coding unit .
[0373] The outputs of this process are as follows. - The motion vectors ctrMvL0 and ctrMvL1, - The prediction list usage flags ctrPredFlagL0, ctrPredFlagL1, - The temporal motion vector tempMv.
[0374] The variable tempMv is set as follows. tempMv[0]=0(8 - 558) tempMv[1]=0(8 - 559)
[0375] The variable currPic defines the current picture.
[0376] If availableFlagA1 is equal to TRUE, the following applies. - If all of the following conditions are true, tempMv is set equal to mvL0A1 . - predFlagL0A1 is equal to 1, - DiffPicOrderCnt(ColPic,RefPicList[0] refIdxL0A1]) is equal to 0, - Otherwise, if all of the following conditions are true, tempMv is set equal to mvL1A1 :. - The slice type is the same as B, - predFlagL1A1 is equal to 1, - DiffPicOrderCnt(ColPic,RefPicList[1] refIdxL1A1]) is equal to 0.
[0377]
Chemical formula
[0378] The array colPredMode is set equal to the prediction mode array CuPredMode[0] of the collocated picture specified by ColPic.
[0379] The motion vectors ctrMvL0, ctrMvL1, and the prediction list usage flags ctrP redFlagL0, ctrPredFlagL1 are derived as follows. - If colPredMode[xColCb][yColCb] is equal to MODE_INTER , the following applies. - The variable currCb is within the current picture and includes (xCtrCb,yCtrCb) Define a luminance quantization block. - The variable colCb defines a luminance quantization block including the modified position given by ((xColCb>>3)<< 3,(yColCb>>3)<<3) inside ColPic. Define. - The luminance position (xColCb,yColCb) is set equal to the top-left luminance sample of the picture located at the location specified by ColPic, at the same position as the top-left sample of the luminance quantization block specified by colCb. For the top-left luminance sample of the picture located at the location specified by ColPic, it is set equal to the top-left sample of the luminance quantization block at the same position specified by colCb. For the top-left luminance sample of the picture located at the location specified by ColPic, it is set equal to the top-left sample of the luminance quantization block at the same position specified by colCb. - The same-position motion vector derivation process defined in item 8.5.2.12 is called by assigning the input cur rCb, colCb, (xColCb,yColCb), 0 equal to refI dxL0, and sbFlag equal to 1, and assigning the output to ctrMvL0 and ctrPredFlagL0. - The same-position motion vector derivation process defined in item 8.5.2.12 is called by assigning the input cur rCb, colCb, (xColCb,yColCb), 0 equal to refI dxL1, and sbFlag equal to 1, and assigning the output to ctrMvL1 and ctrPredFlagL1. - Otherwise, the following applies. ctrPredFlagL0 = 0 (8-563) ctrPredFlagL1 = 0 (8-564)
[0380] 8.5.6.3 Fractional sample interpolation processing
[0381] 8.5.6.3.1 General
[0382] The input to this process is as follows. - The luminance position (xSb, ySb) that defines the top-left sample of the current coded sub-block with respect to the top-left luminance sample of the current picture , - A variable sbWidth that defines the width of the current coded sub-block, - A variable sbHeight that defines the height of the current coded sub-block, - A motion vector offset mvOffset, - A refined motion vector refMvLX, - A selected reference picture sample array refPicLX, - A 1 / 2 sample interpolation filter index hpelIfIdx, - A bidirectional optical flow flag bdofFlag, - A variable cIdx that defines the color component index of the current block.
[0383] The output of this process is as follows. - A (sbWidth + brdExtSize) × (sbHeight + brdExtSize) array predSamplesLX of predicted sample values.
[0384] The prediction block boundary expansion size brdExtSize is derived as follows. brdExtSize = (bdofFlag || (inter_affine_fl ag[xSb][ySb] && sps_affine_prof_enabled_ flag))? 2 : 0 (8 - 752)
[0385] The variable fRefWidth is set equal to PicOutput WidthL of the reference picture in luminance samples.
[0386] The variable fRefHeight is set equal to PicOutpu tHeightL of the reference picture in luminance samples.
[0387] The motion vector mvLX is set equal to (refMvLX - mvOffset). . - If cIdx is equal to 0, the following applies. - The scaling factor and its fixed - point representation are defined as follows. hori_scale_fp = ((fRefWidth << 14)+(PicOut putWidthL >> 1)) / PicOutputWidthL (8 - 753) vert_scale_fp = ((fRefHeight << 14)+(PicOu tputHeightL >> 1)) / PicOutputHeightL (8 - 754 ) - Let (xIntL,yIntL) be the luminance position given in full - sample units, and (xFracL,yFracL) be the offset given in 1 / 16 - sample units. These variables are used only in this section to define the fractional - sample positions within the reference sample array refPicLX. - Set the top - left coordinate of the bounding block (xSbInt L ,ySbIn t L ) for reference sample padding equal to (xSb+(mvLX[0] >> 4),ySb+(mvLX[1] >> 4)). - For each luminance sample position (x L = 0..sbWidth - 1+brdExtSize,y L = 0..sbHeight - L 1+brdExtSize) in the predicted luminance sample array predSamples X, the corresponding predicted luminance sample value predSampl esLX[x L [y L is derived as follows. - (refxSb L , refySb L ), and (refx L , refy L ), are set to the luminance positions pointed to by the motion vectors (refMvLX[0], refMvLX / 16 sample units. The variables refxSb [1]). The variables refxSb L , refx L , refySb L , r efy L are derived as follows. refxSb L = ((xSb << 4) + refMvLX[0]) * hori_sc ale_fp (8 - 755) refx L = ((Sign(refxSb) * ((Abs(refxSb) + 12 8) >> 8) + x L * ((hori_scale_fp + 8) >> 4)) + 32) >> 6 (8 - 756) refySb L = ((ySb << 4) + refMvLX[1]) * vert_sc ale_fp (8 - 757) refyL = ((Sign(refySb) * ((Abs(refySb) + 12 8) >> 8) + yL * ((vert_scale_fp + 8) >> 4)) + 32) >> 6 (8 - 758 ) - The variables xInt L , yInt L , xFrac L , and yFrac L are derived as follows. are derived as follows. xInt L = refx L >> 4 (8 - 759) yInt L = refy L >> 4 (8 - 760) xFrac L =refx L &15 (8-761) yFrac L =refy L &15 (8-762)
[0388]
Chem.
[0389] - If bdofFlag is equal to TRUE (sps_affine_prof_en abled_flag is equal to TRUE, and inter_affine_flag[xS b][ySb] is equal to TRUE), if one or more of the following conditions are true, the predicted luminance sample value predSamplesLX[x L [y L is derived by calling the luminance integer sample extraction process using (xInt +(xFrac L >>3)-1), yInt L +(yFrac L +(yFrac L >>3)-1) and refPicLX as inputs as specified in 8.5.6.3.3. 1. x is equal to 0. L 2. x is equal to sbWidth + 1. L 3. y is equal to 0. L 4. y is equal to sbHeight + 1. L 5.
[0390]
Chem.
[0391] - Otherwise (if cIdx is not equal to 0), the following applies. - Let (xIntC, yIntC) be the chroma position given in full sample units , and (xFracC, yFracC) be the offset given in 1 / 32 sample units . These variables are used only in this section to define the position of the general fractional samples in the reference sample array refPicLX - The top-left coordinate of the reference sample padding bounding block (xSbIntC, ySbIntC) is set equal to ((xSb / SubWidthC)+(mvLX[0]>>5), (ySb / SubHeightC)+(mvLX[1]>>5)) - For each chroma sample position (xC = 0..sbWidth-1, yC = 0..sbHeight-1) in the predicted chroma sample array predSamplesLX , the corresponding predicted chroma sample value predSamplesLX[xC][yC] is derived as follows - Let (refxSb C , refySb C ) and (refx C , refy C ) be the chroma positions pointed to by the motion vector (mvLX[0], mvLX[1]) given in 1 / 32 sample units. The variables refxSb , refySb C , refx C , refy C are C derived as follows refxSb C =((xSb / SubWidthC<<5)+mvLX[0])* hori_scale_fp (8-763) refx C =((Sign(refxSb C )*((Abs(refxSb C )+ 256)>>9) +xC * ((hori_scale_fp + 8) >> 4)) + 16) >> 5 (8 -764) refySb C = ((ySb / SubHeightC << 5) + mvLX[1]) * vert_scale_fp (8 - 765) refy C = ((Sign(refySb C ) * ((Abs(refySb C ) + 256) >> 9) + yC * ((vert_scale_fp + 8) >> 4)) + 16) >> 5 (8 -766) - Variable xInt C 、yInt C 、xFrac C 、yFrac C are derived as follows. Derivation. xInt C = refx C >> 5 (8 - 767) yInt C = refy C >> 5 (8 - 768) xFrac C = refy C & 31 (8 - 769) yFrac C = refy C & 31 (8 - 770) - The predicted sample value predSamplesLX[xC][yC] is (xIntC, yIntC),(xFracC,yFracC),(xSbIntC,ySbIntC) , sbWidth, sbHeight, and refPicLX as inputs, 8.5. It is derived by calling the process specified in 6.3.4.
[0392] 8.5.6.3.2 Luminance Sample Interpolation Filtering Process
[0393] [Chemical formula]
[0394] The output of this process is the predicted luminance sample value predSampleLX L .
[0395] The variables shift1, shift2, and shift3 are derived as follows. - Set the variable shift1 equal to Min(4, BitDepth Y - 8), set the variable s hift2 equal to 6, and set the variable shift3 equal to Max(2, 14 - BitDept h Y ). - Set the variable picW equal to pic_width_in_luma_samples and set the variable picH equal to pic_height_in_luma_samples .
[0396] [Chemical formula]
[0397] For i = 0..7, the luminance i position at the full - sample unit (xInt i , yInt ) is derived as follows. - When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies. xInt i = Clip3(SubPicLeftBoundaryPos, SubP icRightBoundaryPos, xInt L + i - 3) (8 - 771) yInt i = Clip3(SubPicTopBoundaryPos, SubPi cBotBoundaryPos, yInt L +i - 3)(8 - 772) - Otherwise (subpic_treated_as_pic_flag[Sub PicIdx] equals 0), the following applies. xInt i = Clip3(0, picW - 1, sps_ref_wraparoun d_enabled_flag? ClipH((sps_ref_wraparound_offset_minu s1 + 1)*MinCbSizeY, picW, xInt L +i - 3): (8 - 773 ) xInt L +i - 3) yInt i = Clip3(0, picH - 1, yInt L +i - 3)(8 - 774 )
[0398] When i = 0..7, the luminance position in the full - sample unit is further modified as follows. xInt i = Clip3(xSbInt L - 3, xSbIntL + sbWidth + 4, xInt i ) (8 - 775) yInt i = Clip3(ySbInt L - 3, ySbInt L + sbHeight + 4, yInt i ) (8 - 776)
[0399] The predicted luminance sample value predSampleLX L is derived as follows. - xFrac L and yFrac L both equal 0, predSampleL X L The value of is derived as follows. predSampleLX L =refPicLX L [xInt3][yInt3]< <shift3 (8-777) - Otherwise, if xFrac L is not equal to 0 and yFrac L is equal to 0, then predSampleLX L The value of is derived as follows. predSampleLX L =(Σ 7 i=0 f L [xFrac L [i]*refP icLX L [xInt i [yInt3])>>shift1 (8-778) - Otherwise, if xFrac L is equal to 0 and yFrac L is not equal to 0, then p redSampleLX L The value of is derived as follows. predSampleLX L =(Σ 7 i=0 f L [yFrac L [i]*refP icLX L [xInt3][yInt i )>>shift1 (8-779) - Otherwise, if xFrac L is not equal to 0 and yFrac L is not equal to 0 , then the value of predSampleLX L is derived as follows. - The sample array temp[n] for n = 0..7 is derived as follows. temp[n]=(Σ 7 i=0 f L [xFrac L [i]*refPicLX L xInt i [yInt n )>>shift1 (8 - 780) - The predicted luminance sample value predSampleLX L is derived as follows. predSampleLX L =(Σ 7 i=0 f L [yFrac L [i]*temp [i])>>shift2 (8 - 781)
[0400]
Table 11
[0401]
Table 12
[0402] Figure 30 is a flowchart of video processing method 3000. Method 3000 includes, at step 3010, determining whether to add to the merge candidate list based on sub - blocks the maximum number of candidates (ML) and / or the sub - block - based temporal motion vector prediction (SbTMVP) candidates in the merge candidate list based on sub - blocks, whether the temporal temporal motion vector prediction (TMVP) is made effective for use during the conversion, or whether to use the current picture reference (CPR) coding mode for the conversion, and including determining whether to add to the merge candidate list based on sub - blocks. Method 3000 includes, at step 3020, performing the conversion based on the determination
[0403] Method 3000 includes, at step 3020, performing the conversion based on the determination includes performing
[0404] Figure 31 is a flowchart of a video processing method 3100. The method 3100 includes, at step 3110, determining a majority (ML) of candidates in a sub-block-based merge candidate list based on whether time motion vector prediction (TMVP), sub-block-based time motion vector prediction (SbTMVP) tool, and affine coding mode are enabled for conversion between the current block of the video and the bitstream representation of the video.
[0405] The method 3100 includes, at step 3120, performing the conversion based on the determination.
[0406] Figure 32 is a flowchart of a video processing method 3200. The method 3200 includes, at step 3210, determining that the conversion of the sub-block-based motion vector prediction (SbTMVP) mode is disabled because the time motion vector prediction (TMVP) mode at the first video segment level is disabled for conversion between the current block of the first video segment of the video and the bitstream representation of the video.
[0407] The method 3200 includes, at step 3220, performing the conversion based on the determination, and the bitstream representation conforms to a format that defines whether the display of the SbTMVP mode is included and / or the position of the display of the SbTMVP mode relative to the display of the TMVP mode in the merge candidate list.
[0408] Figure 33 is a flowchart of a video processing method 3300. The method 3300 includes, at step In 3310, between the current block of the video encoded using the sub-block based temporal motion vector prediction (SbTMVP) tool or the temporal motion vector prediction (TMVP) tool and the bitstream representation of the video including performing a conversion, the coordinates of the current block or the corresponding position of the current block or the sub-blocks of the current block are selectively masked using a mask based on the compression of the motion vectors associated with the SbTMVP tool or the TMVP tool, and the application of the mask includes a bitwise AND operation between the value of the coordinates and the value of the mask.
[0409] FIG. 34 is a flowchart of a video processing method 3400. Method 3400 includes, at step 3410, determining a valid corresponding region of the current block for applying a sub-block based motion vector prediction (SbTMVP) tool to the current block based on one or more characteristics of the current block of a video segment of the video.
[0410] Method 3400 includes, at step 3420, performing a conversion between the current block and the bitstream representation of the video based on this determination.
[0411] FIG. 35 is a flowchart of a video processing method 3500. Method 3500 includes, at step 3510, determining a default motion vector for the current block of a video to be encoded using a sub-block based temporal motion vector prediction (SbTMVP) tool.
[0412] Method 3500 includes, at step 3520, based on the determination, between the current block and the video including performing conversion between the bitstream representation of the current block and the corresponding position in the associated collocated picture centered at the current block of the current picture, determining a default motion vector when a motion vector cannot be obtained from the block including the corresponding position in the associated collocated picture centered at the current block of the current picture, determining a default motion vector when a motion vector cannot be obtained from the block Figure 36 is a flowchart of a video processing method 3600. Method 3600 includes, at step
[0413] 3610, for the current block of a video segment of a video, inferring that a sub-block based temporal motion vector prediction (SbTMVP) tool or a temporal motion vector prediction ( TMVP) tool is disabled when the current picture is a reference picture with an index set to M in a reference picture list X, where M and X are integers and X = 0 or X = 1. the current picture is a reference picture with an index set to M in a reference picture list X, where M and X are integers and X = 0 or X = 1, inferring that a sub-block based temporal motion vector prediction (SbTMVP) tool or a temporal motion vector prediction ( TMVP) tool is disabled when the current picture is a reference picture with an index set to M in a reference picture list X, where M and X are integers and X = 0 or X = 1. Based on the inference, method 3600 includes, at step 3620, performing conversion between the current block and the bitstream representation of the video. Based on the inference, method 3600 includes, at step 3620, performing conversion between the current block and the bitstream representation of the video.
[0414] Figure 37 is a flowchart of a video processing method 3700. Method 3700 includes, at step 3710, for the current block of a video, determining that the application of a sub-block based temporal motion vector prediction (Sb
[0415] TMVP) tool is enabled when the current picture of the current block is a reference picture having an index set to M in a reference picture list X, where M and X are integers. 3710, for the current block of a video, determining that the application of a sub-block based temporal motion vector prediction (Sb TMVP) tool is enabled when the current picture of the current block is a reference picture having an index set to M in a reference picture list X, where M and X are integers. Based on this determination, method 3700 includes, at step 3720, performing conversion between the current block and the bitstream representation of the video. Based on this determination, method 3700 includes, at step 3720, performing conversion between the current block and the bitstream representation of the video.
[0416] Based on this determination, method 3700 includes, at step 3720, performing conversion between the current block and the bitstream representation of the video. Based on this determination, method 3700 includes, at step 3720, performing conversion between the current block and the bitstream representation of the video.
[0417] Figure 38 is a flowchart of video processing method 3800. Method 3800 includes, at step 3810, performing a conversion between the current block of the video and the bitstream representation of the video, where the current block is encoded using an encoding tool based on sub-blocks, and performing the conversion includes encoding a sub-block-based temporal motion vector prediction (SbTMV P) tool being enabled or disabled using a plurality of bins (N) to encode a sub-block merge index in a unified manner.
[0418] Figure 39 is a flowchart of video processing method 3900. Method 3900 includes, at step 3910, determining a motion vector for defining a corresponding block in a picture different from the current picture including the current block for the current block of the video encoded using a sub-block-based temporal motion vector prediction (SbTMVP) tool using the SbTMVP tool.
[0419] Method 3900 includes, at step 3920, performing a conversion between the current block and the bitstream representation of the video based on the determination.
[0420] Figure 40 is a flowchart of video processing method 4000. Method 4000 includes, at step 4010, determining whether to insert a motion zero affine merge candidate into a sub-block merge candidate list based on whether an affine prediction is enabled for converting the current block for conversion between the current block of the video and the bitstream representation of the video.
[0421] Method 4000 includes, in step 4020, performing the conversion based on the determination. including.
[0422] FIG. 41 is a flowchart of a video processing method 4100. Method 4100 includes, in step 4110, inserting zero motion non-affine padding candidates into a sub-block merge candidate list when the sub-block merge candidate list is not satisfied for conversion between the current block of a video and the bitstream representation of the video using the sub-block merge candidate list. including. including inserting zero motion non-affine padding candidates into a sub-block merge candidate list when the sub-block merge candidate list is not satisfied for conversion between the current block of a video and the bitstream representation of the video using the sub-block merge candidate list. including inserting zero motion non-affine padding candidates into a sub-block merge candidate list when the sub-block merge candidate list is not satisfied for conversion between the current block of a video and the bitstream representation of the video using the sub-block merge candidate list.
[0423] Method 4100 includes, in step 4120, performing the conversion after the insertion.
[0424] FIG. 42 is a flowchart of a video processing method 4200. Method 4200 includes, in step 4210, determining a motion vector using a rule for determining a motion vector from one or more motion vectors of blocks including corresponding positions in a collocated picture for conversion between the current block of a video and the bitstream representation of the video. including determining a motion vector using a rule for determining a motion vector from one or more motion vectors of blocks including corresponding positions in a collocated picture for conversion between the current block of a video and the bitstream representation of the video. including determining a motion vector using a rule for determining a motion vector from one or more motion vectors of blocks including corresponding positions in a collocated picture for conversion between the current block of a video and the bitstream representation of the video. including determining a motion vector using a rule for determining a motion vector from one or more motion vectors of blocks including corresponding positions in a collocated picture for conversion between the current block of a video and the bitstream representation of the video.
[0425] Method 4200 includes, in step 4220, performing the conversion based on the motion vector. including.
[0426] FIG. 43 is a block diagram of a video processing apparatus 4300. Apparatus 4300 may be used to implement one or more of the methods described herein. Apparatus 4300 may be implemented in a smartphone tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 4300 includes one or more processing devices 4302 and one or more memories 430 including one or more processing devices 4302 and one or more memories 430 including one or more processing devices 4302 and one or more memories 430 4 and video processing hardware 4306 may be included. The processing device 4302 may be configured to implement one or more of the methods described in this specification. The memory(ies) 4 304 may be used to store data and code used to implement the methods and techniques described in this specification. The video processing hardware 4306 may be used to implement the techniques described in this specification in hardware circuitry. In some embodiments, the video encoding method may be implemented using an apparatus implemented on a hardware platform, as described with reference to FIG. 43. Some embodiments of the disclosed techniques include determining or deciding to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder uses or implements this tool or mode when processing one video block, but based on the use of this tool or mode, the resulting
[0427] bitstream does not necessarily have to be modified. That is, the conversion from a video block to the bitstream representation of the video uses this video processing tool or mode based on a determination or decision when the video processing tool or mode is enabled. In another example, when a video processing tool or mode is enabled, the decoder recognizes that the bitstream has been modified based on the video processing tool or mode and processes the bitstream. That is, the video processing tool or mode enabled based on a determination or decision is used to perform the conversion from the bitstream representation of the video to the video block.
[0428]
[0429] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In one example, if a video processing tool or mode is disabled, the encoder does not use this tool or mode when converting a block of video into a bitstream representation of the video. In another example, if a video processing tool or mode is disabled, the decoder recognizes that the bitstream has not been modified using a video processing tool or mode enabled based on the decision or determination and processes the bitstream. In one example, when a video processing tool or mode is disabled, the encoder does not use this tool or mode when converting a block of video into a bitstream representation of the video. In one example, when a video processing tool or mode is disabled, the encoder does not use this tool or mode when converting a block of video into a bitstream representation of the video. In another example, if a video processing tool or mode is disabled, the decoder recognizes that the bitstream has not been modified using a video processing tool or mode enabled based on the decision or determination and processes the bitstream. In another example, if a video processing tool or mode is disabled, the decoder recognizes that the bitstream has not been modified using a video processing tool or mode enabled based on the decision or determination and processes the bitstream. In another example, if a video processing tool or mode is disabled, the decoder recognizes that the bitstream has not been modified using a video processing tool or mode enabled based on the decision or determination and processes the bitstream. In another example, if a video processing tool or mode is disabled, the decoder recognizes that the bitstream has not been modified using a video processing tool or mode enabled based on the decision or determination and processes the bitstream.
[0430] FIG. 44 is a block diagram illustrating an exemplary video processing system 4400 in which various technologies disclosed herein may be implemented. Various implementations may include some or all of the modules of system 4400. System 4400 may include an input unit 4402 for receiving video content. The video content may be received in an unprocessed or uncompressed format, such as 8 or 10-bit multi-module pixel values, or in a compressed or encoded format. The input unit 4402 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet®, Passive Optical Network (PON), etc., and wireless interfaces such as Wi-Fi® or cellular interfaces. FIG. 44 is a block diagram illustrating an exemplary video processing system 4400 in which various technologies disclosed herein may be implemented. Various implementations may include some or all of the modules of system 4400. System 4400 may include an input unit 4402 for receiving video content. The video content may be received in an unprocessed or uncompressed format, such as 8 or 10-bit multi-module pixel values, or in a compressed or encoded format. The video content may be received in an unprocessed or uncompressed format, such as 8 or 10-bit multi-module pixel values, or in a compressed or encoded format. The input unit 4402 may represent a network interface, a peripheral bus interface, or a storage interface. The input unit 4402 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet®, Passive Optical Network (PON), etc., and wireless interfaces such as Wi-Fi® or cellular interfaces. Examples of network interfaces include wired interfaces such as Ethernet®, Passive Optical Network (PON), etc., and wireless interfaces such as Wi-Fi® or cellular interfaces. Examples of network interfaces include wired interfaces such as Ethernet®, Passive Optical Network (PON), etc., and wireless interfaces such as Wi-Fi® or cellular interfaces.
[0431] System 4400 implements various encoding or encoding methods described herein. It may include an encoding module 4404 that can. The encoding module 4404 reduces the average bit rate of the video from the input unit 4402 to the output of the encoding module 4404 and may generate an encoded representation of the video. Therefore, this encoding technology is sometimes called video compression or video codec technology. The output of the encoding module 4404 may be stored or transmitted via a connected communication as represented by the module 4406. The bitstream (or encoded) representation of the video received, stored, or communicated in the input unit 4402 is used by the module 4408 to generate pixel values or a viewable video to be transmitted to the display interface unit 1910. The process of generating a video that a user can view from the bitstream representation is sometimes called video decompression. Further, specific video processing operations are called "encoding" operations or tools, but it should be understood that the encoding tools or operations are reversed by the decoder by decoding tools or operations that correspond to the encoder.
[0432] Examples of the peripheral bus interface unit or the display interface unit may include a universal serial bus (USB), a high-definition multimedia interface (HDMI (registered trademark)), a DisplayPort, etc. Examples of the storage interface may include a serial advanced technology attachment (SATA), a PCI, an IDE interface, etc. The technology described in this specification is applicable to mobile phones, notebook computers, smartphones, or any device capable of performing digital data processing and / or video display. It may be implemented in various electronic devices such as other devices.
[0433] In some embodiments, the following technical solutions can be implemented.
[0434] A1. For the conversion between the current block of the video and the bitstream representation of the video , determining whether to add the maximum number of candidates (ML) and / or the sub-block based temporal motion vector prediction (SbTMVP) to the sub-block based merge candidate list based on whether the temporal motion vector prediction (TMVP) is valid for use during the conversion or whether the current picture reference (CPR) coding mode is used for the conversion, and performing the conversion based on the determination, including a video processing method.
[0435] A2. The use of SbTMVP candidates is disabled based on the determination of whether the TMVP tool is disabled or the SbTMVP tool is disabled, as described in solution A1. The method.
[0436] A3. Determining the ML includes excluding SbTMVP candidates from the sub-block based merge candidate list based on whether the SbTMVP tool or the TMVP tool is disabled, as described in solution A2. The method.
[0437] A4. For the conversion between the current block of the video and the bitstream representation of this video , the temporal motion vector prediction (TMVP), the sub-block based temporal motion vector prediction (SbTMVP) tool, and the affine coding mode are used for this conversion and the use for the conversion, Based on whether it is effective, determine the maximum number of candidates (ML) in the merge candidate list based on sub-blocks, and perform this conversion based on the determination, a video processing method. including determining the maximum number of candidates (ML) in the merge candidate list based on sub-blocks and performing this conversion based on the determination A video processing method.
[0438] A5. The method according to solution A4, wherein due to the determination that the affine coding mode is effective, ML is set on-the-fly and signaled in the bitstream representation. The method according to solution A4, wherein due to the determination that the affine coding mode is effective, ML is set on-the-fly and signaled in the bitstream representation.
[0439] A6. The method according to solution A4, wherein due to the determination that the affine coding mode is disabled, ML is predefined. The method according to solution A4, wherein due to the determination that the affine coding mode is disabled, ML is predefined.
[0440] A7. Determining ML includes setting ML to 0 because it is determined that the TMVP tool is disabled, enabling the SbTMVP tool, and disabling the affine coding mode of the current block, according to the method described in solution A2 or A6. including setting ML to 0 because it is determined that the TMVP tool is disabled, enabling the SbTMVP tool, and disabling the affine coding mode of the current block The method according to solution A2 or A6, wherein determining ML includes setting ML to 0 because it is determined that the TMVP tool is disabled, enabling the SbTMVP tool, and disabling the affine coding mode of the current block.
[0441] A8. Determining ML includes setting ML to 1 because it is determined that the SbTMVP tool is effective, enabling the TMVP tool, and disabling the affine coding mode of the current block, according to the method described in solution A2 or A6. including setting ML to 1 because it is determined that the SbTMVP tool is effective, enabling the TMVP tool, and disabling the affine coding mode of the current block The method according to solution A2 or A6, wherein determining ML includes setting ML to 1 because it is determined that the SbTMVP tool is effective, enabling the TMVP tool, and disabling the affine coding mode of the current block.
[0442] A9. The method according to solution A1, wherein the use of SbTMVP candidates is disabled by determining that the SbTMVP tool is disabled or that the collocated reference picture of the current picture of the current block is the current picture. including disabling the use of SbTMVP candidates by determining that the SbTMVP tool is disabled or that the collocated reference picture of the current picture of the current block is the current picture The method according to solution A1, wherein the use of SbTMVP candidates is disabled by determining that the SbTMVP tool is disabled or that the collocated reference picture of the current picture of the current block is the current picture.
[0443] A10. Determining ML includes determining whether the SbTMVP tool is disabled or the current picture Based on whether the collocated reference picture of the kucha is the current picture, the solution includes excluding SbTMVP candidates from the merge candidate list based on sub blocks, the method described in A9. The method described in A9.
[0444] Determining A11.ML includes setting ML to 0 based on the determination that the collocated reference picture of the current picture is the current picture, and the affine coding of the current block being disabled, the method described in solution A9. The method described in solution A9.
[0445] Determining A12.ML includes setting ML to 1 if it is determined that the SbTMVP tool is valid, and thus the collocated reference picture of the current picture is not the current picture, and the affine coding of the current block is disabled, the method described in solution A9. The method described in solution A9. The method described in solution A9.
[0446] A13. Whether the SbTMVP tool is disabled or the reference picture with reference picture index 0 in reference picture list 0 (L 0) is determined to be the current picture of the current block, thereby disabling the use of SbTMVP candidates, the method described in solution A1. The method described in solution A1.
[0447] Determining A14.ML includes excluding SbTMVP candidates from the merge candidate list based on sub-blocks based on whether the SbTMVP tool is disabled or the reference picture with reference picture index 0 in L 0 is the current picture, the method described in solution A13. The method described in solution A13.
[0448] Determining A15.ML implies that the SbTMVP tool is determined to be valid, thus including setting ML to 0, having a reference picture index 0 in L0, where the reference picture having it is the current picture, and disabling the affine encoding of the current block, as described in solution A10 or A13.
[0449] Determining A16.ML implies that the SbTMVP tool is determined to be valid, thus including setting ML to 1, having a reference picture index 0 in L0, where the reference picture having it is not the current picture, and disabling the affine encoding of the current block, as described in A13.
[0450] Based on the determination that the SbTMVP tool is disabled, either the use of SbTMVP candidates is disabled, or the reference picture having a reference picture index 0 in the reference picture list 1 (L1) is the current picture of the current block, as described in solution A1.
[0451] Determining A18.ML includes excluding SbTMVP candidates from the sub-block-based merge candidate list based on whether the SbTMVP tool is determined to be invalid or whether the reference picture having a reference picture index 0 in L1 is the current picture, as described in solution A17.
[0452] Determining A19.ML implies that the SbTMVP tool is determined to be valid, thus setting ML to 0, disabling the reference picture where the reference picture index 0 in L1 is the current picture, and disabling the affine encoding of the current block. The method according to A17, comprising
[0453] Determining A20.ML involves determining that the SbTMVP tool is valid Thus, setting ML to 1, having a reference picture index 0 at L1 The reference picture is not the current picture, and the affine coding of the current block is invalid The method according to A17
[0454] A21. For the conversion between the current block of the first video segment of a video and the bitstream representation of this video Temporal motion vector prediction (TMVP) mode is disabled at the first video segment level Therefore, determining that the motion vector prediction based on one sub-block (SbTMVP) mode is disabled for this conversion And based on this determination, performing the conversion, and the bitstream representation includes whether the display of the SbTMVP mode is included And / or the position of the display of the SbTMVP relative to the display of the TMVP mode in the merge candidate list Conforming to the format that defines Video processing method
[0455] A22. The first video segment is a sequence, slice, tile or picture The method according to solution A21
[0456] A23. The format stipulates omitting the display of the SbTMVP mode by including the display of the TMVP mode at the first video segment level The method according to solution A21
[0457] A24. The format is such that the display of the SbTMVP mode is Subsequently, the method described in Solution A21 that specifies being at the first video segment level in the decryption order. The method.
[0458] A25. Since it is determined that the TMVP mode is invalid in the format, the method described in any one of Solutions A21 to A24 that specifies omitting the display of the SbTMVP mode. The method. The method.
[0459] A26. The method described in Solution A21 that specifies that the display of the SbTMVP mode is included at the video sequence level and omitted at the second video segment level. The method. The method.
[0460] A27. The method described in Solution A26, where the second video segment at the second video segment level is a slice, tile, or picture.
[0461] A28. The method described in any one of Solutions A1 to A27, where the conversion generates the current block from the bitstream representation.
[0462] A29. The method described in any one of Solutions A 1 to A27, where the conversion generates the bitstream representation from the current block.
[0463] A30. Performing the conversion includes the step of syntax-analyzing the bitstream representation based on one or more decryption rules, and the method described in any one of Solutions A1 to A27.
[0464] A31. An apparatus in a video system including a processing device and a non-transitory memory storing instructions, where the instructions executed by the processing device cause the processing device to implement the method described in any one of Solutions A1 to A30. The apparatus in a video system. The apparatus in a video system.
[0465] A32. A computer program product stored on a non-transitory computer-readable medium A program code for executing the method according to any one of solutions A1 to A30. 2. A computer program product, including any
[0466] In some embodiments, the following technical solutions can be implemented.
[0467] B1. Sub-block based temporal motion vector prediction (SbTMVP) tool or The current block of video coded using the Time-domain Motion Vector Prediction (TMVP) tool The SbTMVP tool converts between the bitstream representation of this video and the Use a mask based on the compression of motion vectors associated with the tool or TMVP tool. Then, the coordinates of the position corresponding to the current block or a subblock of this current block are calculated. Selectively masking and applying this mask creates a bilinear relationship between the value of this coordinate and the value of this mask. A video processing method including bit-wise AND operations.
[0468] B2. The coordinates are (xN,yN) and the mask is an integer equal to ~(2M-1). where M is an integer and the masked coordinates (xN', yN') are , xN'=xN&MASK, yN'=yN&MASK, and "~" is a bit "&" is a bitwise NOT operation, and "&" is a bitwise AND operation. Method of posting.
[0469] B3. The method according to solution B2, wherein M=3 or M=4.
[0470] B4. Based on the compression of the motion vectors, a plurality of sub-blocks of size 2K×2K share the same motion information, where K is an integer not equal to M, according to the method described in Solution B2 or B3.
[0471] B5. The method according to Solution B4, where M = K + 1.
[0472] B6. When it is determined that the motion vectors associated with the SbTMVP tool or the TMVP tool are not compressed, the mask is not applied, according to the method described in Solution B1.
[0473] B7. The mask for the SbTMVP tool is the same as the mask for the TMVP tool, according to the method described in any one of Solutions B1 to B6.
[0474] B8. The mask for the ATMVP tool is different from the mask for the TMVP tool, according to the method described in any one of Solutions B1 to B6.
[0475] B9. One type of compression is non - compression, 8×8 compression, or 16×16 compression, according to the method described in Solution B1.
[0476] B10. The type of compression is signaled by a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a slice header, or a tile group header, according to the method described in Solution B9.
[0477] B11. The type of compression is based on the standard profile, level, or layer corresponding to the current block, according to the method described in Solution B9 or B10.
[0478] B12. Based on one or more characteristics of the current block of a video segment of the video, the current Determining a valid corresponding region of a current block for applying a sub-block based motion vector prediction (SbTMVP) tool based on a block, and performing a conversion between the current block and a bitstream representation of a video based on this determination, including: A video processing method.
[0479] B13. The method according to solution B12, wherein one or more features include the height or width of the current block.
[0480] B14. The method according to solution B12, wherein one or more features include the type of compression of motion vectors associated with the current block.
[0481] B15. The method according to solution B14, wherein the valid corresponding region is of a first size when it is determined that the type of compression does not include compression, and the valid corresponding region is of a second size larger than the first size when it is determined that the type of compression includes K×K compression.
[0482] B16. The method according to solution B12, wherein the size of the valid corresponding region is based on a basic region smaller than the size of a coding tree unit (CTU) region of size M×N, and the size of the current block is W×H.
[0483] B17. The method according to solution B16, wherein the size of the CTU region is 128×128, M = 64, and N = 64.
[0484] B18. The method according to solution B16, wherein when it is determined that W≤M and H≤N, the valid corresponding region is the collocation basic region and expansion in the collocated picture. The method described in Decision B16.
[0485] If it is determined that B19.W > M and H > N, the current block is divided into a plurality of parts Each of the plurality of parts includes an individual effective corresponding area for applying the SbTMVP tool The method described in Decision B16.
[0486] B20. Using a sub-block based temporal motion vector prediction (SbTMVP) tool To determine a default motion vector for the current block of the encoded video Based on this determination, perform a conversion between the current block and the bitstream representation of this video And obtain a motion vector from a block including the corresponding position in the collocated picture associated with the center position of the current block If it is determined that there is none, determine the default motion vector, a video processing method.
[0487] B21. The method described in Decision B20, wherein the default motion vector is set to (0, 0).
[0488] B22. The default motion vector is derived from a history based motion vector prediction (HMVP) table, the method described in Decision B20.
[0489] B23. If it is determined that the HMVP table is empty, the default motion vector is set to (0, 0), the method described in Decision B22.
[0490] B24. The default motion vector is predefined, video parameter set (VPS ), sequence parameter set (SPS), picture parameter set (PPS), s The method according to Solution B22, which is signaled to a rice header, a tile group header, a coding tree unit (CTU), or a coding unit (CU).
[0491] Based on the determination that the B25. HMVP table is not empty, set the default motion vector to the first element stored in the HMVP table, as described in Solution B22
[0492] Based on the determination that the B26. HMVP table is not empty, set the default motion vector to the last element stored in the HMVP table, as described in Solution B22
[0493] Based on the determination that the B27. HMVP table is not empty, set the default motion vector to a specific motion vector stored in the HMVP table, Solution B22
[0494] B28. The specific motion vector refers to reference list 0, as described in Solution B27
[0495] B29. The specific motion vector refers to reference list 1, as described in Solution B27
[0496] B30. The specific motion vector refers to a specific reference picture in reference list 0 as described in Solution B27
[0497] B31. The specific motion vector refers to a specific reference picture in reference list 1 as described in Solution B27
[0498] B32. The specific reference picture has index 0, in Solution B30 or B31 The method described in
[0499] B33. A specific motion vector refers to a collocated picture, Solution B The method described in 27.
[0500] B34. In the search process in the HMVP table, if a specific motion vector is not found is determined, set the default motion vector to a predefined default motion vector as described in Solution B22.
[0501] B35. The search process searches only the first element or the last element of the HMVP table as described in Solution B34.
[0502] B36. The search process searches only a subset of the elements of the HMVP table, Solution B 34.
[0503] B37. The default motion vector does not refer to the current picture of the current block as described in Solution B22.
[0504] B38. Based on the determination that the default motion vector does not refer to the collocated picture scale the default motion vector to the collocated picture as described in Solution B22.
[0505] B39. The default motion vector is derived from neighboring blocks, Solution B20 described.
[0506] B40. The upper right corner of the neighboring block (A0) is directly adjacent to the lower left corner of the current block or the lower right corner of the neighboring block (A1) is directly adjacent to the lower left corner of the current block either the bottom left corner of the neighboring block (B0) is directly adjacent to the upper right corner of the current block or the bottom right corner of the neighboring block (B1) is directly adjacent to the upper right corner of the current block or the bottom right corner of the neighboring block (B2) is directly adjacent to the upper left corner of the current block the method described in solution B39.
[0507] B41. The default motion vector is derived from only one of the neighboring blocks A0, A1, B0, B1, B2 the method described in solution B40.
[0508] B42. The default motion vector is derived from one or more of the neighboring blocks A0, A1, B0, B1, B2 the method described in solution B40.
[0509] B43. If it is determined that no valid default motion vector is found in any of the neighboring blocks A0, A1, B0, B1, B2, set the default motion vector to a predefined default motion vector, the method described in solution B40.
[0510] B44. The predefined default motion vector is signaled in the video parameter set (VPS ), sequence parameter set (SPS), picture parameter set (PPS), slice header, tile group header, coding tree unit (CTU), or coding unit (CU), the method described in solution B43.
[0511] B45. The predefined default motion vector is (0,0), the method described in solution B43 or B44.
[0512] The default motion vector is set to a specific motion vector from neighboring blocks, according to the method described in Solution B39.
[0513] The specific motion vector is obtained by referring to reference list 0, according to the method described in Solution B46.
[0514] The specific motion vector is obtained by referring to reference list 1, according to the method described in Solution B46.
[0515] The specific motion vector is obtained by referring to a specific reference picture in reference list 0, according to the method described in Solution B46.
[0516] The specific motion vector is obtained by referring to a specific reference picture in reference list 1, according to the method described in Solution B46.
[0517] The specific reference picture has index 0, according to the method described in Solution B49 or B50.
[0518] The specific motion vector is obtained by referring to the collocated picture, according to the method described in Solution B46.
[0519] If it is determined that the block containing the corresponding position in the collocated picture is intra-coded, the default motion vector is used, according to the method described in Solution B20.
[0520] The derivation method is modified when it is determined that the block containing the corresponding position in the collocated picture is not present, according to the method described in Solution B20.
[0521] Solution B20, wherein the default motion vector candidates are always available Method
[0522] B56. Solution B20, wherein when it is determined that the default motion vector candidates are set to be unavailable the default motion vectors are alternatively derived
[0523] B57. Solution B20, wherein the availability of the default motion vectors is based on syntax information in a bitstream representation associated with a video segment
[0524] B58. Solution B57, wherein the syntax information includes an indication enabling the SbTMVP tool, and the video segment is a slice, tile or picture
[0525] B59. Solution B58, wherein the current picture of the current block is not an intra-random access point (IRAP) reference index picture, and the current picture is not inserted into reference picture list 0 (L0) having reference index 0
[0526] B60. Solution B20, wherein upon determining that the SbTMVP tool is enabled, a fixed index or group of fixed indexes is assigned to candidates associated with the SbTMVP tool, and when it is determined that the SbTMVP tool is disabled, a fixed index or group of fixed indexes is assigned to candidates associated with an encoding tool other than the SbTMVP tool Method
[0527] B61. For the current block of a video segment of a video, wherein the current picture of the current block has a reference index set to M in reference picture list X reference picture where M and X are integers and it is determined that X = 0 or X = 1, and based on this sub-block based temporal motion vector prediction (SbTMVP) tool or temporal motion vector prediction (TMVP) tool is disabled for the video segment, inferring, and based on the inference, performing a conversion between the current block and the bitstream representation of the video, including a video processing method.
[0528] For reference picture list X for the SbTMVP tool or TMVP tool, where M corresponds to the reference picture index that scales the motion information of the temporal block, the method according to solution B61.
[0529] B63. The current picture is an intra-random access point (IRAP) picture, the method according to solution B61.
[0530] B64. For the current block of the video, the current picture of the current block is a reference picture having the index set to M in the reference picture list X, and M and X are determined to be integers, thereby determining that the application of the sub-block based temporal motion vector prediction (SbTMVP) tool is valid, and based on this determination, performing a conversion between the current block and the bitstream representation of the video, including a video processing method. including a video processing method.
[0531] B65. The motion information corresponding to each sub-block of the current block refers to the current picture, the method according to solution B64.
[0532] Motion information for sub - blocks of the current block is derived from one temporal block, and this temporal block is encoded with at least one reference picture that references the current picture of this temporal block, the method according to solution B64. derived, and this temporal block is encoded with at least one reference picture that references the current picture of this temporal block, the method according to solution B64.
[0533] B67. The transformation is the method according to solution B66, excluding a scaling operation.
[0534] B68. Performing a transformation between the current block of a video and the bit - stream representation of the video, where the current block is encoded using an encoding tool based on sub - blocks, and performing this transformation includes encoding a sub - block merge index in a unified manner using a plurality of bins (N) based on whether a sub - block - based temporal motion vector prediction (SbTMVP) tool is determined to be valid or invalid, a video processing method. including encoding a sub - block merge index in a unified manner using a plurality of bins (N) based on whether a sub - block - based temporal motion vector prediction (SbTMVP)
[0535] B69. The first number of bins (L) of the plurality of bins is context - encoded, and the second number of bins (N - L) is bypass - encoded, the method according to solution B68. is context - encoded, and the second number of bins (N - L) is bypass - encoded, the method according to solution B68.
[0536] B70. The method according to solution B69, where L = 1.
[0537] B71. Each of the plurality of bins is context - encoded, the method according to solution B68 .
[0538] B72. The transformation is the method according to any one of solutions B1 to B71, generating the current block from the bit - stream representation.
[0539] B73. The transformation is the method according to any one of solutions B1 to B71, generating the bit - stream representation from the current block.
[0540] Performing the conversion includes syntax-analyzing the bitstream representation based on one or more decoding rules, according to any of the methods described in Solutions B1 to B71. The method according to any of Solutions B1 to B71, wherein performing the conversion includes syntax-analyzing the bitstream representation based on one or more decoding rules.
[0541] An apparatus in a video system, comprising a processing device and a non-transitory memory storing instructions, wherein the instructions executed by the processing device cause the processing device to implement the method according to any of Solutions B1 to B71. The apparatus according to any of Solutions B1 to B71, wherein the instructions executed by the processing device cause the processing device to implement the method described in any one of the solutions. The apparatus is characterized in that the instructions executed by the processing device cause the processing device to implement the method described in any one of Solutions B1 to B71.
[0542] A computer program product stored in a non-transitory computer-readable medium, the computer program product including program code for executing the method according to any one of Solutions B1 to B74. The computer program product according to any one of Solutions B1 to B74, wherein the computer program product includes program code for executing the method described in any one of the solutions. The computer program product includes program code for executing the method described in any one of Solutions B1 to B74.
[0543] In some embodiments, the following technical solutions can be implemented.
[0544] A video processing method, including determining a motion vector used by a sub-block based time motion vector prediction (SbTMVP) tool to locate a corresponding block in a picture different from the current picture containing the current block for the current block of an encoded video, and performing a conversion between the current block and the bitstream representation of the video based on this determination. For the current block of an encoded video, using the SbTMVP tool to determine the motion vector used to locate the corresponding block in a picture different from the current picture containing the current block, and performing a conversion between the current block and the bitstream representation of the video based on this determination. For the current block of an encoded video, using the SbTMVP tool to determine the motion vector used to locate the corresponding block in a picture different from the current picture containing the current block, and performing a conversion between the current block and the bitstream representation of the video based on this determination. For the current block of an encoded video, using the SbTMVP tool to determine the motion vector used to locate the corresponding block in a picture different from the current picture containing the current block, and performing a conversion between the current block and the bitstream representation of the video based on this determination. A video processing method, including determining a motion vector used by a sub-block based time motion vector prediction (SbTMVP) tool to locate a corresponding block in a picture different from the current picture containing the current block for the current block of an encoded video, and performing a conversion between the current block and the bitstream representation of the video based on this determination.
[0545] The method according to Solution C1, wherein the motion vector is set to a default motion vector.
[0546] The method according to Solution C2, wherein the default motion vector is (0, 0).
[0547] C4. The default motion vector is the method described in step C2, which is signaled in a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a slice header, a tile group header, a coding tree unit (CTU), or a coding unit (CU).
[0548] C5. The method described in solution C1, where the motion vector is set to the motion vector stored in the history-based motion vector prediction (HMVP) table.
[0549] C6. The method described in solution C5, where the motion vector is set to the default motion vector based on the determination that the HMVP table is empty.
[0550] C7. The method described in solution C6, where the default motion vector is (0, 0).
[0551] C8. The method described in solution C5, where the motion vector is set to the first motion vector stored in the HMVP table based on the determination that the HMVP table is not empty.
[0552] C9. The method described in solution C5, where the motion vector is set to the last motion vector stored in the HMVP table due to the determination that the HMVP table is not empty.
[0553] C10. The method described in solution C5, where the motion vector is set to a specific motion vector stored in the HMVP table based on the determination that the HMVP table is not empty.
[0554] C11. The method described in solution C10, where the specific motion vector refers to reference list 0.
[0555] C12. A specific motion vector refers to the method described in Solution C10 that refers to reference list 1 .
[0556] C13. A specific motion vector refers to a specific reference picture in reference list 0 , the method described in Solution C10
[0557] C14. A specific motion vector refers to a specific reference picture in reference list 1 , the method described in Solution C10
[0558] C15. A specific reference picture has index 0, in Solution C13 or 14 the method described
[0559] C16. A specific motion vector refers to a collocated picture, Solution C 10
[0560] C17. If it is determined that a specific motion vector is not found in the search process in the HMVP table , set the motion vector to the default motion vector, Solution C5 the method described
[0561] C18. The search process searches only the first element of the HMVP table or only the last element of the HMVP table, the method described in Solution C17
[0562] C19. The search process searches only a subset of the elements of the HMVP table, Solution C 17
[0563] C20. The motion vector stored in the HMVP table does not refer to the current picture , the method described in Solution C5
[0564] The picture in which the motion vectors stored in the C21.HMVP table are collocated is not referenced, so the method described in Solution C5 is used to scale the motion vectors stored in the HMVP table to the collocated picture.
[0565] C22. The motion vectors are set to the motion vectors of specific blocks in a specific neighborhood, the method described in Solution C1.
[0566] C23. The upper right corner of a specific neighboring block (A0) is directly adjacent to the lower left corner of the current block or the lower right corner of a specific neighboring block (A1) is directly adjacent to the lower left corner of the current block or the lower left corner of a specific neighboring block (B0) is directly adjacent to the upper right corner of the current block or the lower right corner of a specific neighboring block (B1) is directly adjacent to the upper right corner of the current block or the lower right corner of a specific neighboring block (B2) is directly adjacent to the upper left corner of the current block or is directly adjacent to the upper left corner of the current block, the method described in Solution C22.
[0567] C24. By determining that there are no specific neighboring blocks, the motion vectors are set to the default motion vectors, the method described in Solution C1.
[0568] C25. By determining that the specific neighboring blocks are not inter-coded the motion vectors are set to the default motion vectors, the method described in Solution C1.
[0569] C26. The specific motion vectors refer to reference list 0, the method described in Solution C22 .
[0570] A specific motion vector refers to the method described in Solution C22 that references Reference List 1 .
[0571] C28. A specific motion vector refers to a specific reference picture in Reference List 0 , the method described in Solution C22.
[0572] C29. A specific motion vector refers to a specific reference picture in Reference List 1 , the method described in Solution C22.
[0573] C30. A specific reference picture has index 0, the method described in Solution C28 or C29 .
[0574] C31. A specific motion vector refers to a collocated picture, the method described in Solution C 22 or C23.
[0575] C32. By determining that a specific neighboring block does not refer to a collocated picture, setting the motion vector to the default motion vector, the method described in Solution C2 2 or C23.
[0576] C33. The default motion vector is (0,0), the method described in any of Solutions C24 to C32 .
[0577] C34. If it is determined that a specific motion vector stored in a specific neighboring block cannot be found, setting this motion vector to the default motion vector, the method described in Solution C1 .
[0578] C35. A specific motion vector is such that the specific motion vector is the collocated pic Determined not to refer to the picture, the method described in Solution C22, which is applied to one collocated picture with a ski -ring.
[0579] C36. A specific motion vector is the method described in Solution C22 that does not refer to the current picture. Method.
[0580] C37. The conversion is the method described in any of Solutions C1 to C36 for generating the current block from the bitstream representation.
[0581] C38. The conversion is the method described in any of Solutions C1 to C36 for generating the bitstream representation from the current block.
[0582] C39. Performing the conversion includes syntax analyzing the bitstream representation based on one or more decoding rules, which is the method described in any of Solutions C1 to C36.
[0583] C40. An apparatus in a video system including a processing device and a non-transitory memory storing instructions, wherein the instructions executed by the processing device cause the processing device to implement the method described in any one of Solutions C1 to C39.
[0584] C41. A computer program product stored in a non-transitory computer-readable medium, wherein the computer program product includes program code for executing the method described in any one of Solutions C1 to C39.
[0585] In some embodiments, the following technical solutions can be implemented.
[0586] D1. For the conversion between the current block of the video and the bitstream representation of the video, Based on whether affine prediction is enabled for the current block conversion, determine whether to insert a motion zero affine merge candidate into the sub-block merge candidate list, and perform this conversion based on this determination, a video processing method. And, based on this determination, perform this conversion, including determining whether to insert a motion zero affine merge candidate into the sub-block merge candidate list. A video processing method.
[0587] D2. The method according to Solution D1, wherein, since it is determined that the affine usage flag in the bitstream representation is off, a motion zero affine merge candidate is not inserted into the sub-block merge candidate list. A solution. The method described in Solution D1.
[0588] D3. The method according to Solution D2, further including inserting a default motion vector candidate, which is a non-affine candidate, into the sub-block merge candidate list because it is determined that the affine usage flag is off. And further including inserting a default motion vector candidate, which is a non-affine candidate, into the sub-block merge candidate list because it is determined that the affine usage flag is off. The method described in Solution D2.
[0589] D4. When it is determined that the sub-block merge candidate list is not satisfied for the conversion between the current block of the video and the bitstream representation of the video using the sub-block merge candidate list, insert a zero motion non-affine padding candidate into the sub-block merge candidate list, and subsequent to this insertion, perform this conversion, a video processing method. And subsequent to this insertion, perform this conversion, including inserting a zero motion non-affine padding candidate into the sub-block merge candidate list when it is determined that the sub-block merge candidate list is not satisfied for the conversion between the current block of the video and the bitstream representation of the video using the sub-block merge candidate list. A video processing method. The method described in Solution D4. A video processing method.
[0590] D5. The method according to Solution D4, further including setting the affine usage flag of the current block to 0. The method described in Solution D4.
[0591] D6. The method according to Solution D4, wherein the insertion step is further based on whether the affine usage flag in the bitstream representation is off. The method described in Solution D4.
[0592] For the conversion between the current block of the video and the bitstream representation of the video, using a rule for determining that the motion vector is derived from one or more motion vectors of a block including corresponding positions in the collocated picture, determine the motion vector, and perform this conversion based on this motion vector, including a video processing method 。
[0593] D8. One or more motion vectors include MV0 and MV1 representing the motion vectors in reference list 0 and reference list 1 respectively, and the motion vector to be derived includes MV0’ and MV1’ representing the motion vectors in reference list 0 and reference list 1, the method according to solution D7.
[0594] D9. Based on the determination that one collocated picture is in reference list 0, the method according to solution D8, where MV0’ and MV1’ are derived based on MV0.
[0595] D10. Based on the determination that one collocated picture is in reference list 1, the method according to solution D8, where MV0’ and MV1’ are derived based on MV1.
[0596] D11. The conversion is the method described in any one of solutions D1 to D10 for generating the current block from the bitstream representation.
[0597] D12. The conversion is the method described in any one of solutions D1 to D16 for generating the bitstream representation from the current block.
[0598] D13. Performing the conversion includes decoding the bitstream representation based on one or more decoding rules The method according to any one of Solutions D1 to D10, including syntax analysis.
[0599] D14. In a video system, an apparatus including a processing device and a non-transitory memory storing instructions wherein the instructions executed by the processing device cause the processing device to implement the method according to any one of Solutions D1 to D13 An apparatus characterized in that the method according to any one of the solutions is implemented.
[0600] D15. A computer program product stored in a non-transitory computer-readable medium wherein the computer program product includes program code for executing the method according to any one of Solutions D1 to D13
[0601] The disclosed and other solutions, examples, embodiments, modules, and functional operations described herein may be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and structural equivalents thereof, or in combinations of one or more of them. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., as one or more modules of computer program instructions encoded on a computer-readable medium for being executed by a data processing apparatus or for controlling the operation of a data processing apparatus. This computer-readable medium may be, for example, a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter that provides a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" includes, for example, a programmable processing device, a computer, or multiple processing devices or computers including all apparatus, devices, and machines for processing data. This apparatus may, in addition to hardware, include code that creates an execution environment for the computer program, such as , processing device firmware, protocol stack, database management system, operating system, or code that constitutes one or more combinations thereof. The propagated signal is an artificially generated signal, such as an electrical, optical, or electromagnetic signal generated by a machine, and is generated to encode information for transmission to a suitable receiving device.
[0602] A computer program (also referred to as a program, software, software application , script, or code) can be described in any form of programming language, including compiled or interpreted languages, and it can be developed in any form, including as a stand-alone program or as a module , component, subroutine, or other unit suitable for use in a computing environment. The computer program does not necessarily correspond to a file in the file system. The program can be recorded as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), or it can be stored in a single file dedicated to the program , or it can be stored in multiple adjustment files (e.g., files that store parts of one or more modules, subprograms, or codes). One computer program can be located on one computer at one site or distributed across multiple sites and communicated via a network connected by a network. to be deployed to be executed on a plurality of computers interconnected by a network is also possible.
[0603] The processes and logic flows described herein operate on input data and perform functions by generating output by executing one or more computer programs for performing the functions on one or more programmable processing devices. The processes and logic flows can also be performed by special-purpose logic circuits, such as, for example, FPGAs (field programmable gate arrays) or ASICs (application specific integrated circuits), and the apparatus can also be implemented as special-purpose logic circuits. A processing device suitable for the execution of a computer program includes, for example, both general-purpose and special-purpose micro
[0604] processing devices, as well as any one or more processing devices of any type of digital computer In general, a processing device receives instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processing device for performing instructions and one or more memory devices for storing the instructions and data In general, a computer may include one or more mass storage devices for storing data, such as, for example, magnetic, magneto-optical disks, or optical disks, or may be operatively coupled to receive data from or transfer data to these mass storage devices However, a computer need not have such devices. A computer-readable medium suitable for storing computer program instructions and data includes any form of non-volatile memory, medium, and memory device, including, for example For example, it includes semiconductor memory devices such as EPROM, EEPROM, flash memory devices, magnetic disks, for example internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processing device and the memory may be complemented by application-specific logic circuits or may be incorporated in application-specific logic circuits.
[0605] This patent specification includes many details, but these should not be construed as limiting the scope of any subject matter or what can be claimed, but rather as descriptions and interpretations of features that may be specific to particular embodiments of a particular technology. Specific features described in the context of separate embodiments in this patent specification may be implemented in combination in one example. Conversely, various features described in the context of one example may be implemented separately or in any suitable sub-combination in multiple embodiments. Further, features may be described and initially claimed above as acting in a particular combination, but one or more features from the claimed combination may, in some cases, be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of sub-combinations. Similarly, operations are shown in the drawings in a particular order, but this should not be understood as requiring that such operations be performed in the particular order shown or in a sequential order to achieve the desired result, nor that all the operations shown be performed. In the context of separate embodiments in this patent specification, specific features described may be implemented in combination in one example. Conversely, various features described in the context of one example may be implemented separately or in any suitable sub-combination in multiple embodiments. Further, features may be described and initially claimed above as acting in a particular combination, but one or more features from the claimed combination may, in some cases, be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of sub-combinations. In the context of separate embodiments in this patent specification, specific features described may be implemented in combination in one example. Conversely, various features described in the context of one example may be implemented separately or in any suitable sub-combination in multiple embodiments. Further, features may be described and initially claimed above as acting in a particular combination, but one or more features from the claimed combination may, in some cases, be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of sub-combinations. In the context of separate embodiments in this patent specification, specific features described may be implemented in combination in one example. Conversely, various features described in the context of one example may be implemented separately or in any suitable sub-combination in multiple embodiments. Further, features may be described and initially claimed above as acting in a particular combination, but one or more features from the claimed combination may, in some cases, be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of sub-combinations. In the context of separate embodiments in this patent specification, specific features described may be implemented in combination in one example. Conversely, various features described in the context of one example may be implemented separately or in any suitable sub-combination in multiple embodiments. Further, features may be described and initially claimed above as acting in a particular combination, but one or more features from the claimed combination may, in some cases, be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of sub-combinations. In the context of separate embodiments in this patent specification, specific features described may be implemented in combination in one example. Conversely, various features described in the context of one example may be implemented separately or in any suitable sub-combination in multiple embodiments. Further, features may be described and initially claimed above as acting in a particular combination, but one or more features from the claimed combination may, in some cases, be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of sub-combinations. In the context of separate embodiments in this patent specification, specific features described may be implemented in combination in one example. Conversely, various features described in the context of one example may be implemented separately or in any suitable sub-combination in multiple embodiments. Further, features may be described and initially claimed above as acting in a particular combination, but one or more features from the claimed combination may, in some cases, be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of sub-combinations. In the context of separate embodiments in this patent specification, specific features described may be implemented in combination in one example. Conversely, various features described in the context of one example may be implemented separately or in any suitable sub-combination in multiple embodiments. Further, features may be described and initially claimed above as acting in a particular combination, but one or more features from the claimed combination may, in some cases, be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of sub-combinations. In the context of separate embodiments in this patent specification, specific features described may be implemented in combination in one example. Conversely, various features described in the context of one example may be implemented separately or in any suitable sub-combination in multiple embodiments. Further, features may be described and initially claimed above as acting in a particular combination, but one or more features from the claimed combination may, in some cases, be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of sub-combinations.
[0606] Similarly, operations are shown in the drawings in a particular order, but this should not be understood as requiring that such operations be performed in the particular order shown or in a sequential order to achieve the desired result, nor that all the operations shown be performed. In the context of separate embodiments in this patent specification, specific features described may be implemented in combination in one example. Conversely, various features described in the context of one example may be implemented separately or in any suitable sub-combination in multiple embodiments. Further, features may be described and initially claimed above as acting in a particular combination, but one or more features from the claimed combination may, in some cases, be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of sub-combinations. In the context of separate embodiments in this patent specification, specific features described may be implemented in combination in one example. Conversely, various features described in the context of one example may be implemented separately or in any suitable sub-combination in multiple embodiments. Further, features may be described and initially claimed above as acting in a particular combination, but one or more features from the claimed combination may, in some cases, be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of sub-combinations. No. Also, the separation of various system modules in the embodiments described in this patent specification should not be understood as requiring such separation in all embodiments. No.
[0607] Only some implementation forms and examples are described, and based on the content described and illustrated in this patent specification, other embodiments, extensions and modifications are possible.
Claims
Claim 1 A method for coding video data, comprising: determining that a time motion vector prediction mode based on sub-blocks is invalidated for the conversion due to the time motion vector prediction mode being invalidated at a first video segment level for conversion between a first video segment of the video and the bitstream of the video; performing the conversion based on the determination; wherein a time motion vector prediction candidate based on sub-blocks derived based on the time motion vector prediction mode based on sub-blocks is used for constructing a sub-block merge candidate list, and a time motion vector prediction candidate derived based on the time motion vector prediction mode is used for constructing a merge candidate list different from the sub-block merge candidate list; the sub-block merge candidate list includes an affine merge candidate and a time motion vector prediction candidate based on the sub-blocks, excluding the time motion vector prediction candidate; the merge candidate list includes a spatial merge candidate, a time motion vector prediction candidate, a motion vector prediction candidate based on history, and a pairwise merge candidate, excluding the time motion vector prediction candidate based on the sub-blocks; a method. Claim 2 The maximum number of candidates (ML) in the merge candidate list based on the sub-blocks is based on whether the time motion vector prediction mode is enabled. The method according to claim 1. Claim 3 The first video segment is a sequence, slice, tile or picture. The method according to claim 1 or 2. Claim 4 The display of the time motion vector prediction mode based on the sub-blocks is at the first video segment level after the display of the time motion vector prediction mode in the bitstream. The method according to any one of claims 1 to 3. Claim 5 The display of the time motion vector prediction mode based on the sub-blocks is omitted when it is displayed that the time motion vector prediction mode is invalidated. The method according to any one of claims 1 to 3. Claim 6 The display of the time motion vector prediction mode based on the sub-blocks is included at the sequence level of the video and omitted at a second video segment level. The method according to any one of claims 1 to 3. Claim 7 The second video segment level is a slice level, a tile level, or a picture level. The method according to claim 6. Claim 8 Constructing a sub-block motion candidate list for blocks within a second video segment of the video, wherein the blocks are coded in a temporal motion vector prediction mode based on the sub-blocks, and constructing; Performing the transform includes performing the transform based on the sub-block motion candidate list. A sub-block merge index is included in the bitstream. The method according to any one of claims 1 to 7. Claim 9 A plurality of bins (N) are used to present the sub-block merge index. The method according to claim 8. Claim 10 A first number of bins (L) of the plurality of bins are context-coded. A second number of bins (N - L) are bypass-coded. The method according to claim 9. Claim 11 L = 1. The method according to claim 10. Claim 12 The transform includes decoding the first video segment from the bitstream. The method according to any one of claims 1 to 11. Claim 13 The transform includes encoding the first video segment into the bitstream. The method according to any one of claims 1 to 11. Claim 14 The transform includes generating the bitstream from the video. The method further includes storing the bitstream in a non-transitory computer-readable recording medium. The method according to any one of claims 1 to 11. Claim 15 An apparatus for processing video data to be coded, comprising a processing device and a non-transitory memory storing instructions. When executed by the processing device, the instructions cause the processing device to Determine that a temporal motion vector prediction mode based on sub-blocks is disabled for the transform due to the temporal motion vector prediction mode being disabled at the first video segment level for the transform between a first video segment of the video and the bitstream of the video; Perform the transform based on the determination. The time motion vector prediction candidates based on sub-blocks derived based on the time motion vector prediction mode based on the sub-blocks are used for constructing a sub-block merge candidate list, and the time motion vector prediction candidates derived based on the time motion vector prediction mode are used for constructing a merge candidate list different from the sub-block merge candidate list. The sub-block merge candidate list includes an affine merge candidate and a time motion vector prediction candidate based on the sub-block, and excludes the time motion vector prediction candidate. The merge candidate list includes a spatial merge candidate, the time motion vector prediction candidate, a motion vector prediction candidate based on history, and a pairwise merge candidate, and excludes the time motion vector prediction candidate based on the sub-block. Device.
16. A method for storing a video bitstream, comprising: Determining that, for a first video segment of the video, the time motion vector prediction mode based on sub-blocks is invalidated for conversion due to the time motion vector prediction mode being invalidated at the first video segment level; Generating the bitstream from the video based on the determination; Storing the bitstream in a non-transitory computer-readable recording medium. Including: The time motion vector prediction candidates based on sub-blocks derived based on the time motion vector prediction mode based on the sub-blocks are used for constructing a sub-block merge candidate list, and the time motion vector prediction candidates derived based on the time motion vector prediction mode are used for constructing a merge candidate list different from the sub-block merge candidate list. The sub-block merge candidate list includes an affine merge candidate and a time motion vector prediction candidate based on the sub-block, and excludes the time motion vector prediction candidate. The merge candidate list includes a spatial merge candidate, the time motion vector prediction candidate, a motion vector prediction candidate based on history, and a pairwise merge candidate, and excludes the time motion vector prediction candidate based on the sub-block. Method.
17. A non-transitory computer-readable storage medium storing instructions, wherein: The instructions cause a processing device to For the conversion of the first video segment of the video and the bitstream of the video, it is determined that the temporal motion vector prediction mode based on sub-blocks is disabled for the conversion due to the temporal motion vector prediction mode being disabled at the first video segment level, Based on the determination, the conversion is performed, The temporal motion vector prediction candidates based on sub-blocks derived based on the temporal motion vector prediction mode based on sub-blocks are used for constructing a sub-block merge candidate list, and the temporal motion vector prediction candidates derived based on the temporal motion vector prediction mode are used for constructing a merge candidate list different from the sub-block merge candidate list, The sub-block merge candidate list includes an affine merge candidate and a temporal motion vector prediction candidate based on the sub-block, and excludes the temporal motion vector prediction candidate, The merge candidate list includes a spatial merge candidate, a temporal motion vector prediction candidate, a motion vector prediction candidate based on history, and a pairwise merge candidate, and excludes the temporal motion vector prediction candidate based on the sub-block, Non-transitory computer-readable storage medium.
Citation Information
Patent Citations
Image converter
JP2013016899A
Block-based altitude residual prediction for 3D video coding
JP2017510127A
Overlapping motion compensation for video coding
JP2018509032A
Simplifying inter-intra combined prediction
JP2022505889A
Weights in inter-intra combined prediction modes
JP2022506119A