Restrictions on the Difference of Motion Vectors
By dynamically adjusting MVD ranges based on codec-specific parameters, the patent addresses inaccuracies in MVD values, enhancing video decoding accuracy and efficiency in standards like VVC.
Patent Information
- Application Number
- JP2023100145
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-04-25
- Filing Date
- 2023-06-19
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-04-26
AI Technical Summary
Existing video coding standards like VVC face inaccuracies in Motion Vector Difference (MVD) values due to bitstream constraints that clip MVD components to a fixed range, especially when resolutions other than 1/4 pixel are used, leading to potential inaccuracies.
Adapt the range of MVD components based on the codec's acceptable resolution/accuracy and coded information, such as AMVR, AMVP mode, and affine inter flag, to ensure accurate MVD values by scaling and clipping them within predefined ranges.
Ensures accurate and efficient MVD representation by adapting the range to the codec's capabilities, reducing inaccuracies and improving video decoding performance.
Smart Images

Figure 0007701406000017 
Figure 0007701406000018 
Figure 0007701406000019
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This This application timely claims the priority and benefits of International Patent Application No. PCT / CN2019 / 084228 filed on April 25, 2019. is a divisional application of Japanese Patent Application No. 2021-562033, which is the national phase of International Patent Application No. PCT / CN2020 / 087068 filed on April 26, 2020. The entire disclosure of the above application is By reference in part as part of the disclosure of this application here is incorporated herein.
[0002] This patent specification relates to video coding technology, devices and systems.
Background Art
[0003] Despite the progress of video compression, digital video still occupies the largest bandwidth usage in the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for the use of digital video is expected to continue to increase.
Summary of the Invention
[0004] This patent specification describes various embodiments and techniques for video encoding or decoding using motion vectors represented using a specified number of bits.
[0005] In one exemplary aspect, a video decoding method is disclosed. The method includes determining, during the conversion between a video region and its bit - stream representation, the range of Motion Vector Difference (MVD) values used for the video region of the video based on the maximum allowable motion vector resolution, the maximum allowable motion vector accuracy, or the priority of the video region, and performing the conversion by restricting the MVD values to be within the range.
[0006] In one exemplary aspect, a video decoding method is disclosed. The method includes determining a range of MVD (Motion Vector Difference) components associated with a first block of video for conversion between the first block of video and a bitstream representation of the first block, where the range of MVD components is [-2 M , 2 M - 1], and M = 17, and suppressing the values of the MVD components to be within the range of the MVD components, and performing the conversion based on the suppressed MVD components.
[0007] In one exemplary aspect, a video decoding method is disclosed. The method includes determining a range of MVD (Motion Vector Difference) components associated with a first block of video for conversion between the first block of video and a bitstream representation of the first block, where the range of MVD components is adapted to a codec's acceptable MVD accuracy and / or acceptable MV (Motion Vector) accuracy, and suppressing the values of the MVD components to be within the range of the MVD components, and performing the conversion based on the suppressed MVD components.
[0008] In one exemplary aspect, a video decoding method is disclosed. The method includes determining a range of MVD (Motion Vector Difference) components associated with a first block of video based on the coded information of the first block for conversion between the first block of video and a bitstream representation of the first block, and suppressing the values of the MVD components to be within the range of the MVD components, and performing the conversion based on the suppressed range of the MVD components.
[0009] In yet another exemplary aspect, a video processing apparatus is disclosed. The apparatus includes a processor configured to execute the method described above.
[0010] In yet another exemplary aspect, a computer-readable medium is disclosed. The medium stores code for implementing the methods described above on a processor.
[0011] These and other aspects are described in this patent specification.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Best Mode for Carrying Out the Invention
[0013] In this specification, section headings are used for ease of understanding, and the embodiments disclosed in one section are not limited to that section only. Further, specific embodiments are described with reference to VVC (Versatile Video Coding) or other specific video codecs, but the disclosed techniques are applicable to other video coding techniques as well. Further, although some embodiments describe video coding steps in detail, it will be understood that the corresponding steps of decoding, which reverse the coding, are performed by a decoder. Further, the term video processing includes video coding or compression, video decoding or decompression, and transcoding of video where the pixels of the video are represented in one compression format to another compression format or at another compression bitrate.
[0014] 1. Summary
[0015] This patent specification relates to video coding technology. Specifically, it relates to the inter-coding process in video coding. It may be applied to existing video coding standards such as HEVC, or may be applied to finalize the Versatile Video Coding standard. The present invention is also applicable to future video coding standards or video codecs.
[0016] 2. Initial negotiation
[0017] Video coding standards have mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and both organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 AVC (Advanced Video Coding), and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure that utilizes temporal prediction and transform coding. To explore future video coding technologies beyond HEVC, in 2015, VCEG and MPEG jointly established the JVET (Joint Video Exploration Team). Since then, many new methods have been adopted by the JVET and incorporated into the reference software called JEM (Joint Exploration Mode). The JVET meetings are held quarterly, and the new coding standard aims to reduce the bitrate by 50% compared to HEVC. At the JVET meeting in April 2018, the new video coding standard was officially named VVC (Versatile Video Coding), and at that time, the first version of the VTM (VVC Test Model) was released. Efforts to contribute to the standardization of VVC continue, and at all JVET meetings, new coding technologies are being adopted into the VVC standard. After each meeting, the VVC working draft and the test model VTM are updated. The VVC project is currently aiming for technical completion (FDIS) at the July 2020 meeting.
[0018] 2.1 Coding Flow of a Typical Video Codec
[0019] Figure 1 shows an example of an encoder block diagram of VVC that includes three in-loop filtering blocks, namely, DF (Deblocking Filter), SAO (Sample Adaptive Offset), and ALF. Different from DF that uses a predefined filter, SAO and ALF utilize the original samples of the current picture, add offsets respectively, and apply an FIR (Finite Inpulse Response) filter to the information on the coded side that signals the offset and filter coefficients, thereby reducing the mean squared error between the original samples and the reconstructed samples. ALF is located at the last processing stage of each picture and can be regarded as a tool to capture and fix the artifacts generated in the previous stage.
[0020] Figure 1 shows an example of an encoder block diagram.
[0021] 2.2 AMVR (Adaptive Motion Vector Resolution)
[0022] In HEVC, when use_integer_mv_flag is equal to 0 in the slice header, MVD (Motion Vector Difference) (the difference between the motion vector of the CU and the predicted motion vector) is signaled in units of 1 / 4 luma samples. In JEM, LAMVR (Locally Adaptive Motion Vector Resolution) is introduced. In VVC, a CU-level AMVR (Adaptive Motion Vector Resolution) scheme is introduced. AMVR enables coding the MVD of the CU with different precisions. Based on the mode for the current CU (normal AMVP mode or affine AVMP mode), the MVD of the current CU can be adaptively selected as follows. - Normal AMVP mode: 1 / 4 luma samples, integer luma samples, or 4 luma samples. - Affine AMVP mode: 1 / 4 luminance sample, integer luminance sample, or 1 / 16 luminance sample.
[0023] If the current CU has at least one non-zero MVD component, the CU-level MVD resolution indication is conditionally signaled. If all MVD components (i.e., both horizontal and vertical MVDs for reference list L0 and reference list L1) are zero, a 1 / 4 luminance sample MVD resolution is inferred.
[0024] For a CU having a component of at least one non-zero MVD component, a first flag is signaled to indicate whether 1 / 4 luminance sample MVD accuracy is used for the CU. If the first flag is 0, no further signaling is required and 1 / 4 luminance sample MVD accuracy is used for the current CU. Otherwise, a second flag is signaled to indicate whether integer luminance sample or 4 luminance sample MVD accuracy is used for a normal AMVP CU. The same second flag is used to indicate whether integer luminance sample or 1 / 16 luminance sample MVD accuracy is used for an affine AMVP CU. To ensure that the reconstructed MV has the intended accuracy (1 / 4 luminance sample, integer luminance sample, or 4 luminance sample), the motion vector predictor of the CU is rounded to the same accuracy as the MVD before being added to the MVD. Round the motion vector predictor towards zero (i.e., round a negative motion vector predictor towards positive infinity and a positive motion vector predictor towards negative infinity).
[0025] The encoder determines the resolution of the motion vector for the current CU using the RD check. To avoid always performing the RD check at the CU level three times for each MVD resolution, in VTM4, the RD check for MVD accuracy other than 1 / 4 luminance samples is called only conditionally. In the normal AVMP mode, first, the RD costs for 1 / 4 luminance sample MVD accuracy and integer luminance sample MV accuracy are calculated. Next, the RD cost of the integer luminance sample MVD accuracy is compared with the RD cost of the 1 / 4 luminance sample MVD accuracy, and it is determined whether it is necessary to further check the RD cost of the 4 luminance sample MVD accuracy. If the RD cost of the 1 / 4 luminance sample MVD accuracy is much smaller than the RD cost of the integer luminance sample MVD accuracy, the RD check for the 4 luminance sample MVD accuracy is omitted. In the affine AMVP mode, after checking the rate-distortion costs of the affine merge / skip mode, the merge / skip mode, the normal AMVP mode with 1 / 4 luminance sampling MVD accuracy, and the affine AMVP mode with 1 / 4 luminance sampling MVD accuracy, if the affine inter mode is not selected, the affine inter mode with 1 / 16 luminance sample MV accuracy and 1 pixel MV accuracy is not checked. Furthermore, as the search starting point in the affine inter mode with 1 / 16 luminance sample and 1 / 4 luminance sample MV accuracy, the affine parameters obtained in the affine inter mode with 1 / 4 luminance sample MV accuracy are used.
[0026] 2.3 Affine AMVP Prediction in VVC
[0027] The Affine AMVP mode can be applied to CUs where both the width and height are 16 or greater. To indicate whether the Affine AMVP mode is used, a CU-level Affine flag is signaled in the bitstream, and then another flag is signaled to indicate whether it is 4-parameter affine or 6-parameter affine. In this mode, the differences between the CPMVs of the current CU and their predictor CPMVs are signaled in the bitstream. The Affine AVMP candidate list size is 2 and is generated in order using the following 4 types of CPVM candidates. 1) Inherited Affine AMVP candidates extrapolated from the CPMVs of neighboring CUs 2) Constructed Affine AMVP candidate CPMVs derived using the translational MVs of neighboring CUs 3) Translational MVs from neighboring CUs 4) Zero MVs
[0028] The checking order of the inherited Affine AMVP candidates is the same as that of the inherited Affine merge candidates. The only difference is that in the case of AVMP candidates, only Affine CUs having the same reference picture as the current block are considered. When inserting an inherited Affine motion predictor into the candidate list, no pruning process is applied.
[0029] The constructed AMVP candidates are derived from the specified spatial neighborhood. Also, the reference picture indices of the neighboring blocks are checked. The first block in the checking order that is inter-coded and has the same reference picture as the current CU is used. There is only one. If the current CU is coded in 4-parameter affine mode and both mv0 and mv1 are available, they are added as one candidate to the Affine AMVP list. If the current CU is coded in 6-parameter affine mode and all 3 CPMVs are available, they are added as one candidate to the Affine AMVP list. Otherwise, the constructed AMVP candidates are set to unavailable.
[0030] After checking the inherited affine AMVP candidates and the constructed AMVP candidates, if the number of affine AMVP list candidates is still less than 2, if available, in the order of mv0, mv1, and mv2, they are added as translational MVs to predict all control point MVs of the current CU. Finally, if the affine AMVP list is not yet fully filled, zero MVs are used to fill the affine AMVP list.
[0031] 2.4 MMVD (Merge mode with MVD) in VVC
[0032] In addition to the merge mode where the implicitly derived motion information is directly used for the prediction sample generation of the current CU, MMVD (Merge mode with Motion Vector Deffirences) is introduced in VVC. Immediately after the skip flag and the merge flag are sent, the MMVD flag is signaled to specify whether the MMVD mode is used for the CU.
[0033] In MMVD, after a merge candidate is selected, it is further refined by the signaled MVD information. The additional information includes the merge candidate flag, an index for specifying the magnitude of the motion, and an index for indicating the direction of the motion. In the MMVD mode, one of the first two candidates in the merge list is selected to be used as the MV base. The merge candidate flag is signaled to specify which one to use.
[0034] The distance index specifies the magnitude information of the motion and indicates a predefined offset from the starting point. The offset is added to either the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is defined in Table 1.
[0035] In VVC, there is an SPS flag sps_fpel_mmvd_enabled_flag for enabling / disabling the fractional MMVD offset at the spss level, and a tile group flag tile_group_fpel_mmvd_enabled_flag for controlling the enabling / disabling of the fractional MMVD offset for "SCC / UHD frames" at the header level of the tile group. When the fractional MVD is enabled, the default distance table in Table 1 is used. Otherwise, all offset elements in the default distance in Table 1 are left-shifted by 2 only.
[0036]
Table 1
[0037] The direction index represents the direction of the MVD with respect to the starting point. The direction index can represent four directions as shown in Table 2. Note that the meaning of the MVD sign may vary according to the information of the starting MV. When the starting MV is an unpredicted MV or a bi-predicted MV and both lists point to the same side of the current picture (i.e., the POCs of both references are both larger than the POC of the current picture, or both are smaller than the POC of the current picture), the sign in Table 2 specifies the sign of the MV offset added to the starting MV. When the starting MV is a bi-predicted MV and the two MVs point to different sides of the current picture (i.e., the POC of one reference is larger than the POC of the current picture and the POC of the other reference is smaller than the POC of the current picture), the sign in Table 2 defines the sign of the MV offset added to the list0 MV component of the starting MV, and the sign of the list1 MV has the opposite value.
[0038]
Table 2
[0039] 2.5 Intra Block Copy (IBC) in VVC
[0040] IBC (Intra Block Copy) is a tool adopted in the HEVC extension of SCC. It is known that this significantly improves the coding efficiency of screen content materials. Since the IBC mode is implemented as a block-level coding mode, BM (Block Matching) is executed in the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector is used to indicate the replacement from the current block to a reference block that has already been reconstructed within the current picture.
[0041] In VVC, the luminance block vector of an IBC-coded CU is of integer precision. The chrominance block vector is rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pixel and 4-pixel motion vector precisions. An IBC-coded CU is treated as a third prediction mode other than the intra prediction mode or the inter prediction mode. The IBC mode is applicable to CUs with both width and height of 64 luminance samples or less.
[0042] The IBC mode is also known as the CPR (Current Picture Reference) mode.
[0043] 2.6 Difference of Motion Vectors in the VVC Specification / Working Draft
[0044] The following text is extracted from the VVC working draft.
[0045] 7.3.6.8 Motion Vector Difference Syntax
[0046]
Table 3
[0047] 7.3.6.7 Merge Data Syntax
[0048]
Table 4
[0049]
Table 5
[0050] 7.4.3.1 Sequence Parameter Set RBSP Semantics When sps_amvr_enabled_flag is equal to 1, it stipulates that adaptive motion vector difference resolution is used for motion vector coding. When amvr_enabled_flag is equal to 0, it stipulates that adaptive motion vector difference resolution is not used for motion vector coding. When sps_affine_amvr_enabled_flag is equal to 1, it stipulates that adaptive motion vector difference resolution is used for motion vector coding in the affine inter mode. When sps_affine_amvr_enabled_flagg is equal to 0, it stipulates that adaptive motion vector difference resolution is not used for motion vector coding in the affine inter mode. When sps_fpel_mmvd_enabled_flag is equal to 1, it stipulates that the merge mode using motion vector difference uses integer sample precision. When sps_fpel_mmvd_enabled_flag is equal to 0, it stipulates that the merge mode using motion vector difference can use fractional sample precision.
[0051] 7.4.5.1 General Tile Group Header Semantics When tile_group_fpel_mmvd_enabled_flag is equal to 1, it stipulates that the merge mode using motion vector difference uses integer sample precision in the current tile group. If tile_group_fpel_mmvd_enabled_flag is equal to 0, it stipulates that the merge mode using the motion vector difference can use the fractional sample precision in the current tile group. If it does not exist, the value of tile_group_fpel_mmvd_enabled_flag is presumed to be 0.
[0052] 7.4.7.5 Coding Unit Semantics amvr_flag[x0][y0] stipulates the resolution of the motion vector difference. The array indices x0, y0 stipulate the position (x0, y0) of the top-left luminance sample of the coding block under consideration, which is related to the top-left luminance sample of the picture. If amvr_flag[x0][y0] is equal to 0, it stipulates that the resolution of the motion vector difference is 1 / 4 of the luminance sample. If amvr_flag[x0][y0] is equal to 1, it stipulates that the resolution of the motion vector difference is further stipulated by amvr_precision_flag[x0][y0]. If amvr_flag[x0][y0] does not exist, it is presumed as follows. - If CuPredMode[x0][y0] is equal to MODE_IBC, amvr_flag[x0][y0] is presumed to be equal to 1. - Otherwise (if CuPredMode[x0][y0] is not equal to MODE_IBC), amvr_flag[x0][y0] is presumed to be 0. When amvr_precision_flag[x0][y0] is equal to 0, the resolution of the motion vector difference is defined as one integer luminance sample when inter_affine_flag[x0][y0] is equal to 0, and 1 / 16 of the luminance sample otherwise. When amvr_precision_flag[x0][y0] is equal to 1, the resolution of the motion vector difference is defined as four luminance samples when inter_affine_flag[x0][y0] is equal to 0, and one integer luminance sample otherwise. The array indices x0, y0 define the position (x0, y0) of the top-left luminance sample of the coding block being considered, which is related to the top-left luminance sample of the picture.
[0053] If amvr_precision_flag[x0][y0] does not exist, it is assumed to be equal to 0. The motion vector difference is modified as follows. - When inter_affine_flag[x0][y0] is equal to 0, the variable MvShift is derived and the variables MvdL0[x0][y0][0], MvdL0[x0][y0][1], MvdL1[x0][y0][0], MvdL1[x0][y0][1] are modified as follows. MvShift=(amvr_flag[x0][y0]+amvr_precision_flag[x0][y0])<<1 (7-98) MvdL0[x0][y0][0]=MvdL0[x0][y0][0]<<(MvShift+2) (7-99) MvdL0[x0][y0][1]=MvdL0[x0][y0][1]<<(MvShift+2) (7-100) MvdL1[x0][y0][0]=MvdL1[x0][y0][0]<<(MvShift+2) (7-101) MvdL1[x0][y0][1]=MvdL1[x0][y0][1]<<(MvShift+2) (7-102) - Otherwise (when inter_afine_flag[x0][y0] is equal to 1), the variable MvShift is derived, and the variables MvdCpL0[x0][y0][0][0], MvdCpL0[x0][y0][0][1], MvdCpL0[x0][y0][1][0], MvdCpL0[x0][y0][1][1], MvdCpL0[x0][y0][2][0], and MvdCpL0[x0][y0][2][1] are modified as follows. MvShift = amvr_precision_flag[x0][y0]? (amvr_precision_flag[x0][y0] << 1) : (-(amvr_flag[x0][y0] << 1))) (7 - 103) MvdCpL0[x0][y0][0][0] = MvdCpL0[x0][y0][0][0] << (MvShift + 2) (7 - 104) MvdCpL1[x0][y0][0][1] = MvdCpL1[x0][y0][0][1] << (MvShift + 2) (7 - 105) MvdCpL0[x0][y0][1][0] = MvdCpL0[x0][y0][1][0] << (MvShift + 2) (7 - 106) MvdCpL1[x0][y0][1][1] = MvdCpL1[x0][y0][1][1] << (MvShift + 2) (7 - 107) MvdCpL0[x0][y0][2][0] = MvdCpL0[x0][y0][2][0] << (MvShift + 2) (7 - 108) MvdCpL1[x0][y0][2][1] = MvdCpL1[x0][y0][2][1] << (MvShift + 2) (7 - 109)
[0054] 7.4.7.7 Merge Data Semantics merge_flag[x0][y0] specifies whether the inter-prediction parameter for the current coding unit is inferred from the neighboring inter-prediction intervals. The array indices x0, y0 specify the position (x0, y0) of the top-left luminance sample of the coding block under consideration, which is related to the top-left luminance sample of the picture. If merge_flag[x0][y0] does not exist, it is inferred as follows. - If cu_skip_flag[x0][y0] is equal to 1, merge_flag[x0][y0] is inferred to be equal to 1. - Otherwise, merge_flag[x0][y0] is inferred to be equal to 0. If mmvd_flag[x0][y0] is equal to 1, it is specified to use the merge mode that uses the motion vector difference to generate the inter-prediction parameter of the current coding unit. The array indices x0, y0 specify the position (x0, y0) of the top-left luminance sample of the coding block under consideration, which is related to the top-left luminance sample of the picture. If mmvd_flag[x0][y0] does not exist, it is inferred to be equal to 0. mmvd_merge_flag[x0][y0] specifies which of the first (0) candidate or the second (1) candidate in the merge candidate list is used with the motion vector difference derived from mmvd_distance_idx[x0][y0] and mmvd_direction_idx[x0][y0]. The array indices x0, y0 specify the position (x0, y0) of the top-left luminance sample of the coding block under consideration, which is related to the top-left luminance sample of the picture. mmvd_distance_idx[x0][y0] specifies the index used to derive MmvdDistance[x0][y0] as defined in Table 7-11. The array indices x0, y0 specify the position (x0, y0) of the top-left luminance sample of the coding block under consideration, which is related to the top-left luminance sample of the picture.
[0055]
Table 6
[0056] mmvd_distance_idx[x0][y0] defines the index used to derive MmvdDistance[x0][y0] as specified in Table 7 - 12. The array indices x0, y0 define the position (x0, y0) of the top - left luminance sample of the coding block under consideration, which is related to the top - left luminance sample of the picture.
[0057]
Table 7
[0058] Both components of the merge + MVD offset MmvdOffset[x0][y0] are derived as follows. MmvdOffset[x0][y0][0]=(MmvdDistance[x0][y0]<<2)*MmvdSign[x0][y0][0] (7 - 112) MmvdOffset[x0][y0][1]=(MmvdDistance[x0][y0]<<2)*MmvdSign[x0][y0][1] (7 - 113) merge_subblock_flag[x0][y0] defines whether the sub - block - based inter - prediction parameter for the current coding unit is inferred from neighboring blocks. The array indices x0, y0 define the position (x0, y0) of the top - left luminance sample of the coding block under consideration, which is related to the top - left luminance sample of the picture. If merge_subblock_flag[x0][y0] does not exist, it is assumed to be equal to 0. merge_subblock_idx[x0][y0] defines the merge candidate index in the sub-block based merge candidate list, where x0, y0 define the position (x0, y0) of the top-left luminance sample of the coding block under consideration, which is related to the top-left luminance sample of the picture. If merge_subblock_idx[x0][y0] does not exist, it is assumed to be equal to 0.
[0059] ciip_flag[x0][y0] defines whether combined inter-picture merge and intra-picture prediction are applied to the current coding unit. The array indices x0, y0 define the position (x0, y0) of the top-left luminance sample of the coding block under consideration, which is related to the top-left luminance sample of the picture. If ciip_flag[x0][y0] does not exist, it is assumed to be equal to 0. The syntax elements ciip_luma_mpm_flag[x0][y0], and ciip_luma_mpm_idx[x0][y0] define the intra prediction mode of the luminance samples used for combined inter-picture merge and intra-picture prediction. The array indices x0, y0 define the position (x0, y0) of the top-left luminance sample of the coding block under consideration, which is related to the top-left luminance sample of the picture. The intra prediction mode is derived according to Section 8.5.6. If ciip_luma_mpm_flag[x0][y0] does not exist, it is assumed as follows. - If cbWidth is greater than 2*cbHeight, or cbHeight is greater than 2*cbWidth, ciip_luma_mpm_flag[x0][y0] is assumed to be equal to 1. - Otherwise, ciip_luma_mpm_flag[x0][y0] is assumed to be equal to 0.
[0060] When merge_triangle_flag[x0][y0] is equal to 1, it is stipulated that for the current coding unit, when decoding the B tile group, motion compensation based on the triangle shape is used to generate the predicted samples of the current coding unit. When merge_triangle_flag[x0][y0] is equal to 0, it is stipulated that the coding unit is not predicted by motion compensation based on the triangle shape. If merge_triangle_flag[x0][y0] does not exist, it is presumed to be equal to 0.
[0061] merge_triangle_split_dir[x0][y0] defines the split direction of the merge triangle mode. The array indices x0, y0 define the position (x0, y0) of the top-left luminance sample of the coding block under consideration, which is related to the top-left luminance sample of the picture. If merge_triangle_split_dir[x0][y0] does not exist, it is presumed to be equal to 0. merge_triangle_idx0[x0][y0] defines the first merge candidate index of the motion compensation candidate list based on the triangle shape, where x0, y0 define the position (x0, y0) of the top-left luminance sample of the coding block under consideration, which is related to the top-left luminance sample of the picture. If merge_triangle_idx0[x0][y0] does not exist, it is presumed to be equal to 0. merge_triangle_idx1[x0][y0] defines the second merge candidate index of the motion compensation candidate list based on the triangle shape, where x0, y0 define the position (x0, y0) of the top-left luminance sample of the coding block under consideration, which is related to the top-left luminance sample of the picture. If merge_triangle_idx1[x0][y0] does not exist, it is presumed to be equal to 0. merge_idx[x0][y0] defines the merge candidate index in the merge candidate list, where x0 and y0 define the position (x0, y0) of the top-left luminance sample of the coding block under consideration, which is related to the top-left luminance sample of the picture. If merge_idx[x0][y0] does not exist, it is inferred as follows. - If mmvd_flag[x0][y0] is equal to 1, merge_idx[x0][y0] is inferred to be equal to mmvd_merge_flag[x0][y0]. Otherwise (if mmvd_flag[x0][y0] is equal to 0), merge_idx[x0][y0] is inferred to be equal to 0.
[0062] 7.4.7.8 Motion Vector Difference Semantics abs_mvd_greater0_flag[compIdx] defines whether the absolute value of the difference of the motion vector components is greater than 0. abs_mvd_greater1_flag[compIdx] defines whether the absolute value of the difference of the motion vector components is greater than 1. If abs_mvd_greater1_flag[compIdx] does not exist, it is inferred to be equal to 0. abs_mvd_minus2[compIdx] + 2 defines the absolute value of the difference of the motion vector components. If abs_mvd_minus2[compIdx] does not exist, it is inferred to be equal to -1. mvd_sign_flag[compIdx] defines the sign of the difference of the motion vector components as follows. - If mvd_sign_flag[compIdx] is equal to 0, the difference of the corresponding motion vector component has a positive value. - Otherwise (if mvd_sign_flag[compIdx] is equal to 1), the difference of the corresponding motion vector component has a negative value. If mvd_sign_flag[compIdx] does not exist, it is inferred to be equal to 0. For compIdx = 0..1, the motion vector difference lMvd[compIdx] is derived as follows.
[0063]
Number
[0064] Based on the value of MotionModelIdc[x][y], the difference of the motion vectors is derived as follows. - When MotionModelIdc[x][y] is equal to 0, the variable MvdLX[x0][y0][compIdx] (X is 0 or 1) defines the difference between the list X vector component to be used and its prediction. The array indices x0, y0 define the position (x0, y0) of the top-left luminance sample of the coding block under consideration, which is related to the top-left luminance sample of the picture. The horizontal motion vector component difference is assigned compIdx = 0, and the vertical motion vector component is assigned compIdx = 1. - When refList is equal to 0, MvdL0[x0][y0][compIdx] is set equal to lMvd[compIdx] for compIdx = 0..1. - Otherwise (when refList is equal to 1), MvdL1[x0][y0][compIdx] is set equal to lMvd[compIdx] for compIdx = 0..1. - Otherwise (when MotionModelIdc[x][y] is not equal to 0), the variable MvdCpLX[x0][y0][cpIdx][compIdx] (X is 0 or 1) defines the difference between the list X vector component to be used and its prediction. The array indices x0, y0 define the position (x0, y0) of the top-left luminance sample of the coding block under consideration, which is related to the top-left luminance sample of the picture, and the array index cpIdx defines the control point index. The horizontal motion vector component difference is assigned compIdx = 0, and the vertical motion vector component is assigned compIdx = 1. - When -refList is equal to 0, for compIdx = 0..1, MvdCpL0[x0][y0][cpIdx][compIdx] is set equal to lMvd[compIdx]. - Otherwise (when -refList is equal to 1), for compIdx = 0..1, MvdCpL1[x0][y0][cpIdx][compIdx] is set equal to lMvd[compIdx].
[0065] 3. Examples of problems to be solved by the embodiments described in this patent specification In some coding standards such as VVC, the MVD (Motion Vector Difference) is not necessarily at a resolution of 1 / 4 pixel (e.g., 1 / 4 luminance sample). However, in existing VVC working drafts, there is a bitstream constraint that always clips the MVD components to a range of -2 15 ~2 15 -1. As a result, especially when an MVD resolution that is not 1 / 4 pixel is used (e.g., the MVD resolution of 1 / 16 luminance samples when affine AMVP is used), the MVD value may become inaccurate.
[0066] 4. Exemplary embodiments and techniques The embodiments listed below should be considered examples for explaining general concepts. These inventions should not be interpreted in a narrow sense. Furthermore, these inventions can be combined in any way. In the following description, the "MVD (Motion Vector Difference) component" refers to either the motion vector difference in the horizontal direction (e.g., along the x-axis) or the motion vector difference in the vertical direction (e.g., along the y-axis). In the case of sub-pixel MV (Motion Vector) representation, the motion vector usually consists of a fractional part and an integer part. The range of the MV is [-2 M ,2 MLet it be [-1], where M is a positive integer value, M = K + L, K represents the range of the integer part of MV, L represents the range of the fractional part of MV, and MV is expressed with a luminance sample accuracy of (1 / 2 L ). For example, in HEVC, K = 13, L = 2, and thus M = K + L = 15. On the other hand, in VVC, K = 13, L = 4, and M = K + L = 17.
[0067] 1. It is proposed that the range of the MVD component may depend on the acceptable MVD resolution / accuracy of the codec. a) In one example, the same range may be applied to all MVD components. i. In one example, the range of the MVD component is the same as the MV range such as [-2 M , 2 M -1], M = 17. b) In one example, all decoded MVD components may first be scaled to a predefined accuracy (1 / 2 L ) luminance samples (e.g., L = 4), and then clipped to a predefined range [-2 M , 2 M -1] (e.g., M = 17). c) In one example, the range of the MVD component may depend on the acceptable MVD / MV resolution in the codec. i. In one example, assuming that the acceptable resolution of MVD is 1 / 16 luminance sample, 1 / 4 luminance sample, 1 luminance sample, or 4 luminance samples, the value of the MVD component may be clipped / suppressed according to the highest resolution (e.g., 1 / 16 luminance sample among these possible resolutions). That is, the value of MVD is in the range of [-2 K+L , 2 K+L -1] with, for example, K = 13, L = 4.
[0068] 2. It is proposed that the range of the MVD component may depend on the coded information of the block. a) In one example, multiple sets of the range of the MVD component may be defined. b) In one example, the range may depend on the MV predictor / MVD / MV accuracy. i. In one example, the MVD accuracy of the MVD component is (1 / 2 L ) luminance samples (e.g., L = 4, 3, 2, 1, 0, -1, -2, -3, -4, etc.), and the value of the MVD may be, for example, [-2 K+L , 2 K+L -1] range may be suppressed and / or clipped. ii. In one example, the range of the MVD component may depend on the variable MvShift, where MvShift may be derived from the affine_inter_flag, amvr_flag, and amvr_precision_flag of VVC. 1. In one example, MvShift may be derived by coded information such as affine_inter_flag, amvr_flag, and / or amvr_precision_flag, and / or sps_fpel_mmvd_enabled_flag, and / or tile_group_fpel_mmvd_enabled_flag, and / or mmvd_distance_idx, and / or CuPredMode. c) In one example, the MVD range may depend on the coding mode, motion model, etc. of the block. i. In one example, the range of the MVD component may depend on the motion model (e.g., MotionModelIdc of the specification), and / or the prediction mode, and / or the affine inter flag of the current block. ii. In one example, when the prediction mode of the current block is MODE_IBC (e.g., when the current block is coded in IBC mode), the value of the MVD may be, for example, [-2 K+L , 2 K+L -1] range. iii. In one example, when the motion model index of the current block (e.g., MotionModelIdc in the specification) is equal to 0 (e.g., when the current block is predicted using a translational motion model), the value of MVD may be in the range of [-2 K+L , 2 K+L -1], for example, with K = 13 and L = 2. 1. Alternatively, when the prediction mode of the current block is MODE_INTER and affine_inter_flag is false (e.g., when the current block is predicted using a translational motion model), the value of MVD may be in the range of [-2 K+L , 2 K+L -1], for example, with K = 13 and L = 2. iv. In one example, when the motion model index of the current block (e.g., MotionModelIdc in the specification) is not equal to 0 (e.g., when the current block is predicted using an affine motion model), the value of MVD may be in the range of [-2 K+L , 2 K+L -1], for example, with K = 13 and L = 4. 1. Alternatively, when the prediction mode of the current block is MODE_INTER and affine_inter_flag is true (e.g., when the current block is predicted using an affine motion model), the value of MVD may be in the range of [-2 K+L , 2 K+L -1], for example, with K = 13 and L = 4. d) Instead of imposing constraints on the decoded MVD component, it is proposed to impose constraints on the rounded MVD value. i. In one example, the compliant bitstream is assumed to satisfy that the rounded integer MVD value is within a given range. 1. In one example, the integer MVD (rounding is required if the decoded MVD is in fractional precision) should be in the range of [-2 K , 2 K -1], for example, with K = 13.
[0069] 3. It is proposed that the values of the decoded MVD components may be explicitly clipped to a range (e.g., the MVD range described above) during semantic interpretation, in addition to using bitstream constraints.
[0070] 5. Embodiments
[0071] 5.1 Embodiment #1 The following embodiments relate to the method of Section 4 Item 1 are. Newly added parts are in italic bold, and parts deleted from the VVC working draft are emphasized with a strikethrough.
[0072] 7.4.7.8 Motion Vector Difference Semantics When compIdx = 0..1, the motion vector difference lMvd[compIdx] is derived as follows.
[0073]
Number
[0074] 5.2 Embodiment #2 The following embodiments relate to the method of Section 4 Item 2 are. Newly added parts are in italic bold, and parts deleted from the VVC working draft are emphasized with a strikethrough.
[0075] 7.4.7.9 Motion Vector Difference Semantics When compIdx = 0..1, the motion vector difference lMvd[compIdx] is derived as follows.
[0076]
Number
[0077] 5.3 Embodiment #3 The following embodiments relate to the method of Section 4 Item 2 are. The newly added parts are in bold italic, and the parts deleted from the VVC working draft are emphasized with green strikethrough.
[0078] 7.4.7.10 Motion Vector Difference Semantics When compIdx = 0..1, the motion vector difference lMvd[compIdx] is derived as follows.
[0079]
Number
[0080] 5.4 Embodiment #4 The following embodiments relate to the method of Section 4 Item 2 and are as follows. The newly added parts are in bold italic, and the parts deleted from the VVC working draft are emphasized with green strikethrough.
[0081] 7.4.7.11 Motion Vector Difference Semantics When compIdx = 0..1, the motion vector difference lMvd[compIdx] is derived as follows.
[0082]
Number
[0083] 5.5 Embodiment #5 The following embodiments relate to the method of Section 4 Item 3 and Item 1 and are as follows. The newly added parts are in bold italic, and the parts deleted from the VVC working draft are emphasized with green strikethrough.
[0084] 7.4.7.12 Motion Vector Difference Semantics When compIdx = 0..1, the motion vector difference lMvd[compIdx] is derived as follows.
[0085]
Number
[0086] 5.6 Embodiment #6 The following embodiments relate to the method of Section 4 Item 3 and Item 2 Newly added parts are in italic bold, and parts deleted from the VVC working draft are emphasized with green strikethrough.
[0087] 7.4.7.13 Motion Vector Difference Semantics When compIdx = 0..1, the motion vector difference lMvd[compIdx] is derived as follows.
[0088]
Number
[0089] Based on the value of MotionModelIdc[x][y], the motion vector difference is derived as follows. - When MotionModelIdc[x][y] is equal to 0, the variable MvdLX[x0][y0][compIdx] (X is 0 or 1) defines the difference between the list X vector component to be used and its prediction. The array indices x0, y0 define the position (x0, y0) of the top-left luminance sample of the coding block under consideration, which is related to the top-left luminance sample of the picture. compIdx = 0 is assigned to the horizontal motion vector component difference, and compIdx = 1 is assigned to the vertical motion vector component.
[0090]
Number
[0091] - Otherwise (when MotionModelIdc[x][y] is not equal to 0), the variable MvdCpLX[x0][y0][cpIdx][compIdx] (where X is 0 or 1) defines the difference between the list X vector component to be used and its prediction. The array indices x0, y0 define the position (x0, y0) of the top-left luminance sample of the coding block under consideration, which is related to the top-left luminance sample of the picture. The array index cpIdx defines the control point index. compIdx = 0 is assigned to the horizontal motion vector component difference, and compIdx = 1 is assigned to the vertical motion vector component.
[0092] [Number]
[0093] FIG. 2 is a block diagram of the video processing apparatus 1000. The apparatus 1000 may be used to implement one or more of the methods described herein. The apparatus 1000 may be implemented by a smartphone, a tablet, a computer, an IoT (Internet of Things) receiver, etc. The apparatus 1000 may include one or more processors 1002, one or more memories 1004, and video processing hardware 1006. The one or more processors 1002 may be configured to implement one or more of the methods described in this patent specification. The one or more memories 1004 may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 1006 may be used to implement the techniques described in this patent specification in a hardware circuit.
[0094] Figure 3 is a flowchart showing an example of a video processing method. Method 300 includes determining (302) a range of values of MVD (Motion Vector Difference) used for a video region of a video during conversion between the video region and a bitstream representation of the video region based on a maximum allowable motion vector resolution, a maximum allowable motion vector accuracy, or an attribute of the video region. Method 300 includes performing (304) the conversion by restricting the MVD value to fall within the range.
[0095] The following list of solutions provides embodiments that can address the technical problems described in this patent specification, among other problems.
[0096] 1. A video processing method, comprising: determining a range of MVD (Motion Vector Difference) values used for a video region of a video during conversion between the video region and a bitstream representation of the video region based on a maximum allowable motion vector resolution, a maximum allowable motion vector accuracy, or a characteristic of the video region; and performing the conversion by restricting the MVD value to be within this range.
[0097] 2. The method according to solution 1, wherein the range is applied during conversion of all video regions of the video.
[0098] 3. The method according to any one of solutions 1 to 2, wherein the range is equal to the range of motion vectors for the video region.
[0099] 4. The method according to any one of solutions 1 to 3, wherein restricting includes scaling the MVD component to the accuracy and clipping the out-of-scale to the range.
[0100] 5. The method according to any one of solutions 1 to 4, wherein the characteristic of the video region includes coded information of the video region.
[0101] 6. The method according to any one of Solutions 1 to 5, wherein the range is selected from a set of a plurality of possible ranges of the video.
[0102] 7. The method according to any one of Solutions 1 to 4, wherein the characteristic of the video region includes the accuracy of the motion vector predictor used for the video region.
[0103] 8. The method according to any one of Solutions 1 to 4, wherein the characteristic of the video region corresponds to the value of MVShift, and MVShift is a variable associated with the video region, and MVShift depends on the affine_inter_flag, or amvr_flag or amvr_precision_flag associated with the video region.
[0104] 9. The method according to Solution 1, wherein the characteristic of the video region corresponds to the coding mode used for the transform.
[0105] 10. The coding mode is an intra block copy mode, and the range corresponds to [-2 K+L , 2 K+L -1], where K and L are integers representing the range of the integer part of the MV (Motion Vector) and the range of the fractional part of the MV, respectively, according to the method described in Solution 9.
[0106] 11. According to the method described in Solution 10, K = 13 and L = 0.
[0107] 12. The method according to Solution 1, wherein the characteristic of the video region corresponds to the motion model used for the transform.
[0108] 13. The characteristic of the video region is that the motion of the video region is modeled using a translational model, and as a result, the range is determined to be [-2 K+L , 2 K+L -1], where K and L are integers representing the range of the integer part of the MV (Motion Vector) and the range of the fractional part of the MV, respectively, according to the method described in Solution 1.
[0109] 14. The method according to Solution 13, wherein K = 13 and L = 2.
[0110] 15. The characteristics of the video region are such that the motion of the video region is modeled using a non-translational model, and as a result, the range is determined to be [-2 K+L , 2 K+L -1], where K and L are integers representing the range of the integer part of the MV (Motion Vector) and the range of the fractional part of the MV, respectively. The method according to Solution 1.
[0111] 16. The method according to Solution 15, wherein K = 13 and L = 4.
[0112] 17. The method according to Solution 1, wherein restricting includes restricting the rounded value of the MVD to the range.
[0113] 18. A video processing method, comprising: determining a range of MVD (Motion Vector Difference) values used for a video region of a video during conversion between the video region and a bitstream representation of the video region; and performing a clipping operation on the MVD values to keep them within the range during semantic interpretation performed during the conversion.
[0114] 19. The method according to any one of Solutions 1 to 18, wherein the video region corresponds to a video block.
[0115] 20. The method according to any one of Solutions 1 to 19, wherein the conversion includes generating pixel values of the video region from the bitstream representation.
[0116] 21. The method according to any one of Solutions 1 to 20, wherein the conversion includes generating a bitstream representation from the pixel values of the video region.
[0117] 22. A video processing apparatus comprising a processor configured to implement one or more of Examples 1 to 21.
[0118] 23. A computer-readable medium storing code that, when executed by a processor, causes the processor to implement the method according to one or more of Examples 1 to 21.
[0119] The items listed in Section 4 provide further variations of the above solutions.
[0120] FIG. 4 is a flowchart illustrating an example of a video processing method 400. The method 400 includes determining (402) a range of MVD (Motion Vector Difference) components associated with a first block for conversion between the first block of video and a bitstream representation of the first block, where the range of the MVD components is [-2 M , 2 M -1] and M = 17, suppressing (404) the values of the MVD components to be within the range of the MVD components, and performing (406) the conversion based on the suppressed range of the MVD components.
[0121] In some examples, the range is adapted to the codec's acceptable MVD accuracy and / or acceptable MV (Motion Vector) accuracy.
[0122] In some examples, the acceptable MVD accuracy and / or acceptable MV (Motion Vector) accuracy is 1 / 16 luminance sample accuracy.
[0123] In some examples, when there are multiple acceptable MVD accuracies and / or MV accuracies in the codec, the range of the MVD components is adapted to the top accuracy of the multiple acceptable MVD accuracies and / or MV accuracies.
[0124] In some examples, when the multiple acceptable MVD accuracies and / or MV accuracies include 1 / 16 luminance sample, 1 / 4 luminance sample, 1 luminance sample, and 4 luminance samples, the range of the MVD components is adapted to the accuracy of 1 / 16 luminance sample.
[0125] In some examples, the range of the MVD component is [-2 M , 2 M - 1], M = K + L, where K indicates the number of bits used to represent the integer part of the MVD component, L indicates the number of bits used to represent the fractional part of the MVD component, the MVD component is represented with 1 / 2 L luminance sample accuracy, and / or the range of the MV component associated with the first block is [-2 M , 2 M - 1], M = K + L, where K indicates the number of bits used to represent the integer part of the MV component, L indicates the number of bits used to represent the fractional part of the MV component, the MV component is represented with 1 / 2 L - luminance sample accuracy, and M, K, L are positive integers.
[0126] In some examples, K = 13, L = 4, and M = 17.
[0127] In some examples, the MVD component is the decoded / signalized MVD component coded in a bitstream, or the transformed MVD component associated with a specific accuracy via an internal shift operation of the decoding process.
[0128] In some examples, the MVD component includes a horizontal MVD component and a vertical MVD component, and the horizontal MVD component and the vertical MVD component have the same range.
[0129] In some examples, the MVD component is represented by integer bits, fractional bits, and sign bits.
[0130] In some examples, the range of the MV associated with the first block is the same as the range of the MVD component.
[0131] FIG. 5 is a flowchart illustrating an example of a video processing method 500. The method 500 includes determining a range of MVD (Motion Vector Difference) components associated with a first block for conversion between the first block of video and a bitstream representation of the first block (502), where the range of MVD components is adapted to the codec's acceptable MVD accuracy and / or acceptable MV (Motion Vector) accuracy, suppressing the values of the MVD components within the range of the MVD components (504), and performing a conversion based on the suppressed range of the MVD components (506).
[0132] In some examples, the MVD component is a decoded / signaled MVD component coded in a bitstream or a transformed MVD component associated with a particular accuracy via an internal shift operation of a decoding process.
[0133] In some examples, the decoded / signaled MVD component is required to be in the range of [-2 M , 2 M -1] and M = 17.
[0134] In some examples, the MVD component is represented by integer bits, fractional bits, and a sign bit.
[0135] In some examples, the range of the MVD component is determined to be [-2 M , 2 M -1], M = K + L, where K indicates the number of bits used to represent the integer part of the MVD component and L indicates the number of bits used to represent the fractional part of the MVD component, and the MVD component is represented with 1 / 2 L luminance sample accuracy, and / or the range of the MV component associated with the first block is determined to be [-2 M , 2 M -1], M = K + L, where K indicates the number of bits used to represent the integer part of the MV component and L indicates the number of bits used to represent the fractional part of the MV component, and the MV component is 1 / 2 LIt is represented by the luminance sample precision, and M, K, and L are positive integers.
[0136] In some examples, the values of all decoded MVD components are first scaled to 1 / 2 L the luminance sample precision, and then the range of the MVD components [-2 M , 2 M -1] is clipped.
[0137] In some examples, if there are multiple acceptable MVD precisions and / or MV precisions in the codec, the range of the MVD components is adapted to the highest precision among the multiple acceptable MVD precisions and / or MV precisions.
[0138] In some examples, if the multiple acceptable MVD precisions and / or MV precisions include 1 / 16 luminance sample precision, 1 / 4 luminance sample precision, 1 luminance sample precision, and 4 luminance sample precision, the range of the MVD components is adapted to 1 / 16 luminance sample precision, and the values of the MVD components are clamped to the range and / or clipped.
[0139] In some examples, K = 13, L = 4, and M = 17.
[0140] In some examples, the MVD components include a horizontal MVD component and a vertical MVD component, and the horizontal MVD component and the vertical MVD component have the same range.
[0141] In some examples, the range of the MV is the same as the range of the MVD components.
[0142] FIG. 6 is a flowchart illustrating an example of a video processing method 600. The method 600 includes determining (602) a range of MVD (Motion Vector Difference) components associated with a first block based on coded information of the first block for conversion between the first block of video and a bitstream representation of the first block, suppressing (604) the values of the MVD components to be within the range of the MVD components, and performing (606) the conversion based on the suppressed MVD components.
[0143] In some examples, the range of MVD components includes a plurality of sets of ranges of MVD components.
[0144] In some examples, the coded information includes at least one of the accuracy of an MV (Motion Vector) predictor, the accuracy of an MVD component, and the accuracy of an MV.
[0145] In some examples, when the MVD accuracy of the MVD component is 1 / 2 L for a luminance sample, the range of the MVD component is determined to be in the range of [-2 K+L , 2 K+L -1], the value of the MVP component is suppressed and / or clipped to be within the range, K indicates the number of bits used to represent the integer part of the MVD component, L indicates the number of bits for representing the fractional part of the MVD component, and K and L are positive integers.
[0146] In some examples, K is 13 and L is one of 4, 3, 2, 1, 0, -1, -2, -3, and -4.
[0147] In some examples, the coded information includes the variable MvShift associated with the MVD, and the derivation of the variable MvShift depends on whether AFFINE is used and / or whether AMVR (Adaptive Motion Vector Resolution) is used and / or the accuracy of AMVR and / or the accuracy of the MVD and / or MMVD (MMerge mode with Motion Vector Difference) information and / or the prediction mode of the first block.
[0148] In some examples, the variable MvShift is derived from one or more syntax elements including inter_affine_flag, amvr_flag, and amvr_precision_idx in the coded information.
[0149] In some examples, the variable MvShift is derived from one or more syntax elements including inter_affine_flag, amvr_flag, amvr_precision_idx, sps_fpel_mmvd_enabled_flag, ph_fpel_mmvd_enabled_flag, mmvd_distance_idx, and CuPredMode in the coded information.
[0150] In some examples, the coded information includes the coding mode, motion mode, and prediction mode of the first block, and one or more variables and / or syntax elements indicating whether AFFIN / AMVR is used in the coding information.
[0151] In some examples, when the prediction mode of the first block is MODE_IBC indicating that the first block is coded in the IBC mode, the range of the MVD components is [-2 K+L , 2 K+LIt is determined that the value is in the range of [-1], and the value of the MVP component is suppressed and / or clipped so as to be within the range. K indicates the number of bits used to represent the integer part of the MVD component, L indicates the number of bits used to represent the fractional part of the MVD component, and K and L are positive integers.
[0152] In some examples, K = 13 and L = 0.
[0153] In some examples, when the index of the motion model of the first block is equal to 0, the range of the MVD component is [-2 K+L , 2 K+L It is determined that the value is in the range of [-1], and the value of the MVP component is suppressed and / or clipped so as to be within the range. K indicates the number of bits used to represent the integer part of the MVD component, L indicates the number of bits used to represent the fractional part of the MVD component, and K and L are positive integers.
[0154] In some examples, K = 13 and L = 2.
[0155] In some examples, when the prediction mode of the first block is such that the first block is MODE_INTER and the variable affine_inter_flag is false, the range of the MVD component is [-2 K+L , 2 K+L It is determined that the value is in the range of [-1], and the value of the MVP component is suppressed and / or clipped so as to be within the range. K indicates the number of bits used to represent the integer part of the MVD component, L indicates the number of bits used to represent the fractional part of the MVD component, and K and L are positive integers.
[0156] In some examples, K = 13 and L = 2.
[0157] In some examples, when the index of the motion model of the first block is not equal to 0, the range of the MVD component is [-2 K+L , 2 K+LIt is determined to be in the range of [-1], and the value of the MVP component is suppressed and / or clipped to be within the range. K indicates the number of bits used to represent the integer part of the MVD component, L indicates the number of bits used to represent the fractional part of the MVD component, and K and L are positive integers.
[0158] In some examples, K = 13 and L = 4.
[0159] In some examples, when the prediction mode of the first block is such that the first block is MODE_INTER and the variable affine_inter_flag is true, the range of the MVD component is [-2 K+L , 2 K+L -1], it is determined to be in the range, and the value of the MVP component is suppressed and / or clipped to be within the range. K indicates the number of bits used to represent the integer part of the MVD component, L indicates the number of bits used to represent the fractional part of the MVD component, and K and L are positive integers.
[0160] In some examples, K = 13 and L = 4.
[0161] In some examples, if the decoded MVD component is in fractional precision, the decoded MVD component is rounded to an integer MVD component.
[0162] In some examples, the rounded integer MVD component is in the range of [-2 K , 2 K -1] and K = 13.
[0163] In some examples, the values of all decoded MVD components are explicitly clipped to the range of the MVD component during semantic interpretation other than using bitstream constraints.
[0164] In some examples, the transformation generates the first block of the video from the bitstream representation.
[0165] In some examples, the transformation generates a bitstream representation from a first block of video.
[0166] In the list of examples in this patent specification, the term transformation may refer to the generation of a bitstream representation of a current video block or the generation of a current video block from a bitstream representation. The bitstream representation need not represent a group of consecutive bits and can be divided into bits included in a header field or a code name representing coded pixel value information.
[0167] In the above examples, the rules of applicability are predefined and may be recognized by the encoder and decoder.
[0168] As described in this patent specification, it will be understood that the disclosed technology can be implemented in a video encoder or decoder and uses techniques including various implementation rules for considerations regarding the use of differential coding modes in intra coding to improve compression efficiency.
[0169] The disclosed and other solutions, examples, embodiments, modules, and implementations of functional operations described in this patent specification may be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed in this patent specification and their structural equivalents, or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., as one or more modules of computer program instructions encoded on a computer-readable medium for being implemented by, or for controlling the operation of, a data processing apparatus. This computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter that provides a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" includes, for example, a programmable processor, a computer, or multiple processors or computers, or all apparatus, devices, and machines for processing data. This apparatus can include, in addition to hardware, code that creates an execution environment for the computer program, for example, processor firmware, protocol stack, database management system, operating system, or code that constitutes a combination of one or more of them. A propagated signal is an artificially generated signal, for example, an electrical, optical, or electromagnetic signal generated by a machine, and is generated for encoding information to be transmitted to a suitable receiving device.
[0170] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, either as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program may be recorded as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), stored in a single file dedicated to the program, or stored in multiple coordinating files (e.g., files that hold one or more modules, subprograms, or portions of code). It is also possible to deploy one computer program to be executed on one computer located at one site, or on multiple computers distributed across multiple sites and interconnected by a communication network.
[0171] The processes and logic flows described in this patent specification can be performed by one or more programmable processors that execute one or more computer programs to function by operating on input data and generating output. The processes and logic flows can also be performed by special-purpose logic circuits, such as FPGAs (Field Programmable Gate Arrays) or ASICs (Application Specific Integrated Circuits), and the apparatus can also be implemented as special-purpose logic circuits.
[0172] Processors suitable for the execution of a computer program include, for example, both general and special purpose microprocessors, as well as any one or more processors of any kind of digital computer. In general, a processor receives instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing the instructions and data. In general, a computer may include one or more mass storage devices for storing data, such as magnetic, magneto-optical disks, or optical disks, or may be operatively coupled to receive data from or transfer data to such mass storage devices. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, EPROM, EEPROM, flash memory devices, magnetic disks, such as internal hard disks or removable disks, magneto-optical disks, and semiconductor memory devices such as CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, application-specific logic circuitry.
[0173] This patent specification contains many details, but these should not be construed as limiting the scope of any subject matter or the claims, but rather as descriptions and interpretations of features that may be specific to particular embodiments of a particular technology. Specific features described in the context of separate embodiments in this patent specification may be implemented in combination in one example. Conversely, the various features described in the context of one example may be implemented separately or in any suitable sub-combination in multiple embodiments. Furthermore, features may be described and initially claimed above as acting in a particular combination, but one or more features from the claimed combination may, in some cases, be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of the sub-combination.
[0174] Similarly, operations are shown in a particular order in the drawings, but this should not be understood as requiring that such operations be performed in the particular order shown or in a sequential order to achieve the desired result, or that all of the operations shown be performed. Also, the separation of the components of the various systems in the examples described in this patent specification should not be understood as requiring such separation in all embodiments.
[0175] Only some implementations and examples are described, and based on the content described and illustrated in this patent document, other embodiments, extensions, and variations are possible.
Claims
1. A method for processing video data, comprising: determining an MVD (Motion Vector Difference) component having a first accuracy associated with the first block for conversion between a first block of video and a bitstream of the video, wherein the first accuracy is due to an accuracy set comprising 1 / 4 luminance samples, 1 luminance sample, and 4 luminance samples; performing the conversion based on the clipped MVD component; wherein the value of the MVD component is clipped to be within the range of the MVD component; the range of the MVD component is [-2M, 2M-1]; M = 17; the range of the MVD component is adapted to an acceptable MVD accuracy or an acceptable MV (Motion Vector) accuracy of the codec; when there are multiple acceptable MVD accuracies or acceptable MV accuracies in the codec, the range of the MVD component is adapted to the top accuracy of 1 / 16 luminance samples; the accuracy of 1 / 16 luminance samples is used for affine coding blocks and does not exist in the accuracy set for the first block.
2. The method according to claim 1, wherein when the multiple acceptable MVD accuracies or the MV accuracies include 1 / 16 luminance samples, 1 / 4 luminance samples, and 1 luminance sample, the range of the MVD component is adapted to the accuracy of 1 / 16 luminance samples.
3. The first block is applied to a first coding mode in which motion information is derived based on a motion vector candidate list and a scaled MVD, the scaled MVD is generated by bit-shifting the MVD, the motion vector candidate list includes inherited candidates, constructed candidates, and candidates with zero MV (Motion Vector) derived from neighboring blocks coded in the first coding mode. The method according to claim 1.
4. The range of the MVD component is determined to be [-2M, 2M-1], M = K + L, K indicates the number of bits used to represent the integer part of the MVD component, L indicates the number of bits used to represent the fractional part of the MVD component, the MVD component is represented at 1 / 2L luminance sample accuracy, or The range of the MV (Motion Vector) component associated with the first block is determined to be [-2M, 2M - 1], where M = K + L, K indicates the number of bits used to represent the integer part of the MV component, L indicates the number of bits used to represent the fractional part of the MV component, the MV component is represented with 1 / 2^L luminance sample accuracy, M, K, and L are positive integers, the method according to any one of claims 1 to 3. **Claim 5** K = 13, L = 4, and M = 17, the method according to claim 4. **Claim 6** The MVD component is a signal-notified MVD component coded in a bitstream or a converted MVD component associated with a specific accuracy through an internal shift operation in the decoding process, the method according to any one of claims 1 to 5. **Claim 7** The MVD component includes a horizontal MVD component and a vertical MVD component, the horizontal MVD component and the vertical MVD component have the same range, the method according to any one of claims 1 to 6. **Claim 8** The MVD component is represented by integer bits, fractional bits, and sign bits, the method according to any one of claims 1 to 7. **Claim 9** The range of the MV (Motion Vector) associated with the first block is the same as the range of the MVD component, the method according to any one of claims 1 to 8. **Claim 10** The conversion includes decoding the first block from the bitstream, the method according to any one of claims 1 to 9. **Claim 11** The conversion includes encoding the first block into the bitstream, the method according to any one of claims 1 to 9. **Claim 12** An apparatus for processing video data, having a processor and a non-transitory memory having instructions, wherein when the instructions are executed by the processor, the processor is caused to determine an MVD (Motion Vector Difference) component having a first accuracy associated with the first block for conversion between a first block of a video and a bitstream of the video, the first accuracy being due to an accuracy set comprising 1 / 4 luminance samples, 1 luminance sample, and 4 luminance samples, execute the conversion based on the suppressed MVD component, to execute, the value of the MVD component is suppressed to be within the range of the MVD component, the range of the MVD component is [-2M, 2M - 1], M = 17, the range of the MVD component is adapted to the acceptable MVD accuracy or acceptable MV (Motion Vector) accuracy of the codec, when there are multiple acceptable MVD accuracies or acceptable MV accuracies in the codec, the range of the MVD component is adapted to the highest accuracy of 1 / 16 luminance samples, the accuracy of the 1 / 16 luminance samples is used for an affine coding block and does not exist in the accuracy set for the first block, device.
13. to a processor, for the conversion between a first block of a video and the bitstream of the video, determining an MVD (Motion Vector Difference) component having a first accuracy associated with the first block, wherein the first accuracy results from an accuracy set comprising 1 / 4 luminance samples, 1 luminance sample, and 4 luminance samples, executing the conversion based on the suppressed MVD component, to execute, the value of the MVD component is suppressed to be within the range of the MVD component, the range of the MVD component is [-2M, 2M - 1], M = 17, the range of the MVD component is adapted to the acceptable MVD accuracy or acceptable MV (Motion Vector) accuracy of the codec, when there are multiple acceptable MVD accuracies or acceptable MV accuracies in the codec, the range of the MVD component is adapted to the highest accuracy of 1 / 16 luminance samples, a non - transitory computer - readable medium storing instructions, wherein the accuracy of the 1 / 16 luminance samples is used for an affine coding block and does not exist in the accuracy set for the first block.
14. A method for storing a bitstream of a video, wherein the method comprises determining an MVD (Motion Vector Difference) component having a first accuracy associated with a first block of the video, wherein the first accuracy results from an accuracy set comprising 1 / 4 luminance samples, 1 luminance sample, and 4 luminance samples, generating the bitstream based on the suppressed MVD component, Storing the bitstream in a non-transitory computer-readable recording medium; having; the value of the MVD component is suppressed to be within the range of the MVD component; the range of the MVD component is [-2M, 2M - 1]; M = 17; the range of the MVD component conforms to an acceptable MVD accuracy or an acceptable MV (Motion Vector) accuracy of the codec; when there are a plurality of acceptable MVD accuracies or acceptable MV accuracies in the codec, the range of the MVD component conforms to the topmost accuracy of 1 / 16 luminance samples; the accuracy of the 1 / 16 luminance samples is used for an affine coding block and does not exist in the accuracy set for the first block, a method.
Citation Information
Patent Citations
JPP7303329B
Encoding method and apparatus therefor, and decoding method and apparatus therefor
US20200236395A1
Encoding method and apparatus therefor, and decoding method and apparatus therefor
WO2019066514A1
Cited By
Restriction related to difference of motion vector
JP2025090654A