Video encoding and decoding
Patent Information
- Application Number
- JP2025124997
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-12-20
- Filing Date
- 2025-07-25
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2039-11-19
AI Technical Summary
【0270】 この単純な評価は、VTM2.0に対して符号化効率の改善を達成する。他の利点は、第7および第8の実施形態と比較して、隣接ブロックのマージインデックスではなくマージフラグのみがチェックされる必要があるので、より低い複雑さである。
Smart Images

Figure 0007923872000007 
Figure 0007923872000008 
Figure 0007923872000009
Abstract
Description
[Technical Field]
[0001] The present invention relates to video encoding and decoding. [Background Art]
[0002] Recently, JVET (Joint Video Experts Team), a joint team formed by MPEG and VCEG of ITU-T Study Group 16, has started research on a new video coding standard called VVC (Versatile Video Coding). The goal of VVC is to provide significant improvement in compression performance over the existing HEVC standard (i.e., typically twice that of the previous standard) and to be completed in 2020. Main target applications and services include, but are not limited to, 360-degree and high dynamic range (HDR) video. Overall, JVET evaluated responses from 32 organizations using formal subjective tests conducted by independent laboratories. Several proposals have demonstrated compression efficiency gains of typically 40% or more compared to the case of using HEVC. The proposals have shown particular effectiveness for ultra-high definition (UHD) video test materials. Therefore, it is expected that the improvement in compression efficiency will far exceed the 50% target for the final standard.
[0003] The JVET Exploration Model (JEM) uses all HEVC tools. An additional tool not present in HEVC is the use of an "affine motion mode" when applying motion compensation. While motion compensation in HEVC is limited to translation, there are actually many types of motion such as zoom-in / zoom-out, rotation, perspective motion, and other irregular motions. When the affine motion mode is used, a more complex transformation is applied to a block to attempt to more accurately predict the formation of such motion. Therefore, it is desirable to be able to use the affine motion mode while reducing complexity while achieving good coding efficiency.
[0004] Another tool not present in HEVC is the use of Alternative Temporal Motion Vector Prediction (ATMVP). ATMVP is a specific motion compensation. Instead of considering only one motion piece for the current block from the temporal reference frame, it considers each motion piece for each collated block. Thus, this temporal motion vector prediction provides segmentation of the current block using the relevant motion piece for each subblock. In the current VTM (VVC Test Model) reference software, ATMVP is signaled as a merge candidate inserted into the list of merge candidates. When ATMVP is enabled at the SPS level, the maximum number of merge candidates increases by one. Therefore, six candidates are considered instead of five, compared to when this mode is disabled.
[0005] These, and other tools described later, introduce coding efficiency and complexity issues related to the coding of indices (e.g., merge indexes) or flags used to indicate which candidate has been selected from a list of candidates (e.g., a list of merge candidates for use with merge mode coding). [Overview of the project]
[0006] Therefore, a solution to at least one of the aforementioned problems is desirable.
[0007] According to a first aspect of the present invention, a method for encoding a motion vector predictor index, Generate a list of motion vector predictor candidates, including ATMVP candidates. Select one of the motion vector predictor candidates from the list above, Using CABAC coding, a motion vector predictor index (merge index) is generated for the selected motion vector predictor candidates, and one or more bits of the motion vector predictor index are bypassed by CABAC coding. A method characterized by the above is provided.
[0008] In one embodiment, all bits of the motion vector predictor index except the first bit are bypassed and CABAC encoded.
[0009] According to a second aspect of the present invention, a method for decoding a motion vector predictor index, Generate a list of motion vector predictor candidates, including ATMVP candidates. The motion vector predictor index is decoded using CABAC decoding, and one or more bits of the motion vector predictor index are bypassed by CABAC decoding. The decoded motion vector predictor index is used to identify one of the motion vector predictor candidates in the list. A method characterized by the above is provided.
[0010] In one embodiment, all bits except the first bit of the motion vector predictor index are bypassed and CABAC decoded.
[0011] According to a third aspect of the present invention, a device for encoding a motion vector predictor index, A means for generating a list of motion vector predictor candidates, including ATMVP candidates, A means for selecting one of the motion vector predictor candidates in the aforementioned list, A means for generating a motion vector predictor index (merge index) of selected motion vector predictor candidates using CABAC coding, wherein one or more bits of the motion vector predictor index are bypassed by CABAC coding. An apparatus is provided that is characterized by comprising the following.
[0012] According to a fourth aspect of the present invention, a device for decoding a motion vector predictor index, A means for generating a list of motion vector predictor candidates, including ATMVP candidates, Means for decoding a motion vector predictor index using CABAC decoding, wherein one or more bits of the motion vector predictor index are bypassed by CABAC decoding, A means for identifying one of the motion vector predictor candidates in the list using the decoded motion vector predictor index. An apparatus is provided that is characterized by comprising the following.
[0013] A fifth aspect of the present invention relates to a method for encoding a motion vector predictor index, Generate a list of motion vector predictor candidates, Select one of the motion vector predictor candidates from the list above, Using CABAC coding, generate motion vector predictor indices for selected motion vector predictor candidates, where two or more bits of the motion vector predictor index share the same context. A method characterized by the above is provided.
[0014] In one embodiment, all bits of the motion vector predictor index share the same context.
[0015] According to a sixth aspect of the present invention, a method for decoding a motion vector predictor index, Generate a list of motion vector predictor candidates, Using CABAC decoding, decode the motion vector predictor index, and if two or more bits of the motion vector predictor index share the same context, The decoded motion vector predictor index is used to identify one of the motion vector predictor candidates in the list. A method characterized by the above is provided.
[0016] In one embodiment, all bits of the motion vector predictor index share the same context.
[0017] According to a seventh aspect of the present invention, there is provided an apparatus for encoding a motion vector predictor index, means for generating a list of motion vector predictor candidates; means for selecting one of the motion vector predictor candidates in the list; means for generating a motion vector predictor index of the selected motion vector predictor candidate using CABAC encoding, wherein two or more bits of the motion vector predictor index share the same context An apparatus characterized by comprising: is provided.
[0018] According to an eighth aspect of the present invention, there is provided an apparatus for decoding a motion vector predictor index, means for generating a list of motion vector predictor candidates; means for decoding a motion vector predictor index using CABAC decoding, wherein two or more bits of the motion vector predictor index share the same context; means for identifying one of the motion vector predictor candidates in the list using the decoded motion vector predictor index An apparatus characterized by comprising: is provided.
[0019] According to a ninth aspect of the present invention, there is provided a method for encoding a motion vector predictor index, generating a list of motion vector predictor candidates, selecting one of the motion vector predictor candidates in the list, generating a motion vector predictor index of the selected motion vector predictor candidate using CABAC encoding, wherein a context variable of at least one bit of the motion vector predictor index of a current block depends on a motion vector predictor index of at least one block adjacent to the current block A method characterized by: is provided.
[0020] In one embodiment, a context variable for at least one bit of the motion vector predictor index depends on the respective motion vector predictor indices of at least two adjacent blocks.
[0021] In another embodiment, a context variable of at least one bit of the motion vector predictor index depends on the motion vector predictor index of the left neighboring block to the left of the current block and the motion vector predictor index of the upper neighboring block above the current block.
[0022] In another embodiment, the left adjacent block is A2 and the upper adjacent block is B3.
[0023] In another embodiment, the left adjacent block is A1 and the upper adjacent block is B1.
[0024] In another embodiment, the context variable has three different possible values.
[0025] Another embodiment includes comparing the motion vector predictor index of at least one adjacent block with the index value of the motion vector predictor index of the current block, and setting the context variable according to the comparison result.
[0026] Another embodiment includes comparing the motion vector predictor index of at least one adjacent block with a parameter representing the bit position of the said or one of the said bits in the motion vector predictor index of the current block, and setting the context variable according to the comparison result.
[0027] Another embodiment includes performing a first comparison, comparing the motion vector predictor index of a first neighboring block with a parameter representing the bit position of the bit of the bit in the motion vector predictor index of the current block; performing a second comparison, comparing the motion vector predictor index of a second neighboring block with the parameter; and setting the context variable according to the results of the first and second comparisons.
[0028] According to a tenth aspect of the present invention, a method for decoding a motion vector predictor index, Generate a list of motion vector predictor candidates, Using CABAC decoding, the motion vector predictor index is decoded, where the context variable of at least one bit of the motion vector predictor index of the current block depends on the motion vector predictor index of at least one block adjacent to the current block. The decoded motion vector predictor index is used to identify one of the motion vector predictor candidates in the list. A method characterized by the above is provided.
[0029] In one embodiment, a context variable for at least one bit of the motion vector predictor index depends on the respective motion vector predictor indices of at least two adjacent blocks.
[0030] In another embodiment, a context variable of at least one bit of the motion vector predictor index depends on the motion vector predictor index of the left neighboring block to the left of the current block and the motion vector predictor index of the upper neighboring block above the current block.
[0031] In another embodiment, the left adjacent block is A2 and the upper adjacent block is B3.
[0032] In another embodiment, the left adjacent block is A1 and the upper adjacent block is B1.
[0033] In another embodiment, the context variable has three different possible values.
[0034] Another embodiment includes comparing the motion vector predictor index of at least one adjacent block with the index value of the motion vector predictor index of the current block, and setting the context variable according to the comparison result.
[0035] Another embodiment includes comparing the motion vector predictor index of at least one adjacent block with a parameter representing the bit position of the said or one of the said bits in the motion vector predictor index of the current block, and setting the context variable according to the comparison result.
[0036] Another embodiment includes performing a first comparison, comparing the motion vector predictor index of a first neighboring block with a parameter representing the bit position of the said or one of the said bits in the motion vector predictor index of the current block; performing a second comparison, comparing the motion vector predictor index of a second neighboring block with the said parameter; and setting the context variable according to the results of the first and second comparisons.
[0037] According to an eleventh aspect of the present invention, a device for encoding a motion vector predictor index, A means for generating a list of motion vector predictor candidates, A means for selecting one of the motion vector predictor candidates in the aforementioned list, Means for generating motion vector predictor indices of selected motion vector predictor candidates using CABAC coding, wherein a context variable of at least one bit of the motion vector predictor index of the current block depends on the motion vector predictor indices of at least one block adjacent to the current block. An apparatus is provided that is characterized by comprising the following.
[0038] According to a twelfth aspect of the present invention, a device for decoding a motion vector predictor index, A means for generating a list of motion vector predictor candidates, Means for decoding a motion vector predictor index using CABAC decoding, wherein a context variable for at least one bit of the motion vector predictor index of the current block depends on the motion vector predictor index of at least one block adjacent to the current block, A means for identifying one of the motion vector predictor candidates in the list using the decoded motion vector predictor index. An apparatus is provided that is characterized by comprising the following.
[0039] According to a thirteenth aspect of the present invention, a method for encoding a motion vector predictor index, Generate a list of motion vector predictor candidates, Select one of the motion vector predictor candidates from the list above, Using CABAC coding, the motion vector predictor index of the selected motion vector predictor candidate is generated, where at least one bit of the context variable of the motion vector predictor index of the current block depends on the skip flag of the current block. A method characterized by the above is provided.
[0040] According to a fourteenth aspect of the present invention, a method for encoding a motion vector predictor index, Generate a list of motion vector predictor candidates, Select one of the motion vector predictor candidates from the list above, Using CABAC coding, a motion vector predictor index is generated for the selected motion vector predictor candidate, where at least one bit of the motion vector predictor index of the current block is a context variable that depends on another parameter or syntax element of the current block available before decoding the motion vector predictor index. A method characterized by the above is provided.
[0041] According to a 15th aspect of the present invention, a method for encoding a motion vector predictor index, Generate a list of motion vector predictor candidates, Select one of the motion vector predictor candidates from the list above, Using CABAC coding, a motion vector predictor index is generated for the selected motion vector predictor candidate, where at least one bit of the motion vector predictor index of the current block is a context variable that depends on another parameter or syntax element of the current block, which is an indicator of the complexity of motion within the current block. A method characterized by the above is provided.
[0042] According to a sixteenth aspect of the present invention, a method for decoding a motion vector predictor index, Generate a list of motion vector predictor candidates, Using CABAC decoding, the motion vector predictor index is decoded, where the context variable for at least one bit of the motion vector predictor index of the current block depends on the skip flag of the current block. The decoded motion vector predictor index is used to identify one of the motion vector predictor candidates in the list. A method characterized by the above is provided.
[0043] According to a 17th aspect of the present invention, a method for decoding a motion vector predictor index, Generate a list of motion vector predictor candidates, Using CABAC decoding, the motion vector predictor index is decoded, where the context variable of at least one bit of the motion vector predictor index of the current block depends on another parameter or syntax element of the current block that is available before decoding the motion vector predictor index. The decoded motion vector predictor index is used to identify one of the motion vector predictor candidates in the list. A method characterized by the above is provided.
[0044] According to an eighteenth aspect of the present invention, a method for decoding a motion vector predictor index, Generate a list of motion vector predictor candidates, Using CABAC decoding, the motion vector predictor index is decoded, where at least one bit of the motion vector predictor index of the current block is a context variable that depends on another parameter or syntax element of the current block, which is an indicator of the complexity of motion in the current block. The decoded motion vector predictor index is used to identify one of the motion vector predictor candidates in the list. A method characterized by the above is provided.
[0045] According to a 19th aspect of the present invention, a device for encoding a motion vector predictor index, A means for generating a list of motion vector predictor candidates, A means for selecting one of the motion vector predictor candidates in the aforementioned list, Means for generating motion vector predictor indices of selected motion vector predictor candidates using CABAC coding, wherein a context variable for at least one bit of the motion vector predictor index of the current block depends on the skip flag of the current block. An apparatus is provided that is characterized by comprising the following.
[0046] According to a 20th aspect of the present invention, a device for encoding a motion vector predictor index, A means for generating a list of motion vector predictor candidates, A means for selecting one of the motion vector predictor candidates in the aforementioned list, Means for generating a motion vector predictor index of a selected motion vector predictor candidate using CABAC coding, wherein a context variable of at least one bit of the motion vector predictor index of the current block depends on another parameter or syntax element of the current block that is available before decoding the motion vector predictor index. An apparatus is provided that is characterized by comprising the following.
[0047] According to a 21st aspect of the present invention, a device for encoding a motion vector predictor index, A means for generating a list of motion vector predictor candidates, A means for selecting one of the motion vector predictor candidates in the aforementioned list, Means for generating motion vector predictor indices of selected motion vector predictor candidates using CABAC coding, wherein a context variable of at least one bit of the motion vector predictor index of the current block depends on another parameter or syntax element of the current block which is an indicator of the complexity of motion in the current block. An apparatus is provided that is characterized by comprising the following.
[0048] According to a 22nd aspect of the present invention, a device for decoding a motion vector predictor index, A means for generating a list of motion vector predictor candidates, Means for decoding a motion vector predictor index using CABAC decoding, wherein a context variable for at least one bit of the motion vector predictor index of the current block depends on the skip flag of the current block, A means for identifying one of the motion vector predictor candidates in the list using the decoded motion vector predictor index. An apparatus is provided that is characterized by comprising the following.
[0049] According to a 23rd aspect of the present invention, a device for decoding a motion vector predictor index, A means for generating a list of motion vector predictor candidates, Means for decoding a motion vector predictor index using CABAC decoding, wherein a context variable for at least one bit of the motion vector predictor index of the current block depends on another parameter or syntax element of the current block that is available before decoding the motion vector predictor index. A means for identifying one of the motion vector predictor candidates in the list using the decoded motion vector predictor index. An apparatus is provided that is characterized by comprising the following.
[0050] According to a 24th aspect of the present invention, a device for decoding a motion vector predictor index, A means for generating a list of motion vector predictor candidates, Means for decoding a motion vector predictor index using CABAC decoding, wherein a context variable for at least one bit of the motion vector predictor index of the current block depends on another parameter or syntax element of the current block which is an indicator of the complexity of motion in the current block. A means for identifying one of the motion vector predictor candidates in the list using the decoded motion vector predictor index. An apparatus is provided that is characterized by comprising the following.
[0051] A 25th aspect of the present invention provides a method for encoding information about a motion information predictor, wherein one of a plurality of motion information predictor candidates is selected, and CABAC coding is used to encode information for identifying the selected motion information predictor candidate, wherein the CABAC coding uses the same context variable used for another interprediction mode when one or both of the triangle merge mode or motion vector difference (MMVD) merge mode are used for at least one bit of the information.
[0052] A method for decoding information about a motion information predictor is provided, comprising decoding information for identifying one of a plurality of motion information predictor candidates using CABAC decoding, selecting one of the plurality of motion information predictor candidates using the decoded information, wherein the CABAC decoding uses the same context variable used for another interprediction mode when one or both of the triangle merge mode or motion vector difference (MMVD) merge mode merge is used for at least one bit of the information.
[0053] With respect to the 25th or 26th aspect of the present invention, the following features may be provided according to the embodiment.
[0054] Preferably, all bits of the information except the first bit are bypassed CABAC encoded or bypassed CABAC decoded. Preferably, the first bit is CABAC encoded or CABAC decoded. Preferably, another interprediction mode includes one or both of a merge mode or an affine merge mode. Preferably, another interprediction mode includes an MHII (Multi-Hypothesis Intra Inter) merge mode. Preferably, multiple motion information predictor candidates for another interprediction mode include ATMVP candidates. Preferably, CABAC encoding or CABAC decoding includes using the same context variables for both when a triangle merge mode is used and an MMVD merge mode is used. Preferably, at least one bit of the information is CABAC encoded or CABAC decoded when a skip mode is used. Preferably, the skip mode includes one or more of a merge skip mode, an affine merge skip mode, a triangle merge skip mode, or a Motion Vector Difference (MMVD) merge skip mode.
[0055] A 27th aspect of the present invention provides a method for encoding information about a motion information predictor, wherein one of a plurality of motion information predictor candidates is selected, information for identifying the selected motion information predictor candidate is encoded, and the encoding of the information is performed by bypassing CABAC encoding at least one bit of the information when one or both of the triangle merge mode or motion vector difference (MMVD) merge mode merge is used.
[0056] According to a 28th aspect of the present invention, a method is provided for decoding information about a motion information predictor, wherein the decoding of information for identifying one of a plurality of motion information predictor candidates, the decoding of information for selecting one of the plurality of motion information predictor candidates using the decoded information, and the decoding of information by bypassing CABAC decoding at least one bit of the information when one or both of the triangle merge mode or motion vector difference (MMVD) merge mode merge is used.
[0057] With respect to the 27th or 28th aspect of the present invention, the following features may be provided according to the embodiment.
[0058] Preferably, all bits of the information except the first bit are bypassed CABAC encoded or bypassed CABAC decoded. Preferably, the first bit is CABAC encoded or CABAC decoded. Preferably, all bits of the information are bypassed CABAC encoded or bypassed CABAC decoded when either or both of triangle merge mode or MMVD merge mode are used. Preferably, all bits of the information are bypassed CABAC encoded or bypassed CABAC decoded.
[0059] Suitablely, at least one bit of the information is CABAC encoded or CABAC decoded when affine merge mode is used. Suitablely, all bits of the information are bypassed CABAC encoded or bypassed CABAC decoded, except when affine merge mode is used.
[0060] Preferably, at least one bit of the information is CABAC encoded or decoded when either or both of merge mode or MHII (Multi-Hypothesis Intra Inter) merge mode are used. Preferably, all bits of the information are bypassed CABAC encoded or decoded, except when either or both of merge mode or MHII (Multi-Hypothesis Intra Inter) merge mode are used.
[0061] Preferably, at least one bit of the information is CABAC encoded or decoded if the multiple motion information predictor candidates include ATMVP candidates. Preferably, all bits of the information are bypassed CABAC encoded or decoded unless the multiple motion information predictor candidates include ATMVP candidates.
[0062] Preferably, at least one bit of the information is CABAC encoded or CABAC decoded when the skip mode is used. Preferably, all bits of the information are bypass CABAC encoded or bypass CABAC decoded except when the skip mode is used. Preferably, the skip mode includes one or more merges of merge skip modes, affine merge skip modes, triangle merge skip modes, or motion vector difference (MMVD) merge skip modes.
[0063] With respect to the 25th, 26th, 27th, or 28th aspect of the present invention, the following features may be provided according to the embodiment.
[0064] Preferably, at least one bit includes the first bit of the information. Preferably, the information includes a motion information prediction index or flag. Preferably, the motion information predictor candidate includes information for obtaining a motion vector.
[0065] With respect to the 25th or 27th aspect of the present invention, the following features may be provided according to the embodiment.
[0066] Preferably, the method further includes information to indicate the use of one of the following modes in the bitstream: triangle merge mode, MMVD merge mode, merge mode, affine merge mode, or MHII (Multi-Hypothesis Intra Inter) merge mode. Preferably, the method further includes information to determine the maximum number of motion information predictor candidates that may be included in a plurality of motion information predictor candidates in the bitstream.
[0067] With respect to the 26th or 28th aspect of the present invention, the following features may be provided according to the embodiment.
[0068] Preferably, the method further includes obtaining information from the bitstream to indicate the use of one of the following modes: triangle merge mode, MMVD merge mode, merge mode, affine merge mode, or MHII (Multi-Hypothesis Intra Inter) merge mode. Preferably, the method further includes obtaining information from the bitstream to determine the maximum number of motion predictor candidates that may be included in a plurality of motion predictor candidates.
[0069] According to a 29th aspect of the present invention, an apparatus is provided for encoding information about a motion information predictor, comprising means for selecting one of a plurality of motion information predictor candidates, and means for encoding information for identifying the selected motion information predictor candidate using CABAC coding, wherein the CABAC coding uses the same context variable used for another interprediction mode when one or both of the triangle merge mode or motion vector difference (MMVD) merge mode merge are used for at least one bit of the information. Preferably, the apparatus comprises means for performing a method for encoding information about a motion information predictor according to a 25th or 27th aspect of the present invention.
[0070] According to a 30th aspect of the present invention, an apparatus is provided for encoding information about a motion information predictor, comprising means for selecting one of a plurality of motion information predictor candidates, and means for encoding information for identifying the selected motion information predictor candidate, wherein encoding the information comprises bypassing CABAC encoding at least one bit of the information when one or both of a triangle merge mode or a motion vector difference (MMVD) merge mode merge are used. Preferably, the apparatus comprises means for performing a method for encoding information about a motion information predictor according to a 25th or 27th aspect of the present invention.
[0071] According to a 31st aspect of the present invention, a device for decoding information relating to a motion information predictor is provided, comprising means for decoding information for identifying one of a plurality of motion information predictor candidates using CABAC decoding, and means for selecting one of the plurality of motion information predictor candidates using the decoded information, wherein the CABAC decoding uses the same context variable used for another interprediction mode when one or both of the triangle merge mode or motion vector difference (MMVD) merge mode are used for at least one bit of the information. Preferably, the device comprises means for performing a method for decoding information relating to a motion information predictor according to a 26th or 28th aspect of the present invention.
[0072] According to a 32nd aspect of the present invention, a device for decoding information relating to a motion information predictor is provided, comprising means for decoding information for identifying one of a plurality of motion information predictor candidates, and means for selecting one of the plurality of motion information predictor candidates using the decoded information, wherein decoding the information comprises bypassing CABAC decoding at least one bit of the information when one or both of a triangle merge mode or a motion vector difference (MMVD) merge mode merge is used. Preferably, the device comprises means for performing a method for decoding information relating to a motion information predictor according to a 26th or 28th aspect of the present invention.
[0073] A 33rd aspect of the present invention provides a method for encoding a motion vector predictor index, comprising: generating a list of motion vector predictor candidates; selecting one of the motion vector predictor candidates in the list; generating a motion vector predictor index for the selected motion vector predictor candidate using CABAC encoding; and deriving a context variable of at least one bit of the motion vector predictor index of the current block from at least one context variable of the skip flag and affine flag of the current block.
[0074] A 34th aspect of the present invention provides a method for decoding a motion vector predictor index, comprising: generating a list of motion vector predictor candidates; decoding the motion vector predictor index using CABAC decoding; a context variable of at least one bit of the motion vector predictor index of the current block being derived from at least one context variable of the skip flag and affine flag of the current block; and using the decoded motion vector predictor index to identify one of the motion vector predictor candidates in the list.
[0075] According to a 35th aspect of the present invention, there is a device for encoding a motion vector predictor index, comprising means for generating a list of motion vector predictor candidates, means for selecting one of the motion vector predictor candidates in the list, and means for generating a motion vector predictor index of the selected motion vector predictor candidate using CABAC encoding, wherein a context variable of at least one bit of the motion vector predictor index of the current block is derived from at least one context variable of the skip flag and affine flag of the current block.
[0076] According to a 36th aspect of the present invention, there is a device for decoding a motion vector predictor index, comprising: means for generating a list of motion vector predictor candidates; means for decoding the motion vector predictor index using CABAC decoding, wherein a context variable of at least one bit of the motion vector predictor index of the current block is derived from at least one context variable of the skip flag and affine flag of the current block, and means for identifying one of the motion vector predictor candidates in the list using the decoded motion vector predictor index.
[0077] A 37th aspect of the present invention provides a method for encoding a motion vector predictor index, comprising: generating a list of motion vector predictor candidates; selecting one of the motion vector predictor candidates in the list; generating a motion vector predictor index for the selected motion vector predictor candidate using CABAC encoding, wherein a context variable of at least one bit of the motion vector predictor index of the current block has only two distinct possible values.
[0078] A 38th aspect of the present invention provides a method for decoding a motion vector predictor index, comprising generating a list of motion vector predictor candidates, decoding the motion vector predictor index using CABAC decoding, wherein a context variable for at least one bit of the motion vector predictor index of the current block has only two distinct possible values, and using the decoded motion vector predictor index, one of the motion vector predictor candidates in the list is identified.
[0079] According to a 39th aspect of the present invention, there is a device for encoding a motion vector predictor index, comprising means for generating a list of motion vector predictor candidates, means for selecting one of the motion vector predictor candidates in the list, and means for generating a motion vector predictor index for the selected motion vector predictor candidate using CABAC encoding, wherein a context variable of at least one bit of the motion vector predictor index of the current block has only two different possible values.
[0080] According to a 40th aspect of the present invention, there is a device for decoding a motion vector predictor index, comprising means for generating a list of motion vector predictor candidates; means for decoding the motion vector predictor index using CABAC decoding, wherein a context variable for at least one bit of the motion vector predictor index of the current block has only two distinct possible values, and means for identifying one of the motion vector predictor candidates in the list using the decoded motion vector predictor index.
[0081] According to a forty-first aspect of the present invention, there is a method for encoding a motion information predictor index, characterized in that a list of motion information predictor candidates is generated, one of the motion information predictor candidates in the list is selected as an affine merge mode predictor when affine merge mode is used, one of the motion information predictor candidates in the list is selected as a non-affine merge mode predictor when non-affine merge mode is used, a motion information predictor index of the selected motion information predictor candidate is generated using CABAC encoding, and one or more bits of the motion information predictor index are bypassed by CABAC encoding.
[0082] Suitablely, the CABAC coding comprises using the same context variable for at least one bit of the motion information predictor index of the current block, whether affine merge mode or non-affine merge mode is used. Alternatively, the CABAC coding comprises using a first context variable for at least one bit of the motion information predictor index of the current block, if affine merge mode is used, or a second context variable if non-affine merge mode is used, and the method further comprises including data in the bitstream indicating the use of affine merge mode when affine merge mode is used.
[0083] Preferably, the method further includes data for determining the maximum number of motion predictor candidates that may be included in the generated list of motion predictor candidates in the bitstream. Preferably, all bits of the motion predictor index except the first bit are bypassed CABAC encoded. Preferably, the first bit is CABAC encoded. Preferably, the motion predictor index of the selected motion predictor candidates is encoded using the same syntax elements when affine merge mode and when non-affine merge mode are used.
[0084] A forty-second aspect of the present invention provides a method for decoding a motion information predictor index, characterized by generating a list of motion information predictor candidates, decoding the motion information predictor index using CABAC decoding, one or more bits of the motion information predictor index being bypassed by CABAC decoding, and, if affine merge mode is used, using the decoded motion information predictor index to identify one of the motion information predictor candidates in the list as an affine merge mode predictor, and, if non-affine merge mode is used, using the decoded motion information predictor index to identify one of the motion information predictor candidates in the list as a non-affine merge mode predictor.
[0085] Suitablely, CABAC decoding includes using the same context variable for at least one bit of the motion information predictor index of the current block, whether affine merge mode or non-affine merge mode is used. Alternatively, the method further includes obtaining data from the bitstream indicating the use of affine merge mode, and CABAC decoding includes using a first context variable for at least one bit of the motion information predictor index of the current block if the obtained data indicates the use of affine merge mode, and using a second context variable if the obtained data indicates the use of non-affine merge mode.
[0086] Appropriately, the method further includes obtaining data from the bitstream indicating the use of affine merge mode, and the generated list of motion information predictor candidates includes affine merge mode predictor candidates if the obtained data indicates the use of affine merge mode, and non-affine merge mode predictor candidates if the obtained data indicates the use of non-affine merge mode.
[0087] Preferably, the method further includes obtaining data from the bitstream to determine the maximum number of motion predictor candidates that may be included in the generated list of motion predictor candidates. Preferably, all bits of the motion predictor index except the first bit are bypassed CABAC decoded. Preferably, the first bit is CABAC decoded. Preferably, decoding the motion predictor index includes parsing the same syntax elements from the bitstream, when affine merge mode and non-affine merge mode are used. Preferably, the motion predictor candidates include information for obtaining motion vectors. Preferably, the generated list of motion predictor candidates includes ATMVP candidates. Preferably, the generated list of motion predictor candidates has the same maximum number of motion predictor candidates that may be included therein, when affine merge mode and non-affine merge mode are used.
[0088] According to a 43rd aspect of the present invention, there is a device for encoding a motion information predictor index, comprising: means for generating a list of motion information predictor candidates; means for selecting one of the motion information predictor candidates in the list as an affine merge mode predictor when an affine merge mode is used; means for selecting one of the motion information predictor candidates in the list as a non-affine merge mode predictor when a non-affine merge mode is used; and means for generating a motion information predictor index of the selected motion information predictor candidates using CABAC encoding, wherein one or more bits of the motion information predictor index are bypassed by CABAC encoding. Preferably, the device comprises means for performing a method for encoding a motion information predictor index according to the 41st aspect.
[0089] According to a forty-fourth aspect of the present invention, there is a device for decoding a motion information predictor index, comprising: means for generating a list of motion information predictor candidates; means for decoding the motion information predictor index using CABAC decoding, wherein one or more bits of the motion information predictor index are bypassed by CABAC decoding; and, when affine merge mode is used, means for using the decoded motion information predictor index to identify one of the motion information predictor candidates in the list as an affine merge mode predictor; and, when non-affine merge mode is used, means for using the decoded motion information predictor index to identify one of the motion information predictor candidates in the list as a non-affine merge mode predictor. Preferably, the device comprises means for performing a method for decoding a motion information predictor index according to the forty-second aspect.
[0090] According to a forty-fifth aspect of the present invention, there is a method for encoding an affine merge mode motion information predictor index, characterized in that a list of motion information predictor candidates is generated, one of the motion information predictor candidates in the list is selected as an affine merge mode predictor, a motion information predictor index of the selected motion information predictor candidate is generated using CABAC encoding, and one or more bits of the motion information predictor index are bypassed by CABAC encoding.
[0091] Preferably, when non-affine merge mode is used, the method further includes selecting one of the motion information predictor candidates in the list as the non-affine merge mode predictor. Preferably, the CABAC coding comprises using a first context variable for at least one bit of the motion information predictor index of the current block when affine merge mode is used, or using a second context variable when non-affine merge mode is used, and the method further includes including data in the bitstream indicating the use of affine merge mode when affine merge mode is used. Alternatively, the CABAC coding comprises using the same context variable for at least one bit of the motion information predictor index of the current block when affine merge mode is used and when non-affine merge mode is used.
[0092] Preferably, the method further includes data for determining the maximum number of motion information predictor candidates that may be included in the generated list of motion information predictor candidates in the bitstream.
[0093] Preferably, all bits of the motion information predictor index except the first bit are bypassed CABAC encoded. Appropriately, the first bit is CABAC encoded. Appropriately, the motion information predictor index of the selected motion information predictor candidate is encoded using the same syntax elements, both when affine merge mode and non-affine merge mode are used.
[0094] A forty-sixth aspect of the present invention provides a method for decoding an affine merge mode motion information predictor index, characterized by generating a list of motion information predictor candidates, decoding the motion information predictor index using CABAC decoding, and, if one or more bits of the motion information predictor index are bypassed by CABAC decoding and affine merge mode is used, using the decoded motion information predictor index to identify one of the motion information predictor candidates in the list as an affine merge mode predictor.
[0095] Preferably, when non-affine merge mode is used, the method further includes using the decoded motion information predictor index to identify one of the motion information predictor candidates in the list as a non-affine merge mode predictor. Preferably, the method further includes obtaining data from the bitstream indicating the use of affine merge mode, and the CABAC decoding includes using a first context variable for at least one bit of the motion information predictor index of the current block if the obtained data indicates the use of affine merge mode, and using a second context variable if the obtained data indicates the use of non-affine merge mode. Alternatively, the CABAC decoding includes using the same context variables for at least one bit of the motion information predictor index of the current block, both when affine merge mode and when non-affine merge mode are used.
[0096] Appropriately, the method further includes obtaining data from the bitstream indicating the use of affine merge mode, and the generated list of motion information predictor candidates includes affine merge mode predictor candidates if the obtained data indicates the use of affine merge mode, and non-affine merge mode predictor candidates if the obtained data indicates the use of non-affine merge mode.
[0097] Preferably, decoding the motion information predictor index involves parsing the same syntax elements from the bitstream, when affine merge mode and non-affine merge mode are used. Preferably, the method further includes obtaining data from the bitstream to determine the maximum number of motion information predictor candidates that may be included in the generated list of motion information predictor candidates. Preferably, all bits of the motion information predictor index except the first bit are bypassed CABAC decoded. Preferably, the first bit is CABAC decoded. Preferably, the motion information predictor candidates include information for obtaining motion vectors. Preferably, the generated list of motion information predictor candidates includes ATMVP candidates. Preferably, the generated list of motion information predictor candidates has the same maximum number of motion information predictor candidates that may be included, when affine merge mode and non-affine merge mode are used.
[0098] According to a 47th aspect of the present invention, there is a device for encoding an affine merge mode motion information predictor index, comprising means for generating a list of motion information predictor candidates, means for selecting one of the motion information predictor candidates in the list as an affine merge mode predictor, and means for generating a motion information predictor index of the selected motion information predictor candidate using CABAC encoding, wherein one or more bits of the motion information predictor index are bypassed by CABAC encoding. Preferably, the device comprises means for performing a method for encoding a motion information predictor index according to a 45th aspect.
[0099] According to a forty-eighth aspect of the present invention, there is a device for decoding an affine merge mode motion information predictor index, comprising: means for generating a list of motion information predictor candidates; means for decoding the motion information predictor index using CABAC decoding, wherein one or more bits of the motion information predictor index are bypassed by CABAC decoding; and, when affine merge mode is used, means for using the decoded motion information predictor index to identify one of the motion information predictor candidates in the list as an affine merge mode predictor. Preferably, the device comprises means for performing a method for decoding a motion information predictor index according to a forty-sixth aspect.
[0100] A further aspect of the present invention relates to a program that, when executed by a computer or processor, causes the computer or processor to perform any of the methods of the preceding aspects. The program may be provided by itself, or it may be carried on, by, or within a transport medium. The transport medium may be non-temporary, for example, a storage medium, in particular a computer-readable storage medium. The transport medium may also be temporary, for example, a signal or other transmission medium. The signal may be transmitted over any suitable network, including the Internet.
[0101] A further aspect of the present invention relates to a camera comprising an apparatus according to any of the apparatus embodiments described above. In one embodiment, the camera further comprises a zooming means. In one embodiment, the camera is adapted to indicate when the zooming means is operational and to signal an interpredictive mode in response to the indication that the zooming means is operational. In another embodiment, the camera further comprises a panning means. In another embodiment, the camera is adapted to indicate when the panning means is operational and to signal an interpredictive mode in response to the indication that the panning means is operational.
[0102] According to yet another aspect of the present invention, a mobile device is provided comprising a camera embodying any of the above camera configurations. In one embodiment, the mobile device further comprises at least one position sensor adapted to sense a change in the orientation of the mobile device. In one embodiment, the mobile device is adapted to transmit an interpredictive mode signal depending on the sensing of the change in the orientation of the mobile device.
[0103] Further features of the present invention are characterized by other independent and dependent claims.
[0104] Any feature in one aspect of the present invention may be applied to other aspects of the present invention in any suitable combination. In particular, a method aspect may be applied to an apparatus aspect, and vice versa. Furthermore, a feature implemented in hardware may be implemented in software, and vice versa. Accordingly, any reference in this specification to software and hardware features should be interpreted as such. Any apparatus feature described herein may also be provided as a method feature, and vice versa. As used herein, means-plus-function features may be expressed alternatively with respect to their corresponding structures, such as a appropriately programmed processor and associated memory.
[0105] Furthermore, it should be understood that specific combinations of the various features described and defined in any embodiment of the present invention can be implemented and / or supplied and / or used independently. [Brief explanation of the drawing]
[0106] For example, please refer to the attached drawing. [Figure 1] Figure 1 is a diagram used to explain the coding structure used in HEVC. [Figure 2] Figure 2 is a schematic block diagram showing a data communication system that can implement one or more embodiments of the present invention. [Figure 3] Figure 3 is a block diagram showing the components of a processing apparatus that can carry out one or more embodiments of the present invention. [Figure 4] Figure 4 is a flowchart showing the steps of the encoding method according to an embodiment of the present invention. [Figure 5] Figure 5 is a flowchart showing the steps of the decoding method according to an embodiment of the present invention. [Figure 6a] Figure 6a shows spatial and temporal blocks that can be used to generate motion vector predictors. [Figure 6b] Figure 6b shows spatial and temporal blocks that can be used to generate motion vector predictors. [Figure 7] Figure 7 shows the simplified steps of the AMVP predictor set derivation process. [Figure 8] Figure 8 is a schematic diagram of the motion vector derivation process in merge mode. [Figure 9] Figure 9 shows the current block segmentation and time motion vector prediction. [Figure 10] Figure 10(a) shows the encoding of the merge index for HEVC, or when ATMVP is not enabled at the SPS level. Figure 10(b) shows the encoding of the merge index when ATMVP is enabled at the SPS level. [Figure 11] Figure 11(a) shows a simple affine motion field. Figure 11(b) shows a more complex affine motion field. [Figure 12] Figure 12 is a flowchart of the partial decoding process for several syntax elements related to the encoding mode. [Figure 13] Figure 13 is a flowchart showing the derivation of merge candidates. [Figure 14] Figure 14 shows the encoding of a merge index according to the first embodiment of the present invention. [Figure 15]Figure 15 is a flowchart of the partial decoding process of several syntax elements related to the coding mode in the twelfth embodiment of the present invention. [Figure 16] Figure 16 is a flowchart showing the generation of a list of merge candidates in a twelfth embodiment of the present invention. [Figure 17] Figure 17 is a block diagram used to illustrate a CABAC encoder suitable for use in embodiments of the present invention. [Figure 18] Figure 18 is a schematic block diagram of a communication system for implementing one or more embodiments of the present invention. [Figure 19] Figure 19 is a schematic block diagram of the computing device. [Figure 20] Figure 20 shows a network camera system. [Figure 21] Figure 21 shows a smartphone. [Figure 22] Figure 22 is a flowchart of the partial decoding process of several syntax elements related to the coding mode according to the 16th embodiment. [Figure 23] Figure 23 is a flowchart showing the use of a single-index signaling scheme for both merge mode and affine merge mode according to the embodiment. [Figure 24] Figure 24 is a flowchart showing the affine merge candidate derivation process in the affine merge mode according to the embodiment. [Figure 25] Figures 25(a) and 25(b) show the predictor derivation process for the triangle merge mode according to one embodiment. [Figure 26] Figure 26 is a flowchart of the decoding process for the interprediction mode for the current coding unit according to one embodiment. [Figure 27] Figure 27(a) shows the encoding of flags for merging in motion vector difference (MMVD) merge mode according to one embodiment. Figure 27(b) shows the encoding of an index for triangle merge mode according to one embodiment. [Figure 28] Figure 28 is a flowchart showing the affine merge candidate derivation process for the affine merge mode using ATMVP candidates according to one embodiment. [Figure 29] Figure 29 is a flowchart of the decoding process in interprediction mode according to the 18th embodiment. [Figure 30] Figure 30(a) shows the encoding of flags for merging in motion vector difference (MMVD) merge mode according to the 19th embodiment. Figure 30(b) shows the encoding of indexes for triangle merge mode according to the 19th embodiment. Figure 30(c) shows the encoding of indexes for affine merge mode or merge mode according to the 19th embodiment. [Figure 31] Figure 31 is a flowchart of the decoding process in the interprediction mode according to the 19th embodiment. [Modes for carrying out the invention]
[0107] The embodiments of the present invention described below relate to improving the encoding and decoding of indexes / flags / information / data using CABAC. It should be understood that alternative embodiments of the present invention may also be implemented to improve other context-based arithmetic coding schemes functionally similar to CABAC. Before describing the embodiments, video encoding and decoding techniques, as well as related encoders and decoders, will be discussed.
[0108] In this specification, “signaling” may mean inserting (providing / including / encoding) or extracting / retrieving (decoding) bitstream information relating to one or more syntax elements that represent use, deactivation, activation or deactivation of a mode (e.g., interpredictive mode) or other information (such as information about a selection).
[0109] Figure 1 illustrates the coding structure used in the High Efficiency Video Coding (HEVC) video standard. Video sequence 1 consists of a series of digital images i. Each such digital image is represented by one or more matrices. The matrix coefficients represent pixels.
[0110] Image 2 in this sequence is divided into slice 3. A slice may constitute the entire image. These slices are divided into non-overlapping coding tree units (CTUs). A coding tree unit (CTU) is a fundamental processing unit of the High Efficiency Video Coding (HEVC) video standard, and conceptually, its structure corresponds to the macroblock units used in some earlier video standards. A CTU is sometimes also called a Large Coding Unit (LCU). A CTU has luminance and chrominance component parts, each of which is called a coding tree block (CTB). These different color components are not shown in Figure 1.
[0111] A CTU is typically 64 pixels x 64 pixels for HEVC, but for VVC, this size can be 128 pixels x 128 pixels. Each CTU may be sequentially and iteratively divided into smaller variable-size coding units (CUs) 5 using quadtree decomposition.
[0112] A coding unit is a fundamental coding element and consists of two types of subunits called prediction units (PUs) and transformation units (TUs). The maximum size of a PU or TU is equal to the CU size. A prediction unit corresponds to a partition of a CU for predicting pixel values. Various different partitions of a CU into PUs are possible, including partitions into four square PUs and two different partitions into two rectangular PUs, as shown in 6. A transformation unit is a fundamental unit that performs spatial transformations using DCTs. A CU can be partitioned into TUs based on a quadtree representation 7. Thus, a slice, tile, CTU / LCU, CTB, CU, PU, TU, or block of pixels / samples is sometimes called an image portion, i.e., a part of the image 2 of a sequence.
[0113] Each slice is embedded in a single Network Abstraction Layer (NAL) unit. Furthermore, the encoding parameters of a video sequence are stored in a dedicated NAL unit called a parameter set. HEVC and H.264 / AVC use two types of parameter set NAL units: firstly, the Sequence Parameter Set (SPS) NAL unit, which collects all parameters that do not change throughout the entire video sequence. Typically, it handles the encoding profile, video frame size, and other parameters. Secondly, the Picture Parameter Set (PPS) NAL unit contains parameters that can change from one image (or frame) to another in the sequence. HEVC also includes a Video Parameter Set (VPS) NAL unit, which contains parameters that describe the overall structure of the bitstream. VPS is a new type of parameter set defined in HEVC and applies to all layers of the bitstream. A layer can contain multiple temporal sublayers, while all version 1 bitstreams are limited to one layer. HEVC has specific layer extensions for extensibility and multi-view, which enable multiple layers with backward-compatible version 1 base layers.
[0114] Figures 2 and 18 illustrate a data communication system that can implement one or more embodiments of the present invention. The data communication system includes a transmitting device, for example, a server 201 in Figure 2 or a content provider 150 in Figure 18, which is capable of transmitting data packets of a data stream 204 (or bitstream 101 in Figure 18) over a data communication network 200 to a receiving device, for example, a client terminal 202 in Figure 2 or a content consumer 100 in Figure 18. The data communication network 200 may be a wide area network (WAN) or a local area network (LAN). Such a network may be, for example, a wireless network (Wifi / 802.11a or b or g), an Ethernet network, an Internet network, or a mixed network consisting of several different networks. In a particular embodiment of the present invention, the data communication system may be a digital television broadcasting system in which a server 201 (or content provider 150 in Figure 18) transmits the same data content to multiple clients (or content consumers).
[0115] The data stream 204 (or bitstream 101) provided by the server 201 (or content provider 150) may consist of multimedia data representing video and audio data. In some embodiments of the present invention, the audio and video data streams may be captured by the server 201 (or content provider 150) using a microphone and a camera, respectively. In some embodiments, the data streams may be stored on the server 201 (or content provider 150), received by the server 201 (or content provider 150) from another data provider, or generated by the server 201 (or content provider 150). The server 201 (or content provider 150) particularly includes an encoder for encoding the video and audio streams (e.g., the original sequence 151 of the images in Figure 18) in order to provide a compressed bitstream 204,101 for transmission, which is a more compact representation of the data presented as input to the encoder.
[0116] To obtain a better ratio of the quality to the amount of transmitted data, video data compression may follow, for example, the HEVC format, or the H.264 / AVC format, or the VVC format.
[0117] The client 202 (or content consumer 100) receives the transmitted bitstream, decodes the reconstructed bitstream, and plays the video image (e.g., the video signal 109 in Figure 18) on the display device and plays the audio data through the speaker.
[0118] While the examples in Figure 2 or Figure 18 consider a streaming scenario, it will be understood that in some embodiments of the present invention, data communication between the encoder and decoder may be performed using a storage medium such as an optical disc.
[0119] In one or more embodiments of the present invention, the video image may be transmitted along with data representing a compensation offset to be applied to the reconstructed pixels of the image to provide filtered pixels in the final image.
[0120] Figure 3 schematically shows a processing apparatus 300 configured to carry out at least one embodiment of the present invention. The processing apparatus 300 may be a device such as a microcomputer, a workstation, or a light portable device. The device 300 is -Central processing unit 311 such as a microprocessor indicated by CPU - A read-only memory 307, referred to as ROM, for storing a computer program for carrying out the present invention. - Random access memory 312, indicated by RAM, which stores registers configured to record variables and parameters necessary to implement the executable code for the method of an embodiment of the present invention, and the method for encoding a sequence of digital images and / or decoding a bitstream according to an embodiment of the present invention. - A communication interface 302 connected to a communication network 303 through which digital data to be processed is sent and received. It is equipped with a communication bus 313 connected to it.
[0121] Optionally, the device 300 may also include the following components:
[0122] - A computer program for carrying out one or more embodiments of the present invention, and a data storage means such as a hard disk for storing data used or generated during the implementation of one or more embodiments of the present invention. - A disk drive 305 for disk 306, the disk drive is configured to read data from disk 306 or write data to disk. - Screen 309, which displays data and / or functions as a graphical interface with the user, using a keyboard 310 or any other pointing / input means. The device 300 can be connected to various peripheral devices, such as a digital camera 320 or a microphone 308, and each peripheral device is connected to an input / output card (not shown) to supply multimedia data to the device 300.
[0123] The communication bus provides communication and interoperability between various elements included in or connected to the device 300. The representation of the bus is not limited, and in particular, the central processing unit can operate to communicate commands to any element of the device 300, either directly or through another element of the device 300.
[0124] The disk 306 can be replaced with any information medium, such as a compact disc (CD-ROM), rewritable or non-rewritable ZIP disk or memory card, and more generally, it can be replaced with information storage means that can be read by a microcomputer or microprocessor, and can be integrated or unintegrated into the device, and, if possible, removable, and can be configured to store one or more programs that enable execution of a method for encoding a sequence of digital images and / or a method for decoding a bitstream according to the present invention.
[0125] The executable code can be stored in read-only memory 307, hard disk 304, or a removable digital medium such as disk 306 as described earlier. In a modified example, the program's executable code can be received by the communication network 303 via interface 302 to be stored in one of the storage means of device 300, such as hard disk 304, before execution.
[0126] The central processing unit 311 is configured to control and direct the execution of instructions or parts of the program or software code of the program according to the present invention using instructions stored in one of the aforementioned storage means. When power is turned on, the program or program stored in non-volatile memory, for example, on the hard disk 304 or disk 306 or in read-only memory 307, is transferred to the random access memory 312, and the random access memory 213 includes registers for storing the program or executable code of the program, as well as variables and parameters necessary to carry out the present invention.
[0127] In this embodiment, the device is a programmable device that uses software to carry out the present invention. However, alternatively, the present invention may be carried out in hardware (for example, in the form of an application-specific integrated circuit or ASIC).
[0128] Figure 4 shows a block diagram of an encoder according to at least one embodiment of the present invention. The encoder is represented by connected modules, each module adapted to perform at least one corresponding step of a method for implementing at least one embodiment of the present invention for encoding an image of a sequence of images according to one or more embodiments of the present invention, in the form of program instructions to be executed, for example, by the CPU 311 of device 300.
[0129] The original sequence of digital images i0 to in401 is received as input by encoder 400. Each digital image is represented by a set of samples, sometimes also called pixels (hereinafter referred to as pixels).
[0130] The bitstream 410 is output by the encoder 400 after the encoding process has been performed. The bitstream 410 comprises a plurality of encoding units or slices, each slice comprising a slice header for transmitting encoded values of encoding parameters used to encode the slice, and a slice body containing encoded video data.
[0131] The input digital images i0 to in401 are divided into blocks of pixels by module 402. The blocks correspond to parts of the image and may be of variable size (e.g., 4x4, 8x8, 16x16, 32x32, 64x64, 128x128 pixels, and several rectangular block sizes can also be considered). An encoding mode is selected for each input block. Two families of encoding modes are provided: those based on spatial prediction coding (intra-prediction) and those based on temporal prediction (inter-coding, merge, SKIP). Possible encoding modes are tested.
[0132] Module 403 performs an intra-prediction process in which a given block to be encoded is predicted by predictors calculated from neighboring pixels of the block to be encoded. The indication of the selected intra-predictor, and the difference between the given block and its predictor, are encoded to provide a residual if intra-coding is selected.
[0133] Temporal prediction is performed by motion estimation module 404 and motion compensation module 405. First, a reference image is selected from the set of reference images 416, and the portion of the reference image, also called the reference region or image portion, that is closest to a given block to be encoded (closest in terms of pixel value similarity) is selected by the motion estimation module 404. Next, the motion compensation module 405 uses the selected region to predict the block to be encoded. The difference between the selected reference region and the given block, also called the residual block, is calculated by the motion compensation module 405. The selected reference region is represented using a motion vector.
[0134] Therefore, in both cases (spatial and temporal prediction), the residual is calculated by subtracting the predictor from the original block when the original block is not in SKIP mode.
[0135] In the INTRA prediction performed by module 403, the prediction direction is encoded. In the inter prediction performed by modules 404, 405, 416, 418, and 417, at least one motion vector or data is encoded for temporal prediction to identify such motion vectors.
[0136] If interpretation is selected, information related to the motion vector and residual block is encoded. To further reduce the bitrate, assuming uniform motion, the motion vector is encoded by the difference with respect to the motion vector predictor. The motion vector predictor from a pair of motion information predictor candidates is obtained from the motion vector field 418 by the motion vector prediction coding module 417.
[0137] The encoder 400 further includes a selection module 406 for selecting an encoding mode by applying encoding cost criteria such as rate distortion criteria. To further reduce redundancy, a transformation module 407 applies a transformation (such as DCT) to the residual block, the resulting transformed data is quantized by a quantization module 408, and entropy encoded by an entropy encoding module 409. Finally, when the encoded residual block of the currently encoded block is not in SKIP mode, it is inserted into the bitstream 410, and the mode requires the residual block to be encoded in the bitstream.
[0138] Furthermore, encoder 400 performs decoding of the encoded image to generate reference images (e.g., those in reference image / picture 416) for motion estimation of subsequent images. This allows the encoder and decoder receiving the bitstream to have the same reference frame (reconstructed image or image portion is used). The dequantization module 411 performs dequantization of the quantized data, followed by inverse transformation by the inverse transformation module 412. The intra-prediction module 413 uses the prediction information to determine which predictor to use for a given block, and the motion compensation module 414 actually adds the residuals obtained by module 412 to the reference region obtained from the set of reference images 416.
[0139] Subsequently, post-filtering is applied by module 415, filtering the reconstructed frame (image or portion of an image) of pixels. In embodiments of the present invention, an SAO loop filter is used in which a compensation offset is added to the pixel values of the reconstructed pixels in the reconstructed image. It is understood that post-filtering is not necessarily required. In addition, any other type of post-filtering can be performed in addition to or instead of SAO loop filtering.
[0140] Figure 5 shows a block diagram of a decoder 60 that may be used to receive data from an encoder, according to one embodiment of the present invention. The decoder is represented by connected modules, each module configured to implement a corresponding step of the method implemented by the decoder 60, for example, by forming program instructions executed by the CPU 311 of device 300.
[0141] Decoder 60 receives a bitstream 61 containing encoding units (e.g., data corresponding to image portions, blocks, or encoding units), each encoding unit consisting of a header containing information about encoding parameters and a body containing encoded video data. As described with reference to Figure 4, the encoded video data is entropy encoded, and the indices of the motion vector predictors are encoded with a predetermined number of bits for a given image portion (e.g., a block or CU). The received encoded video data is entropy decoded by module 62. The residual data is then inversely quantized by module 63, and then an inverse transform is applied by module 64 to obtain pixel values.
[0142] Mode data indicating the encoding mode is also entropy-decoded, and based on this mode, INTRA-type decoding or INTER-type decoding is performed on the encoded blocks (units / sets / groups) of the image data.
[0143] In INTRA mode, the INTRA predictor is determined by the INTRA prediction module 65 based on the INTRA prediction mode specified in the bitstream.
[0144] When the mode is INTER, motion prediction information is extracted from the bitstream to find (identify) the reference region used by the encoder. The motion prediction information includes the reference frame index and the motion vector residual. The motion vector predictor is added to the motion vector residual by the motion vector decoding module 70 to obtain the motion vector.
[0145] The motion vector decoding module 70 applies motion vector decoding to each image portion (e.g., the current block or CU) encoded by motion prediction. Once the index of the motion vector predictor for the current block is obtained, the actual value of the motion vector associated with the image portion (e.g., the current block or CU) is decoded and can be used by module 66 to apply motion compensation. The reference image portion indicated by the decoded motion vector is extracted from the reference image 68 and motion compensation 66 is applied. The motion vector field data 71 is updated with the decoded motion vector for use in predicting subsequent decoded motion vectors.
[0146] Finally, the decoded block is obtained. If appropriate, post-filtering is applied by the post-filtering module 67. The decoded video signal 69 is finally obtained and given by the decoder 60.
[0147] CABAC HEVC uses several types of entropy coding, including CABAC (Context-based Adaptive Binary Arithmetic Coding), Golomb-rice coding, or a simple binary representation called Fixed Length Coding. In most cases, binary coding is performed to represent different syntax elements. This binary coding is also very specific and depends on different syntax elements. Arithmetic coding represents syntax elements according to their current probabilities. CABAC is an extension of arithmetic coding that separates the probabilities of syntax elements according to the "context" defined by a context variable. This is equivalent to conditional probability. The context variable can be derived from the current syntax values of the upper left block (A2 in Figure 6b, which will be explained in detail below) and the upper left block (B3 in Figure 6b), which have already been coded.
[0148] CABAC is adopted as a base part of the H.264 / AVC and H.265 / HEVC standards. In H.264 / AVC, it is one of two alternative methods of entropy coding. The other method specified in H.264 / AVC is a low-complexity entropy coding technique based on the use of context-adaptively switched sets of variable-length codes, so-called context-adaptive variable-length coding (CAVLC). Compared to CABAC, CAVLC offers reduced implementation costs at the expense of lower compression efficiency. For standard-definition or high-definition resolution TV signals, CABAC typically offers a 10-20% bitrate saving compared to CAVLC at the same objective video quality. In HEVC, CABAC is one of the entropy coding methods used. Many bits are also bypassed with CABAC coding (also expressed as CABAC bypass coding). Furthermore, some syntax elements are coded with unary codes or Golomb codes, which are other types of entropy codes.
[0149] Figure 17 shows the main block of the CABAC encoder.
[0150] Input syntax elements, which are non-binary values, are binarized by the binaryr 1701. CABAC's coding strategy is based on the discovery that highly efficient coding of syntax element values in a hybrid block-based video encoder, such as motion vector difference or transformation coefficient level value components, can be achieved by using a binarization scheme as a kind of preprocessing unit for the subsequent stages of context modeling and binary arithmetic coding. In general, a binarization scheme defines a unique mapping of a sequence of binary decisions of syntax element values to so-called bins, which can be "bits" and therefore can also be interpreted in terms of a binary code tree. The design of the binarization scheme in CABAC is based on a few basic prototypes that allow its structure to be easily computed online and applied to several suitable model-probability distributions.
[0151] Each bin can be processed in one of two basic ways, depending on the setting of switch 1702. When the switch is in the "regular" setting, the bin is fed to the context modeler 1703 and the regular encoding engine 1704. When the switch is in the "bypass" setting, the context modeler is bypassed and the bin is fed to the bypass encoding engine 1705. Another switch 1706 has "regular" and "bypass" settings similar to switch 1702, and as a result, the bin encoded by one of the applicable encoding engines 1704 and 1705 can form a bitstream as the output of the CABAC encoder.
[0152] It is understood that other switches 1706 may be used with storage to group some of the bins encoded by encoding engine 1705 (e.g., bins for encoding image portions such as blocks or encoding units) to provide blocks of bypass-encoded data in the bitstream, and to group some of the bins encoded by encoding engine 1704 (e.g., bins for encoding blocks or encoding units) to provide another block of "canonical" (or arithmetic)-encoded data in the bitstream. This separate grouping of bypass-encoded data and canonical-encoded data can result in improved throughput during decoding (since bypass-encoded data can be processed first / in parallel with canonical CABAC-encoded data).
[0153] By decomposing each syntax element value into a sequence of bins, further processing of each bin value in CABAC depends on the associated coding mode decision, which can be selected as either normal mode or bypass mode. The latter is selected for lower-significant bins that are assumed to be uniformly distributed and therefore simply bypass the entire normal binary arithmetic coding process. In normal coding mode, each bin value is coded using a normal binary arithmetic coding engine, and the associated probabilistic model is determined by a fixed selection without contextual modeling or adaptively selected depending on the associated contextual model. As a key design decision, in the latter case, it is generally applied only to the most frequently observed bins, while other bins that are not usually observed are handled using joints, typically zero-order probabilistic models. In this way, CABAC enables selective contextual modeling at the sub-symbol level and thus provides an efficient means of leveraging inter-symbol redundancy with a significant reduction in overall modeling or learning costs. For specific selections of contextual models, four basic design types were adopted in CABAC, two of which were applied to coding at the transformation coefficient level only. The designs of these four prototypes are based on a priori knowledge of the typical characteristics of the source data being modeled, and reflect the objective of finding a good compromise between the conflicting goals of avoiding unnecessary modeling overhead and making significant use of statistical dependencies.
[0154] In CABAC's lowest level of processing, each bin value enters a binary arithmetic encoder in either normal or bypass coding mode. In the latter case, the fast branch of the coding engine, which has significantly reduced complexity, is used, while in the former coding mode, the coding of a given bin value depends on the actual state of the associated adaptive probabilistic model passed to the M coder along with the bin value, a term chosen for the table-based adaptive binary arithmetic coding engine in CABAC.
[0155] Next, the corresponding CABAC decoder receives the bitstream output from the CABAC encoder and processes the bypass-encoded data and the regular CABAC-encoded data accordingly. Once the CABAC decoder processes the regular CABAC-encoded data, the context modeler (and its probabilistic model) is updated, allowing it to correctly decode / process (e.g., inverse binarization) the bins that make up the bitstream to obtain the syntax elements.
[0156] Intercoding HEVC uses three distinct intermodes: intermode (Advanced Motion Vector Prediction (AMVP) which signals motion information differences), "classical" merge mode (also known as "non-affine merge mode" or "normal" merge mode which does not signal motion information differences), and "classical" merge-skip mode (also known as "non-affine merge-skip" mode or "normal" merge-skip mode which does not signal motion information differences and also does not signal residual data of sample values). The main difference between these modes is the data signaling in the bitstream. For motion vector coding, the current HEVC standard includes a competition-based scheme for motion vector prediction that was not present in previous versions of the standard. This means that for each inter-coding mode (AMVP) or merge mode (i.e., "classical / normal" merge mode or "classical / normal" merge-skip mode), several candidates compete with the encoder-side rate-distortion criterion to find the best motion vector predictor or best motion information. An index or flag corresponding to the best candidate for the best predictor or motion information is then inserted into the bitstream. The decoder can derive the same set of predictors or candidates and use the best one according to the decoded index / flag. In HEVC's Screen Content Extension, a new coding tool called Intra-Block Copy (IBC) is signaled as one of these three intermodes, and the difference between IBC and equivalent intermodes is done by checking whether the reference frame is the current one. IBC is also known as Current Image Reference (CPR). This can be done, for example, by checking the reference index in list L0 and assuming that this is an intra-block copy if it is the last frame in the list. Another way is to compare the picture order count of the current frame and the reference frame, and if they are equal, this is an intra-block copy.
[0157] The design of predictor and candidate derivations is crucial for achieving the best coding efficiency without disproportionately impacting complexity. HEVC uses two motion vector derivations: one for intermode (Advanced Motion Vector Prediction (AMVP)) and one for merge mode (Merge derivation process - for the classical Merge mode and the classical Merge Skip mode). These processes are described below.
[0158] Figures 6a and 6b show spatial and temporal blocks that can be used, for example, to generate motion vector predictors in the advanced motion vector prediction (AMVP) and merge modes of an HEVC coding and decoding system, and Figure 7 shows simplified steps of the process for deriving an AMVP predictor set.
[0159] Two spatial predictors, namely two spatial motion vectors for AMVP mode, are selected from the motion vectors of the upper block (indicated by the letter "B") and left block (indicated by the letter "A"), including the upper corner block (block B2) and the left corner block (block A0), and one temporal predictor is selected from the motion vectors of the lower right block (H) and the center block (Center) of the collated block, as shown in Figure 6a.
[0160] Table 1 below outlines the naming convention used when referring to blocks relative to the current block, as shown in Figures 6a and 6b. While this naming convention is used for brevity, it should be understood that other labeling systems may be used, particularly in future versions of the standard.
[0161] [Table 1]
[0162] It should be noted that the "current block" is variable in size, such as 4x4, 16x16, 32x32, 64x64, 128x128, or any size in between. The dimensions of the block are preferably multiples of 2 (i.e., 2^n × 2^m, where n and m are positive integers), which results in more efficient use of bits when using binary coding. The current block does not have to be square, but this is often a preferred embodiment due to the complexity of the coding.
[0163] Referring to Figure 7, the first step aims to select a first spatial predictor (Cand1, 706) from blocks A0 and A1 in the lower left, whose spatial location is shown in Figure 6a. To this end, these blocks are selected one after another in a given (i.e., predetermined / preset) order (700, 702), and for each selected block, the following conditions are evaluated in a given order (704), and the first block that satisfies the conditions is set as the predictor.
[0164] - Motion vector from the same reference image and the same reference list - Motion vectors from the same reference image and other reference lists - Scaled motion vectors from different reference images and the same reference list - Scaled motion vectors from different reference images and other reference lists If no value is found, the left predictor is considered unavailable. This indicates that the associated blocks are intra-encoded or that those blocks do not exist.
[0165] The subsequent steps aim to select a second spatial predictor (Cand2, 716) from the upper right block B0, the upper block B1, and the upper left (upper left) block B2, whose spatial positions are shown in Figure 6a. To this end, these blocks are selected one after another in a given order (708, 710, 712), and for each selected block, the above conditions are evaluated in a given order (714), and the first block that satisfies the above conditions is set as the predictor.
[0166] Again, if the value is not found, the predictor above is considered unavailable. In this case, it indicates that the associated blocks are intra-encoded or that those blocks do not exist.
[0167] In the next step (718), the two predictors are compared to each other if both are available, and if they are equal (i.e., same motion vector value, same reference list, same reference index, and same direction type), then one of them is eliminated. If only one spatial predictor is available, the algorithm searches for a temporal predictor in a later step.
[0168] The temporal motion predictor (Cand3,726) is derived as follows: The lower right (H,720) position of the collated block in the previous / referenced frame is first considered in the availability check module 722. If it does not exist, or if the motion vector predictor is not available, the center (center,724) of the collated block is selected to be checked. These temporal positions (center and H) are shown in Figure 6a. In any case, scaling 723 is applied to these candidates to match the temporal distance between the current frame and the first frame in the reference list.
[0169] Next, the motion predictor value is added to the set of predictors. Then, the number of predictors (Nb_Cand) is compared to the maximum number of predictors (Max_Cand) (728). As mentioned above, the maximum number of predictors (Max_Cand) for motion vector predictors that the AMVP derivation process needs to generate is 2 in the current version of the HEVC standard.
[0170] If this maximum number is reached, the final list or set (732) of AMVP predictors is constructed. Otherwise, a zero predictor is added to the list (730). A zero predictor is a motion vector equal to (0,0).
[0171] As shown in Figure 7, the final list or set (732) of AMVP predictors is constructed from a subset of candidate spatial motion predictors (700-712) and a subset of candidate temporal motion predictors (720, 724).
[0172] As described above, motion predictor candidates in classical merge mode or classical merge skip mode can represent all the necessary motion information: direction, list, reference frame index, and motion vector (or any subset thereof for performing the prediction). An indexed list of several candidates is generated by the merge derivation process. In the current HEVC design, the maximum number of candidates for both merge modes (i.e., classical merge mode and classical merge skip mode) is equal to 5 (four spatial candidates and one temporal candidate).
[0173] Figure 8 is a schematic diagram of the motion vector derivation process for merge modes (classical merge mode and classical merge skip mode). In the first step of the derivation process, five block locations are considered (800-808). These locations are the spatial locations shown in Figure 6a with reference numbers A1, B1, B0, A0, and B2. In a later step, the availability of spatial motion vectors is checked, and at most five motion vectors are selected / acquired for consideration (810). If a predictor exists and the block is not intra-encoded, the predictor is considered available. Therefore, the selection of motion vectors corresponding to the five blocks as candidates is performed according to the following conditions.
[0174] If the “left” A1 motion vector (800) is available (810), i.e., it exists and this block is not intra-encoded, the motion vector for the “left” block is selected and used as the first candidate in the candidate list (814).
[0175] If the "up" B1 motion vector (802) is available (810), the candidate "up" block motion vector is compared to the "left" A1 motion vector, if it exists (812). If the B1 motion vector is equal to the A1 motion vector, B1 is not added to the list of spatial candidates (814). Conversely, if the B1 motion vector is not equal to the A1 motion vector, B1 is added to the list of spatial candidates (814).
[0176] If the "upper right" B0 motion vector (804) is available (810), the "upper right" motion vector is compared to the B1 motion vector (812). If the B0 motion vector is equal to the B1 motion vector, the B0 motion vector is not added to the list of spatial candidates (814). Conversely, if the B0 motion vector is not equal to the B1 motion vector, the B0 motion vector is added to the list of spatial candidates (814).
[0177] If the "bottom left" A0 motion vector (806) is available (810), the "bottom left" motion vector is compared to the A1 motion vector (812). If the A0 motion vector is equal to the A1 motion vector, the A0 motion vector is not added to the list of spatial candidates (814). Conversely, if the A0 motion vector is not equal to the A1 motion vector, the A0 motion vector is added to the list of spatial candidates (814).
[0178] If the list of spatial candidates does not contain four candidates, the availability of the "top-left" B2 motion vector (808) is checked (810). If available, it is compared with the A1 motion vector and the B1 motion vector. If the B2 motion vector is equal to either the A1 motion vector or the B1 motion vector, the B2 motion vector is not added to the list of spatial candidates (814). Conversely, if the B2 motion vector is not equal to either the A1 motion vector or the B1 motion vector, the B2 motion vector is added to the list of spatial candidates (814).
[0179] At the end of this stage, the list of spatial candidates will contain up to four candidates.
[0180] Two temporal candidates can be used: the lower right position of the collated block (816, H shown in Figure 6a) and the center of the collated block (818). These positions are shown in Figure 6a.
[0181] As described in relation to Figure 7 for the temporal motion predictor of the AMVP motion vector derivation process, the first step is to check the availability of the block at position H (820). Next, if it is not available, the availability of the block at the central position is checked (820). If at least one motion vector is available for these positions, the temporal motion vector may be scaled to a reference frame with index 0 for both lists L0 and L1, if necessary, to create a temporal candidate (824) that is added to the list of merge motion vector predictor candidates (822). This is placed after the spatial candidates in the list. Lists L0 and L1 are two reference frame lists containing zero, one or more reference frames.
[0182] A combined candidate is generated (828) if the number of candidates (Nb_Cand) is strictly less than the maximum number of candidates (826) (Max_Cand information for which the value is signaled in the bitstream slice header and is determined to be equal to 5 in the current HEVC design), and the current frame is of type B. The combined candidate is generated based on the available candidates in the list of merge motion vector predictor candidates. This mainly consists of combining (pairing) the motion information of one candidate in list L0 with the motion information of one candidate in list L1.
[0183] If the number of candidates (Nb_Cand) is strictly less than the maximum number of candidates (Max_Cand) (830), zero motion candidates are generated until the number of candidates in the list of merge motion vector predictor candidates reaches the maximum number of candidates (832).
[0184] At the end of this process, a list or set of merge motion vector predictor candidates is constructed (i.e., a list or set of merge mode candidates, which are classical merge modes and classical merge skip modes) (834). As shown in Figure 8, the list or set of merge motion vector predictor candidates is constructed from a subset of spatial candidates (800-808) and a subset of temporal candidates (816, 818) (834).
[0185] Alternative Time-Motion Vector Prediction (ATMVP) Alternative Temporal Motion Vector Prediction (ATMVP) is a special type of motion compensation. Instead of considering only one motion information for the current block from the temporal reference frame, it considers each motion information for each collated block. Thus, this temporal motion vector prediction gives the segmentation of the current block using the relevant motion information for each subblock, as shown in Figure 9.
[0186] In the VTM reference software, ATMVP is signaled as a merge candidate inserted into the list of merge candidates (i.e., a list or set of candidates for merge modes, which are the classic merge mode and the classic merge skip mode). When ATMVP is enabled at the SPS level, the maximum number of merge candidates is increased by 1. Thus, 6 candidates are considered instead of 5, which was the case when this ATMVP mode was disabled. According to one embodiment of the present invention, it is understood that ATMVP may be signaled as an affine merge candidate inserted into the list of affine merge candidates (e.g., ATMVP candidates) (i.e., a separate list or set of candidates for affine merge modes, which will be described in more detail below).
[0187] Furthermore, when this prediction is enabled at the SPS level, all bins of the merge index (i.e., identifiers, indices, or information for identifying candidates from the list of merge candidates) are context-encoded by CABAC. While in HEVC, or when ATMVP is not enabled at the SPS level in JEM, only the first bin is context-encoded, and the remaining bins are context-bypass encoded (i.e., bypass CABAC encoded).
[0188] Figure 10(a) shows the encoding of the merge index when ATMVP is not enabled at the SPS level of HEVC or JEM. This corresponds to the unary maximal code. Furthermore, in Figure 10(a), the first bit is CABAC encoded, and the other bits are bypassed CABAC encoded.
[0189] Figure 10(b) shows the encoding of the merge index when ATMVP is enabled at the SPS level. All bits are CABAC encoded (bits 1 through 5). Note that each bit used to encode the index has its own context, in other words, their probabilities used in CABAC encoding are isolated.
[0190] Affine Mode HEVC only applied translational motion models for motion compensation prediction (MCP). In contrast, the real world has many types of motion, including zoom in / zoom out, rotation, perspective motion, and other irregular motions.
[0191] In JEM, a simplified affine transform motion compensation prediction is applied, and the general principles of affine modes are described below, based on excerpts from the JVET-G1001 document presented at the JVET conference in Torino from July 13-21, 2017. This entire document is incorporated herein by reference to the extent that it describes other algorithms used in JEM.
[0192] As shown in Figure 11(a), the affine motion field of a block in this document is described by two control point motion vectors (it should be understood that other affine models, such as those with more control point motion vectors, can also be used according to embodiments of the present invention).
[0193] The motion vector field (MVF) of a block is described by the following formula:
[0194]
number
[0195] Here, (v0x, v0y) is the motion vector of the control point in the upper left corner, and (v1x, v1y) is the motion vector of the control point in the upper right corner. Also, w is the width of block Cur (the current block).
[0196] To further simplify motion compensation prediction, we applied subblock-based affine transformation prediction. The subblock size MxN is derived as shown in Equation 2, where MvPre is the fractional precision of the motion vector (1 / 16 in JEM) and (v2x,v2y) is the motion vector of the upper-left control point calculated according to Equation 1.
[0197]
number
[0198] After being derived by Equation 2, M and N may be adjusted downwards if necessary so that they become divisors of w and h, respectively, where h is the height of the current block Cur (the current block).
[0199] To derive the motion vector for each M×N subblock, the motion vector of the central sample of each subblock is calculated according to Equation 1 and rounded to a fractional precision of 1 / 16, as shown in Figure 11(b). Next, a motion compensation interpolation filter is applied to generate predictions for each subblock with the derived motion vectors.
[0200] Affine mode is a motion compensation mode similar to intermode (AMVP, "classical" merge, or "classical" merge skip). Its principle is to generate one motion information per pixel according to two or three adjacent motion information. In JEM, affine mode derives one motion information for each 4x4 block as shown in Figure 11(b) (each square is a 4x4 block, and the entire block in Figure 11(b) is a 16x16 block divided into 16 such 4x4 squares, and each 4x4 square block has its associated motion vector). In embodiments of the present invention, it should be understood that affine mode can drive one motion information for blocks of different sizes or shapes, as long as it can derive one motion information.
[0201] According to one embodiment, this mode becomes available for AMVP mode and merge mode (i.e., classical merge mode, also called "non-affine merge mode," and classical merge skip mode, also called "non-affine merge skip mode") by enabling affine mode with a flag. This flag is CABAC encoded. In one embodiment, the context depends on the sum of the affine flags in the left block (position A2 in Figure 6b) and the upper left block (position B3 in Figure 6b).
[0202] Therefore, in JEM, there are three possible context variables (0, 1, or 2) for the affine flag given in the following expression.
[0203] Ctx=IsAffine(A2)+IsAffine(B3) Here, IsAffine(block) is a function that returns 0 if the block is not an affine block, and 1 if the block is affine.
[0204] Derivation of Affine Merge Candidates In JEM, the affine merge mode (or affine merge skip mode), also known as the subblock (merge) mode, derives motion information for the current block from the first neighboring block (i.e., the first neighboring block encoded using the affine mode) that is affine within the block at positions A1, B1, B0, A0, B2. These positions are shown in Figures 6a and 6b. However, how the affine parameters are derived is not fully defined, and the present invention aims to improve at least this aspect by, for example, defining the affine parameters of the affine merge mode, thereby enabling a wider selection of affine merge candidates (i.e., at least one other candidate is available for selection using identifiers such as an index, in addition to the first neighboring block that is affine).
[0205] For example, according to some embodiments of the present invention, an affine merge mode having its own list of affine merge candidates (candidates for deriving / acquiring motion information for an affine mode) and an affine merge index (for identifying one affine merge candidate from the list of affine merge candidates) are used to encode or decode a block.
[0206] Affine Merge Signaling Figure 12 is a flowchart of the partial decoding process for several syntax elements related to the coding mode to signal the use of affine merge mode. In this figure, the skip flag (1201), prediction mode (1211), merge flag (1203), merge index (1208), and affine flag (1206) can be decoded.
[0207] For all CUs in the interslice, the skip flag is decoded (1201). If the CU is not a skip (1202), the prediction mode is decoded (1211). This syntax element indicates whether the current CU is encoded (decoded) in inter-mode or intra-mode. Note that if the CU is a skip (1202), its current mode is inter-mode. If the CU is not a skip (1202: No), the CU is encoded in AMVP mode or merge mode. If the CU is inter (1212), the merge flag is decoded (1203). If the CU is merge (1204), or if the CU is a skip (1202: Yes), it is verified / checked whether the affine flag (1206) needs to be decoded, i.e., (1205), whether it is possible that the current CU was encoded in affine mode (1205). This flag is decoded if the current CU is a 2N×2N CU, meaning that in the current VVC, the height and width of the CUs are equal. Furthermore, at least one adjacent CU A1 or B1 or B0 or A0 or B2 must be encoded in affine mode (either affine merge mode or AMVP mode with affine mode enabled). Finally, the current CU is not a 4x4 CU, and by default, 4x4 CUs are disabled in the VTM reference software. If this condition (1205) is false, it is certain that the current CU is encoded in classical merge mode (or classical merge skip mode) as specified in HEVC, and the merge index is decoded (1208). If the affine flag (1206) is set to equal 1 (1207), then the CU is either a merged affine CU (i.e., a CU encoded in affine merge mode) or a merge-skip affine CU (i.e., a CU encoded in affine merge-skip mode), and the merge index (1208) does not need to be decoded (because affine merge mode is used, i.e., the CU is decoded using affine mode with an affine first adjacent block).Otherwise, the current CU is a classical (basic) merge or merge-skip CU (i.e., a CU encoded in classical merge or merge-skip mode), and the merge candidate index (1208) is decoded.
[0208] Derivation of merge candidates Figure 13 is a flowchart illustrating the derivation of merge candidates (i.e., candidates for classical merge mode or classical merge skip mode) according to one embodiment. This derivation is built upon the merge mode motion vector derivation process shown in Figure 8 (i.e., HEVC merge candidate list derivation). The main changes compared to HEVC are the addition of ATMVP candidates (1319, 1321, 1323), a full duplicate check of candidates (1325), and a new order of candidates. ATMVP predictions are set as dedicated candidates because they represent some motion information of the current CU. The value of the first sub-block (top left) is compared with the temporal candidates, and if the temporal candidates are equal, they are not added to the list of merge candidates (1320). ATMVP candidates are not compared with other spatial candidates. This is in contrast to temporal candidates, which are compared with each spatial candidate already in the list (1325) and are not added to the merge candidate list if they are duplicate candidates.
[0209] When a spatial candidate is added to the list, it is compared to other spatial candidates in the list (1312), which is not the case in the final version of HEVC.
[0210] In the current VTM version, the list of merge candidates is set in the following order when it is determined that they provide the best results for the encoding test conditions. A1 B1 B0 A0 ATMVP B2 · Temporal · combination · Zero_MV It is important to note that spatial candidate B2 is set after the ATMVP candidate.
[0211] Furthermore, when ATMVP is enabled at the slice level, the maximum number in the list of candidates is 6, not 5 for HEVC.
[0212] Other Interpretation Modes In the first few embodiments described below (up to the 16th embodiment), the description explains the encoding or decoding of indices for (canonical) merge mode and affine merge mode. In recent versions of the VVC standard under development, additional interprediction modes are also considered in addition to (canonical) merge mode and affine merge mode. Such additional interprediction modes currently considered are the Multi-Hypothesis Intra Inter (MHII) merge mode, triangle merge mode, and motion vector difference merge (MMVD) merge mode, which are described below.
[0213] It should be understood that, according to some of these first embodiments, in addition to or instead of merge mode or affine merge mode, one or more of additional interprediction modes can be used, and indices (or flags or information) for one or more of the additional interprediction modes can be signaled (encoded or decoded) using the same techniques as for either one of the merge mode or affine merge mode.
[0214] MHII (Multi-Hypothesis Intra Inter) Merge Mode The MHII (Multi-Hypothesis Intra Inter) merge mode is a hybrid combining the regular merge mode and the intra mode. The block predictor in this mode is obtained as the average between the (regular) merge mode block predictor and the intra mode block predictor. The resulting block predictor is added to the residual of the current block to obtain the reconstructed block. To obtain this merge mode block predictor, the MHII merge mode uses the same number of candidates and the same merge candidate derivation process as the regular merge mode. Therefore, index signaling for the MHII merge mode can use the same techniques as index signaling for the regular merge mode. Furthermore, this mode is only valid for blocks encoded / decoded in non-skip mode. Therefore, if the current CU is encoded / decoded in skip mode, MHII cannot be used in the encoding / decoding process.
[0215] Triangle Merge Mode The triangle merge mode is a type of dual prediction mode that uses motion compensation based on triangle shape. Figures 25(a) and 25(b) show different partition configurations used for its block predictor generation. Block predictors are derived from a first triangle (first block predictor 2501 or 2511) and a second triangle (second block predictor 2502 or 2512) within a block. Two different configurations are used for this block predictor generation. For the first, the partition / split between the triangular parts / regions (to which the two block predictor candidates are associated) is from the upper left corner to the lower right corner, as shown in Figure 25(a). For the second, the partition / split between the triangular regions (to which the two block predictor candidates are associated) is from the upper right corner to the lower left corner, as shown in Figure 25(b). Furthermore, samples around the boundary between the triangular regions are filtered by a weighted average whose weights depend on the sample location (e.g., distance from the boundary). An independent list of triangle merge candidates is generated, and index signaling for triangle merge mode can use techniques that are modified accordingly compared to the techniques for index signaling in merge mode or affine merge mode.
[0216] Merge Motion Vector Difference (MMVD) Merge Mode The MMVD merge mode is a special type of regular merge mode candidate derivation that generates an independent list of MMVD merge candidates. The selected MMVD merge candidate for the current CU is obtained by adding an offset value to one of the motion vector components (mvx or mvy) of the MMVD merge candidate. The offset value is added to the motion vector component from the first list L0 or the second list L1, depending on the configuration of these reference frames (e.g., both backward, forward, or both forward and backward). The selected MMVD merge candidate is signaled using an index. The offset value is signaled using a distance index between eight possible preset distances (1 / 4-pel, 1 / 2-pel, 1-pel, 2-pel, 4-pel, 8-pel, 16-pel, 32-pel) and a direction index that gives the x or y axis and sign of the offset. Thus, index signaling for the MMVD merge mode can use the same techniques as index signaling for the merge mode or affine merge mode.
[0217] Embodiment Embodiments of the present invention will be described with reference to the remaining drawings. Unless otherwise specified, embodiments may be combined, and for example, certain combinations of embodiments may improve coding efficiency at the expense of complexity, although this may be acceptable in specific use cases.
[0218] First Embodiment As described above, the VTM reference software signals ATMVP as a merge candidate that has been inserted into the list of merge candidates. ATMVP can be enabled or disabled for the entire sequence (at the SPS level). When ATMVP is disabled, the maximum number of merge candidates is 5. When ATMVP is enabled, the maximum number of merge candidates increases by 1, from 5 to 6.
[0219] The encoder generates a list of merge candidates using the method shown in Figure 13. One merge candidate is selected from the list of merge candidates based, for example, on a rate distortion criterion. The selected merge candidate is signaled to the decoder in the bitstream using a syntax element called a merge index.
[0220] In current VTM reference software, the method of encoding merge indices differs depending on whether ATMVP is enabled or disabled.
[0221] Figure 10(a) shows the encoding of the merge index when ATMVP is not enabled at the SPS level. The five merge candidates Cand0, Cand1, Cand2, Cand3, and Cand4 are encoded to 0, 10, 110, 1110, and 1111, respectively. This corresponds to unary maximal encoding. Furthermore, the first bit is encoded by CABAC using a single context, while the other bits are bypass encoded.
[0222] Figure 10(b) shows the encoding of the merge index when ATMVP is enabled. The six merge candidates Cand0, Cand1, Cand2, Cand3, Cand4, and Cand5 are encoded as 0, 10, 110, 1110, 11110, and 11111, respectively. In this case, all bits of the merge index (bits 1 through 5) are context-encoded by CABAC. Each bit has its own context, and there are separate probabilistic models for different bits.
[0223] In the first embodiment of the present invention, as shown in Figure 14, if ATMVP is included as a merge candidate in the list of merge candidates (for example, if ATMVP is enabled at the SPS level), the encoding of the merge index is modified so that only the first bit of the merge index is encoded by CABAC using a single context. The context is set in the same way as the current VTM reference software if ATMVP is not enabled at the SPS level, i.e., the other bits (second through fifth) are bypass-encoded. If ATMVP is not included as a merge candidate in the list of merge candidates (for example, if ATMVP is disabled at the SPS level), there are five merge candidates. Only the first bit of the merge index is encoded by CABAC using a single context. The context is set in the same way as the current VTM reference software if ATMVP is not enabled at the SPS level. The other bits (second through fourth) are bypass-decoded.
[0224] The decoder generates the same list of merge candidates as the encoder. This can be achieved by using the method shown in Figure 13. If ATMVP is not included as a merge candidate in the list of merge candidates (e.g., ATMVP is disabled at the SPS level), there are five merge candidates. Only the first bit of the merge index is decoded by CABAC using a single context. The other bits (bits 2 through 4) are bypassed. In contrast to the current reference software, if ATMVP is included as a merge candidate in the list of merge candidates (e.g., ATMVP is enabled at the SPS level), only the first bit of the merge index is decoded by CABAC using a single context in the decoding of the merge index. The other bits (bits 2 through 5) are bypassed. The decoded merge index is used to identify the merge candidate selected by the encoder from the list of merge candidates.
[0225] Compared with VTM2.0 reference software, the advantage of this embodiment is that the complexity of merge index decoding and decoder design (and encoder design) is reduced without affecting coding efficiency. In fact, in this embodiment, only one CABAC state instead of 5 is required for the merge index for current VTM merge index encoding / decoding. Furthermore, other bits are CABAC bypass encoded, which reduces the number of operations compared to encoding all bits with CABAC, thus reducing the worst-case complexity.
[0226] Second Embodiment In the second embodiment, all bits of the merge index are CABAC encoded, and all of them share the same context. In this case, there may be a single context shared between bits, as in the first embodiment. As a result, when ATMVP is included as a merge candidate in the merge candidate list (e.g., when ATMVP is enabled at the SPS level), only one context is used, compared with 5 in VTM2.0 reference software. Compared with VTM2.0 reference software, the advantage of this embodiment is that the complexity of merge index decoding and decoder design (and encoder design) is reduced without affecting coding efficiency.
[0227] Alternatively, as described below in connection with the third to sixteenth embodiments, context variables may be shared between bits such that although two or more contexts are available, the current context is shared by the bits.
[0228] When ATMVP is disabled, the same context is still used for all bits.
[0229] This embodiment and all subsequent embodiments can be applied even when ATMVP is not in an available mode or is disabled.
[0230] In a variation of the second embodiment, any two or more bits of a merge index are CABAC encoded and share the same context. The other bits of the merge index are bypass coded. For example, the first N bits of the merge index may be CABAC encoded, where N is 2 or greater.
[0231] Third Embodiment In the first embodiment, the first bit of the merge index is CABAC encoded using a single context.
[0232] In the third embodiment, the context variable for the bits of the merge index depends on the value of the merge index of an adjacent block. This allows multiple contexts for a target bit, where each context corresponds to a different value of the context variable.
[0233] The adjacent block may be any block that has already been decoded, such that its merge index is available to the decoder by the time the current block is being decoded. For example, the adjacent block may be any one of blocks A0, A1, A2, B0, B1, B2, and B3 shown in FIG. 6b.
[0234] In a first variation, only the first bit is CABAC encoded using this context variable.
[0235] In a second variation, the first N bits of the merge index, where N is 2 or greater, are CABAC encoded, and the context variable is shared among these N bits.
[0236] In the third variation, any N bits of the merge index, where N is greater than or equal to 2, are CABAC encoded, and a context variable is shared among these N bits.
[0237] In the fourth variation, the first N bits of the merge index, where N is 2 or greater, are CABAC encoded, and N context variables are used for these N bits. Assuming that the context variables have K values, KxN CABAC states are used. For example, in this embodiment, using one adjacent block, the context variables can conveniently have two values, e.g., 0 and 1. In other words, 2N CABAC states are used.
[0238] In the fifth variation, any N bits of the merge index, where N is greater than or equal to 2, are adaptively PM encoded, and N context variables are used for these N bits.
[0239] Similar modifications are applicable to the fourth to sixteenth embodiments described below.
[0240] Fourth Embodiment In the fourth embodiment, the context variable of the merge index bits depends on the respective merge index values of two or more neighboring blocks. For example, the first neighboring block is the left block A0, A1, or A2, and the second neighboring block is the upper block B0, B1, B2, or B3. The method of combining two or more merge index values is not particularly limited. An example is shown below.
[0241] For convenience, since there are two adjacent blocks, the context variable can have three different values in this case, for example, 0, 1, and 2. Therefore, if the fourth modification described in relation to the third embodiment is applied to this embodiment having three different values, K is 3 instead of 2. In other words, 3N CABAC states are used.
[0242] Fifth Embodiment In the fifth embodiment, the context variable of the merge index bits depends on the respective merge index values of adjacent blocks A2 and B3.
[0243] Sixth Embodiment In the sixth embodiment, the context variable of the merge index bits depends on the respective merge index values of adjacent blocks A1 and B1. The advantage of this modification is its alignment with the merge candidate derivation. As a result, some decoder and encoder implementations can achieve a reduction in memory access.
[0244] Seventh Embodiment In the seventh embodiment, a bit context variable having bit position idx_num in the merge index of the current block is obtained according to the following formula:
[0245] ctxIdx=(Merge_index_left==idx_num)+(Merge_index_up==idx_num) Here, Merge_index_left is the merge index of the left block, Merge_index_up is the merge index of the block above, and symbol== is the equivalent symbol.
[0246] For example, if there are 6 merge candidates, then 0 <= idx_num <= 5.
[0247] The left block is block A1, and the upper block is block B1 (same as in the sixth embodiment). Alternatively, the left block may be block A2, and the upper block may be block B3 (same as in the fifth embodiment).
[0248] If the merge index of the left block is equal to idx_num, then the expression (Merge_index_left==idx_num) is equal to 1. The following table shows the results of this expression (Merge_index_left==idx_num).
[0249]
Table 2
[0250] Of course, for the formula, the tables where Merge_index_up==idx_num are the same.
[0251] The following table shows the maximum unary code for each merge index value and the relative bit position of each bit. This table corresponds to FIG. 10(b).
[0252]
Table 3
[0253] If the left block is not a merge block or an affine merge block (that is, it is encoded using affine merge mode), the left block is considered unavailable. The same condition applies to the upper block.
[0254] For example, if only the first bit is CABAC encoded, the context variable ctxIdx is 0 if the top-left block does not have a merge index, or if the left block merge index is not the first index (i.e., not 0), and the upper block merge index is not the first index (i.e., not 0) 1 if one but not the other of the left block and the upper block has a merge index equal to the first index 2 if the merge index of each of the left block and the upper block is equal to the first index is set equal to
[0255] More generally, for the target bit at CABAC-encoded position idx_num, the context variable ctxIdx is If the top-left block does not have a merge index, or if the left block merge index is not the i-th index (i=idx_num), and if the top block merge index is not the i-th index, then the value is 0. If one of the left block and the upper block, but not the other, has a merge index equal to the i-th index, then 1 If the merge index for each of the left and top blocks is equal to the i-th index, then 2 It is set to be equal to . Here, the i-th index means the first index if i=0, the second index if i=1, and so on.
[0256] Eighth Embodiment In the eighth embodiment, a bit context variable having bit position idx_num in the merge index of the current block is obtained according to the following formula:
[0257] Ctx = (Merge_index_left > idx_num) + (Merge_index_up > idx_num) where Merge_index_left is the merge index of the left block, Merge_index_up is the merge index of the block above, and the symbol > means "greater than".
[0258] For example, if there are 6 merge candidates, then 0 <= idx_num <= 5.
[0259] The left block is block A1, and the upper block is block B1 (same as in the sixth embodiment). Alternatively, the left block may be block A2, and the upper block may be block B3 (same as in the fifth embodiment).
[0260] If the merge index of the left block is greater than idx_num, the expression (Merge_index_left>idx_num) is equal to 1. If the left block is not a merge block or an affine merge block (i.e., encoded using affine merge mode), the left block is considered unavailable. The same conditions apply to the upper block.
[0261] The following table shows the result of this formula (Merge_index_left>idx_num).
[0262] [Table 4]
[0263] For example, if only the first bit is CABAC encoded, the context variable ctxIdx is: If the top-left block does not have a merge index, or if the left block merge index is less than or equal to the first index (i.e., not 0), and if the top block merge index is less than or equal to the first index (i.e., not 0), then it is 0. If one of the left block and the upper block, but not the other, has a merge index greater than the first index, then 1 If the merge index for each of the left and top blocks is greater than the first index, then 2 It is set to be equal to.
[0264] More generally, for the target bit of the CABAC-encoded position idx_num, the context variable ctxIdx is: If the top-left block does not have a merge index, or if the left block merge index is less than the i-th index (i=idx_num), and if the top block merge index is less than or equal to the i-th index, then the value is 0. If one of the left block and the upper block, but not the other, has a merge index greater than the i-th index, then 1 If the merge index for each of the left and top blocks is greater than the i-th index, then 2 It is set to be equal to.
[0265] The eighth embodiment further improves coding efficiency compared to the seventh embodiment.
[0266] Ninth Embodiment In the fourth to eighth embodiments, the context variable of the current block's merge index bits depended on the respective values of the merge indices of two or more adjacent blocks.
[0267] In the ninth embodiment, the context variable of the bits of the merge index of the current block depends on the merge flags of two or more adjacent blocks. For example, the first adjacent block is the left block A0, A1, or A2, and the second adjacent block is the upper block B0, B1, B2, or B3.
[0268] The merge flag is set to 1 if the block is encoded using merge mode, and to 0 if other modes such as skip mode or affine merge mode are used. Note that in VMT2.0, affine merge is a separate mode from basic mode or "classic" merge mode. Affine merge mode can be signaled using a dedicated affine flag. Alternatively, the list of merge candidates may include affine merge candidates, in which case affine merge mode may be selected and signaled using a merge index.
[0269] Subsequently, the context variable is, If neither the left adjacent block nor the top adjacent block has the merge flag set to 1, then it is set to 0. If one of the adjacent blocks to the left and above, but not the other, has its merge flag set to 1, then 1 If each of the adjacent blocks to the left and above has its merge flag set to 1, then 2 It will be set to this.
[0270] This simple evaluation achieves an improvement in encoding efficiency compared to VTM2.0. Another advantage is lower complexity, as only the merge flag, rather than the merge index of adjacent blocks, needs to be checked, compared to the seventh and eighth embodiments.
[0271] In a modified version, the context variable for the current block's merge index bits depends on the merge flag of a single adjacent block.
[0272] Tenth Embodiment In the third to ninth embodiments, the context variable of the current block's merge index bits depended on the merge index values or merge flags of one or more adjacent blocks.
[0273] In the tenth embodiment, the context variable of the bits of the merge index of the current block depends on the value of the skip flag for the current block (the current coding unit, i.e., CU). The skip flag is equal to 1 if the current block uses merge skip mode, and equal to 0 otherwise.
[0274] A skip flag is a first example of another variable or syntax element that has already been decoded or parsed for the current block. This other variable or syntax element is preferably an indicator of the complexity of the motion information in the current block. Since the occurrence of merge index values depends on the complexity of the motion information, variables or syntax elements like skip flags generally correlate with merge index values.
[0275] More specifically, the merge-skip mode is generally chosen for static scenes or scenes with constant motion. As a result, the merge index value is generally lower in the merge-skip mode than in the classical merge mode used to encode interpredictions including block residuals. This generally occurs for more complex motion. However, the choice between these modes is often also related to quantization and / or RD criteria.
[0276] This simple evaluation improves encoding efficiency compared to VTM2.0. Furthermore, it is very easy to implement because it does not involve checking adjacent block or merge index values.
[0277] In the first modification, the context variable of the bits of the merge index of the current block is simply set to equal the skip flag of the current block. The bits can be only the first bit; the other bits are bypass-encoded as in the first embodiment.
[0278] In the second variation, all bits of the merge index are CABAC encoded, and each of them has its own context variable depending on the merge flag. This requires 10 probabilistic states when there are 5 CABAC encoded bits in the merge index (corresponding to 6 merge candidates).
[0279] In the third variation, to limit the number of states, only N bits of the merge index are CABAC encoded, where N is 2 or greater, for example, the first N bits. This requires 2N states. For example, if the first 2 bits are CABAC encoded, then 4 states are required.
[0280] Generally, instead of a skip flag, it is possible to use any other variable or syntax element that has already been decoded or parsed for the current block and serves as an indicator of the complexity of the motion information in the current block.
[0281] 11th Embodiment The eleventh embodiment relates to the affine merge signaling described above with reference to Figures 11(a), 11(b), and 12.
[0282] In the eleventh embodiment, the context variable of the CABAC-encoded bits of the merge index of the current block (current CU) depends, if any, on the affine merge candidate in the list of merge candidates. This bit may be only the first bit of the merge index, or the first N bits, where N is 2 or greater, or any N bits. The other bits are bypass-encoded.
[0283] Affine prediction is designed to compensate for complex movements. Therefore, in the case of complex movements, the merge index generally has a higher value than in the case of less complex movements. As a result, if the first Affine Merge candidate is quite far down the list, or if there are no Affine Merge candidates at all, the merge index of the current CU may have a small value.
[0284] Therefore, it is reasonable for the context variable to depend on the presence and / or position of at least one affine merge candidate in the list.
[0285] For example, the context variable is set to 1 if A1 is affine, 2 if B1 is affine, 3 if B0 is affine, 4 if A0 is affine, 5 if B2 is affine, and 0 if the adjacent block is not affine.
[0286] When the merge index of the current block is decrypted or parsed, the affine flags for the merge candidates at these locations have already been checked. Therefore, no further memory access is required to derive the context of the merge index of the current block.
[0287] This embodiment improves encoding efficiency compared to VTM2.0. Since step 1205 already includes checking the adjacent CU affine mode, no additional memory access is required.
[0288] In the first variation, to limit the number of states, the context variable is: It is set to 0 if the adjacent block is not affine, or if A1 or B1 is affine, and equal to 1 if B0, A0, or B2 is affine.
[0289] In the second variation, to limit the number of states, the context variable is: It is set to 0 if the adjacent block is not affine, 1 if A1 or B1 is affine, and 2 if B0, A0, or B2 is affine.
[0290] In the third variation, the context variable is: The value is set to 1 if A1 is affine, 2 if B1 is affine, 3 if B0 is affine, 4 if A0 or B2 is affine, and 0 if the adjacent block is not affine.
[0291] It should be noted that these locations are already checked when the merge index is decoded or parsed, as affine flag decoding depends on these locations. Therefore, no additional memory access is required to derive the merge index context that is encoded after the affine flag.
[0292] Twelfth Embodiment In the twelfth embodiment, signaling an affine mode includes inserting an affine mode as a candidate motion predictor.
[0293] In one example of the twelfth embodiment, affine merge (and affine merge skip) is signaled as a merge candidate (i.e., as one of the merge candidates for use with the classical merge mode or classical merge skip mode). In this case, modules 1205, 1206, and 1207 in Figure 12 are removed. Furthermore, the maximum number of possible merge candidates is incremented so as not to affect the coding efficiency of the merge mode. For example, in the current VTM version, this value is set to equal 6, and therefore, when this embodiment is applied to the current version of VTM, the value becomes 7.
[0294] The advantage is the simplification of the syntax element design in merge mode, as fewer syntax elements need to be decoded. In some situations, improvements / changes in coding efficiency may be observed.
[0295] Two possibilities for implementing this example are described below.
[0296] The merge index of an affine merge candidate always occupies the same position in the list, regardless of the values of other merge MVs. The position of a candidate motion predictor indicates its likelihood of being selected; therefore, the higher it is placed in the list (lower index value), the more likely that motion vector predictor is selected.
[0297] In the first example, the merge index of an affine merge candidate always occupies the same position within the list of merge candidates. This means it has a fixed "merge idx" value. For example, since the affine merge mode should represent a complex movement that is not the most probable content, this value can be set to equal 5. An additional advantage of this embodiment is that the current block can be set as an affine block when it is parsed (decoding the data itself, not just decoding / reading the syntax elements). As a result, this value can be used to determine the CABAC context of the affine flag used in AMVP. Thus, the conditional probability should be improved for this affine flag, and the coding efficiency should be better.
[0298] In the second example, the affine merge candidate is derived along with other merge candidates. In this example, the new affine merge candidate is added to the list of merge candidates (in classical merge mode or classical merge skip mode). Figure 16 shows this example. Compared to Figure 13, the affine merge candidate is the first affine neighboring block from A1, B1, B0, A0, and B2 (1917). If the same conditions as 1205 in Figure 12 are valid (1927), a motion vector field is generated using the affine parameters, and the affine merge candidate is obtained (1929). The initial list of merge candidates can have 4, 5, 6, or 7 candidates, depending on the use of ATMVP, temporal, and affine merge candidates.
[0299] The order among all these candidates is important, as the more likely candidate should be processed first to ensure that the most likely candidate is more likely to make the cut of the motion vector candidate. The preferred order is as follows:
[0300] A1, B1, B0, A0, Affine Merge, ATMVP, B2, Temporal, Combination, Zero_MV It is important to note that the affine merge candidate is placed before the ATMVP candidate but after the four major adjacency blocks. The advantage of placing the affine merge candidate before the ATMVP candidate is that coding efficiency increases compared to placing it after the ATMVP and temporal predictor candidates. This improvement in coding efficiency depends on the GOP (group of pictures) structure and the QP (Quantization Parameter) setting for each picture within the GOP. However, with the most commonly used GOP and QP settings, this order results in increased coding efficiency.
[0301] A further advantage of this solution is the clean design of the classic merge and classic merge skip modes (i.e., merge modes with additional candidates such as ATMVP or affine merge candidates) for both syntax and derivation processing. Furthermore, the merge index of the affine merge candidate can be modified according to the availability or value (duplicate check) of previous candidates in the list of merge candidates. As a result, efficient signaling can be obtained.
[0302] In a further example, the merge index for an affine merge candidate is variable according to one or more conditions.
[0303] For example, the merge index or position in the list associated with an affine merge candidate changes according to a criterion. The principle is to set a lower value for the merge index corresponding to the affine merge candidate if the probability of selecting the affine merge candidate is high (and a higher value if the probability of selection is low).
[0304] In the twelfth embodiment, the affine merge candidate has a merge index value. To improve the coding efficiency of the merge index, it is useful to make the context variable of the merge index bits dependent on the affine flags of the adjacent blocks and / or the current block.
[0305] For example, context variables can be determined using the following expression.
[0306] ctxIdx=IsAffine(A1)+IsAffine(B1)+IsAffine(B0)+IsAffine(A0)+IsAffine(B2) The resulting context value can have a value of 0, 1, 2, 3, 4, or 5.
[0307] The affine flag improves encoding efficiency.
[0308] In the first variation, ctxIdx = IsAffine(A1) + IsAffine(B1) is used to include fewer adjacent blocks. The resulting context value can have a value of 0, 1, or 2.
[0309] Furthermore, in the second variation, to include fewer adjacent blocks, ctxIdx = IsAffine(A2) + IsAffine(B3). In this case as well, the resulting context value can have a value of 0, 1, or 2.
[0310] In the third variation, ctxIdx = IsAffine(current block) is used to avoid including adjacent blocks. The resulting context value can be 0 or 1.
[0311] Figure 15 is a flowchart of the partial decoding process for several syntax elements related to the coding mode according to a third modification. In this figure, the skip flag (1601), prediction mode (1611), merge flag (1603), merge index (1608), and affine flag (1606) can be decoded. This flowchart is similar to the flowchart in Figure 12 described earlier, so a detailed explanation is omitted. The difference is that the merge index decoding process takes the affine flag into consideration so that when obtaining the context variable for the merge index, the affine flag that is decoded before the merge index can be used. This is not the case in VTM2.0. In VTM2.0, the affine flag of the current block always has the same value "0", so it cannot be used to obtain the context variable for the merge index.
[0312] 13th Embodiment In the tenth embodiment, the context variable for the bits of the merge index of the current block depends on the value of the skip flag for the current block (the current coding unit, i.e., CU). In the thirteenth embodiment, instead of using the skip flag value directly to derive the context variable for the target bit of the merge index, the context value of the target bit is derived from the context variable used to encode the skip flag of the current CU. This is possible because the skip flag itself is CABAC encoded and therefore has a context variable. Preferably, the context variable for the target bit of the merge index of the current CU is set to (copied from) the context variable used to encode the skip flag of the current CU. The target bit can be only the first bit; the other bits may be bypass encoded as in the first embodiment.
[0313] The context variable for the current CU's skip flag is derived in the manner specified in VTM2.0. The advantage of this embodiment compared to the VTM2.0 reference software is that the complexity of the merge index decoding and decoder design (and encoder design) is reduced without affecting encoding efficiency. In fact, in this embodiment, instead of 5 for current VTM merge index coding (encoding / decoding), only one CABAC state is required to encode the merge index. Furthermore, the other bits are CABAC bypass encoded, reducing the number of operations compared to encoding all bits with CABAC, thus reducing the complexity in the worst case.
[0314] Embodiment 14 In the 13th embodiment, the context variable / value of the target bit is derived from the context variable of the skip flag of the current CU. In the 14th embodiment, the context value of the target bit is derived from the context variable of the affine flag of the current CU.
[0315] This is possible because the affine flag itself is CABAC encoded and therefore has a context variable. Preferably, the context variable for the target bit of the merge index of the current CU is set equal to (copied from) the context variable for the affine flag of the current CU. The target bit can be only the first bit; the other bits are bypass encoded as in the first embodiment.
[0316] The context variable for the current CU's affine flag is derived in the manner specified in VTM2.0.
[0317] The advantage of this embodiment compared to the VTM2.0 reference software is that it reduces the complexity of the merge index decoding and decoder design (and encoder design) without affecting encoding efficiency. In fact, in this embodiment, a minimum of one CABAC state is required for the merge index, instead of five for current VTM merge index coding (encoding / decoding). Furthermore, the other bits are CABAC bypass coded, reducing the number of operations compared to coding all bits with CABAC, thus reducing the complexity in the worst case.
[0318] Embodiment 15 In some of the embodiments described above, the context variable had more than two values, for example, three values: 0, 1, and 2. However, to reduce complexity and the number of states to be processed, it is possible to limit the number of allowed context variable values to two, for example, 0 and 1. This can be achieved, for example, by changing any initial context variable that has the value 2 to 1. In practice, this simplification has little to no effect on coding efficiency.
[0319] Embodiments and combinations of other embodiments Any two or more of the embodiments described above may be combined.
[0320] The above description has focused on the encoding and decoding of merge indices. For example, the first embodiment includes generating a list of merge candidates, including ATMVP candidates (in the case of classical merge mode or classical merge skip mode, i.e., non-affine merge mode or non-affine merge skip mode), selecting one of the merge candidates in the list, and generating a merge index for the selected merge candidate using CABAC encoding, where one or more bits of the merge index are bypass CABAC encoded. In principle, the present invention can be applied to modes other than merge modes (e.g., affine merge mode), which include generating a list of motion information predictor candidates (e.g., a list of affine merge candidates or motion vector predictor (MVP) candidates), selecting one of the motion information predictor candidates (e.g., MVP candidates) in the list, and generating an identifier or index for the selected motion information predictor candidate (e.g., a selected affine merge candidate or a selected MVP candidate for predicting the motion vector of the current block) in the list. Therefore, the present invention is not limited to merge modes (i.e., classical merge mode and classical merge skip mode), and the index to be encoded or decoded is not limited to the merge index. For example, in the development of VVC, the technology of the embodiments described above may be applied to (or extended to) modes other than merge mode, such as the AMVP mode of HEVC, or its equivalent mode in VVC, or the affine merge mode. The attached claims should be interpreted accordingly.
[0321] As described above, in the embodiments described above, the affine merge mode (affine merge or affine merge skip mode) and / or one or more motion information candidates (e.g., motion vectors) of one or more affine parameters are obtained from first adjacent blocks that are affine-encoded between spatially adjacent blocks (e.g., positions A1, B1, B0, A0, B2) or temporally related blocks (e.g., the “center” block with collated blocks, or its spatially adjacent blocks such as “H”). These positions are shown in Figures 6a and 6b. To enable this retrieval (e.g., deriving, sharing, or "merging") of affine parameters between one or more motion information and / or the current block (or the current CU, a group of sample / pixel values currently being encoded / decoded) and an adjacent block (which is spatially adjacent to or temporally related to the current block), one or more affine merge candidates are added to the list of merge candidates (i.e., classical merge mode candidates), and as a result, if the selected merge candidate (e.g., signaled using the merge index with a syntax element such as "merge_idx" in HEVC or its functionally equivalent syntax element) is an affine merge candidate, the current CU / block is encoded / decoded using the affine merge mode together with the affine merge candidate.
[0322] As described above, one or more affine merge candidates for obtaining (e.g., deriving or sharing) one or more motion information and / or affine parameters of an affine merge mode can also be signaled using a separate list (or set) of affine merge candidates (which may be the same as or different from the list of merge candidates used for the classical merge mode).
[0323] According to one embodiment of the present invention, when the techniques of the above-described embodiment are applied to the affine merge mode, the list of affine merge candidates is shown in Figure 8 and can be generated using the same techniques as the motion vector derivation process for the classical merge mode described in relation thereto, or shown in Figure 13 and can be generated using the same techniques as the merge candidate derivation process described in relation thereto. The advantage of sharing the same techniques to generate / compile this list of affine merge candidates (for the affine merge mode or affine merge skip mode) and the list of merge candidates (for the classical merge mode or classical merge skip mode) is that the complexity of the encoding / decoding process is reduced compared to having separate techniques.
[0324] To achieve similar advantages, it should be understood that, according to other embodiments, similar techniques can be applied to other interprediction modes that require signaling a motion information predictor selected (from multiple candidates).
[0325] According to another embodiment, a list of affine merge candidates can be generated / compiled using a separate technique shown below in relation to Figure 24.
[0326] Figure 24 is a flowchart of the affine merge candidate derivation process for affine merge modes (affine merge mode and affine merge skip mode). In the first step of the derivation process, five block positions are considered (2401-2405) to obtain / derive a spatial affine merge candidate 2413. These positions are the spatial positions shown in Figure 6a (and Figure 6b) with reference numbers A1, B1, B0, A0, and B2. In the next step, the availability of spatial motion vectors is checked to determine whether each of the intermode encoded blocks associated with each position A1, B1, B0, A0, and B2 is encoded in affine mode (e.g., using one of affine merge, affine merge skip, or affine AMVP mode) (2410). At most five motion vectors (i.e., spatial affine merge candidates) are selected / obtained / derivated. A predictor is considered available if a predictor exists (for example, if there is information to obtain / derive the motion vector associated with its position), and if the block is not intra-encoded, and if the block is affine (i.e., encoded using affine mode).
[0327] Next, affine motion information is derived / obtained for each available block position (2411)(2410). This derivation is performed for the current block based on the affine model of the block position (and its affine model parameters, for example, as described in relation to Figures 11(a) and 11(b)). Then, a pruning process (2412) is applied to remove candidates that previously added to the list that give each other the same affine motion compensation (or have the same affine model parameters).
[0328] At the end of this stage, the list of spatial affine merge candidates may include up to five candidates.
[0329] If the number of candidates (Nb_Cand) is strictly less than the maximum number of candidates (2426) (where Max_Cand is a value signaled in the bitstream slice header, which is equal to 5 in affine merge mode, but may vary / be variable depending on the implementation).
[0330] Next, constructed affine merge candidates are generated (i.e., additional affine merge candidates that serve a similar role to the combined bipredictive merge candidates in HEVC, for example, not only approaching the target number but also providing some diversity) (2428). These constructed affine merge candidates are based on motion vectors related to the adjacent spatial and temporal positions of the current block. First, control points are defined to generate motion information for generating the affine model (2418, 2419, 2420, 2421). Two of these control points correspond to v0 and v1 in Figures 11(a) and 11(b), for example. These four control points correspond to the four corners of the current block.
[0331] The motion information for the upper left (2418) of the control point is obtained from the motion information for the block position at position B2 (2405) (for example, by making it equal to it), if it exists and this block is encoded in intermode (2414). Otherwise, the motion information for the upper left (2418) of the control point is obtained from the motion information for the block position at position B3 (2406) (for example, by making it equal to it), if it does exist and this block is encoded in intermode (2414), and if this is not the case, the motion information for the upper left (2418) of the control point is obtained from the motion information for the block position at position A2 (2407), (for example, by making it equal to it), if it exists and this block is encoded in intermode (2414). If there is no available block at this control point, it is considered unavailable.
[0332] The motion information for the upper right control point (2419) is obtained from (e.g., equal to) the motion information for the block at position B1 (2402), if it exists and this block is encoded in intermode (2415). Otherwise, the motion information for the upper right control point (2419) is obtained from (e.g., equal to) the motion information for the block at position B0 (2403), if it exists and this block is encoded in intermode (2415). If there is no available block at this control point, it is considered unavailable.
[0333] The motion information for the lower left control point (2420) is obtained from the motion information for the block at position A1 (2401) if it exists and this block is encoded in intermode (2416) (for example, they are equal). Otherwise, the motion information for the lower left control point (2420) is obtained from the motion information for the block at position A0 (2404) if it exists and this block is encoded in intermode (2416) (for example, they are equal). If there is no available block at this control point, it is considered unavailable.
[0334] The motion information for the lower right control point (2421) is obtained from (e.g., equal to) the temporal candidate motion information, e.g., the collated block position at position H (2408) (as shown in Figure 6a), if it exists and this block is encoded in intermode (2417). If there is no available block at this control point, it is considered unavailable.
[0335] Based on these control points, up to 10 constructed affine merge candidates can be generated (2428). These candidates are generated based on affine models having 4, 3, or 2 control points. For example, the first constructed affine merge candidate may be generated using 4 control points. Next, the following 4 constructed affine merge candidates are 4 possibilities that can be generated using 4 different sets of 3 control points (i.e., 4 different possible combinations of sets containing 3 of the 4 available control points). Then, the other constructed affine merge candidates are generated using 2 different sets of 2 control points (i.e., 2 different possible combinations of sets containing 2 of the 4 control points).
[0336] If the number of candidates (Nb_Cand) remains strictly less than the maximum number of candidates (Max_Cand) after adding these additional (constructed) affine merge candidates (2430), other additional virtual motion information candidates, such as zero motion vector candidates (or combined bipredictive merge candidates, if applicable), are added / generated until the number of candidates in the list of affine merge candidates reaches the target number (e.g., the maximum number of candidates) (2432).
[0337] At the end of this process, a list or set of affine merge mode candidates (i.e., a list or set of candidates for affine merge modes, which are affine merge modes and affine merge skip modes) is generated / constructed (2434). As shown in Figure 24, the list or set of affine merge (motion vector predictor) candidates is constructed / generated from a subset of spatial candidates (2401-2407) and temporal candidates (2408) (2434). According to embodiments of the present invention, it should be understood that other affine merge candidate derivation processes having a different order for checking availability, pruning processes, or the number / type of potential candidates (for example, ATMVP candidates can also be added in a similar manner to the merge candidate list derivation process in Figure 13 or Figure 16) may also be used to generate the list / set of affine merge candidates.
[0338] The following embodiments demonstrate how a list (or set) of affine merge candidates can be used to signal (e.g., encode or decode) a selected affine merge candidate (which can be signaled using the merge index used in the merge mode, or a separate affine merge index used in particular in the affine merge mode).
[0339] In the following embodiments, a merge mode (i.e., a merge mode other than the affine merge mode as defined later, in other words, a classical non-affine merge mode or a classical non-affine merge skip mode) is a type of merge mode in which motion information of either a spatially adjacent block or a temporally related block is obtained for (or derived for or shared with) the current block; a merge mode predictor candidate (i.e., a merge candidate) is information about one or more spatially adjacent or temporally related blocks from which the current block can obtain / derive motion information in a merge mode; a merge mode predictor is a selected merge mode predictor candidate, the information of which is used when predicting motion information of the current block and during signaling in a merge mode (e.g., encoding or decoding) process; and an index (e.g., a merge index) that identifies the merge mode predictor from a list (or set) of merge mode predictor candidates is The signaled affine merge mode is a type of merge mode in which motion information of either a spatially adjacent block or a temporally related block is acquired (derived for or shared with the current block) for the current block so that motion information of the current block and / or affine parameters for affine mode processing (or affine motion model processing) can use this acquired / derived / shared motion information; an affine merge mode predictor candidate (i.e., an affine merge candidate) is information about one or more spatially adjacent or temporally related blocks from which the current block can acquire / derive motion information in affine merge mode; and an affine merge mode predictor is a selected affine merge mode predictor candidate whose information is available in the affine motion model when predicting motion information of the current block and while signaling in affine merge mode (e.g., encoding or decoding) processing.An index (e.g., an affine merge index) is signaled to identify an affine merge mode predictor from a list (or set) of affine merge mode predictor candidates. In the following embodiments, it is understood that an affine merge mode is a merge mode that has its own affine merge index (an identifier, which is a variable) for identifying one affine merge mode predictor candidate from a list / set of candidates (also known as the "affine merge list" or "subblock merge list"), and has a single index value associated with it, whereas the affine merge index is signaled to identify that particular affine merge mode predictor candidate.
[0340] In the following embodiments, “Merge Mode” refers to either the classic merge-skip mode or the classic merge mode in HEVC / JEM / VTM, or any functionally equivalent mode, provided that the acquisition of motion information (e.g., derivation or sharing) and signaling of the merge index as described above are used in the above mode. It should be understood that “Affine Merge Mode” also refers to either the affine merge mode or the affine merge-skip mode (using such acquisition / derivation, if present), or any other functionally equivalent mode, provided that the same features are used in the above mode.
[0341] Embodiment 16 In the 16th embodiment, a motion information predictor index for identifying affine merge mode predictors (candidates) from a list of affine merge candidates is signaled using CABAC coding, and one or more bits of the motion information predictor index are bypassed by CABAC coding.
[0342] According to a first modification of the embodiment, in the encoder, the motion information predictor index for affine merge mode is encoded by generating a list of motion information predictor candidates, selecting one of the motion information predictor candidates in the list as the affine merge mode predictor, and generating a motion information predictor index for the selected motion information predictor candidate using CABAC coding, wherein one or more bits of the motion information predictor index are bypass CABAC coded. Next, data indicating the index for this selected motion information predictor candidate is included in the bitstream. Next, the decoder decodes the motion information predictor index for affine merge mode by generating a list of motion information predictor candidates from the bitstream containing this data, decoding the motion information predictor index using CABAC decoding, wherein one or more bits of the motion information predictor index are bypass CABAC coded, and when affine merge mode is used, the decoded motion information predictor index is used to identify one of the motion information predictor candidates in the list as the affine merge mode predictor.
[0343] According to a further modification of the first modification, one or more motion information predictor candidates in a list are also selectable as merge mode predictors when merge mode is used, so that when merge mode is used, the decoder can use a decoded motion information predictor index (e.g., a merge index) to identify one of the motion information predictor candidates in a list as a merge mode predictor. In this further modification, an affine merge index is used to signal the affine merge mode predictor(s), and the signaling of the affine merge index is implemented using merge index signaling according to any one of the first to fifteenth embodiments, or index signaling similar to the merge index signaling used in current VTM or HEVC.
[0344] In this modification, when merge mode is used, signaling the merge index can be performed using merge index signaling according to any one of the 1st to 15th embodiments, or using the merge index signaling used in the current VTM or HEVC. In this modification, signaling the affine merge index and signaling the merge index can use different index signaling schemes. The advantage of this modification is that better coding efficiency is achieved by using efficient index coding / signaling for both affine merge mode and merge mode. Furthermore, in this modification, separate syntax elements can be used for the merge index (e.g., HEVC's "Merge_idx[][]" or its functional equivalent) and the affine merge index (e.g., "A_Merge_idx[][]"). This allows for separate signaling (coding / decoding) of the merge index and the affine merge index.
[0345] In yet another further variation, if merge mode is used and one of the motion information predictor candidates in the list is also selectable as a merge mode predictor, CABAC coding uses the same context variable for at least one bit of the motion information predictor index of the current block (e.g., merge index or affine merge index) for both modes, i.e., when affine merge mode is used and when merge mode is used, so that at least one bit of the affine merge index and merge index share the same context variable. The decoder then uses the decoded motion information predictor index when merge mode is used to identify one of the motion information predictor candidates in the list as a merge mode predictor, and CABAC decoding uses the same context variable for at least one bit of the motion information predictor index of the current block for both modes, i.e., when affine merge mode is used and when merge mode is used.
[0346] According to a second modification of the embodiment, in the encoder, the motion information predictor index is encoded by generating a list of motion information predictor candidates, selecting one of the motion information predictor candidates in the list as the affine merge mode predictor when affine merge mode is used, selecting one of the motion information predictor candidates in the list as the merge mode predictor when merge mode is used, and generating a motion information predictor index for the selected motion information predictor candidate using CABAC coding, wherein one or more bits of the motion information predictor index are bypassed CABAC coded. Next, data indicating the index for this selected motion information predictor candidate is included in the bitstream. Next, the decoder decodes the motion information predictor index from the bitstream, generates a list of motion information predictor candidates, decodes the motion information predictor index using CABAC decoding, one or more bits of the motion information predictor index are bypassed by CABAC decoding, and if affine merge mode is used, the decoded motion information predictor index is used to identify one of the motion information predictor candidates in the list as an affine merge mode predictor, and if merge mode is used, the decoded motion information predictor index is used to identify one of the motion information predictor candidates in the list as a merge mode predictor.
[0347] According to a further modification of the second modification, the affine merge index signaling and merge index signaling use the same index signaling scheme as any one of the first to fifteenth embodiments, or the merge index signaling used in the current VTM or HEVC. The advantage of this further modification is its simple design in practice, which can also lead to less complexity. In this modification, when the affine merge mode is used, the encoder's CABAC coding includes using a context variable for at least one bit of the motion information predictor index (affine merge index) of the current block, the context variable being separable from another context variable for at least one bit of the motion information predictor index (merge index) when the merge mode is used, and the data indicating the use of the affine merge mode is included in the bitstream so that the context variables for the affine merge mode and the merge mode can be distinguished (clearly identified) for the CABAC decoding process. Next, the decoder retrieves data from the bitstream to indicate the use of affine merge mode in the bitstream, and when affine merge mode is used, CABAC decoding uses this data to distinguish between context variables for the affine merge index and the merge index. Furthermore, the decoder can also use the data indicating the use of affine merge mode to generate a list (or set) of affine merge mode predictor candidates when the retrieved data indicates the use of affine merge mode, or to generate a list (or set) of merge mode predictor candidates when the retrieved data indicates the use of merge mode.
[0348] This variation allows both merged and affine merged indices to be signaled using the same index signaling scheme, while merged and affine merged indices are still encoded / decoded independently of each other (e.g., by using separate context variables).
[0349] One way to use the same index signaling scheme is to use the same syntax elements for both affine merge indices and merge indices, meaning that whether affine merge mode or merge mode is used, the motion predictor index of the selected motion predictor candidate is encoded using the same syntax elements in both cases. The decoder then decodes the motion predictor index by parsing the same syntax elements from the bitstream, regardless of whether the current block was encoded (and decoded) using affine merge mode or merge mode.
[0350] Figure 22 shows the partial decoding of several syntax elements related to the coding mode (i.e., the same index signaling scheme) by this modification of the 16th embodiment. This figure shows the signaling of the affine merge index (2255 - "merge idx affine") for the affine merge mode (2257: Yes) and the merge index (2258 - "merge idx") for the merge mode (2257: No) having the same index signaling scheme. It should be understood that in some modifications, the affine merge candidate list may include ATMVP candidates, as in the merge candidate list of the current VTM. The coding of the affine merge index is similar to the coding of the merge index for the merge mode, as shown in Figure 10(a), Figure 10(b), or Figure 14. In some variations, even if ATMVP merge candidates are not defined in the affine merge candidate derivation, if ATMVP is enabled for merge mode with up to five other candidates (i.e., a total of six candidates) such that the maximum number of candidates in the affine merge candidate list matches the maximum number of candidates in the merge candidate list, the affine merge index is encoded as described in Figure 10(b). Thus, each bit of the affine merge index has its own context. All context variables used for bits in merge index signaling are independent of the context variables used for bits in affine merge index signaling.
[0351] In a further modification, this same index signaling scheme, shared by merge index and affine merge index signaling, uses CABAC coding only for the first bin, similar to the first embodiment. That is, all bits except the first bit of the motion information predictor index are bypassed with CABAC coding. In this further modification of the 16th embodiment, if ATMVP is included as a candidate in one of the lists of merge candidates or affine merge candidates (for example, if ATMVP is enabled at the SPS level), the coding for each index (i.e., merge index or affine merge index) is modified so that only the first bit of the index is coded by CABAC using a single context variable as shown in Figure 14. This single context is set in the same way as in the current VTM reference software when ATMVP is not enabled at the SPS level. The other bits (the second through fifth bits or the fourth bit, if only five candidates are present in the list) are bypassed. If ATMVP is not included as a candidate in the merge candidate list (for example, if ATMVP is disabled at the SPS level), there are five merge candidates and five affine merge candidates available. Only the first bit of the merge index for merge mode is encoded by CABAC using a first single context variable. And only the first bit of the affine merge index for affine merge mode is encoded by CABAC using a second single context variable. These first and second context variables are set in the same way as in the current VTM reference software if ATMVP is not enabled at the SPS level for both the merge index and the affine merge index. The other bits (bits 2 through 4) are bypass-decoded.
[0352] The decoder generates the same list of merge candidates and the same list of affine merge candidates as the encoder. This is achieved, for example, by using the method shown in Figure 24. The same index signaling scheme is used for both merge mode and affine merge mode, but the affine flag (2256) is used to determine whether the data currently being decoded is for a merge index or an affine merge index, and as a result the first and second context variables are separable (or distinguishable) from each other for the CABAC decoding process. That is, the affine flag (2256) is used during the index decoding process (i.e., used in step 2257) to determine whether to decode "merge idx2258" or "merge idx affine 2255". If ATMVP is not included as a candidate in the list of merge candidates (for example, if ATMVP is disabled at the SPS level), there are five merge candidates in both lists of candidates (for merge mode and affine merge mode). Only the first bit of the merge index is decoded by CABAC using the first single context variable. Then, only the first bit of the affine merge index is decoded by CABAC using a second single context variable. All other bits (bits 2 through 4) are bypass-decoded. In contrast to the current reference software, if ATMVP is included as a candidate in the list of merge candidates (for example, if ATMVP is enabled at the SPS level), only the first bit of the merge index is decoded by CABAC, using a first single context variable in the decodement of the merge index and a second single context variable in the decodement of the affine merge index. The other bits (bits 2 through 5 or bit 4) are bypass-decoded. The decoded index is then used to identify the candidate selected by the encoder from the corresponding list of candidates (i.e., merge candidates or affine merge candidates).
[0353] The advantage of this modification is that by using the same index signaling scheme for both merged and affine merged indices, the complexity of the index decoding and decoder design (and encoder design) for implementing these two different modes is reduced without significantly impacting encoding efficiency. In fact, in this variable, only two CABAC states (one for each of the first and second single-context variables) are required for index signaling, instead of 9 or 10 when all bits of the merged index and all bits of the affine merged index are CABAC encoded / decoded. Furthermore, since all other bits (except the first bit) are CABAC bypass encoded, the worst-case complexity is reduced, and the number of operations required during the CABAC encoding / decoding process is reduced compared to encoding all bits by CABAC.
[0354] In yet another further variation, CABAC encoding or decoding uses the same context variable for at least one bit of the motion information predictor index of the current block, both when affine merge mode is used and when merge mode is used. In this further variation, the context variable used for the first bit of the merge index and the first bit of the affine merge index is independent of which index is being encoded or decoded; that is, the first and second single context variables (from the previous variation) are not distinguished / separated but are the same single context variable. Thus, in contrast to the previous variation, the merge index and the affine merge index share one context variable during CABAC processing. As shown in Figure 23, the index signaling scheme is the same for both the merge index and the affine merge index; that is, only one type of index, "merge idx(2308)", is encoded or decoded for both modes. As far as CABAC decoders are concerned, the same syntax elements are used for both the merge index and the affine merge index, and there is no need to distinguish them when considering the context variable. Therefore, as in step (2257) in Figure 22, there is no need to use the affine flag (2306) to determine whether the current block is encoded (decoded) in affine merge mode, and there is no branch after step 2306 in Figure 23, since only one index ("merge idx") requires decoding. The affine flag is used to perform motion information prediction in affine merge mode, i.e., during the prediction process after the CABAC decoder has decoded the index ("merge idx"). Furthermore, only the first bit of this index (i.e., the merge index and the affine merge index) is encoded by CABAC using one single context, while the other bits are bypass encoded as described in the first embodiment.Therefore, in this further modification, one context variable of the first bit of the merge index and the affine merge index is shared by both merge index and affine merge index signaling. If the size of the candidate list differs between the merge index and the affine merge index, the maximum number of bits for signaling the relevant index for each case may also differ, i.e., they are independent of each other. Thus, the number of bypass-encoded bits can be adjusted as needed, according to the value of the affine flag (2306), for example, to enable parsing of data for the relevant index from the bitstream.
[0355] The advantage of this modification is that it reduces the complexity of the merge index and affine merge index decoding process and decoder design (and encoder design) without significantly impacting encoding efficiency. In fact, in this further modification, when signaling both merge index and affine merge index, only one CABAC state is required instead of the previous modification or 9 or 10 CABAC states. Furthermore, since all other bits (except the first bit) are CABAC bypass encoded, the worst-case complexity is reduced, and the number of operations required during CABAC encoding / decoding is reduced compared to encoding all bits by CABAC.
[0356] In the aforementioned modifications of this embodiment, affine merge index signaling and merge index signaling can reduce the number of contents and / or share one or more contexts, as described in any of the first to fifteenth embodiments. The advantage of this is the reduction in complexity due to the decrease in the number of contexts required to encode or decode these indices.
[0357] In the aforementioned modifications of this embodiment, the motion information predictor candidate comprises information for obtaining (or deriving) one or more of the following: direction, list ID, reference frame index, and motion vector. Preferably, the motion information predictor candidate comprises information for obtaining a motion vector predictor candidate. In a preferred modification, a motion information predictor index (e.g., an affine merge index) is used to signal an affine merge mode predictor candidate, and the affine merge index signaling is implemented using merge index signaling according to any one of the first to fifteenth embodiments, or index signaling similar to the merge index signaling used in current VTM or HEVC (with affine merge mode motion information predictor candidates as merge candidates).
[0358] In the aforementioned modifications of this embodiment, the generated list of motion information predictor candidates includes ATMVP candidates, as in the first embodiment or as in some of the other modifications of the second to fifteenth embodiments described above. ATMVP candidates may be included in either or both of the merge candidate list and the affine merge candidate list. Alternatively, the generated list of motion information predictor candidates does not include ATMVP candidates.
[0359] In the aforementioned modified embodiment, the maximum number of candidates that can be included in the list of merge index and affine merge index candidates is fixed. The maximum number of candidates that can be included in the list of merge index and affine merge index candidates may be the same. Data for determining (or indicating) the maximum number (or target number) of motion information predictor candidates that can be included in the generated list of motion information predictor candidates is included in the bitstream by the encoder, and the decoder obtains data from the bitstream for determining the maximum number (or target number) of motion information predictor candidates that can be included in the generated list of motion information predictor candidates. This allows the data for decoding the merge index or affine merge index to be analyzed from the bitstream. This data for determining (or indicating) the maximum number (or target number) may be the maximum number (or target number) itself when decoded, or the decoder may be able to determine this maximum / target number in relation to other parameters / syntax elements, for example, "five_minus_max_num_merge_cand" or "MaxNumMergeCand-1" or a functionally equivalent parameter used in HEVC.
[0360] Alternatively, if the maximum number of candidates (or target number) in the list of candidates for the merge index and affine merge index can change or differ (for example, because the use of ATMVP candidates or any other candidate is enabled or disabled for one list but not for the other, or because the lists use different candidate list generation / derivation processes), then when affine merge mode is used, and when merge mode is used, the maximum number of motion predictor candidates (or target number) that can be included in the generated list of motion predictor candidates can be determined separately, and the encoder includes data in the bitstream to determine the maximum / target number. The decoder then retrieves the data from the bitstream to determine the maximum / target number and uses the retrieved data to parse or decode the motion predictor index. Then, an affine flag can be used to switch, for example, between parsing or decoding the merge index and the affine merge index.
[0361] As described above, one or more of the additional interprediction modes (such as MHII merge mode, triangle merge mode, and MMVD merge mode) can be used in addition to or instead of the merge mode or affine merge mode, and an index (or flag or information) for one or more of the additional interprediction modes can be signaled (encoded or decoded). The following embodiments relate to the signaling of information (such as an index) for the additional interprediction modes.
[0362] Embodiment 17 Signaling for all interprediction modes (including merge mode, affine merge mode, MHII merge mode, triangle merge mode, and MMVD merge mode) These multiple interprediction "merge" modes are signaled using data provided in the bitstream along with their associated syntax (elements) according to the 17th embodiment. Figure 26 shows the decoding process for the interprediction mode for the current CU (image portion or block) according to one embodiment of the present invention. As described in relation to Figure 12 (and the skip flag in 1201), the first CU skip flag is extracted from the bitstream (2601). If the CU is not skip (2602), i.e., the current CU is not processed in skip mode, the pred mode flag (2603) and / or merge flag (2606) are decoded to determine whether the current CU is a merge CU. If the current CU is processed as a merge skip (2602) or merge CU (2607), the MMVD_Skip_Flag or MMVD_Merge_Flag is decoded (2608). If this flag is equal to 1 (2609), the current CU is decoded using MMVD merge mode (i.e., in MMVD merge mode or in MMVD merge mode), resulting in the decoding of the MMVD merge index (2610), followed by the decoding of the MMVD distance index (2611) and the MMVD direction index (2612). If the CU is not an MMVD merge CU (2609), the merge subblock flag is decoded (2613). This flag was also referred to as the "affine flag" in the previous description. If the current CU is processed in affine merge mode (also known as "subblock merge" mode) (2614), the merge subblock index (i.e., the affine merge index) is decoded (2615). If the current CU is not processed in affine merge mode (2614) and not in skip mode (2616), the MHII merge flag is decoded (2620). If this block is processed in MHII merge mode (2621), the regular merge index (2619) is decoded using the associated intra-prediction mode (2622) for MHII merge mode. Note that MHII merge mode is only available for non-skip "merge" mode and not for skip mode.If the MHII merge flag is equal to 0 (2621), or if the current CU is not processed in skip mode (2616) or affine merge mode (2614), the triangle merge flag is decoded (2617). If this CU is processed in triangle merge mode (2618), the triangle merge index is decoded (2623). If the current CU is not processed in triangle merge mode (2618), the current CU is a regular merge mode CU and the merge index is decoded.
[0363] Signaling for each merge candidate MMVD Merge Flag / Index Signaling In the first modification of the 17th embodiment, only two initial candidates are available for use / selection in MMVD merge mode. However, if eight possible values for the distance index and four possible values for the direction index are also signaled along with the bitstream, the number of potential candidates for use in MMVD merge mode in the decoder is 64 (2 candidates × 8 distance indices × 4 direction indices), and each potential candidate is different from another (i.e., unique) if the initial candidates are different. These 64 potential candidates can be evaluated / compared for MMVD merge mode on the encoder side, and then the MMVD merge index (2610) for the selected initial candidate is signaled with a unary max code. Since only two initial candidates are used, this MMVD merge index (2610) corresponds to a flag. Figure 27(a) shows the encoding of this flag, which is CABAC encoded using one context variable. In another variation, it should be understood that a different number of initial candidates, distance index values, and / or direction index values can be used instead, and the signaling of the MMVD merge index will be adapted accordingly (for example, at least one bit will be CABAC encoded using one context variable).
[0364] Triangle Merge Index Signaling In the first modification of the 17th embodiment, the triangle merge index is signaled differently compared to the index signaling for other interprediction modes. In the triangle merge mode, 40 possible substitutions of candidates are available, corresponding to combinations of five initial candidates and two possible types of triangles (see Figures 25(a) and 25(b), where two possible first block predictors (2501 or 2511) and second block predictors (2502 or 2512) are for each type of triangle). Figure 27(b) shows the encoding of the index for the triangle merge mode, i.e., the signaling of these candidates. The first bit (i.e., the first bin) is CABAC decoded in one context. If this first bit is equal to 0, the second bit (i.e., the second bin) is CABAC bypass decoded. If this second bit is equal to 0, the index corresponds to the first candidate in the list, i.e., index 0 (Cand0). Otherwise (when the second bit is equal to 1), the index corresponds to the second candidate in the list, i.e., index 1 (Cand1). If the first bit is equal to 1, the Exponential-Golomb code is extracted from the bitstream, and the Exponential-Golomb code represents the index of the selected candidate in the list. That is, it is selected from index 2 (Cand2) to index 39 (Cand39).
[0365] In another variation, a different number of initial candidates may be used instead, and the signaling of the triangle merge index may be adapted accordingly (for example, at least one bit may be CABAC encoded using one context variable).
[0366] Affine Merger List ATMVP In a second modification of the 17th embodiment, ATMVP is available as a candidate for the affine merge candidate list (i.e., in affine merge mode, also known as the "subblock merge" mode). Figure 28 shows a list of affine merge list derivations with this additional ATMVP candidate (2848). This figure is similar to Figure 24 (described earlier), but since this additional ATMVP candidate (2848) has been added to the list, a repetition of the detailed explanation is omitted here. In another modification, a different number of initial candidates may be used instead, and the signaling of the triangle merge index is adapted accordingly (e.g., at least one bit is CABAC encoded using one context variable).
[0367] In another variation, it is understood that ATMVP candidates may be added to a list of candidates for another interprediction mode, and the signaling of that index is adapted accordingly (for example, at least one bit is CABAC encoded using one context variable).
[0368] Figure 26, in another modification, provides a complete overview of the signaling for all interpredictive modes (i.e., merge mode, affine merge mode, MHII merge mode, triangle merge mode, and MMVD merge mode), although it should be understood that only a subset of interpredictive modes may be used instead.
[0369] Embodiment 18 According to the 18th embodiment, either or both of the triangle merge mode or the MMVD merge mode are available for use in the encoding or decoding process, and either or both of these interprediction modes share context variables (used with CABAC encoding) with another interprediction mode when signaling its index / flags.
[0370] In this embodiment or further variations of the embodiments below, it is understood that one or more interprediction modes may use two or more context variables when signaling their index / flags (for example, an affine merge mode may use four or five context variables, depending on whether ATMVP candidates can also be included in the list for its affine merge index coding / decoding process).
[0371] For example, before this embodiment or any modification of the following embodiment is implemented, the total number of context variables for signaling all bits of the index / flags for all interprediction modes may be 7: (normal) merge = 1 (as shown in Figure 10(a)); affine merge = 4 (as shown in Figure 10(b), however, there is one less candidate, e.g., no ATMVP candidate); triangle = MMVD = 1; and MHII (if available) = 0 (shared with normal merge). Then, by implementing the modification, the total number of context variables for signaling all bits of the index / flags for all interprediction modes can be reduced to 5: (normal) merge = 1 (as shown in Figure 10(a)); affine merge = 4 (as shown in Figure 10(b), however, there is one less candidate, e.g., no ATMVP candidate); and triangle = MMVD = MHII (if available) = 0 (shared with normal merge).
[0372] In another example, before this modification is implemented, the total number of context variables for signaling all bits of the index / flag for all interprediction modes can be 4:(normal)merge=affinemerge=triangle=MMVD=1 (as shown in Figure 10(a)); and MHII(if available)=0 (shared with normalmerge). Then, by implementing this modification, the total number of context variables for signaling all bits of the index / flag for all interprediction modes is reduced to 2:(normal)merge=affinemerge=1 (as shown in Figure 10(a)); and triangle=MMVD=MHII(if available)=0 (shared with normalmerge).
[0373] For the sake of simplicity in the following explanation, note that we will describe whether or not a single context variable (e.g., only the first bit) is shared. This means that in the following explanation, we will often see a simple case where a context variable is used to signal only the first bit for each interprediction mode, and this context variable is either 1 (a separate / independent context variable is used) or 0 (this bit is either bypassed CABAC encoded or shares the same context variable as another interprediction mode, so no separate / independent context variable exists). It will be understood that the context variables of other bits, or in fact all bits, may be shared / not shared / bypassed CABAC encoded in the same way, and that this embodiment and the different variations of the embodiments below are not limited thereto.
[0374] In the first modification of the 18th embodiment, all interprediction modes available for use in the encoding or decoding process share at least some CABAC contexts.
[0375] In this modification, the relevant parameters for index coding and interprediction modes (e.g., the number of (initial) candidates) can be set to be the same or similar as long as possible / compatible. For example, to simplify signaling, the number of candidates for affine merge mode and merge mode can be set to 5 and 6, respectively, the initial number of candidates for MMVD merge mode can be set to 2, and the maximum number of candidates for triangle merge mode can be 40. Also, triangle merge index is not signaled using unary maximal code like the other interprediction modes. In this triangle merge mode, only the first bit of the context variable (in the case of triangle merge index) can be shared with the other interprediction modes. The advantage of this modification is that the encoder and decoder design is simplified.
[0376] In a further variation, the CABAC content for all merge interpretation mode indices is shared. This means that only one CABAC context variable is required for the first bit of all indices. In yet another variation, if an index contains two or more bits to be CABAC encoded, the encoding of the additional bits (all CABAC encoded bits away from the first bit) is treated as a separate part (i.e., as if it were for a different syntax element as far as the CABAC encoding process is concerned), and if two or more indices have two or more bits to be CABAC encoded, one same context variable is shared for these CABAC encoded "additional" bits. The advantage of this variation is a reduction in the amount of CABAC context. This reduces the memory requirements for the context state that needs to be stored on the encoder and decoder sides without significantly impacting the encoding efficiency for most sequences processed by video codecs that implement the variation.
[0377] Figure 29 shows the decoding process for another further modification for interprediction mode. This figure is similar to Figure 26 but includes an implementation of this modification. In this figure, when the current CU is processed in MMVD merge mode, its MMVD merge index is decoded as the same index as the merge index in normal merge mode (i.e., the “merge index” (2919)). However, unlike normal merge mode, in MMVD merge mode there are only two initial candidates to choose from, not six. Since there are only two possibilities, this “shared” index used in MMVD merge mode is essentially a flag. Since the same index is shared, the CABAC context variable is the same for this flag in MMVD merge mode and for the first bit of the merge index in merge mode. Next, if it is determined that the current CU should be processed in MMVD merge mode (2925), the distance index (2911) and direction index (2912) are decoded. If the current CU is determined to be processed in affine merge mode (2914), its affine merge index is decoded as the same index as the merge index in normal merge mode (i.e., the "merge index" (2919)). However, in affine merge mode, unlike normal merge mode, the maximum number of candidates (i.e., the maximum number of indices) is 5, not 6. If the current CU is determined to be processed in triangle merge mode (2918), the first bit is decoded as a shared index (2919), and as a result, the same CABAC context variable is shared with normal merge mode. If this CU is processed in triangle merge mode (2926), the remaining bits related to the triangle merge index are decoded (2923).
[0378] Therefore, for example, when processing these indices / flags during CABAC coding, the number of separate (independent) context variables used for the first bit of the index / flag for each interprediction mode is: (Regular) Merge = 1; MHII = Affine Merge = Triangle = MMVD = 0 (Shared with regular merge) That is the case.
[0379] In the second variation, when either or both of the triangle merge mode or the MMVD merge mode are used (i.e., information regarding the motion information predictor selection of the current CU is processed / encoded / decoded in the relevant interprediction mode), its / their index signaling shares context variables with the merge mode index signaling. In this variation, the CABAC context of the triangle merge index and / or the CABAC context of the MMVD merge index / flag share the same CABAC context as the merge index (in the (canonical) merge mode). This means that at least one CABAC state is required for these three modes.
[0380] In a further variation of the second variation, the CABAC context of the triangle merge index and / or the CABAC context of the MMVD merge index / flag share the same first CABAC context variable of the merge index, for example, the same context variable of the first bit of the merge index.
[0381] Therefore, for example, when processing these indices / flags during CABAC encoding, the number of separate (independent) context variables used for the first bit of the index / flag is: (Regular) Merge = 1; MHII (if available) = Affine merge (if available) = 0 (shared with regular merge) or 1 depending on the implementation; Triangle = MMVD = 0 (Shared with regular merge) That is the case.
[0382] In yet another variation of the second variation, if two or more context variables are used for triangle merge index CABAC encoding / decoding, or if two or more context variables are used for MMVD merge index CABAC encoding / decoding, they can all be shared with, or at least partially shared with, two or more CABAC context variables used for merge index CABAC encoding / decoding, or at least partially shared whenever they are compatible.
[0383] The advantage of this second variation is that the amount of context that needs to be stored is reduced, and as a result, the amount of state that needs to be stored on the encoder and decoder side is reduced without significantly affecting the encoding efficiency of most of the sequences processed by the video codecs that implement them.
[0384] In a third variation, when either or both of the triangle merge mode or the MMVD merge mode are used (i.e., information regarding the motion information predictor selection of the current CU is processed / encoded / decoded in the relevant interprediction mode), its / their index signaling shares context variables with the index signaling of the affine merge mode. In this variation, the CABAC context of the triangle merge index and / or the CABAC context of the MMVD merge index / flag share the same CABAC context as the affine merge index (for the affine merge mode).
[0385] In a further variation of the third variation, the CABAC context of the triangle merge index and / or the CABAC context of the MMVD merge index / flag share the same first CABAC context variable as the affine merge index, for example, the same context variable as the first bit of the affine merge index.
[0386] Therefore, for example, when processing these indices / flags during CABAC encoding, the number of separate (independent) context variables used for the first bit of the index / flag is: (Regular) merge (if available) = 0 (shared with affine merge) or 1 depending on the implementation; MHII(if available) = 0 (regular merge and sharing); Affine merge = 1; Triangle = MMVD = 0 (shared with affine merge) That is the case.
[0387] In yet another variation of the third variation, if two or more context variables are used for triangle merge index CABAC encoding / decoding, or if two or more context variables are used for MMVD merge index CABAC encoding / decoding, they are all shared, or can be at least partially shared if they are compatible with two or more CABAC context variables used for affine merge index CABAC encoding / decoding.
[0388] In the fourth variation, when MMVD merge mode is used (i.e., information regarding the motion information predictor selection of the current CU is processed / encoded / decoded in MMVD merge mode), its index signaling shares context variables with the index signaling of merge mode or affine merge mode. In this variation, the CABAC context of the MMVD merge index / flag is the same CABAC context as the merge index or the same CABAC context as the affine merge index.
[0389] Therefore, for example, when processing these indices / flags during CABAC encoding, the number of separate (independent) context variables used for the first bit of the index / flag is: (Regular) Merge = 1; MHII(if available) = 0 (regular merge and sharing); Affine merge (if available) = 0 (shared with regular merge) or 1 depending on the implementation; MMVD=0 (Regular merge and shared) or (Regular) merge (if available) = 0 (shared with affine merge) or 1 depending on the implementation; MHII(if available) = 0 (regular merge and sharing); Affine merge = 1; MMVD=0 (shared with affine merge) That is the case.
[0390] In the fifth variation, when triangle merge mode is used (i.e., information regarding the motion information predictor selection of the current CU is processed / encoded / decoded in triangle merge mode), its index signaling shares context variables with the index signaling for merge mode or affine merge mode. In this variation, the CABAC context of the triangle merge index is the same CABAC context of the merge index or the same CABAC context of the affine merge index.
[0391] Therefore, for example, when processing these indices / flags during CABAC encoding, the number of separate (independent) context variables used for the first bit of the index / flag is: (Regular) Merge = 1; MHII(if available) = 0 (regular merge and sharing); Affine merge (if available) = 0 (shared with regular merge) or 1 depending on the implementation; Triangle = 0 (Regular merge and sharing) or (Regular) merge (if available) = 0 (shared with affine merge) or 1 depending on the implementation; MHII(if available) = 0 (regular merge and sharing); Affine merge = 1; Triangle = 0 (shared with affine merge) That is the case.
[0392] In the sixth variation, when triangle merge mode is used (i.e., information regarding the motion information predictor selection of the current CU is processed / encoded / decoded in triangle merge mode), its index signaling shares context variables with the index signaling of MMVD merge mode. In this variation, the CABAC context of the triangle merge index is the same CABAC context as the MMVD merge index. Therefore, for example, when processing these indices / flags during CABAC encoding, the number of separate (independent) context variables used for the first bit of the index / flag is: MMVD=1; Triangle = 0 (shared with MMVD); (Regular) merge or MHII or Affine merge = depending on whether it is implemented and available. That is the case.
[0393] In the seventh variation, if either or both of the triangle merge mode or the MMVD merge mode are used (i.e., information regarding the motion information predictor selection of the current CU is processed / encoded / decoded in the relevant inter-prediction mode), then its / their index signaling shares context variables with the index signaling for the inter-prediction mode, which may include ATMVP predictor candidates in its list of candidates, i.e., the inter-prediction mode may have ATMVP predictor candidates as one of the available candidates. In this variation, the CABAC context of the triangle merge index and / or the MMVD merge index / flags share the same CABAC context of the index of the inter-prediction mode which may use ATMVP predictors.
[0394] In further variations, the CABAC context variable for the triangle merge index and / or MMVD merge index / flag is shared with the containable ATMVP candidate for the merge index in merge mode, or for the affine merge index in affine merge mode.
[0395] In yet another variation, if two or more context variables are used for triangle merge index CABAC encoding / decoding, or if two or more context variables are used for MMVD merge index CABAC encoding / decoding, they can all be shared, or at least partially shared, with two or more CABAC context variables used for affine merge index or merge index in affine merge mode that have inclusive ATMVP candidates.
[0396] Since ATMVP (predictor) candidates are the predictors that benefit most from CABAC adaptation compared to other types of predictors, the advantage of these modifications is improved coding efficiency.
[0397] Embodiment 19 According to the 19th embodiment, either or both of the triangle merge mode or the MMVD merge mode are available for use in the encoding or decoding process, and the index / flags for either or both of these interprediction modes are CABAC bypass encoded when signaling the index / flags.
[0398] In the first modification of the 19th embodiment, all interprediction modes available for use in the encoding or decoding process signal the index / flags by CABAC bypass encoding / decoding their indices / flags. In this modification, all indices of all interprediction modes are encoded without using CABAC context variables (e.g., by the bypass encoding engine 1705 in Figure 17). This means that all bits of the merge index (2619), affine merge index (2615), MMVD merge index (2610), and triangle merge index (2623) in Figure 26 are CABAC bypass encoded. Figures 30(a) to 30(c) show the encoding of indices / flags according to this embodiment. Figure 30(a) shows the MMVD merge index encoding of the initial MMVD merge candidate. Figure 30(b) shows the triangle merge index encoding. Figure 30(c) shows the affine merge index encoding, which can also be easily used for merge index encoding.
[0399] Therefore, for example, when processing these indices / flags during CABAC encoding, the number of separate (independent) context variables used for the indices / flags is: (Normal) Merge (if available) = MHII (if available) = Affine Merge (if available) = Triangle (if available) = MMVD (if available) = 0 (all bypass codes). The advantage of this modification is a reduction in the amount of context that needs to be stored, and consequently a reduction in the amount of state that needs to be stored on the encoder and decoder sides, with only a slight impact on the encoding efficiency of most sequences processed by video codecs that implement the modification. However, it should be noted that the loss may be significant when used to encode screen content. This modification represents another compromise between encoding efficiency and complexity when compared to other modifications / embodiments. The impact on encoding efficiency is often small. In fact, with the many available interprediction modes, the average amount of data required to signal an index for each interprediction mode is smaller than the average amount of data required to signal a merge index when only merge modes are available (this comparison is for the same sequence and the same encoding efficiency compromise). This means that the efficiency of CABAC encoding / decoding from context-based bin probability adaptation may be inefficient.
[0400] In a second variation, when either or both of the triangle merge mode or the MMVD merge mode are used (i.e., information regarding the motion information predictor selection of the current CU is processed / encoded / decoded in the relevant interprediction mode), their indices / flags are signaled by CABAC bypass encoding / decoding the indices / flags. In this variation, the MMVD merge index and / or triangle merge index are CABAC bypass encoded. Depending on the implementation, i.e., when merge mode and affine merge mode are available, the merge index and affine merge index have their own context. In yet another variation, the context of the affine merge index and the merge index are shared. Therefore, for example, when processing these indices / flags during CABAC encoding, the number of separate (independent) context variables used for the first bit of the index / flag is: (Regular) merge = Affine merge = 0 or 1 depending on the implementation; MHII (if available) = 0 (shared with regular merge); Triangle = MMVD = 0 (bypass encoding) That is the case.
[0401] The advantage of these modifications is that they offer improved coding efficiency compared to previous modifications, as there is yet another compromise between reducing the CABAC context and coding efficiency. In fact, the triangle merge mode is not often selected. Consequently, when its context is removed, i.e., when the triangle merge mode uses CABAC bypass coding, the impact on coding efficiency is small. The MMVD merge mode tends to be selected more frequently than the triangle merge mode, but the probability of selecting the first and second candidates for the MMVD merge mode tends to be equal to or greater than that of other interprediction modes such as merge mode or affine merge mode, and therefore the advantage gained from using the context of CABAC coding in the MMVD merge mode is not so great. Another advantage of these modifications is the small impact on coding efficiency for screen content sequences, since the merge mode is the most influential interprediction mode for screen content.
[0402] In the third variation, when merge mode, triangle merge mode, or MMVD merge mode is used (i.e., information regarding the motion information predictor selection of the current CU is processed / encoded / decoded in the relevant interprediction mode), its / their indices / flags are signaled by CABAC bypass encoding / decoding the indices / flags. In this variation, the MMVD merge index, triangle merge index, and merge index are CABAC bypass encoded. Therefore, for example, when processing these indices / flags during CABAC encoding, the number of separate (independent) context variables used for the first bit of the index / flag is: (Normal) Merge = Triangle (if available) = MMVD (if available) = 0 (Bypass Encoding); MHII (if available) = 0 (same as a regular merge); and Affine merge = 1 That is the case.
[0403] This modification offers an alternative compromise compared to other modifications, for example, it results in a greater encoding efficiency reduction for screen content sequences than the previous modification.
[0404] In the fourth variation, when affine merge mode, triangle merge mode, or MMVD merge mode is used (i.e., information regarding the motion information predictor selection of the current CU is processed / encoded / decoded in the relevant interprediction mode), its / their indices / flags are signaled by CABAC bypass encoding / decoding the indices / flags. In this variation, the affine merge index, MMVD merge index, and triangle merge index are CABAC bypass encoded, and the merge index is encoded in one or more CABAC contexts. Thus, for example, when processing these indices / flags during the CABAC encoding process, the number of separate (independent) context variables used for the first bit of the index / flag is: (Regular) Merge = 1; Affine merge = Triangle (if available) = MMVD (if available) = 0 (bypass encoding); MHII (if available) = 0 (regular merge and share). The advantage of this variant compared to previous variants is the increased encoding efficiency for screen content sequences.
[0405] In the fifth variation, an interprediction mode available for use in the encoding or decoding process encodes / decodes the index / flag via CABAC bypass in order to signal the index / flag, except that the interprediction mode can include ATMVP predictor candidates in its list of candidates, i.e., the interprediction mode can have ATMVP predictor candidates as one of the available candidates. In this variation, all indices of all interprediction modes are CABAC bypass encoded, except that the interprediction mode can have ATMVP predictor candidates. Therefore, for example, when processing these indices / flags during the CABAC encoding process, the number of separate (independent) context variables used for the first bit of the index / flag is: (Canonical) merge with inclusive ATMVP candidates = 1; Affine Merge (if available) = Triangle (if available) = MMVD (if available) = 0 (Bypass Encoding); Depending on the implementation, MHII (if available) = 1 (shared with regular merge) or 0 or Affine merge = 1 with an inclusive ATMVP candidate; (Regular) Merge (if available) = Triangle (if available) = MMVD (if available) = 0 (Bypass Encoding); MHII (if available) = 0 (same as regular MERGE) That is the case.
[0406] This modification introduces other complexity / encoding efficiency compromises to most natural sequences. However, it should be noted that for screen content sequences, it may be preferable to have ATMVP predictor candidates within the canonical merge candidate list.
[0407] In the sixth variation, the interprediction mode available for use in the encoding or decoding process is not a skip mode (e.g., not one of the normal merge skip mode, affine merge skip mode, triangle merge skip mode, or MMVD merge skip mode), and the index / flag is CABAC bypass encoded / decoded to signal the index / flag. In this variation, all indices are CABAC bypass encoded for any CUs that are not processed in skip mode, i.e., not skipped. For skipped CUs (i.e., CUs processed in skip mode), the index may be processed using one of the CABAC encoding techniques described in relation to the above embodiments / variations (e.g., only the first bit, or two or more bits, have a context variable, which may or may not be shared).
[0408] Figure 31 is a flowchart of the decoding process in interprediction mode illustrating this modification. The process in Figure 31 is similar to Figure 29, except that it has an additional “skip mode” determination / check step (3127), after which the index / flag ("Merge_idx") is decoded using either CABAC decoding (3119) or CABAC bypass decoding (3128). According to yet another modification, as a result of the determination / check performed in the previous step, the CU is skip (2902 / 3102), MMVD skip (2908 / 3108), and it is understood that the CU performs the “skip mode” determination / check instead of the additional “skip mode” determination / check step (3127) using the skip (2916 / 3116) in Figure 29 or Figure 31.
[0409] This variation has little impact on coding efficiency because skip modes are generally selected more frequently than non-skip modes (i.e., non-skip interprediction modes such as normal merge mode, MHII merge mode, affine merge mode, triangle merge mode, or MMVD merge mode), and the selection of the first candidate is more likely in skip modes than in non-skip modes. Since SKIP modes are designed for more predictable movements, their indices must also be more predictable. Therefore, the probability of utilizing CABAC coding / decoding is likely to be more useful for skip modes. However, non-skip modes are more likely to be used when movements are less predictable, as this increases the likelihood of more random selection from predictor candidates. Therefore, CABAC coding / decoding is less likely to be efficient in non-skip modes.
[0410] 20th Embodiment According to the 20th embodiment, the data is provided as a bitstream and is used to determine whether an index / flag should be signaled for one or more of the interpredictive modes by using CABAC bypass coding / decoding, CABAC coding / decoding with separate context variables, or CABAC coding / decoding with one or more shared context variables. For example, such data could be flags for enabling or disabling the use of one or more independent contexts for index coding / decoding of the interpredictive modes. Using such data, it is possible to control the use or non-use of context sharing in CABAC coding / decoding or CABAC bypass coding / decoding.
[0411] In a variation of the 20th embodiment, the CABAC context shared between two or more indices of two or more interprediction modes depends on data transmitted in a bitstream at a level higher than, for example, the CU level (e.g., at the level of image portions greater than the smallest CU, such as sequence, frame, slice, tile, or CTU level). For example, this data may indicate that for any CU within a particular image portion, the CABAC context of the merge index of a merge mode is shared (or not shared) with one or more other CABAC contexts of another interprediction mode.
[0412] In another variation, one or more indices are CABAC bypass encoded in the bitstream in response to data transmitted at a level higher than the CU level (e.g., at the slice level). For example, this data might indicate that for any CU within a particular image portion, the index for a particular interprediction mode should be CABAC bypass encoded.
[0413] In one variation, to further improve coding efficiency, the encoder may select a value for this data to indicate the sharing of context of one or more indices of one or more interprediction modes, or CABAC bypass coding / decoding may be selected based on how frequently one or more interprediction modes are used in previously coded frames. An alternative is to select a value for this data based on the type of sequence being processed or the type of application in which the variation is implemented.
[0414] The advantage of this embodiment is the increased controlled coding efficiency compared to the previous embodiment / modification.
[0415] Embodiments of the present invention One or more of the embodiments described above are implemented by the processor 311 of the processing device 300 in Figure 3, or the corresponding functional module / unit of the decoder 60 in Figure 5, the CABAC coder in Figure 17, the encoder 400 in Figure 4, or the corresponding CABAC decoder, and perform one or more method steps of the embodiments described above.
[0416] Figure 19 is a schematic block diagram of a computing device 2000 for implementing one or more embodiments of the present invention. The computing device 2000 may be a device such as a microcomputer, a workstation, or a light portable device. The computing device 2000 includes: - a central processing unit (CPU) 2001 such as a microprocessor; - random access memory (RAM) 2002 for storing executable code of the method of an embodiment of the present invention and registers for recording variables and parameters necessary to implement a method for encoding or decoding at least a portion of an image according to an embodiment of the present invention, the memory capacity of which can be expanded, for example, by optional RAM connected to an expansion port; - read-only memory (ROM) 2003 for storing a computer program for implementing an embodiment of the present invention; and - a communication bus connected to a network interface (NET) 2004, which is typically connected to a communication network on which digital data to be processed is transmitted or received. The network interface (NET) 2004 may be a single network interface or may consist of a set of different network interfaces (e.g., wired and wireless interfaces, or different types of wired or wireless interfaces). Data packets are written to the network interface for transmission or read from the network interface for reception, under the control of a software application running on CPU2001. - A user interface (UI) 2005 may be used to receive user input or display information to the user. - A hard disk (HD) 2006 may be provided as mass storage. - An input / output module (IO) 2007 may be used to send and receive data to and from external devices such as video sources or displays. Executable code can be stored in ROM2003, HD2006, or any removable digital medium such as a disk.In a modified version, the executable code of a program can be received via a communication network through NET2004 so that it is stored in one of the storage means of a communication device 2000, such as HD2006, before execution. The CPU2001 is adapted to control and direct the execution of instructions or parts of the software code of a program or program according to an embodiment of the present invention, where the instructions are stored in one of the aforementioned storage means. After power-on, the CPU2001 can execute instructions relating to a software application from the main RAM memory 2002, for example, after these instructions have been loaded from the program ROM 2003 or HD2006. When such a software application is executed by the CPU2001, it causes the steps of the method according to the present invention to be performed.
[0417] Furthermore, according to other embodiments of the present invention, the decoder according to the above embodiments may be provided to a user terminal such as a computer, a mobile phone, a tablet, or any other type of device capable of providing / displaying content to a user (e.g., a display device). In yet another embodiment, the encoder according to the above embodiments may be provided in an image capture device that also includes a camera, video camera, or network camera (e.g., a closed-circuit television or video surveillance camera) that captures and provides content for the encoder to encode. Two such examples are provided below with reference to Figures 20 and 21.
[0418] Figure 20 shows a network camera system 2100, which includes a network camera 2102 and a client device 2104.
[0419] The network camera 2102 includes an imaging unit 2106, an encoding unit 2108, a communication unit 2110, and a control unit 2112.
[0420] The network camera 2102 and the client device 2104 are interconnected via the network 200 so that they can communicate with each other.
[0421] The imaging unit 2106 includes a lens and an image sensor (e.g., a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS)) and captures an image of an object and generates image data based on that image. This image may be a still image or a video image. The imaging unit may also include zoom and / or pan means adapted to zoom or pan (optically or digitally).
[0422] The encoding unit 2108 encodes the image data using the encoding method described in one or more of the embodiments described above. The encoding unit 2108 uses at least one of the encoding methods described in the embodiments described above. In other examples, the encoding unit 2108 can use a combination of the encoding methods described in the embodiments described above.
[0423] The communication unit 2110 of the network camera 2102 transmits the encoded image data encoded by the encoding unit 2108 to the client device 2104. The communication unit 2110 also receives commands from the client device 2104. These commands include commands to set the encoding parameters for the encoding unit 2108.
[0424] The control unit 2112 controls other units within the network camera 2102 according to commands received by the communication unit 2110.
[0425] The client device 2104 includes a communication unit 2114, a decoding unit 2116, and a control unit 2118. The communication unit 2114 of the client device 2104 transmits commands to the network camera 2102. The communication unit 2114 of the client device 2104 also receives encoded image data from the network camera 2102.
[0426] The decoding unit 2116 decodes the encoded image data using the decoding method described in one or more of the embodiments described above. In other examples, the decoding unit 2116 can use a combination of the decoding methods described in the embodiments described above.
[0427] The control unit 2118 of the client device 2104 controls other units within the client device 2104 according to user operations and commands received by the communication unit 2114. The control unit 2118 of the client device 2104 controls the display device 2120 to display the image decoded by the decoding unit 2116. The control unit 2118 of the client device 2104 also controls the display device 2120 to display a GUI (Graphical User Interface) and specifies the parameter values of the network camera 2102, including the encoding parameters of the encoding unit 2108.
[0428] Furthermore, the control unit 2118 of the client device 2104 controls other units within the client device 2104 in response to user input to the GUI displayed by the display device 2120. The control unit 2118 of the client device 2104 controls the communication unit 2114 of the client device 2104 to send a command to the network camera 2102 specifying the parameter values for the network camera 2102 in response to user input to the GUI displayed by the display device 2120.
[0429] The network camera system 2100 can determine whether camera 2102 is using zoom or pan while recording video, and such information can be used when encoding the video stream, as zoom or pan during shooting can benefit from the use of affine modes, which are well suited to encoding complex movements such as zoom, rotation, and / or stretching (which can be a side effect of panning, especially if the lens is a "fisheye" lens).
[0430] Figure 21 shows a smartphone 2200.
[0431] The smartphone 2200 comprises a communication unit 2202, a decoding / encoding unit 2204, a control unit 2206, and a display unit 2208.
[0432] The communication unit 2202 receives encoded image data via the network 200.
[0433] The decoding / encoding unit 2204 decodes the encoded image data received by the communication unit 2202. The decoding / encoding unit 2204 decodes the encoded image data using the decoding method described in one or more of the embodiments described above. The decoding / encoding unit 2204 can use at least one of the decoding methods described in the embodiments described above. In other examples, the decoding / encoding unit 2204 can use a combination of the decoding or encoding methods described in the embodiments described above.
[0434] The control unit 2206 controls other units within the smartphone 2200 in response to user operations or commands received by the communication unit 2202 or via the input unit. For example, the control unit 2206 controls the display device 2208 to display the image decoded by the decoding unit 2204.
[0435] The smartphone may further include an image recording device 2210 (e.g., a digital camera and associated circuitry) for recording images or videos. Such recorded images or videos may be encoded by a decoding / encoding unit 2204 under the direction of a control unit 2206. The smartphone may further include a sensor 2212 configured to sense the orientation of the mobile device. Such a sensor may include an accelerometer, gyroscope, compass, global positioning (GPS) unit, or similar position sensor. Such a sensor 2212 can determine whether the smartphone is changing orientation, and such information is used when encoding the video stream as a change in orientation during recording, benefiting from the use of affine modes, which are well suited to encoding complex movements such as rotation.
[0436] Alternatives and changes The object of the present invention is to ensure that affine modes are utilized in the most efficient manner, and the particular examples described above will be understood to relate to signaling the use of affine modes, depending on the likelihood that affine modes are perceived as useful. Further examples of this can be applied to encoders when it is known that complex motion (where affine transforms may be particularly efficient) is being encoded. An example of such a case is, a) Camera zoom in / out b) A portable camera (e.g., a mobile phone) that changes direction during shooting (i.e., rotational motion). c) Panning with a "fisheye" lens camera (e.g., stretching / distorting part of the image) Includes.
[0437] Therefore, complex motion instructions can be increased during the recording process, making it more likely that affine mode will be used for slices, frame sequences, or even the entire video stream.
[0438] In further examples, affine mode is more likely to be used depending on the characteristics or functionality of the device used to record video. For example, mobile devices are more likely to change orientation than (e.g.) fixed security cameras, so affine mode may be more suitable for encoding video from the former. Examples of characteristics or functionality include the presence / use of zoom means, the presence / use of position sensors, the presence / use of pan means, whether the device is portable or not, or user selection on the device.
[0439] While the present invention has been described with reference to embodiments, it should be understood that the present invention is not limited to the disclosed embodiments. Those skilled in the art will understand that various changes and modifications can be made without departing from the scope of the invention, as defined in the appended claims. All features disclosed herein (including any appended claims, abstract, and drawings) and / or all steps of any method or process disclosed herein may be combined in any combination, except for any combination in which at least some of such features and / or steps are mutually exclusive. Each feature disclosed herein (including any appended claims, abstract, and drawings) may be replaced by an alternative feature serving the same, equivalent, or similar purpose unless otherwise specified. Therefore, unless otherwise specified, each feature disclosed is merely an example of a general series of equivalent or similar features.
[0440] Furthermore, it is understood that any result of the comparison, determination, evaluation, selection, execution, or consideration described above, such as a selection made during an encoding or filtering process, may be indicated in or determinable / inferred from data in the bitstream, such as flags or data indicating the result, and that the indicated or determined / inferred result may be used in the process, for example, during a decoding process, instead of actually performing the comparison, determination, evaluation, selection, execution, or consideration.
[0441] In the claims, the word “having” does not preclude other elements or steps, and the indefinite article “a” or “an” does not preclude the plural. The mere fact that different features are described in different dependent claims does not imply that combinations of these features cannot be used advantageously.
[0442] The reference numerals used in the claims are for illustrative purposes only and do not limit the scope of the claims.
[0443] In the embodiments described above, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as one or more instructions or codes on or transmitted through a computer-readable medium and executed by a hardware-based processing unit.
[0444] Computer-readable media may include computer-readable storage media corresponding to tangible media such as data storage media, or communication media including any media that facilitates the transfer of computer programs from one place to another, for example, according to a communication protocol. Thus, computer-readable media can generally correspond to (1) non-transient tangible computer-readable storage media, or (2) communication media such as signals or carrier waves. Data storage media may be any available media accessible by one or more computers or one or more processors for retrieving instructions, code and / or data structures for implementing the technology described herein. Computer program products may include computer-readable media.
[0445] Such computer-readable storage media can include, but are not limited to, any other media that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer, such as RAM, ROM, EEPROM, CD-ROM or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other media that can be used to store desired program code in the form of instructions or data structures. Also, any connection is appropriately called a computer-readable medium. For example, if instructions are transmitted from a website, server or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of a medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carriers, signals, or other temporary media, but instead refer to non-temporary, tangible storage media. As used herein, the terms "Disk" and "Disc" include Compact Disc (CD), LaserDisc, Optical Disc, Digital Multipurpose Disc (DVD), Floppy Disk (registered trademark), and Blu-ray Disc, where a Disc typically reproduces data magnetically, and a Disc reproduces data optically using a laser. The above combinations should also be included within the scope of computer-readable media.
[0446] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate / logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, the term “processor” as used herein may refer to any of the aforementioned structures or any other structure suitable for implementing the technology described herein. Furthermore, in some embodiments, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. The technology can also be fully implemented with one or more circuits or logic elements.
Claims
1. A method for encoding information about a motion information predictor, Selecting one of several motion information predictor candidates, The method involves encoding one of several indices, including a first index and a second index, for identifying the selected motion information predictor candidate, using Context-based Adaptive Binary Arithmetic Coding (CABAC) coding, Includes, The first index is used for a first merge mode in which block predictors can be obtained from a first block predictor associated with a first region in a block and a second block predictor associated with a second region in the block that is different from the first region. The second index is used for a second merge mode of interprediction mode, which is different from the first merge mode. The CABAC coding of the first bit of the first index for the first merge mode uses the same context variables as the CABAC coding of the first bit of the second index for the second merge mode, All bits of the first index except the first bit of the first index are bypass encoded, and all bits of the second index except the first bit of the second index are bypass encoded. Based on the aforementioned second merge mode, and using the average of the intrablock predictor and the interblock predictor, the block predictor can be obtained. A spatial merge candidate is available for the second merge mode. A method characterized by the following:
2. The method according to claim 1, characterized in that each of the first merge mode and the second merge mode is a merge mode independent of the merge mode that uses affine motion information.
3. The method according to claim 1, characterized in that each of the first and second regions has a shape different from a rectangle.
4. The first region in the block does not include the lower left vertex of the block, and includes the upper right vertex of the block. The method according to 3, characterized in that the second region in the block does not include the upper right vertex of the block, but includes the lower left vertex of the block.
5. The method according to claim 1, characterized in that a weighted average is applied to the region between the first region and the second region within the block.
6. A method for decoding information about a motion information predictor, To identify one of several motion information predictor candidates, one of several indices, including the first and second indices, is decoded using Context-based Adaptive Binary Arithmetic Coding (CABAC) decoding, Using the decoded index, select one of the multiple motion information predictor candidates, Includes, The first index is used for a first merge mode in which block predictors can be obtained from a first block predictor associated with a first region in a block and a second block predictor associated with a second region in the block that is a region different from the first region. The second index is used for a second merge mode of interprediction mode, which is different from the first merge mode. CABAC decoding of the first bit of the first index for the first merge mode uses the same context variables as CABAC decoding of the first bit of the second index for the second merge mode in interprediction mode, All bits of the first index except the first bit of the first index are bypass-decoded, and all bits of the second index except the first bit of the second index are bypass-decoded. Based on the aforementioned second merge mode, and using the average of the intrablock predictor and the interblock predictor, the block predictor can be obtained. A spatial merge candidate is available for the second merge mode. A method characterized by the following:
7. The method according to 6, characterized in that each of the first merge mode and the second merge mode is a merge mode independent of the merge mode that uses affine motion information.
8. The method according to 6, characterized in that each of the first and second regions has a shape different from a rectangle.
9. The first region in the block does not include the lower left vertex of the block, and includes the upper right vertex of the block. The method according to 8, characterized in that the second region in the block does not include the upper right vertex of the block, but includes the lower left vertex of the block.
10. The method according to 6, characterized in that a weighted average is applied to the region between the first region and the second region within the block.
11. A device for encoding information about motion information predictors, A means for selecting one of several motion information predictor candidates, Means for encoding one of a plurality of indices, including a first index and a second index for identifying the selected motion information predictor candidate, using Context-based Adaptive Binary Arithmetic Coding (CABAC) coding, Includes, The first index is used for a first merge mode in which block predictors can be obtained from a first block predictor associated with a first region in a block and a second block predictor associated with a second region in the block that is different from the first region. The second index is used for a second merge mode of interprediction mode, which is different from the first merge mode. The CABAC coding of the first bit of the first index for the first merge mode uses the same context variables as the CABAC coding of the first bit of the second index for the second merge mode, All bits of the first index except the first bit of the first index are bypass encoded, and all bits of the second index except the first bit of the second index are bypass encoded. Based on the aforementioned second merge mode, and using the average of the intrablock predictor and the interblock predictor, the block predictor can be obtained. A spatial merge candidate is available for the second merge mode. A device characterized by the following features.
12. A device for decoding information about motion information predictors, A means for decrypting one of several indices, including a first index and a second index, using Context-based Adaptive Binary Arithmetic Coding (CABAC) decoding in order to identify one of several motion information predictor candidates, A means for selecting one of the multiple motion information predictor candidates using the decoded index, Includes, The first index is used for a first merge mode in which block predictors can be obtained from a first block predictor associated with a first region in a block and a second block predictor associated with a second region in the block that is a region different from the first region. The second index is used for a second merge mode of interprediction mode, which is different from the first merge mode. CABAC decoding of the first bit of the first index for the first merge mode uses the same context variables as CABAC decoding of the first bit of the second index for the second merge mode in interprediction mode, All bits of the first index except the first bit of the first index are bypass-decoded, and all bits of the second index except the first bit of the second index are bypass-decoded. Based on the aforementioned second merge mode, and using the average of the intrablock predictor and the interblock predictor, the block predictor can be obtained. A spatial merge candidate is available for the second merge mode. A device characterized by the following features.
13. A computer program for causing a computer to perform the method described in claim 1.
14. A computer program for causing a computer to perform the method described in claim 6.
Citation Information
Patent Citations
Methods and apparatus for context-adaptive binary arithmetic coding of syntactic elements
JP2014531819A
Split Based Motion Vector Operation Reduction
US20180352223A1