Video encoding and decoding
By employing CABAC-based methods for encoding motion vector predictor indices with strategies like ATMVP and affine motion mode, the challenges of encoding complex motions in video coding are addressed, resulting in enhanced efficiency and reduced complexity.
Patent Information
- Application Number
- JP2025124997
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-12-20
- Filing Date
- 2025-07-25
- Publication Date
- 2025-10-22
AI Technical Summary
The existing video coding standards, such as HEVC, face challenges in efficiently encoding complex motions like zooming in/out, rotation, and perspective motion due to limited motion compensation tools, leading to increased coding complexity and reduced efficiency.
Implementing methods for encoding and decoding motion vector predictor indices using Context Adaptive Binary Arithmetic Coding (CABAC) with techniques like ATMVP and affine motion mode, including strategies to bypass or share contexts for bits of the motion vector predictor index, and utilizing context variables based on neighboring blocks for improved encoding efficiency.
Enhances coding efficiency by reducing complexity while maintaining good prediction accuracy for complex motions, achieving improved compression performance in video encoding.
Smart Images

Figure 2025160338000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to video encoding and decoding. [Background technology]
[0002] Recently, the Joint Video Experts Team (JVET), a joint team formed by MPEG and ITU-T Study Group 16's VCEG, initiated research into a new video coding standard called Versatile Video Coding (VVC). The goal of VVC is to provide significant improvements in compression performance over the existing HEVC standard (i.e., typically double the previous standard) and to be completed in 2020. Primary target applications and services include, but are not limited to, 360-degree and high dynamic range (HDR) video. Overall, JVET evaluated responses from 32 organizations using formal subjective testing conducted by an independent testing laboratory. Several proposals demonstrated compression efficiency gains of typically 40% or more compared to using HEVC. This was particularly evident for ultra-high-definition (UHD) video test material. Therefore, compression efficiency gains are expected to far exceed the 50% target for the final standard.
[0003] The JVET Search Model (JEM) uses all of the HEVC tools. An additional tool not present in HEVC is the use of "affine motion mode" when applying motion compensation. While motion compensation in HEVC is limited to translation, in practice there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. When utilizing affine motion mode, more complex transformations are applied to blocks to attempt to more accurately predict the formation of such motion. Therefore, it is desirable to be able to use affine motion mode with reduced complexity while achieving good coding efficiency.
[0004] Another tool not present in HEVC is the use of Alternate Temporal Motion Vector Prediction (ATMVP). Alternate Temporal Motion Vector Prediction (ATMVP) is a specific type of motion compensation. Instead of considering only one motion information for the current block from a temporal reference frame, each motion information of each co-located block is considered. This temporal motion vector prediction therefore provides a segmentation of the current block using the associated motion information of each sub-block. In the current VTM (VVC Test Model) reference software, ATMVP is signaled as a merge candidate inserted into the list of merge candidates. When ATMVP is enabled at the SPS level, the maximum number of merge candidates is increased by one. Thus, six candidates are considered instead of five from when this mode is disabled.
[0005] These and other tools described below raise concerns about the coding efficiency and complexity of encoding an index (e.g., merge index) or flag used to signal which candidate was selected from among a list of candidates (e.g., from a list of merge candidates for use with merge mode encoding). Summary of the Invention
[0006] Therefore, a solution to at least one of the aforementioned problems is desirable.
[0007] According to a first aspect of the present invention, there is provided a method for encoding a motion vector predictor index, the method comprising the steps of: generating a list of motion vector predictor candidates including ATMVP candidates; selecting one of the motion vector predictor candidates in the list; Generate a motion vector predictor index (merge index) for the selected motion vector predictor candidate using CABAC coding, and one or more bits of the motion vector predictor index are bypass CABAC coded. A method is provided, characterized in that:
[0008] In one embodiment, all bits except the first bit of the motion vector predictor index are bypass CABAC coded.
[0009] According to a second aspect of the present invention, there is provided a method for decoding a motion vector predictor index, the method comprising the steps of: generating a list of motion vector predictor candidates including ATMVP candidates; decoding a motion vector predictor index using CABAC decoding, wherein one or more bits of the motion vector predictor index are bypass CABAC decoded; Using the decoded motion vector predictor index to identify one of the motion vector predictor candidates in the list. A method is provided, characterized in that:
[0010] In one embodiment, all bits except the first bit of the motion vector predictor index are bypass CABAC decoded.
[0011] According to a third aspect of the present invention, there is provided an apparatus for encoding a motion vector predictor index, the apparatus comprising: means for generating a list of motion vector predictor candidates, including ATMVP candidates; means for selecting one of the motion vector predictor candidates in the list; means for generating a motion vector predictor index (merge index) of a selected motion vector predictor candidate using CABAC coding, wherein one or more bits of the motion vector predictor index are bypass CABAC coded; An apparatus is provided comprising:
[0012] According to a fourth aspect of the present invention, there is provided an apparatus for decoding a motion vector predictor index, the apparatus comprising: means for generating a list of motion vector predictor candidates, including ATMVP candidates; means for decoding a motion vector predictor index using CABAC decoding, wherein one or more bits of the motion vector predictor index are bypass CABAC decoded; means for identifying one of the candidate motion vector predictors in the list using the decoded motion vector predictor index; An apparatus is provided comprising:
[0013] According to a fifth aspect of the present invention, there is provided a method for encoding a motion vector predictor index, the method comprising the steps of: generating a list of candidate motion vector predictors; selecting one of the motion vector predictor candidates in the list; CABAC coding is used to generate a motion vector predictor index for a selected motion vector predictor candidate, and two or more bits of the motion vector predictor index share the same context. A method is provided, characterized in that:
[0014] In one embodiment, all bits of the motion vector predictor index share the same context.
[0015] According to a sixth aspect of the present invention, there is provided a method for decoding a motion vector predictor index, the method comprising the steps of: generating a list of candidate motion vector predictors; Using CABAC decoding, decode the motion vector predictor index, and two or more bits of the motion vector predictor index share the same context; Using the decoded motion vector predictor index to identify one of the motion vector predictor candidates in the list. A method is provided, characterized in that:
[0016] In one embodiment, all bits of the motion vector predictor index share the same context.
[0017] According to a seventh aspect of the present invention, there is provided an apparatus for encoding a motion vector predictor index, comprising: means for generating a list of candidate motion vector predictors; means for selecting one of the motion vector predictor candidates in the list; means for generating a motion vector predictor index for a selected motion vector predictor candidate using CABAC coding, wherein two or more bits of the motion vector predictor index share the same context; An apparatus is provided comprising:
[0018] According to an eighth aspect of the present invention, there is provided an apparatus for decoding a motion vector predictor index, the apparatus comprising: means for generating a list of candidate motion vector predictors; means for decoding a motion vector predictor index using CABAC decoding, wherein two or more bits of the motion vector predictor index share the same context; means for identifying one of the candidate motion vector predictors in the list using the decoded motion vector predictor index; An apparatus is provided comprising:
[0019] According to a ninth aspect of the present invention, there is provided a method for encoding a motion vector predictor index, the method comprising the steps of: generating a list of candidate motion vector predictors; selecting one of the motion vector predictor candidates in the list; A motion vector predictor index of a selected motion vector predictor candidate is generated using CABAC coding, and a context variable of at least one bit of the motion vector predictor index of a current block depends on the motion vector predictor index of at least one block adjacent to the current block. A method is provided, characterized in that:
[0020] In one embodiment, the context variable for at least one bit of the motion vector predictor index depends on the motion vector predictor index of each of the at least two neighboring blocks.
[0021] In another embodiment, the context variable of at least one bit of the motion vector predictor index depends on the motion vector predictor index of the left neighboring block to the left of the current block and the motion vector predictor index of the upper neighboring block above the current block.
[0022] In another embodiment, the left adjacent block is A2 and the above adjacent block is B3.
[0023] In another embodiment, the left adjacent block is A1 and the above adjacent block is B1.
[0024] In another embodiment, the context variable has three different possible values.
[0025] Another embodiment includes comparing a motion vector predictor index of at least one neighboring block with an index value of a motion vector predictor index of a current block, and setting said context variable according to a comparison result.
[0026] Another embodiment includes comparing a motion vector predictor index of at least one neighboring block with a parameter representing a bit position of the or one of the bits in the motion vector predictor index of the current block, and setting the context variable according to the comparison result.
[0027] Yet another embodiment includes performing a first comparison, comparing a motion vector predictor index of a first neighboring block with a parameter representing a bit position of the or one of the bits in the motion vector predictor index of the current block; performing a second comparison, comparing a motion vector predictor index of a second neighboring block with the parameter; and setting the context variable according to results of the first and second comparisons.
[0028] According to a tenth aspect of the present invention, there is provided a method for decoding a motion vector predictor index, the method comprising the steps of: generating a list of candidate motion vector predictors; Decode a motion vector predictor index using CABAC decoding, where a context variable of at least one bit of the motion vector predictor index of the current block depends on the motion vector predictor index of at least one block adjacent to the current block; Using the decoded motion vector predictor index to identify one of the motion vector predictor candidates in the list. A method is provided, characterized in that:
[0029] In one embodiment, a context variable for at least one bit of the motion vector predictor index depends on the motion vector predictor index of each of the at least two neighboring blocks.
[0030] In another embodiment, the context variable of at least one bit of the motion vector predictor index depends on the motion vector predictor index of the left neighboring block to the left of the current block and the motion vector predictor index of the upper neighboring block above the current block.
[0031] In another embodiment, the left adjacent block is A2 and the above adjacent block is B3.
[0032] In another embodiment, the left adjacent block is A1 and the above adjacent block is B1.
[0033] In another embodiment, the context variable has three different possible values.
[0034] Another embodiment includes comparing a motion vector predictor index of at least one neighboring block with an index value of a motion vector predictor index of a current block, and setting said context variable according to a comparison result.
[0035] Another embodiment includes comparing a motion vector predictor index of at least one neighboring block with a parameter representing a bit position of the or one of the bits in the motion vector predictor index of the current block, and setting the context variable according to the comparison result.
[0036] Yet another embodiment includes performing a first comparison, comparing a motion vector predictor index of a first neighboring block with a parameter representing a bit position of the or one of the bits in the motion vector predictor index of the current block, performing a second comparison, comparing a motion vector predictor index of a second neighboring block with the parameter, and setting the context variable depending on results of the first and second comparisons.
[0037] According to an eleventh aspect of the present invention, there is provided an apparatus for encoding a motion vector predictor index, the apparatus comprising: means for generating a list of candidate motion vector predictors; means for selecting one of the motion vector predictor candidates in the list; means for generating a motion vector predictor index for a selected motion vector predictor candidate using CABAC coding, wherein a context variable of at least one bit of the motion vector predictor index for a current block depends on the motion vector predictor index of at least one block adjacent to the current block; An apparatus is provided comprising:
[0038] According to a twelfth aspect of the present invention, there is provided an apparatus for decoding a motion vector predictor index, the apparatus comprising: means for generating a list of candidate motion vector predictors; means for decoding a motion vector predictor index using CABAC decoding, wherein a context variable for at least one bit of the motion vector predictor index of a current block depends on a motion vector predictor index of at least one block neighboring the current block; means for identifying one of the candidate motion vector predictors in the list using the decoded motion vector predictor index; An apparatus is provided comprising:
[0039] According to a thirteenth aspect of the present invention, there is provided a method for encoding a motion vector predictor index, the method comprising the steps of: generating a list of candidate motion vector predictors; selecting one of the motion vector predictor candidates in the list; generating a motion vector predictor index for a selected motion vector predictor candidate using CABAC coding, wherein a context variable of at least one bit of the motion vector predictor index for a current block depends on a skip flag of the current block; A method is provided, characterized in that:
[0040] According to a fourteenth aspect of the present invention, there is provided a method for encoding a motion vector predictor index, the method comprising the steps of: generating a list of candidate motion vector predictors; selecting one of the motion vector predictor candidates in the list; Generate a motion vector predictor index for a selected motion vector predictor candidate using CABAC coding, where a context variable of at least one bit of the motion vector predictor index for a current block depends on another parameter or syntax element of the current block that is available before decoding the motion vector predictor index. A method is provided, characterized in that:
[0041] According to a fifteenth aspect of the present invention, there is provided a method for encoding a motion vector predictor index, the method comprising the steps of: generating a list of candidate motion vector predictors; selecting one of the motion vector predictor candidates in the list; Generate a motion vector predictor index for a selected motion vector predictor candidate using CABAC coding, wherein a context variable of at least one bit of the motion vector predictor index for a current block depends on another parameter or syntax element of the current block that is an indicator of motion complexity within the current block. A method is provided, characterized in that:
[0042] According to a sixteenth aspect of the present invention, there is provided a method for decoding a motion vector predictor index, the method comprising the steps of: generating a list of candidate motion vector predictors; Decode a motion vector predictor index using CABAC decoding, where a context variable for at least one bit of the motion vector predictor index of a current block depends on a skip flag of the current block; Using the decoded motion vector predictor index to identify one of the motion vector predictor candidates in the list. A method is provided, characterized in that:
[0043] According to a seventeenth aspect of the present invention, there is provided a method for decoding a motion vector predictor index, the method comprising the steps of: generating a list of candidate motion vector predictors; Decode a motion vector predictor index using CABAC decoding, wherein a context variable of at least one bit of the motion vector predictor index of a current block depends on another parameter or syntax element of the current block that is available before decoding the motion vector predictor index; Using the decoded motion vector predictor index to identify one of the motion vector predictor candidates in the list. A method is provided, characterized in that:
[0044] According to an eighteenth aspect of the present invention, there is provided a method for decoding a motion vector predictor index, the method comprising the steps of: generating a list of candidate motion vector predictors; Decode a motion vector predictor index using CABAC decoding, wherein a context variable of at least one bit of the motion vector predictor index of a current block depends on another parameter or syntax element of the current block that is an indicator of motion complexity in the current block; Using the decoded motion vector predictor index to identify one of the motion vector predictor candidates in the list. A method is provided, characterized in that:
[0045] According to a nineteenth aspect of the present invention, there is provided an apparatus for encoding a motion vector predictor index, the apparatus comprising: means for generating a list of candidate motion vector predictors; means for selecting one of the motion vector predictor candidates in the list; means for generating a motion vector predictor index for a selected motion vector predictor candidate using CABAC coding, wherein a context variable for at least one bit of the motion vector predictor index for a current block depends on a skip flag for the current block; An apparatus is provided comprising:
[0046] According to a twentieth aspect of the present invention, there is provided an apparatus for encoding a motion vector predictor index, the apparatus comprising: means for generating a list of candidate motion vector predictors; means for selecting one of the motion vector predictor candidates in the list; means for generating a motion vector predictor index for a selected motion vector predictor candidate using CABAC coding, wherein a context variable of at least one bit of the motion vector predictor index of a current block depends on another parameter or syntax element of the current block that is available before decoding of the motion vector predictor index; An apparatus is provided comprising:
[0047] According to a twenty-first aspect of the present invention, there is provided an apparatus for encoding a motion vector predictor index, the apparatus comprising: means for generating a list of candidate motion vector predictors; means for selecting one of the motion vector predictor candidates in the list; means for generating a motion vector predictor index for a selected motion vector predictor candidate using CABAC coding, wherein a context variable of at least one bit of the motion vector predictor index for a current block depends on another parameter or syntax element of the current block that is an indicator of motion complexity within the current block; An apparatus is provided comprising:
[0048] According to a 22nd aspect of the present invention, there is provided an apparatus for decoding a motion vector predictor index, the apparatus comprising: means for generating a list of candidate motion vector predictors; means for decoding a motion vector predictor index using CABAC decoding, wherein a context variable for at least one bit of the motion vector predictor index of a current block depends on a skip flag of the current block; means for identifying one of the candidate motion vector predictors in the list using the decoded motion vector predictor index; An apparatus is provided comprising:
[0049] According to a 23rd aspect of the present invention, there is provided an apparatus for decoding a motion vector predictor index, the apparatus comprising: means for generating a list of candidate motion vector predictors; means for decoding a motion vector predictor index using CABAC decoding, wherein a context variable for at least one bit of the motion vector predictor index of a current block depends on another parameter or syntax element of the current block that is available before decoding the motion vector predictor index; means for identifying one of the candidate motion vector predictors in the list using the decoded motion vector predictor index; An apparatus is provided comprising:
[0050] According to a 24th aspect of the present invention, there is provided an apparatus for decoding a motion vector predictor index, the apparatus comprising: means for generating a list of candidate motion vector predictors; means for decoding a motion vector predictor index using CABAC decoding, wherein a context variable for at least one bit of the motion vector predictor index of a current block depends on another parameter or syntax element of the current block that is an indicator of motion complexity in the current block; means for identifying one of the candidate motion vector predictors in the list using the decoded motion vector predictor index; An apparatus is provided comprising:
[0051] According to a 25th aspect of the present invention, there is provided a method for encoding information related to a motion information predictor, comprising selecting one of a plurality of motion information predictor candidates and encoding information for identifying the selected motion information predictor candidate using CABAC encoding, wherein the CABAC encoding uses, for at least one bit of the information, the same context variable used for another inter prediction mode when one or both of a triangle merge mode or a motion vector difference (MMVD) merge mode is used.
[0052] According to a 26th aspect of the present invention, there is provided a method for decoding information relating to a motion information predictor, the method comprising: using CABAC decoding to decode information for identifying one of a plurality of motion information predictor candidates; and using the decoded information to select one of the plurality of motion information predictor candidates, wherein for at least one bit of the information, the CABAC decoding uses the same context variable used for another inter prediction mode when one or both of a triangle merge mode or a motion vector difference (MMVD) merge mode is used.
[0053] Regarding the twenty-fifth or twenty-sixth aspects of the present invention, the following features may be provided according to the embodiments thereof.
[0054] Suitably, all bits except the first bit of the information are bypass CABAC encoded or bypass CABAC decoded. Suitably, the first bit is CABAC encoded or CABAC decoded. Suitably, the other inter prediction mode includes one or both of a merge mode or an affine merge mode. Suitably, the other inter prediction mode includes a Multi-Hypothesis Intra Inter (MHII) merge mode. Suitably, the multiple motion information predictor candidates for the other inter prediction mode include ATMVP candidates. Suitably, the CABAC encoding or CABAC decoding includes using the same context variable for both a triangle merge mode and an MMVD merge mode when both are used. Suitably, at least one bit of the information is CABAC encoded or CABAC decoded when a skip mode is used. Suitably, the skip mode includes one or more of a merge skip mode, an affine merge skip mode, a triangle merge skip mode, or a merge of motion vector difference (MMVD) merge skip mode.
[0055] According to a 27th aspect of the present invention, there is provided a method for encoding information related to a motion information predictor, comprising: selecting one of a plurality of motion information predictor candidates; encoding information for identifying the selected motion information predictor candidate; and, encoding the information by bypass CABAC encoding of at least one bit of the information when one or both of a triangle merge mode or a motion vector difference (MMVD) merge mode is used.
[0056] According to a 28th aspect of the present invention, there is provided a method for decoding information relating to a motion information predictor, the method comprising: decoding information to identify one of a plurality of motion information predictor candidates; selecting one of the plurality of motion information predictor candidates using the decoded information; and, decoding the information comprises bypass CABAC decoding of at least one bit of the information when one or both of a triangle merge mode or a motion vector difference (MMVD) merge mode merging are used.
[0057] Regarding the twenty-seventh or twenty-eighth aspects of the present invention, the following features may be provided according to the embodiments thereof.
[0058] Suitably, all bits of information except the first bit are bypass CABAC encoded or bypass CABAC decoded. Suitably, the first bit is CABAC encoded or CABAC decoded. Preferably, all bits of said information are bypass CABAC encoded or bypass CABAC decoded when one or both of triangle merge mode or MMVD merge mode are used. Preferably, all bits of said information are bypass CABAC encoded or bypass CABAC decoded.
[0059] Suitably, at least one bit of said information is CABAC encoded or CABAC decoded when affine merge mode is used, and suitably, all bits of said information are bypass CABAC encoded or bypass CABAC decoded except when affine merge mode is used.
[0060] Preferably, at least one bit of the information is CABAC encoded or CABAC decoded when one or both of merge mode and MHII (Multi-Hypothesis Intra Inter) merge mode are used. Preferably, all bits of the information are bypass CABAC encoded or bypass CABAC decoded except when one or both of merge mode and MHII (Multi-Hypothesis Intra Inter) merge mode are used.
[0061] Preferably, at least one bit of the information is CABAC encoded or CABAC decoded when the plurality of motion information predictor candidates includes an ATMVP candidate. Preferably, all bits of the information are bypass CABAC encoded or bypass CABAC decoded except when the plurality of motion information predictor candidates includes an ATMVP candidate.
[0062] Preferably, at least one bit of the information is CABAC encoded or CABAC decoded when a skip mode is used. Preferably, all bits of the information are bypass CABAC encoded or bypass CABAC decoded except when a skip mode is used. Preferably, the skip mode includes one or more of a merge skip mode, an affine merge skip mode, a triangle merge skip mode, or a merge of motion vector difference (MMVD) merge skip mode.
[0063] Regarding the 25th, 26th, 27th or 28th aspects of the present invention, the following features may be provided according to the embodiments thereof.
[0064] Preferably, the at least one bit includes the first bit of the information. Preferably, the information includes a motion information prediction index or a flag. Preferably, the motion information predictor candidate includes information for obtaining a motion vector.
[0065] Regarding the twenty-fifth or twenty-seventh aspects of the present invention, the following features may be provided according to the embodiments thereof.
[0066] Preferably, the method further includes, in the bitstream, information for indicating use of one of a triangle merge mode, an MMVD merge mode, a merge mode, an affine merge mode, or an MHII (Multi-Hypothesis Intra Inter) merge mode. Preferably, the method further includes, in the bitstream, information for determining a maximum number of motion information predictor candidates that may be included in the plurality of motion information predictor candidates.
[0067] Regarding the 26th or 28th aspects of the present invention, the following features may be provided according to the embodiments thereof.
[0068] Suitably, the method further includes obtaining, from the bitstream, information for indicating use of one of a triangle merge mode, an MMVD merge mode, a merge mode, an affine merge mode, or an MHII (Multi-Hypothesis Intra Inter) merge mode. Preferably, the method further includes obtaining, from the bitstream, information for determining a maximum number of motion information predictor candidates that may be included in the plurality of motion information predictor candidates.
[0069] According to a 29th aspect of the present invention, there is provided an apparatus for encoding information related to a motion information predictor, comprising: means for selecting one of a plurality of motion information predictor candidates; and means for encoding information for identifying the selected motion information predictor candidate using CABAC encoding, wherein the CABAC encoding uses, for at least one bit of the information, the same context variable used for another inter-prediction mode when one or both of a triangle merge mode or a motion vector difference (MMVD) merge mode are used. Suitably, the apparatus comprises means for performing the method for encoding information related to a motion information predictor according to the 25th or 27th aspect of the present invention.
[0070] According to a 30th aspect of the present invention, there is provided an apparatus for encoding information related to a motion information predictor, comprising: means for selecting one of a plurality of motion information predictor candidates; and means for encoding information identifying the selected motion information predictor candidate, wherein encoding the information comprises bypass CABAC encoding of at least one bit of the information when one or both of a triangle merge mode or a motion vector difference (MMVD) merge mode are used. Suitably, the apparatus comprises means for performing the method for encoding information related to a motion information predictor according to the 25th or 27th aspects of the present invention.
[0071] According to a 31st aspect of the present invention, there is provided an apparatus for decoding information relating to a motion information predictor, the apparatus comprising: means for decoding information for identifying one of a plurality of motion information predictor candidates using CABAC decoding; and means for selecting one of the plurality of motion information predictor candidates using the decoded information, wherein the CABAC decoding comprises, for at least one bit of the information, using the same context variable as used for another inter-prediction mode when one or both of a triangle merge mode or a motion vector difference (MMVD) merge mode are used. Suitably, the apparatus comprises means for performing the method for decoding information relating to a motion information predictor according to the 26th or 28th aspect of the present invention.
[0072] According to a 32nd aspect of the present invention, there is provided an apparatus for decoding information relating to a motion information predictor, the apparatus comprising: means for decoding information to identify one of a plurality of motion information predictor candidates; and means for using the decoded information to select one of the plurality of motion information predictor candidates, wherein decoding the information comprises bypass CABAC decoding of at least one bit of the information when one or both of a triangle merge mode or a motion vector difference (MMVD) merge mode are used. Suitably, the apparatus comprises means for performing the method for decoding information relating to a motion information predictor according to the 26th or 28th aspect of the present invention.
[0073] According to a thirty-third aspect of the present invention, there is provided a method for encoding a motion vector predictor index, the method comprising: generating a list of motion vector predictor candidates; selecting one of the motion vector predictor candidates in the list; and generating a motion vector predictor index for the selected motion vector predictor candidate using CABAC coding, wherein a context variable for at least one bit of the motion vector predictor index of a current block is derived from at least one context variable of a skip flag and an affine flag of the current block.
[0074] According to a thirty-fourth aspect of the present invention, there is provided a method for decoding a motion vector predictor index, the method comprising: generating a list of motion vector predictor candidates; decoding the motion vector predictor index using CABAC decoding; a context variable of at least one bit of the motion vector predictor index of a current block is derived from a context variable of at least one of a skip flag and an affine flag of the current block; and identifying one of the motion vector predictor candidates in the list using the decoded motion vector predictor index.
[0075] According to a 35th aspect of the present invention, there is provided an apparatus for encoding a motion vector predictor index, the apparatus comprising: means for generating a list of motion vector predictor candidates; means for selecting one of the motion vector predictor candidates in the list; and means for generating a motion vector predictor index for the selected motion vector predictor candidate using CABAC encoding, wherein a context variable for at least one bit of the motion vector predictor index of a current block is derived from a context variable of at least one of a skip flag and an affine flag of the current block.
[0076] According to a 36th aspect of the present invention, there is provided an apparatus for decoding a motion vector predictor index, the apparatus comprising: means for generating a list of motion vector predictor candidates; means for decoding the motion vector predictor index using CABAC decoding; and means for identifying one of the motion vector predictor candidates in the list using the decoded motion vector predictor index, wherein a context variable of at least one bit of the motion vector predictor index of a current block is derived from a context variable of at least one of a skip flag and an affine flag of the current block.
[0077] According to a 37th aspect of the present invention, there is provided a method for encoding a motion vector predictor index, comprising generating a list of motion vector predictor candidates, selecting one of the motion vector predictor candidates in the list, and generating a motion vector predictor index for the selected motion vector predictor candidate using CABAC coding, wherein a context variable of at least one bit of the motion vector predictor index of a current block has only two different possible values.
[0078] According to a 38th aspect of the present invention, there is provided a method for decoding a motion vector predictor index, comprising generating a list of motion vector predictor candidates and decoding the motion vector predictor index using CABAC decoding, wherein a context variable for at least one bit of the motion vector predictor index of a current block has only two different possible values, and wherein the decoded motion vector predictor index is used to identify one of the motion vector predictor candidates in the list.
[0079] According to a thirty-ninth aspect of the present invention, there is provided an apparatus for encoding a motion vector predictor index, the apparatus comprising: means for generating a list of motion vector predictor candidates; means for selecting one of the motion vector predictor candidates in the list; and means for generating a motion vector predictor index for the selected motion vector predictor candidate using CABAC coding, wherein a context variable of at least one bit of the motion vector predictor index of a current block has only two different possible values.
[0080] According to a fortieth aspect of the present invention, there is provided an apparatus for decoding a motion vector predictor index, the apparatus comprising: means for generating a list of motion vector predictor candidates; means for decoding the motion vector predictor index using CABAC decoding; and means for identifying one of the motion vector predictor candidates in the list using the decoded motion vector predictor index, wherein a context variable for at least one bit of the motion vector predictor index of a current block has only two different possible values.
[0081] According to a forty-first aspect of the present invention, there is provided a method for encoding a motion information predictor index, comprising generating a list of motion information predictor candidates, selecting one of the motion information predictor candidates in the list as an affine merge mode predictor when an affine merge mode is used, and selecting one of the motion information predictor candidates in the list as a non-affine merge mode predictor when a non-affine merge mode is used, and generating a motion information predictor index for the selected motion information predictor candidate using CABAC coding, wherein one or more bits of the motion information predictor index are bypass CABAC coded.
[0082] Suitably, the CABAC encoding comprises using the same context variable for at least one bit of a motion information predictor index of the current block when an affine merge mode is used and when a non-affine merge mode is used, or the CABAC encoding comprises using a first context variable when an affine merge mode is used or a second context variable when a non-affine merge mode is used, for at least one bit of a motion information predictor index of the current block, and the method further comprises including data indicating use of the affine merge mode in the bitstream when the affine merge mode is used.
[0083] Preferably, the method further includes data for determining a maximum number of motion information predictor candidates that can be included in the generated list of motion information predictor candidates in the bitstream. Preferably, all bits except the first bit of the motion information predictor index are bypass CABAC coded. Suitably, the first bit is CABAC coded. Suitably, the motion information predictor index of the selected motion information predictor candidate is coded using the same syntax element when an affine merge mode and a non-affine merge mode are used.
[0084] According to a forty-second aspect of the present invention, there is provided a method for decoding a motion information predictor index, the method comprising: generating a list of motion information predictor candidates; decoding the motion information predictor index using CABAC decoding, wherein one or more bits of the motion information predictor index are bypass CABAC decoded; and, if an affine merge mode is used, using the decoded motion information predictor index to identify one of the motion information predictor candidates in the list as an affine merge mode predictor; and, if a non-affine merge mode is used, using the decoded motion information predictor index to identify one of the motion information predictor candidates in the list as a non-affine merge mode predictor.
[0085] Suitably, the CABAC decoding includes using the same context variable for at least one bit of the motion information predictor index of the current block when the affine merge mode is used and when the non-affine merge mode is used. Alternatively, the method further includes obtaining data from the bitstream indicating use of the affine merge mode, and the CABAC decoding includes using a first context variable for at least one bit of the motion information predictor index of the current block when the obtained data indicates use of the affine merge mode, and using a second context variable when the obtained data indicates use of the non-affine merge mode.
[0086] Suitably, the method further comprises obtaining data from the bitstream indicating the use of an affine merge mode, and the generated list of motion information predictor candidates includes an affine merge mode predictor candidate if the obtained data indicates the use of an affine merge mode, and a non-affine merge mode predictor candidate if the obtained data indicates the use of a non-affine merge mode.
[0087] Suitably, the method further includes obtaining data from the bitstream for determining a maximum number of motion information predictor candidates that can be included in the generated list of motion information predictor candidates. Suitably, all bits except the first bit of the motion information predictor index are bypass CABAC decoded. Suitably, the first bit is CABAC decoded. Suitably, decoding the motion information predictor index includes parsing the same syntax element from the bitstream when the affine merge mode is used and when the non-affine merge mode is used. Suitably, the motion information predictor candidate includes information for obtaining a motion vector. Suitably, the generated list of motion information predictor candidates includes ATMVP candidates. Suitably, the generated list of motion information predictor candidates has the same maximum number of motion information predictor candidates that can be included therein when the affine merge mode is used and when the non-affine merge mode is used.
[0088] According to a 43rd aspect of the present invention, there is provided an apparatus for encoding a motion information predictor index, the apparatus comprising: means for generating a list of motion information predictor candidates; means for selecting one of the motion information predictor candidates in the list as an affine merge mode predictor when an affine merge mode is used; means for selecting one of the motion information predictor candidates in the list as a non-affine merge mode predictor when a non-affine merge mode is used; and means for generating a motion information predictor index for the selected motion information predictor candidate using CABAC coding, wherein one or more bits of the motion information predictor index are bypass CABAC coded. Preferably, the apparatus comprises means for performing the method for encoding a motion information predictor index according to the 41st aspect.
[0089] According to a 44th aspect of the present invention, there is provided an apparatus for decoding a motion information predictor index, the apparatus comprising: means for generating a list of motion information predictor candidates; means for decoding the motion information predictor index using CABAC decoding, where one or more bits of the motion information predictor index are bypass CABAC decoded; means for identifying one of the motion information predictor candidates in the list as an affine merge mode predictor using the decoded motion information predictor index if an affine merge mode is used; and means for identifying one of the motion information predictor candidates in the list as a non-affine merge mode predictor using the decoded motion information predictor index if a non-affine merge mode is used. Preferably, the apparatus comprises means for performing the method for decoding a motion information predictor index according to the 42nd aspect.
[0090] According to a forty-fifth aspect of the present invention, there is provided a method for encoding a motion information predictor index for an affine merge mode, comprising generating a list of motion information predictor candidates, selecting one of the motion information predictor candidates in the list as an affine merge mode predictor, and generating a motion information predictor index for the selected motion information predictor candidate using CABAC coding, wherein one or more bits of the motion information predictor index are bypass CABAC coded.
[0091] Suitably, when the non-affine merge mode is used, the method further includes selecting one of the motion information predictor candidates in the list as the non-affine merge mode predictor. Suitably, the CABAC encoding comprises using a first context variable when the affine merge mode is used or using a second context variable when the non-affine merge mode is used for at least one bit of the motion information predictor index of the current block, and the method further comprises, when the affine merge mode is used, including data indicating the use of the affine merge mode in the bitstream. Alternatively, the CABAC encoding comprises using the same context variable for at least one bit of the motion information predictor index of the current block when the affine merge mode and the non-affine merge mode are used.
[0092] Preferably, the method further comprises data for determining a maximum number of motion information predictor candidates that may be included in the generated list of motion information predictor candidates in the bitstream.
[0093] Preferably, all bits except the first bit of the motion information predictor index are bypass CABAC coded. Suitably, the first bit is CABAC coded. Suitably, the motion information predictor index of the selected motion information predictor candidate is coded using the same syntax element when the affine merge mode and the non-affine merge mode are used.
[0094] According to a forty-sixth aspect of the present invention, there is provided a method for decoding a motion information predictor index for an affine merge mode, comprising: generating a list of motion information predictor candidates; decoding the motion information predictor index using CABAC decoding; and, if one or more bits of the motion information predictor index are bypass CABAC decoded and affine merge mode is used, using the decoded motion information predictor index to identify one of the motion information predictor candidates in the list as an affine merge mode predictor.
[0095] Suitably, when the non-affine merge mode is used, the method further includes identifying one of the motion information predictor candidates in the list as a non-affine merge mode predictor using the decoded motion information predictor index. Suitably, the method further includes obtaining data indicating use of the affine merge mode from the bitstream, and the CABAC decoding includes using, for at least one bit of the motion information predictor index of the current block, a first context variable if the obtained data indicates use of the affine merge mode, and using a second context variable if the obtained data indicates use of the non-affine merge mode. Alternatively, the CABAC decoding includes using the same context variable for at least one bit of the motion information predictor index of the current block when the affine merge mode and when the non-affine merge mode are used.
[0096] Suitably, the method further includes obtaining data from the bitstream indicating the use of an affine merge mode, and the generated list of motion information predictor candidates includes an affine merge mode predictor candidate if the obtained data indicates the use of an affine merge mode, and a non-affine merge mode predictor candidate if the obtained data indicates the use of a non-affine merge mode.
[0097] Suitably, decoding the motion information predictor index includes parsing the same syntax element from the bitstream when the affine merge mode and the non-affine merge mode are used. Suitably, the method further includes obtaining data from the bitstream for determining a maximum number of motion information predictor candidates that can be included in the generated list of motion information predictor candidates. Suitably, all bits except the first bit of the motion information predictor index are bypass CABAC decoded. Suitably, the first bit is CABAC decoded. Suitably, the motion information predictor candidate includes information for obtaining a motion vector. Suitably, the generated list of motion information predictor candidates includes ATMVP candidates. Suitably, the generated list of motion information predictor candidates has the same maximum number of motion information predictor candidates that can be included therein when the affine merge mode and the non-affine merge mode are used.
[0098] According to a 47th aspect of the present invention, there is provided an apparatus for encoding a motion information predictor index for an affine merge mode, the apparatus comprising: means for generating a list of motion information predictor candidates; means for selecting one of the motion information predictor candidates in the list as an affine merge mode predictor; and means for generating a motion information predictor index for the selected motion information predictor candidate using CABAC coding, wherein one or more bits of the motion information predictor index are bypass CABAC coded. Preferably, the apparatus comprises means for performing the method for encoding a motion information predictor index according to the 45th aspect.
[0099] According to a 48th aspect of the present invention, there is provided an apparatus for decoding a motion information predictor index for affine merge mode, the apparatus comprising: means for generating a list of motion information predictor candidates; means for decoding the motion information predictor index using CABAC decoding, where one or more bits of the motion information predictor index are bypass CABAC decoded; and means for identifying one of the motion information predictor candidates in the list as an affine merge mode predictor using the decoded motion information predictor index if affine merge mode is used. Preferably, the apparatus comprises means for performing the method for decoding a motion information predictor index according to the 46th aspect.
[0100] Yet another aspect of the present invention relates to a program that, when executed by a computer or processor, causes the computer or processor to perform any of the methods of the preceding aspects. The program may be provided on its own or may be carried on, by, or within a carrier medium. The carrier medium may be non-transitory, for example, a storage medium, in particular a computer-readable storage medium. The carrier medium may also be transitory, for example, a signal or other transmission medium. The signal may be transmitted over any suitable network, including the Internet.
[0101] Yet another aspect of the invention relates to a camera comprising an apparatus according to any of the aforementioned apparatus aspects. In one embodiment, the camera further comprises zooming means. In one embodiment, the camera indicates when said zooming means is operable and is adapted to signal an inter-prediction mode in response to said indication that the zooming means is operable. In another embodiment, the camera further comprises panning means. In another embodiment, the camera indicates when said panning means is operable and is adapted to signal an inter-prediction mode in response to said indication that the panning means is operable.
[0102] According to yet another aspect of the present invention, there is provided a mobile device comprising a camera embodying any of the above camera aspects. In one embodiment, the mobile device further comprises at least one position sensor adapted to sense a change in orientation of the mobile device. In one embodiment, the mobile device is adapted to signal an inter-prediction mode dependent on said sensing a change in orientation of the mobile device.
[0103] Further features of the invention are characterized by the other independent and dependent claims.
[0104] Any feature in one aspect of the invention may be applied to other aspects of the invention in any appropriate combination. In particular, method aspects may be applied to apparatus aspects, and vice versa. Furthermore, features implemented in hardware may be implemented in software, and vice versa. Any references herein to software and hardware features should be interpreted accordingly. Any apparatus features as described herein may also be provided as method features, and vice versa. As used herein, means-plus-function features may alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and associated memory.
[0105] It is also to be understood that specific combinations of the various features described and defined in any embodiment of the present invention can be implemented and / or provided and / or used independently. [Brief explanation of the drawings]
[0106] Reference will now be made, by way of example, to the accompanying drawings in which: [Figure 1] FIG. 1 is a diagram used to explain the coding structure used in HEVC. [Figure 2] FIG. 2 is a block diagram that schematically illustrates a data communications system in which one or more embodiments of the present invention may be implemented. [Figure 3] FIG. 3 is a block diagram illustrating components of a processing device capable of implementing one or more embodiments of the present invention. [Figure 4] FIG. 4 is a flow chart illustrating steps of an encoding method according to an embodiment of the invention. [Figure 5] FIG. 5 is a flow chart illustrating the steps of a decoding method according to an embodiment of the present invention. [Figure 6a] FIG. 6a shows spatial and temporal blocks that can be used to generate a motion vector predictor. [Figure 6b] FIG. 6b shows spatial and temporal blocks that can be used to generate a motion vector predictor. [Figure 7] FIG. 7 shows simplified steps in the process of AMVP predictor set derivation. [Figure 8] FIG. 8 is a schematic diagram of a motion vector derivation process in merge mode. [Figure 9] FIG. 9 shows the segmentation and temporal motion vector prediction of the current block. [Figure 10] Figure 10(a) shows the encoding of the merge index for HEVC or when ATMVP is not enabled at the SPS level, and Figure 10(b) shows the encoding of the merge index when ATMVP is enabled at the SPS level. [Figure 11] Figure 11(a) shows a simple affine motion field, and Figure 11(b) shows a more complex affine motion field. [Figure 12] FIG. 12 is a flowchart of a partial decoding process for some syntax elements related to coding modes. [Figure 13] FIG. 13 is a flowchart showing deriving merge candidates. [Figure 14] FIG. 14 shows the encoding of the merge index according to the first embodiment of the present invention. [Figure 15]FIG. 15 is a flowchart of a partial decoding process of some syntax elements related to encoding modes in the twelfth embodiment of the present invention. [Figure 16] FIG. 16 is a flowchart showing generation of a list of merge candidates in the twelfth embodiment of the present invention. [Figure 17] FIG. 17 is a block diagram for use in describing a CABAC encoder suitable for use in embodiments of the present invention. [Figure 18] FIG. 18 is a schematic block diagram of a communication system for the implementation of one or more embodiments of the present invention. [Figure 19] FIG. 19 is a schematic block diagram of a computing device. [Figure 20] FIG. 20 is a diagram showing a network camera system. [Figure 21] FIG. 21 is a diagram illustrating a smartphone. [Figure 22] FIG. 22 is a flowchart of a partial decoding process of some syntax elements related to coding modes according to the sixteenth embodiment. [Figure 23] FIG. 23 is a flowchart illustrating the use of a single index signaling scheme for both merge mode and affine merge mode according to an embodiment. [Figure 24] FIG. 24 is a flowchart showing the affine merge candidate derivation process in the affine merge mode according to this embodiment. [Figure 25] 25(a) and 25(b) illustrate the predictor derivation process for triangle merge mode according to one embodiment. [Figure 26] FIG. 26 is a flowchart of a decoding process for inter prediction mode for a current coding unit according to one embodiment. [Figure 27] Figure 27(a) shows the encoding of flags for merging in the motion vector differential (MMVD) merge mode, and Figure 27(b) shows the encoding of indices for the triangle merge mode, according to one embodiment. [Figure 28] FIG. 28 is a flowchart illustrating an affine merge candidate derivation process for the affine merge mode using ATMVP candidates, according to one embodiment. [Figure 29] FIG. 29 is a flowchart of the decoding process in the inter prediction mode according to the eighteenth embodiment. [Figure 30] Figure 30(a) shows encoding of flags for merging in motion vector differential (MMVD) merge mode according to the 19th embodiment, Figure 30(b) shows encoding of indices for triangle merge mode according to the 19th embodiment, and Figure 30(c) shows encoding of indices for affine merge mode or merge mode according to the 19th embodiment. [Figure 31] FIG. 31 is a flowchart of a decoding process in the inter prediction mode according to the nineteenth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0107] The embodiments of the present invention described below relate to improving the encoding and decoding of indexes, flags, information, and data using CABAC. It should be understood that alternative embodiments of the present invention may also be implemented to improve other context-based arithmetic coding schemes that are functionally similar to CABAC. Before describing the embodiments, video encoding and decoding techniques and related encoders and decoders will be described.
[0108] In this specification, "signaling" may refer to inserting (providing / including / during encoding) or extracting / obtaining (decoding) bitstream information regarding one or more syntax elements that represent the use, disuse, enabling or disabling of a mode (e.g., inter-prediction mode) or other information (such as information regarding a selection).
[0109] Figure 1 relates to the coding structure used in the High Efficiency Video Coding (HEVC) video standard. A video sequence 1 consists of a sequence of digital images i. Each such digital image is represented by one or more matrices. The matrix coefficients represent pixels.
[0110] Image 2 of this sequence is divided into slices 3, which in some cases constitute the entire image. These slices are divided into non-overlapping coding tree units (CTUs). A coding tree unit (CTU) is the basic processing unit of the High Efficiency Video Coding (HEVC) video standard and conceptually corresponds in structure to the macroblock unit used in some previous video standards. A CTU is sometimes also called a maximal coding unit (LCU). A CTU has luma and chroma component parts, each of which is called a coding tree block (CTB). These different color components are not shown in FIG. 1.
[0111] A CTU is typically sized 64 pixels by 64 pixels for HEVC, but for VVC the size may be 128 pixels by 128 pixels. Each CTU, in turn, may be iteratively divided into smaller variable-size coding units (CUs) 5 using quadtree decomposition.
[0112] A coding unit is a basic coding element and consists of two types of subunits called prediction units (PUs) and transform units (TUs). The maximum size of a PU or TU is equal to the CU size. A prediction unit corresponds to the partitioning of a CU for pixel value prediction. As shown in Figure 6, various different partitions of a CU into PUs are possible, including a partition into four square PUs and two different partitions into two rectangular PUs. A transform unit is a basic unit that performs spatial transformations using DCT. A CU can be partitioned into TUs based on a quadtree representation. Thus, a slice, tile, CTU / LCU, CTB, CU, PU, TU, or block of pixels / samples may be referred to as an image portion, i.e., part of an image 2 of a sequence.
[0113] Each slice is embedded in one network abstraction layer (NAL) unit. Furthermore, the coding parameters of a video sequence are stored in a dedicated NAL unit called a parameter set. HEVC and H.264 / AVC use two types of parameter set NAL units: the sequence parameter set (SPS) NAL unit, which collects all parameters that do not change during the entire video sequence. Typically, it handles coding profiles, video frame sizes, and other parameters. The picture parameter set (PPS) NAL unit contains parameters that can change from one picture (or frame) of a sequence to another. HEVC also includes the video parameter set (VPS) NAL unit, which contains parameters that describe the overall structure of the bitstream. The VPS is a new type of parameter set defined in HEVC that applies to all layers of the bitstream. A layer can contain multiple temporal sublayers; all Version 1 bitstreams are limited to one layer. HEVC has specific layer extensions for scalability and multiview, which allow multiple layers with a backward-compatible Version 1 base layer.
[0114] 2 and 18 illustrate data communication systems in which one or more embodiments of the present invention may be implemented. The data communication system includes a transmitting device, e.g., server 201 of FIG. 2 or content provider 150 of FIG. 18, operable to transmit data packets of a data stream 204 (or bitstream 101 of FIG. 18) to a receiving device, e.g., client terminal 202 of FIG. 2 or content consumer 100 of FIG. 18, via a data communication network 200. Data communication network 200 may be a wide area network (WAN) or a local area network (LAN). Such a network may be, for example, a wireless network (Wifi / 802.11a or b or g), an Ethernet network, an Internet network, or a hybrid network consisting of several different networks. In certain embodiments of the present invention, the data communication system may be a digital television broadcasting system in which server 201 (or content provider 150 of FIG. 18) transmits the same data content to multiple clients (or content consumers).
[0115] The data stream 204 (or bitstream 101) provided by the server 201 (or content provider 150) may be composed of multimedia data representing video and audio data. The audio and video data streams may, in some embodiments of the present invention, be captured by the server 201 (or content provider 150) using a microphone and a camera, respectively. In some embodiments, the data streams may be stored on the server 201 (or content provider 150), received by the server 201 (or content provider 150) from another data provider, or generated at the server 201 (or content provider 150). The server 201 (or content provider 150) particularly comprises an encoder for encoding the video and audio streams (e.g., the original sequence 151 of images in FIG. 18) to provide a compressed bitstream 204, 101 for transmission, which is a more compact representation of the data presented as input to the encoder.
[0116] In order to obtain a better ratio between the quality of the transmitted data and the amount of the transmitted data, the compression of the video data may for example be according to the HEVC format, or the H.264 / AVC format, or the VVC format.
[0117] The client 202 (or content consumer 100) receives the transmitted bitstream, decodes the reconstructed bitstream, and plays the video image (e.g., video signal 109 in FIG. 18) on a display device and the audio data through speakers.
[0118] Although the examples of Figure 2 or Figure 18 consider a streaming scenario, it will be appreciated that in some embodiments of the present invention, data communication between the encoder and decoder may be performed using a media storage device, such as an optical disc, for example.
[0119] In one or more embodiments of the present invention, a video image may be transmitted along with data representing a compensation offset to be applied to the reconstructed pixels of the image to provide filtered pixels in the final image.
[0120] 3 shows a schematic diagram of a processing device 300 configured to implement at least one embodiment of the present invention. The processing device 300 may be a device such as a microcomputer, a workstation, or a light portable device. The device 300 may include: - a central processing unit 311 such as a microprocessor indicated by CPU a read-only memory 307, denoted ROM, for storing the computer program for implementing the invention; a random access memory 312, denoted RAM, for storing the executable code of the methods of the present invention, as well as registers configured to record variables and parameters necessary for implementing the methods of encoding a sequence of digital images and / or decoding a bitstream according to embodiments of the present invention; a communications interface 302 connected to a communications network 303 over which the digital data to be processed is sent and received; The communication bus 313 is connected to the
[0121] Optionally, the device 300 may also include the following components:
[0122] - data storage means 304, such as a hard disk, for storing computer programs for implementing the methods of one or more embodiments of the present invention and data used or generated during the implementation of one or more embodiments of the present invention; a disk drive 305 for a disk 306, the disk drive configured to read data from or write data to the disk 306; A screen 309 that displays data and / or acts as a graphical interface with the user using a keyboard 310 or any other pointing / input means. The device 300 can be connected to various peripheral devices, such as a digital camera 320 or a microphone 308, each connected to an input / output card (not shown) to provide multimedia data to the device 300.
[0123] The communication bus provides communication and interoperability between the various elements included in or connected to the device 300. The representation of the bus is not limiting, and in particular a central processing unit is operable to communicate instructions to any element of the device 300, either directly or by means of another element of the device 300.
[0124] The disk 306 may be replaced by any information carrier, such as a compact disk (CD-ROM), rewritable or not, a ZIP disk or a memory card, and generally speaking by any information storage means readable by a microcomputer or microprocessor, integrated or not integrated into the device, possibly removable, and configured to store one or more programs whose execution enables the method of encoding a sequence of digital images and / or the method of decoding a bitstream according to the invention to be carried out.
[0125] The executable code can be stored either in the read-only memory 307, on the hard disk 304 or on a removable digital medium as previously described, such as for example the disk 306. According to a variant, the executable code of the program can be received by the communication network 303, via the interface 302, to be stored in one of the storage means of the device 300 before being executed, such as the hard disk 304.
[0126] The central processing unit 311 is configured to control and direct the execution of instructions or parts of the software code of the program or program according to the invention with instructions stored in one of the aforementioned storage means. On power-up, the program or programs stored in non-volatile memory, for example on the hard disk 304 or disk 306 or in read-only memory 307, are transferred to the random access memory 312, which contains registers for storing the executable code of the program or program, as well as variables and parameters necessary to implement the invention.
[0127] In this embodiment, the device is a programmable device that uses software to implement the invention, however, the invention may alternatively be implemented in hardware (e.g., in the form of an application specific integrated circuit or ASIC).
[0128] 4 shows a block diagram of an encoder according to at least one embodiment of the present invention, the encoder being represented by connected modules, each module adapted to perform, e.g., in the form of program instructions to be executed by the CPU 311 of the device 300, at least one corresponding step of a method for implementing at least one embodiment of encoding images of a sequence of images according to one or more embodiments of the present invention.
[0129] An original sequence of digital images i0 to in401 is received as input by the encoder 400. Each digital image is represented by a set of samples, sometimes also called picture elements (hereafter referred to as pixels).
[0130] A bitstream 410 is output by the encoder 400 after the encoding process has been performed. The bitstream 410 comprises a number of coding units or slices, each of which comprises a slice header for transmitting coded values of coding parameters used to code the slice, and a slice body comprising coded video data.
[0131] The input digital image i0~in401 is divided into blocks of pixels by module 402. The blocks correspond to image portions and may be of variable size (for example, 4x4, 8x8, 16x16, 32x32, 64x64, 128x128 pixels, and several rectangular block sizes can also be considered). A coding mode is selected for each input block. Two families of coding modes are provided: coding modes based on spatial predictive coding (intra prediction) and coding modes based on temporal prediction (inter coding, merge, SKIP). Possible coding modes are tested.
[0132] The module 403 performs an intra prediction process in which a given block to be coded is predicted by a predictor calculated from pixels neighboring said block to be coded. An indication of the selected intra predictor and the difference between the given block and its predictor are coded to provide a residual when intra coding is selected.
[0133] Temporal prediction is performed by a motion estimation module 404 and a motion compensation module 405. First, a reference image is selected from a set of reference images 416, and the portion of the reference image, also called a reference region or image portion, that is the closest region (closest in terms of pixel value similarity) to a given block to be coded is selected by the motion estimation module 404. Then, the motion compensation module 405 uses the selected region to predict the block to be coded. The difference between the selected reference region and the given block, also called a residual block, is calculated by the motion compensation module 405. The selected reference region is indicated using a motion vector.
[0134] Therefore, in both cases (spatial prediction and temporal prediction), the residual is calculated by subtracting the predictor from the original block when the original block is not in SKIP mode.
[0135] In the INTRA prediction performed by module 403, the prediction direction is coded. In the INTER prediction performed by modules 404, 405, 416, 418, 417, at least one motion vector or data for identifying such a motion vector is coded for the temporal prediction.
[0136] If inter-prediction is selected, information related to the motion vector and the residual block is coded. To further reduce the bit rate, assuming uniform motion, the motion vector is coded by the difference with respect to the motion vector predictor. The motion vector predictor from a set of motion information predictor candidates is obtained from the motion vector field 418 by the motion vector predictive coding module 417.
[0137] The encoder 400 further comprises a selection module 406 for selecting a coding mode by applying a coding cost criterion, such as a rate-distortion criterion. To further reduce redundancy, a transform (e.g., a DCT) is applied to the residual block by a transform module 407, and the resulting transformed data is quantized by a quantization module 408 and entropy coded by an entropy coding module 409. Finally, the coded residual block of the current block being coded is inserted into the bitstream 410 when not in SKIP mode, a mode requiring the residual block to be coded in the bitstream.
[0138] The encoder 400 also performs decoding of the coded images to generate reference images for motion estimation of subsequent images (e.g., those in the reference images / pictures 416). This allows the encoder and decoder receiving the bitstream to have the same reference frame (reconstructed images or image portions are used). An inverse quantization ("dequantization") module 411 performs inverse quantization ("dequantization") of the quantized data, followed by an inverse transform by an inverse transform module 412. An intra prediction module 413 uses the prediction information to decide which predictor to use for a given block, and a motion compensation module 414 actually adds the residual obtained by module 412 to a reference region obtained from the set of reference images 416.
[0139] Post-filtering is then applied by module 415 to filter the reconstructed frame of pixels (image or image portion). In an embodiment of the present invention, an SAO loop filter is used in which a compensation offset is added to the pixel values of the reconstructed pixels of the reconstructed image. It will be understood that post-filtering does not necessarily have to be performed. Also, any other type of post-filtering can be performed in addition to or instead of SAO loop filtering.
[0140] 5 shows a block diagram of a decoder 60 that may be used to receive data from an encoder, according to one embodiment of the present invention. The decoder is represented by connected modules, each configured to implement a corresponding step of a method implemented by the decoder 60, e.g., in the form of program instructions executed by the CPU 311 of the device 300.
[0141] Decoder 60 receives a bitstream 61 containing coding units (e.g., data corresponding to image portions, blocks, or coding units), each of which consists of a header containing information about coding parameters and a body containing the coded video data. As described with respect to Figure 4, the coded video data is entropy coded, and a motion vector predictor index is coded with a predetermined number of bits for a given image portion (e.g., block or CU). The received coded video data is entropy decoded by module 62. The residual data is then inverse quantized by module 63, and then an inverse transform is applied by module 64 to obtain pixel values.
[0142] Mode data indicating the coding mode is also entropy decoded, and based on this mode, INTRA type decoding or INTER type decoding is performed on the coded block (unit / set / group) of image data.
[0143] For INTRA mode, the INTRA predictor is determined by the intra prediction module 65 based on the intra prediction mode specified in the bitstream.
[0144] If the mode is INTER, motion prediction information is extracted from the bitstream to find (identify) the reference region used by the encoder. The motion prediction information includes a reference frame index and a motion vector residual. The motion vector predictor is added to the motion vector residual by the motion vector decoding module 70 to obtain a motion vector.
[0145] Motion vector decoding module 70 applies motion vector decoding for each image portion (e.g., current block or CU) coded with motion prediction. Once the motion vector predictor index for the current block is obtained, the actual value of the motion vector associated with the image portion (e.g., current block or CU) can be decoded and used to apply motion compensation by module 66. The reference image portion indicated by the decoded motion vector is extracted from reference image 68, and motion compensation 66 is applied. Motion vector field data 71 is updated with the decoded motion vector for use in predicting subsequent decoded motion vectors.
[0146] Finally, decoded blocks are obtained, where appropriate post-filtering is applied by a post-filtering module 67. A decoded video signal 69 is finally obtained and provided by the decoder 60.
[0147] CABAC HEVC uses several types of entropy coding, including context-based adaptive binary arithmetic coding (CABAC), Golomb-Rice coding, or a simple binary representation called fixed-length coding. In most cases, a binary coding process is performed to represent different syntax elements. This binary coding process is also very specific and depends on the different syntax elements. Arithmetic coding represents syntax elements according to their current probability. CABAC is an extension of arithmetic coding that separates the probability of syntax elements according to a "context" defined by a context variable. This corresponds to conditional probability. The context variable can be derived from the values of the currently decoded syntax elements of the top-left block (A2 in Figure 6b, described in detail below) and the top-left block (B3 in Figure 6b).
[0148] CABAC has been adopted as a reference part of the H.264 / AVC and H.265 / HEVC standards. In H.264 / AVC, it is one of two alternative methods for entropy coding. The other method specified in H.264 / AVC is a low-complexity entropy coding technique based on the use of a context-adaptively switched set of variable-length codes, known as context-adaptive variable-length coding (CAVLC). Compared to CABAC, CAVLC offers reduced implementation costs at the expense of lower compression efficiency. For standard-definition or high-definition TV signals, CABAC typically offers 10–20% bitrate savings over CAVLC for the same objective video quality. In HEVC, CABAC is one of the entropy coding methods used. Many bits are also bypassed (also referred to as CABAC bypass coding). Additionally, some syntax elements are coded with unary or Golomb codes, which are other types of entropy codes.
[0149] Figure 17 shows the main blocks of the CABAC encoder.
[0150] Input syntax elements that are non-binary values are binarized by a binarizer 1701. CABAC's coding strategy is based on the discovery that highly efficient coding of syntax element values in hybrid block-based video coders, such as components of motion vector differences or transform coefficient level values, can be achieved by using a binarization scheme as a kind of preprocessing unit for subsequent stages of context modeling and binary arithmetic coding. In general, a binarization scheme defines a unique mapping of syntax element values to a sequence of binary decisions, so-called bins, which can be "bits," and thus can also be interpreted in terms of a binary code tree. The design of the binarization scheme in CABAC is based on a few basic prototypes whose structure allows for simple online computation and which are applied to several suitable model-probability distributions.
[0151] Each bin can be processed in one of two basic ways, depending on the setting of switch 1702. When the switch is in the "regular" setting, the bin is fed to a context modeler 1703 and a regular encoding engine 1704. When the switch is in the "bypass" setting, the context modeler is bypassed and the bin is fed to a bypass encoding engine 1705. Another switch 1706 has similar "regular" and "bypass" settings as switch 1702, so that the bins encoded by the applicable one of encoding engines 1704 and 1705 can form a bitstream as the output of the CABAC encoder.
[0152] It will be appreciated that other switch 1706 may be used in conjunction with storage to group some of the bins encoded by coding engine 1705 (e.g., bins for encoding image portions such as blocks or coding units) to provide blocks of bypass-coded data in the bitstream, and to group some of the bins encoded by coding engine 1704 (e.g., bins for encoding blocks or coding units) to provide other blocks of “regular” (or arithmetically) coded data in the bitstream. This separate grouping of bypass-coded data and regular-coded data can result in improved throughput during the decoding process (since the bypass-coded data can be processed first / in parallel with the regular CABAC-coded data).
[0153] By decomposing each syntax element value into a sequence of bins, further processing of each bin value in CABAC depends on an associated coding mode decision, which can be selected as either normal mode or bypass mode. The latter is selected for bins associated with code information or for lower-level significant bins, which are assumed to be uniformly distributed and, as a result, simply bypass the entire normal binary arithmetic coding process. In normal coding mode, each bin value is coded using a normal binary arithmetic coding engine, and the associated probability model is either determined by a fixed selection without context modeling or adaptively selected depending on the associated context model. A key design decision is that the latter is generally applied only to the most frequently observed bins, while other, usually less frequently observed bins are processed using a joint, typically zero-order, probability model. In this way, CABAC enables selective context modeling at the sub-symbol level, thus providing an efficient means for exploiting inter-symbol redundancy with significantly reduced overall modeling or training costs. For a particular choice of context model, four basic design types are adopted in CABAC, two of which are applied to coding only at the transform coefficient level. The design of these four prototypes is based on a priori knowledge of the typical characteristics of the source data to be modeled and reflects the aim of finding a good compromise between the conflicting objectives of avoiding unnecessary modeling cost overhead and exploiting statistical dependencies to a large extent.
[0154] At the lowest level of processing in CABAC, each bin value enters a binary arithmetic coder in either normal or bypass coding mode. In the latter case, a fast branch of the coding engine with significantly reduced complexity is used, while in the former coding mode, the coding of a given bin value depends on the actual state of an associated adaptive probability model that is passed along with the bin value to the M-coder, the term chosen for the table-based adaptive binary arithmetic coding engine in CABAC.
[0155] A corresponding CABAC decoder then receives the bitstream output from the CABAC encoder and processes the bypass-encoded data and the regular CABAC-encoded data accordingly. Once the CABAC decoder processes the regular CABAC-encoded data, the context modeler (and its probability model) is updated so that it can correctly decode / process (e.g., debinarize) the bins that form the bitstream to obtain the syntax elements.
[0156] Inter-coding HEVC uses three different inter modes: inter mode (Advanced Motion Vector Prediction (AMVP) with motion information difference signaling), "classical" merge mode (i.e., also known as "non-affine merge mode" or "regular" merge mode with no motion information difference signaling), and "classical" merge skip mode (i.e., also known as "non-affine merge skip" mode or "regular" merge skip mode with no motion information difference signaling and no residual data for sample values). The main difference between these modes is the data signaling in the bitstream. For motion vector coding, the current HEVC standard includes a contention-based scheme for motion vector prediction that was not present in previous versions of the standard. This means that for each inter coding mode (AMVP) or merge mode (i.e., "classical / regular" merge mode or "classical / regular" merge skip mode), several candidates compete against the encoder's rate-distortion criterion to find the best motion vector predictor or best motion information. An index or flag corresponding to the best predictor or best candidate for motion information is then inserted into the bitstream. The decoder can derive the same set of predictors or candidates and use the best one according to the decoded index / flag. In the HEVC screen content extension, a new coding tool called Intra Block Copy (IBC) is signaled as one of these three inter modes. The difference between IBC and its equivalent inter mode is done by checking whether the reference frame is the current one. IBC is also known as Current Picture Reference (CPR). This can be implemented, for example, by checking the reference index of list L0 and inferring that if this is the last frame in the list, it is an intra block copy. Another way is to compare the picture order counts of the current and reference frames; if they are equal, it is an intra block copy.
[0157] The design of predictor and candidate derivation is important to achieve the best coding efficiency without disproportionately impacting complexity. In HEVC, two motion vector derivations are used: one for inter-mode (Advanced Motion Vector Prediction (AMVP)) and one for merge-mode (merge derivation process - for the classical merge mode and the classical merge skip mode). These processes are described below.
[0158] Figures 6a and 6b show spatial and temporal blocks that can be used to generate motion vector predictors, for example, in Advanced Motion Vector Prediction (AMVP) and merge modes of an HEVC encoding and decoding system, and Figure 7 shows simplified steps in the process of AMVP predictor set derivation.
[0159] Two spatial predictors, i.e., two spatial motion vectors for AMVP mode, are selected from the motion vectors of the top block (indicated by the letter "B") and the left block (indicated by the letter "A"), including the top corner block (block B2) and the left corner block (block A0), and one temporal predictor is selected from the motion vectors of the bottom right block (H) and the center block (Center) of the co-located blocks, as shown in Figure 6a.
[0160] Table 1 below outlines the nomenclature used when referencing blocks relative to the current block, as shown in Figures 6a and 6b. This nomenclature is used for simplicity, but it should be understood that other labeling systems may be used, particularly in future versions of the standard.
[0161] [Table 1]
[0162] It should be noted that the "current block" can be variable in size, such as 4x4, 16x16, 32x32, 64x64, 128x128, or any size in between. The block dimensions are preferably a multiple of 2 (i.e., 2^n x 2^m, where n and m are positive integers), which results in more efficient use of bits when using binary encoding. The current block does not have to be square, although this is often the preferred embodiment due to encoding complexity.
[0163] 7, the first step aims to select a first spatial predictor (Cand1, 706) from among the bottom-left blocks A0 and A1, whose spatial locations are shown in FIG. 6a. To do so, these blocks are selected one after the other in a given (i.e., predetermined / preset) order (700, 702), and for each selected block, the following conditions are evaluated in the given order (704), and the first block for which the conditions are satisfied is set as a predictor:
[0164] - Motion vectors from the same reference image and from the same reference list -Motion vectors from the same reference image and other reference lists - Scaled motion vectors from different reference images and the same reference list - Scaled motion vectors from different reference pictures and other reference lists If the value is missing, the left predictor is considered to be unavailable, which indicates that the associated blocks are intra-coded or do not exist.
[0165] The next step aims to select a second spatial predictor (Cand2, 716) from among the upper right block B0, the upper block B1, and the upper left (upper left) block B2, whose spatial locations are shown in Figure 6a. To do this, these blocks are selected one after the other in a given order (708, 710, 712), and for each selected block, the above-mentioned conditions are evaluated in a given order (714), and the first block for which the above-mentioned conditions are met is set as the predictor.
[0166] Again, if a value is missing, the above predictor is considered unavailable, indicating that the associated blocks are intra-coded or do not exist.
[0167] In the next step (718), the two predictors, if both are available, are compared to each other in order to eliminate one of them if they are equal (i.e., same motion vector value, same reference list, same reference index, and same direction type). If only one spatial predictor is available, the algorithm looks for a temporal predictor in a later step.
[0168] The temporal motion predictor (Cand3, 726) is derived as follows: The bottom right (H, 720) location of the co-located block in the previous / reference frame is first considered in the availability check module 722. If it does not exist, or if a motion vector predictor is not available, the center (Center, 724) of the co-located block is selected to be checked. These temporal locations (Center and H) are shown in Figure 6a. In either case, scaling 723 is applied to these candidates to match the temporal distance between the current frame and the first frame in the reference list.
[0169] The motion predictor value is then added to the set of predictors. The number of predictors (Nb_Cand) is then compared to the maximum number of predictors (Max_Cand) (728). As mentioned above, the maximum number of predictors (Max_Cand) for the motion vector predictors that the AMVP derivation process must generate is 2 in the current version of the HEVC standard.
[0170] If this maximum number is reached, a final list or set of AMVP predictors (732) is constructed. Otherwise, a zero predictor is added to the list (730). A zero predictor is a motion vector equal to (0,0).
[0171] As shown in FIG. 7, a final list or set of AMVP predictors (732) is constructed from a subset of candidate spatial motion predictors (700-712) and a subset of candidate temporal motion predictors (720, 724).
[0172] As mentioned above, a motion predictor candidate for classical merge mode or classical merge skip mode can represent all necessary motion information: direction, list, reference frame index, and motion vector (or any subset thereof to perform prediction). An indexed list of several candidates is generated by the merge derivation process. In the current HEVC design, the maximum number of candidates for both merge modes (i.e., classical merge mode and classical merge skip mode) is equal to 5 (four spatial candidates and one temporal candidate).
[0173] FIG. 8 is a schematic diagram of the motion vector derivation process for merge modes (classical merge mode and classical merge skip mode). In the first step of the derivation process, five block positions are considered (800-808). These positions are the spatial positions indicated in FIG. 6a by reference numerals A1, B1, B0, A0, and B2. In a later step, the availability of spatial motion vectors is checked, and at most five motion vectors are selected / obtained for consideration (810). If a predictor exists and the block is not intra-coded, the predictor is considered available. Therefore, the selection of motion vectors corresponding to the five blocks as candidates is performed according to the following conditions:
[0174] If a "left" A1 motion vector (800) is available (810), i.e., if it exists and this block is not intra-coded, the motion vector of the "left" block is selected and used as the first candidate in the candidate list (814).
[0175] If a "top" B1 motion vector (802) is available (810), the candidate "top" block motion vector is compared to the "left" A1 motion vector, if present (812). If the B1 motion vector is equal to the A1 motion vector, B1 is not added to the list of spatial candidates (814). Conversely, if the B1 motion vector is not equal to the A1 motion vector, B1 is added to the list of spatial candidates (814).
[0176] If an "upper right" B0 motion vector (804) is available (810), the "upper right" motion vector is compared to the B1 motion vector (812). If the B0 motion vector is equal to the B1 motion vector, the B0 motion vector is not added to the list of spatial candidates (814). Conversely, if the B0 motion vector is not equal to the B1 motion vector, the B0 motion vector is added to the list of spatial candidates (814).
[0177] If a "bottom left" A0 motion vector (806) is available (810), the "bottom left" motion vector is compared to the A1 motion vector (812). If the A0 motion vector is equal to the A1 motion vector, the A0 motion vector is not added to the list of spatial candidates (814). Conversely, if the A0 motion vector is not equal to the A1 motion vector, the A0 motion vector is added to the list of spatial candidates (814).
[0178] If the list of spatial candidates does not contain four candidates, the availability of the "top left" B2 motion vector (808) is checked (810). If available, it is compared with the A1 and B1 motion vectors. If the B2 motion vector is equal to the A1 or B1 motion vector, the B2 motion vector is not added to the list of spatial candidates (814). Conversely, if the B2 motion vector is not equal to the A1 or B1 motion vector, the B2 motion vector is added to the list of spatial candidates (814).
[0179] At the end of this stage, the list of spatial candidates contains up to four candidates.
[0180] For temporal candidates, two positions can be used: the bottom right position of the co-located block (816, shown as H in Figure 6a) and the center of the co-located block (818). These positions are shown in Figure 6a.
[0181] As described in connection with FIG. 7 for the temporal motion predictor of the AMVP motion vector derivation process, the first step is to check the availability of a block in the H position (820). Next, if it is not available, the availability of a block in the center position is checked (820). If at least one motion vector in these positions is available, the temporal motion vector may be scaled to the reference frame with index 0, if necessary, for both lists L0 and L1 (822) to create a temporal candidate that is added to the list of merge motion vector predictor candidates (824). This is placed after the spatial candidate in the list. Lists L0 and L1 are two reference frame lists that may contain zero, one, or multiple reference frames.
[0182] If the number of candidates (Nb_Cand) is strictly less than the maximum number of candidates (826) (the Max_Cand information for determining the value is signaled in the bitstream slice header and is equal to 5 in the current HEVC design), and if the current frame is of type B, a combined candidate is generated (828). The combined candidate is generated based on the available candidates in the list of merge motion vector predictor candidates. This mainly consists of combining (pairing) the motion information of one candidate in list L0 with the motion information of one candidate in list L1.
[0183] If the number of candidates (Nb_Cand) is strictly less than the maximum number of candidates (Max_Cand) (830), zero motion candidates are generated (832) until the number of candidates in the list of merge motion vector predictor candidates reaches the maximum number of candidates.
[0184] At the end of this process, a list or set of merge motion vector predictor candidates (i.e., a list or set of candidates for the classic merge mode and the classic merge skip mode) is constructed (834). As shown in Figure 8, the list or set of merge motion vector predictor candidates is constructed (834) from a subset of spatial candidates (800-808) and a subset of temporal candidates (816, 818).
[0185] Alternative Temporal Motion Vector Prediction (ATMVP) Alternative temporal motion vector prediction (ATMVP) is a special type of motion compensation. Instead of considering only one motion information for the current block from the temporal reference frame, each motion information of each co-located block is considered. Therefore, this temporal motion vector prediction provides a segmentation of the current block using the associated motion information of each sub-block, as shown in Figure 9.
[0186] In the VTM reference software, the ATMVP is signaled as a merge candidate inserted into a list of merge candidates (i.e., a list or set of candidates for merge modes that are classical merge mode and classical merge skip mode). When ATMVP is enabled at the SPS level, the maximum number of merge candidates is increased by one. Thus, six candidates are considered instead of five, as would be the case if this ATMVP mode were disabled. It is understood that, according to one embodiment of the present invention, the ATMVP can be signaled as an affine merge candidate (e.g., an ATMVP candidate) inserted into a list of affine merge candidates (i.e., a separate list or set of candidates for affine merge mode, described in more detail below).
[0187] Furthermore, when this prediction is enabled at the SPS level, all bins of the merge index (i.e., identifiers or indexes or information for identifying a candidate from a list of merge candidates) are context coded by CABAC. While in HEVC or when ATMVP is not enabled at the SPS level in JEM, only the first bin is context coded, and the remaining bins are context bypass coded (i.e., bypass CABAC coded).
[0188] Figure 10(a) shows the encoding of the merge index when ATMVP is not enabled at the SPS level of HEVC or JEM. This corresponds to a unary max code. Furthermore, in Figure 10(a), the first bit is CABAC encoded and the other bits are bypass CABAC encoded.
[0189] Figure 10(b) shows the encoding of the merge index when ATMVP is enabled at the SPS level. All bits are CABAC encoded (from the 1st to the 5th bits). Note that each bit for encoding the index has its own context, in other words, their probabilities for use in CABAC encoding are separated.
[0190] Affine Mode HEVC only applies the translational motion model for motion compensated prediction (MCP), whereas in the real world there are many types of motion, such as zoom-in / zoom-out, rotation, perspective motion, and other irregular motions.
[0191] JEM applies simple affine transformation motion compensation prediction, and the general principles of affine mode are described below based on an excerpt from the JVET-G1001 document presented at the JVET conference in Torino, July 13-21, 2017. This document is incorporated herein in its entirety by reference insofar as it describes other algorithms used in JEM.
[0192] As shown in Figure 11(a), the affine motion field of a block in this document is described by two control point motion vectors (it should be understood that other affine models, such as those with more control point motion vectors, may also be used according to embodiments of the present invention).
[0193] The motion vector field (MVF) of a block is described by the following equation:
[0194]
number
[0195] where (v0x, v0y) is the motion vector of the control point in the upper left corner, (v1x, v1y) is the motion vector of the control point in the upper right corner, and w is the width of block Cur (the current block).
[0196] To further simplify the motion compensation prediction, we apply subblock-based affine transformation prediction. The subblock size MxN is derived as in Equation 2, where MvPre is the motion vector fractional precision (1 / 16 in JEM), and (v2x, v2y) is the motion vector of the top-left control point calculated according to Equation 1.
[0197]
number
[0198] After being derived by Equation 2, M and N may be adjusted downwards if necessary to be divisors of w and h, respectively, where h is the height of the current block Cur.
[0199] To derive the motion vector for each M × N sub-block, the motion vector of the center sample of each sub-block is calculated according to Equation 1 and rounded to 1 / 16 fractional precision, as shown in Figure 11(b). Then, a motion-compensated interpolation filter is applied to generate a prediction for each sub-block with the derived motion vector.
[0200] Affine mode is a motion compensation mode like inter mode (AMVP, "classical" merge, or "classical" merge skip). Its principle is to generate one motion information per pixel according to the motion information of two or three neighboring pixels. In JEM, the affine mode derives one motion information for each 4x4 block, as shown in Figure 11(b) (each square is a 4x4 block, and the entire block in Figure 11(b) is a 16x16 block divided into 16 such square blocks of 4x4 size, and each 4x4 square block has a motion vector associated with it). It should be understood that in embodiments of the present invention, the affine mode can drive one motion information for blocks of different sizes or shapes, as long as one motion information can be derived.
[0201] According to one embodiment, this mode is made available for AMVP mode and merge mode (i.e., classical merge mode, also called "non-affine merge mode," and classical merge skip mode, also called "non-affine merge skip mode") by enabling affine mode with a flag. This flag is CABAC coded. In one embodiment, the context depends on the sum of the affine flags of the left block (position A2 in Figure 6b) and the top-left block (position B3 in Figure 6b).
[0202] Thus, in JEM, there are three possible context variables (0, 1, or 2) for the affine flag given in the later formula.
[0203] Ctx=IsAffine(A2)+IsAffine(B3) Here, IsAffine(block) is a function that returns 0 if the block is not an affine block, and returns 1 if the block is affine.
[0204] Affine merge candidate derivation In JEM, affine merge mode (or affine merge skip mode), also known as subblock (merge) mode, derives motion information for the current block from the first affine neighboring block (i.e., the first neighboring block coded using affine mode) among the blocks at positions A1, B1, B0, A0, and B2. These positions are shown in Figures 6a and 6b. However, how the affine parameters are derived is not fully defined. The present invention aims to improve at least this aspect, for example, by defining affine parameters for the affine merge mode, thereby enabling a wider selection of affine merge candidates (i.e., not only the affine first neighboring block, but at least one other candidate is available for selection using an identifier such as an index).
[0205] For example, according to some embodiments of the present invention, an affine merge mode having its own list of affine merge candidates (candidates for deriving / obtaining motion information for the affine mode) and an affine merge index (for identifying one affine merge candidate from the list of affine merge candidates) are used to encode or decode a block.
[0206] Affine Merge Signaling Figure 12 is a flowchart of a partial decoding process for some syntax elements related to coding modes for signaling the use of affine merge mode. In this figure, the skip flag (1201), prediction mode (1211), merge flag (1203), merge index (1208), and affine flag (1206) can be decoded.
[0207] For all CUs in the inter slice, the skip flag is decoded (1201). If the CU is not a skip (1202), the prediction mode (prediction_mode) is decoded (1211). This syntax element indicates whether the current CU is coded (decoded) in inter mode or intra mode. Note that if the CU is a skip (1202), its current mode is inter mode. If the CU is not a skip (1202: No), the CU is coded in AMVP mode or merge mode. If the CU is inter (1212), the merge flag is decoded (1203). If the CU is a merge (1204) or if the CU is a skip (1202: Yes), it is verified / checked (1205) whether the affine flag (1206) needs to be decoded, i.e., a determination is made as to whether the current CU could have been coded in affine mode. This flag is decoded if the current CU is a 2Nx2N CU, which means that in the current VVC, the height and width of the CU are equal. Furthermore, at least one neighboring CU A1 or B1 or B0 or A0 or B2 must be coded in affine mode (either affine merge mode or AMVP mode with affine mode enabled). Finally, the current CU is not a 4x4 CU, and by default, CU 4x4 is disabled in the VTM reference software. If this condition (1205) is false, it is certain that the current CU is coded in classical merge mode (or classical merge skip mode) as specified in HEVC, and the merge index is decoded (1208). If the affine flag (1206) is set equal to 1 (1207), the CU is a merge affine CU (i.e., a CU coded in affine merge mode) or a merge skip affine CU (i.e., a CU coded in affine merge skip mode), and the merge index (1208) does not need to be decoded (because affine merge mode is used, i.e., the CU is decoded using affine mode with its first neighboring block being affine).Otherwise, the current CU is a classic (basic) merge or merge skip CU (i.e., a CU coded in classic merge or merge skip mode), and the merge candidate index (1208) is decoded.
[0208] Merge candidate derivation FIG. 13 is a flowchart illustrating merge candidate derivation (i.e., candidates for classical merge mode or classical merge skip mode) according to one embodiment. This derivation builds on the merge mode motion vector derivation process (i.e., merge candidate list derivation for HEVC) shown in FIG. 8. The main changes compared to HEVC are the addition of ATMVP candidates (1319, 1321, 1323), a full overlap check for candidates (1325), and a new candidate ordering. ATMVP prediction is set as a dedicated candidate because it represents some motion information of the current CU. The value of the first sub-block (top left) is compared with the temporal candidate, and if they are equal, the temporal candidate is not added to the merge candidate list (1320). The ATMVP candidate is not compared with other spatial candidates. This is in contrast to temporal candidates, which are compared with each spatial candidate already in the list (1325) and are not added to the merge candidate list if they are overlapping candidates.
[0209] When a spatial candidate is added to the list, it is compared 1312 with other spatial candidates in the list, which is not the case in the final version of HEVC.
[0210] In the current VTM version, the list of merge candidates is set in the following order as determined to provide the best results for the coding test criteria: A1 B1 B0 A0 ATMVP B2 · Temporal · combination · Zero_MV It is important to note that spatial candidate B2 is set after the ATMVP candidate.
[0211] Furthermore, when ATMVP is enabled at the slice level, the maximum number in the list of candidates is 6 instead of 5 in HEVC.
[0212] Other Inter Prediction Modes In the first few embodiments (up to the sixteenth embodiment) described below, the description describes encoding or decoding of indices for the (regular) merge mode and the affine merge mode. In recent versions of the VVC standard under development, in addition to the (regular) merge mode and the affine merge mode, additional inter prediction modes are also considered. Such additional inter prediction modes currently considered are the Multi-Hypothesis Intra Inter (MHII) merge mode, the Triangle merge mode, and the Motion Vector Difference Merge (MMVD) merge mode, which are described below.
[0213] According to variations of these first several embodiments, it will be appreciated that one or more of the additional inter prediction modes may be used in addition to or instead of the merge mode or the affine merge mode, and that an index (or flag or information) for one or more of the additional inter prediction modes may be signaled (encoded or decoded) using the same techniques as any one of the merge mode or the affine merge mode.
[0214] MHII (Multi-Hypothesis Intra Inter) merge mode The Multi-Hypothesis Intra-Inter (MHII) merge mode is a hybrid that combines the regular merge mode and the intra-mode. The block predictor in this mode is obtained as the average between the (regular) merge mode block predictor and the intra-mode block predictor. The obtained block predictor is added to the residual of the current block to obtain a reconstructed block. To obtain this merge mode block predictor, the MHII merge mode uses the same number of candidates and the same merge candidate derivation process as the merge mode. Therefore, index signaling for the MHII merge mode can use the same technique as for the merge mode. Also, this mode is only enabled for blocks coded / decoded in non-skip mode. Therefore, if the current CU is coded / decoded in skip mode, the MHII is no longer available in the coding / decoding process.
[0215] Triangle Merge Mode The triangle merge mode is a type of bi-prediction mode that uses motion compensation based on triangle shapes. Figures 25(a) and 25(b) show different partition configurations used for its block predictor generation. The block predictor is obtained from the first triangle (first block predictor 2501 or 2511) and the second triangle (second block predictor 2502 or 2512) in a block. Two different configurations are used for generating this block predictor. In the first one, the division / partition between triangular parts / regions (to which two block predictor candidates are associated) is from the upper left corner to the lower right corner, as shown in Figure 25(a). In the second one, the division / partition between triangular regions (to which two block predictor candidates are associated) is from the upper right corner to the lower left corner, as shown in Figure 25(b). Furthermore, samples around the boundary between triangular regions are filtered with a weighted average, where the weight depends on the sample position (e.g., distance from the boundary). A separate triangle merge candidate list is generated, and index signaling for the triangle merge mode can use techniques modified accordingly relative to the techniques for index signaling in the merge mode or affine merge mode.
[0216] Merge Motion Vector Difference (MMVD) merge mode MMVD merge mode is a special type of regular merge mode candidate derivation that generates an independent MMVD merge candidate list. The selected MMVD merge candidate for the current CU is obtained by adding an offset value to the motion vector component (mvx or mvy) of one of the MMVD merge candidates. The offset value is added to the motion vector component from the first list L0 or the second list L1 depending on the configuration of these reference frames (e.g., both backward, forward, or both forward and backward). The selected MMVD merge candidate is signaled using an index. The offset value is signaled using a distance index among eight possible preset distances (1 / 4-pel, 1 / 2-pel, 1-pel, 2-pel, 4-pel, 8-pel, 16-pel, 32-pel) and a direction index that indicates the x-axis or y-axis and the sign of the offset. Therefore, index signaling for MMVD merge mode can use the same technique as index signaling for merge mode or affine merge mode.
[0217] Embodiment Embodiments of the present invention will now be described with reference to the remaining figures. It should be noted that embodiments may be combined unless otherwise stated, for example, certain combinations of embodiments may improve coding efficiency at the expense of increased complexity, which may be acceptable in certain use cases.
[0218] First embodiment As mentioned above, in the VTM reference software, the ATMVP is signaled as a merge candidate inserted into the list of merge candidates. The ATMVP can be enabled or disabled for the entire sequence (at the SPS level). When the ATMVP is disabled, the maximum number of merge candidates is 5. When the ATMVP is enabled, the maximum number of merge candidates is increased by 1, from 5 to 6.
[0219] At the encoder, a list of merge candidates is generated using the method of Figure 13. One merge candidate is selected from the list of merge candidates based, for example, on a rate-distortion criterion. The selected merge candidate is signaled to the decoder in the bitstream using a syntax element called a merge index.
[0220] The current VTM reference software encodes the merge index differently depending on whether ATMVP is enabled or disabled.
[0221] Figure 10(a) shows the encoding of merge indices when ATMVP is not enabled at the SPS level. Five merge candidates, Cand0, Cand1, Cand2, Cand3, and Cand4, are encoded as 0, 10, 110, 1110, and 1111, respectively. This corresponds to unary maximal encoding. Furthermore, the first bit is encoded by CABAC using a single context, and the other bits are bypass encoded.
[0222] Figure 10(b) shows the encoding of the merge index when ATMVP is enabled. Six merge candidates, Cand0, Cand1, Cand2, Cand3, Cand4, and Cand5, are encoded as 0, 10, 110, 1110, 11110, and 11111, respectively. In this case, all bits (from the first to the fifth bits) of the merge index are context-coded by CABAC. Each bit has its own context, and there are separate probability models for different bits.
[0223] In the first embodiment of the present invention, as shown in FIG. 14 , when the list of merge candidates includes ATMVP as a merge candidate (e.g., when ATMVP is enabled at the SPS level), the encoding of the merge index is modified so that only the first bit of the merge index is encoded by CABAC using a single context. The context is set in the same way as in the current VTM reference software when ATMVP is not enabled at the SPS level, i.e., the other bits (second to fifth) are bypass-coded. When the list of merge candidates does not include ATMVP as a merge candidate (e.g., when ATMVP is disabled at the SPS level), five merge candidates exist. Only the first bit of the merge index is encoded by CABAC using a single context. The context is set in the same way as in the current VTM reference software when ATMVP is not enabled at the SPS level. The other bits (second to fourth bits) are bypass-decoded.
[0224] The decoder generates the same list of merge candidates as the encoder. This can be achieved by using the method in Figure 13. If ATMVP is not included as a merge candidate in the list of merge candidates (e.g., if ATMVP is disabled at the SPS level), there are five merge candidates. Only the first bit of the merge index is decoded by CABAC using a single context. The other bits (the second through fourth bits) are bypass-decoded. In contrast to the current reference software, if ATMVP is included as a merge candidate in the list of merge candidates (e.g., if ATMVP is enabled at the SPS level), only the first bit of the merge index is decoded by CABAC using a single context in the decoding of the merge index. The other bits (the second through fifth bits) are bypass-decoded. The decoded merge index is used to identify the merge candidate selected by the encoder from the list of merge candidates.
[0225] The advantage of this embodiment compared to the VTM 2.0 reference software is that the complexity of merge index decoding and decoder design (and encoder design) is reduced without affecting coding efficiency. Indeed, in this embodiment, only one CABAC state is needed for the merge index, instead of five for the current VTM merge index encoding / decoding. Furthermore, other bits are CABAC bypass coded, reducing the number of operations compared to coding all bits with CABAC, thus reducing worst-case complexity.
[0226] Second embodiment In the second embodiment, all bits of the merge index are CABAC coded, but they all share the same context. In this case, there can be a single context, as in the first embodiment, shared between the bits. As a result, when the list of merge candidates includes ATMVP as a merge candidate (e.g., when ATMVP is enabled at the SPS level), only one context is used, compared to five in the VTM2.0 reference software. The advantage of this embodiment compared to the VTM2.0 reference software is that the complexity of merge index decoding and decoder design (and encoder design) is reduced without impacting coding efficiency.
[0227] Alternatively, as described below in connection with the third through sixteenth embodiments, context variables may be shared between bits such that more than one context is available but the current context is shared by the bits.
[0228] When ATMVP is disabled, the same context is still used for all bits.
[0229] This embodiment and all subsequent embodiments are applicable even if ATMVP is not an available mode or is disabled.
[0230] In a variation of the second embodiment, any two or more bits of the merge index are CABAC coded and share the same context. Other bits of the merge index are bypass coded. For example, the first N bits of the merge index may be CABAC coded, where N is 2 or greater.
[0231] Third embodiment In the first embodiment, the first bit of the merge index was CABAC coded using a single context.
[0232] In a third embodiment, the context variable for a bit of a merge index depends on the value of the merge index of the neighboring block, which allows multiple contexts for the target bit, with each context corresponding to a different value of the context variable.
[0233] A neighboring block may be any block that has already been decoded, so that its merge index is available to the decoder by the time the current block is decoded. For example, a neighboring block may be any of blocks A0, A1, A2, B0, B1, B2, and B3 shown in Figure 6b.
[0234] In a first variant, only the first bit is CABAC coded using this context variable.
[0235] In a second variant, the first N bits of the merge index, where N is 2 or greater, are CABAC coded and the context variables are shared among these N bits.
[0236] In a third variant, any N bits of the merge index, where N is 2 or greater, are CABAC coded, and the context variables are shared among these N bits.
[0237] In a fourth variant, the first N bits of the merge index, where N is 2 or greater, are CABAC coded, and N context variables are used for these N bits. Assuming that the context variable has K values, KxN CABAC states are used. For example, in this embodiment, with one adjacent block, the context variable can conveniently have two values, e.g., 0 and 1. In other words, 2N CABAC states are used.
[0238] In a fifth variant, any N bits of the merge index, where N is 2 or greater, are adaptive PM coded and N context variables are used for these N bits.
[0239] Similar modifications can be applied to the fourth to sixteenth embodiments described below.
[0240] Fourth embodiment In the fourth embodiment, the context variable of a bit of a merge index depends on the respective values of the merge indexes of two or more adjacent blocks. For example, the first adjacent block is the left block A0, A1, or A2, and the second adjacent block is the top block B0, B1, B2, or B3. The method of combining two or more merge index values is not particularly limited. An example is shown below.
[0241] The context variable can conveniently have three different values in this case, for example, 0, 1, and 2, since there are two neighboring blocks. Therefore, if the fourth variant described in relation to the third embodiment is applied to this embodiment with three different values, K is 3 instead of 2. In other words, 3N CABAC states are used.
[0242] Fifth embodiment In the fifth embodiment, the context variable of a bit of a merge index depends on the respective values of the merge index of neighboring blocks A2 and B3.
[0243] Sixth embodiment In the sixth embodiment, the context variables of the merge index bits depend on the respective values of the merge indexes of the neighboring blocks A1 and B1. The advantage of this variant is alignment with the merge candidate derivation. As a result, some decoder and encoder implementations can achieve reduced memory accesses.
[0244] Seventh embodiment In the seventh embodiment, the context variable of the bit having the bit position idx_num in the merge index of the current block is obtained according to the following formula:
[0245] ctxIdx=(Merge_index_left==idx_num)+(Merge_index_up==idx_num) Here, Merge_index_left is the merge index of the left block, Merge_index_up is the merge index of the top block, and the symbol == is the equality symbol.
[0246] For example, if there are 6 merge candidates, then 0<=idx_num<=5.
[0247] The left block is block A1 and the top block is block B1 (similar to the sixth embodiment), or the left block may be block A2 and the top block may be block B3 (similar to the fifth embodiment).
[0248] If the merge index of the left block is equal to idx_num, then the formula (Merge_index_left==idx_num) is equal to 1. The following table shows the result of this formula (Merge_index_left==idx_num).
[0249] [Table 2]
[0250] Of course, the formula table (Merge_index_up==idx_num) is the same.
[0251] The following table shows the unary maximum code for each merge index value and the relative bit position of each bit: This table corresponds to Figure 10(b).
[0252] [Table 3]
[0253] If the left block is not a merge block or an affine merge block (i.e., it is coded using the affine merge mode), then the left block is considered unavailable. Similar conditions apply for the top block.
[0254] For example, if only the first bit is CABAC coded, the context variable ctxIdx is 0 if the top-left block does not have a merge index, or if the left block merge index is not the first index (i.e., not 0), and if the top block merge index is not the first index (i.e., not 0) 1 if one of the left and top blocks but not the other has a merge index equal to the first index 2 if the merge index is equal to the first index for each of the left and top blocks is set equal to
[0255] More generally, for a target bit at CABAC coded position idx_num, the context variable ctxIdx is 0 if the top-left block has no merge index, or if the left block merge index is not the i-th index (if i=idx_num), and if the top block merge index is not the i-th index 1 if one of the left and top blocks, but not the other, has a merge index equal to the i-th index 2 if the merge index is equal to the i-th index for each of the left and top blocks where the i-th index means the first index if i=0, the second index if i=1, and so on.
[0256] Eighth embodiment In the eighth embodiment, the context variable of the bit having the bit position idx_num in the merge index of the current block is obtained according to the following formula:
[0257] Ctx=(Merge_index_left>idx_num)+(Merge_index_up>idx_num) where Merge_index_left is the merge index of the left block, Merge_index_up is the merge index of the top block, and the symbol > means "greater than".
[0258] For example, if there are 6 merge candidates, then 0<=idx_num<=5.
[0259] The left block is block A1 and the top block is block B1 (similar to the sixth embodiment), or the left block may be block A2 and the top block may be block B3 (similar to the fifth embodiment).
[0260] If the merge index of the left block is greater than idx_num, then the formula (Merge_index_left>idx_num) is equal to 1. If the left block is not a merge block or an affine merge block (i.e., it is coded using the affine merge mode), then the left block is considered unavailable. Similar conditions apply for the top block.
[0261] The following table shows the result of this formula (Merge_index_left>idx_num).
[0262] [Table 4]
[0263] For example, if only the first bit is CABAC coded, the context variable ctxIdx is 0 if the top-left block has no merge index, or if the left block merge index is less than or equal to the first index (i.e., not 0), and if the top block merge index is less than or equal to the first index (i.e., not 0) 1 if one of the left and top blocks but not the other has a merge index greater than the first index 2 if the merge index is greater than the first index for each of the left and top blocks is set equal to
[0264] More generally, for a target bit at CABAC coded position idx_num, the context variable ctxIdx is 0 if the top-left block has no merge index, or if the left block merge index is less than the i-th index (if i=idx_num), and if the top block merge index is less than or equal to the i-th index 1 if one of the left and top blocks, but not the other, has a merge index greater than the i-th index 2 if the merge index is greater than the i-th index for each of the left and top blocks is set equal to
[0265] The eighth embodiment further improves the coding efficiency compared to the seventh embodiment.
[0266] Ninth embodiment In the fourth to eighth embodiments, the context variables of the bits of the merge index of the current block depend on the respective values of the merge indexes of two or more adjacent blocks.
[0267] In the ninth embodiment, the context variable of the bit of the merge index of the current block depends on the merge flags of each of two or more neighboring blocks, for example, the first neighboring block is the left block A0, A1 or A2, and the second neighboring block is the top block B0, B1, B2 or B3.
[0268] The merge flag is set to 1 if the block is coded using merge mode, and to 0 if other modes such as skip mode or affine merge mode are used. Note that in VMT2.0, affine merge is a separate mode from basic or "classical" merge mode. Affine merge mode can be signaled using a dedicated affine flag. Alternatively, the list of merge candidates may contain affine merge candidates, in which case affine merge mode may be selected and signaled using the merge index.
[0269] The context variable is then 0 if neither the left adjacent block nor the top adjacent block has its merge flag set to 1 1 if one of the left and top neighbors but not the other has its merge flag set to 1 2 if each of the left and top neighboring blocks has its merge flag set to 1 is set to
[0270] This simple evaluation achieves improved coding efficiency over VTM 2.0. Another advantage is lower complexity compared to the seventh and eighth embodiments, since only the merge flags need to be checked, rather than the merge indices of neighboring blocks.
[0271] In a variant, the context variable of the merge index bit of the current block depends on the merge flag of a single neighboring block.
[0272] Tenth embodiment In the third to ninth embodiments, the context variables of the merge index bits of the current block depended on the merge index values or merge flags of one or more neighboring blocks.
[0273] In a tenth embodiment, the context variable for the merge index bits of the current block depends on the value of the skip flag for the current block (current coding unit, or CU): the skip flag is equal to 1 if the current block uses merge skip mode, and is equal to 0 otherwise.
[0274] The skip flag is the first example of another variable or syntax element that has already been decoded or parsed for the current block. This other variable or syntax element is preferably an indicator of the complexity of the motion information in the current block. Because the occurrence of the merge index value depends on the complexity of the motion information, variables or syntax elements such as the skip flag generally correlate with the merge index value.
[0275] More specifically, merge skip mode is generally selected for static scenes or scenes with constant motion. As a result, merge index values are generally lower in merge skip mode than in classical merge mode, which is used to encode inter-prediction including block residuals. This generally occurs for more complex motion. However, the choice between these modes is often also related to quantization and / or RD criteria.
[0276] This simple evaluation improves coding efficiency over VTM2.0, and is very easy to implement since it does not involve checking neighboring blocks or merge index values.
[0277] In a first variant, the context variable of the bit of the merge index of the current block is simply set equal to the skip flag of the current block. This bit can be only the first bit. The other bits are bypass coded as in the first embodiment.
[0278] In the second variant, all bits of the merge index are CABAC coded, each with its own context variable depending on the merge flag, which requires 10 probability states when there are 5 CABAC coded bits in the merge index (corresponding to 6 merge candidates).
[0279] In a third variant, to limit the number of states, only N bits of the merge index are CABAC coded, where N is 2 or more, e.g., the first N bits. This requires 2N states. For example, if the first two bits are CABAC coded, four states are required.
[0280] In general, instead of the skip flag, it is possible to use any other variable or syntax element that has already been decoded or analyzed for the current block and that is an indicator of the complexity of the motion information in the current block.
[0281] Eleventh embodiment The eleventh embodiment relates to affine merge signaling as described above with reference to FIGS. 11(a), 11(b) and 12.
[0282] In an eleventh embodiment, the context variable of the CABAC coded bit of the merge index of the current block (current CU) depends on the affine merge candidate, if any, in the list of merge candidates. This bit can be only the first bit of the merge index, or the first N bits, where N is 2 or more, or any N bits. The other bits are bypass coded.
[0283] Affine prediction is designed to compensate for complex motion. Therefore, for complex motion, the merge index generally has a higher value than for less complex motion. As a result, if the first Affine Merge candidate is far down the list, or if there is no Affine Merge candidate at all, the merge index of the current CU may have a small value.
[0284] Therefore, the context variable may usefully depend on the presence and / or position of at least one affine merge candidate in the list.
[0285] For example, the context variable is set equal to 1 if A1 is affine, 2 if B1 is affine, 3 if B0 is affine, 4 if A0 is affine, 5 if B2 is affine, and 0 if the neighboring block is not affine.
[0286] When the merge index of the current block is decoded or parsed, the affine flags of the merge candidates at these locations have already been checked, so no further memory accesses are required to derive the context of the merge index of the current block.
[0287] This embodiment improves coding efficiency over VTM 2.0: since step 1205 already involves checking neighboring CU affine modes, no additional memory accesses are required.
[0288] In a first variant, to limit the number of states, the context variables are: Set equal to 0 if the neighboring block is not affine or if A1 or B1 is affine, or 1 if B0, A0, or B2 is affine.
[0289] In a second variant, to limit the number of states, the context variables are: It is set equal to 0 if the neighboring block is not affine, 1 if A1 or B1 is affine, and 2 if B0, A0, or B2 is affine.
[0290] In a third variant, the context variables are: Set equal to 1 if A1 is affine, 2 if B1 is affine, 3 if B0 is affine, 4 if A0 or B2 are affine, 0 if the neighboring block is not affine.
[0291] Note that these locations are already checked when the merge index is decoded or parsed, since the affine flag decoding depends on these locations, so no additional memory accesses are required to derive the merge index context that is coded after the affine flag.
[0292] Twelfth embodiment In a twelfth embodiment, signaling the affine mode includes inserting the affine mode as a candidate motion predictor.
[0293] In one example of the twelfth embodiment, affine merge (and affine merge skip) is signaled as a merge candidate (i.e., as one of the merge candidates for use with the classical merge mode or the classical merge skip mode). In this case, modules 1205, 1206, and 1207 of FIG. 12 are removed. Furthermore, the maximum possible number of merge candidates is incremented so as not to affect the coding efficiency of the merge mode. For example, in current VTM versions, this value is set equal to 6, so when applying this embodiment to the current version of VTM, the value becomes 7.
[0294] The advantage is a simplified design of syntax elements for merge mode, since fewer syntax elements need to be decoded. In some situations, an improvement / change in coding efficiency may be observed.
[0295] Two possibilities for implementing this example are described below.
[0296] The merge index of an affine merge candidate always has the same position in the list, whatever the values of the other merge MVs. The position of a candidate motion predictor indicates its likelihood of being selected, so the higher it is placed in the list (lower index value), the more likely that motion vector predictor is selected.
[0297] In the first example, the merge index of an affine merge candidate always has the same position in the list of merge candidates. This means that it has a fixed "merge idx" value. For example, this value can be set equal to 5, since the affine merge mode should represent complex motion, which is not the most probable content. An additional advantage of this embodiment is that when the current block is parsed (decoding / reading syntax elements, not just decoding the data itself), the current block can be set as an affine block. As a result, this value can be used to determine the CABAC context of the affine flag used for AMVP. Therefore, the conditional probability should be improved for this affine flag, and the coding efficiency should be better.
[0298] In the second example, an affine merge candidate is derived along with other merge candidates. In this example, a new affine merge candidate is added to the list of merge candidates (for classic merge mode or classic merge skip mode). Figure 16 illustrates this example. Compared to Figure 13, the affine merge candidate is the first affine neighboring block from A1, B1, B0, A0, and B2 (1917). If the same conditions as 1205 in Figure 12 are valid (1927), a motion vector field generated using affine parameters is generated, resulting in an affine merge candidate (1929). The initial merge candidate list can have 4, 5, 6, or 7 candidates depending on the ATMVP, temporal, and affine merge candidate usage.
[0299] The order among all these candidates is important as the more likely candidates should be processed first to ensure that they are more likely to make the cut in the motion vector candidates, the preferred order is as follows:
[0300] A1, B1, B0, A0, Affine merge, ATMVP, B2, Temporal, Combination, Zero_MV It is important to note that the affine merge candidate is placed before the ATMVP candidate but after the four major neighboring blocks. The advantage of placing the affine merge candidate before the ATMVP candidate is increased coding efficiency compared to placing it after the ATMVP and temporal predictor candidates. This coding efficiency improvement depends on the GOP (group of pictures) structure and the QP (quantization parameter) setting of each picture within the GOP. However, for most commonly used GOP and QP settings, this order results in increased coding efficiency.
[0301] A further advantage of this solution is the clean design of the classical merge and classical merge skip modes (i.e., merge modes with additional candidates such as ATMVP or affine merge candidates) for both syntax and derivation processing. Furthermore, the merge index of an affine merge candidate can be modified according to the availability or value (duplicate check) of previous candidates in the list of merge candidates. As a result, efficient signaling can be obtained.
[0302] In a further example, the merge index for the affine merge candidates is variable according to one or several conditions.
[0303] For example, the merge index or position in the list associated with an affine merge candidate varies according to a criterion: the principle is to set a low value for the merge index corresponding to an affine merge candidate if the probability of the affine merge candidate being selected is high (and a higher value if the probability of selection is low).
[0304] In a twelfth embodiment, the affine merge candidates have a merge index value. To improve the coding efficiency of the merge index, it is effective to make the context variables of the bits of the merge index dependent on the affine flags of the neighboring blocks and / or the current block.
[0305] For example, the context variable may be determined using the following formula:
[0306] ctxIdx=IsAffine(A1)+IsAffine(B1)+IsAffine(B0)+IsAffine(A0)+IsAffine(B2) The resulting context value can have the values 0, 1, 2, 3, 4, or 5.
[0307] The affine flag increases coding efficiency.
[0308] In a first variant, to include fewer neighboring blocks, ctxIdx=IsAffine(A1)+IsAffine(B1). The resulting context value can have the values 0, 1, or 2.
[0309] Also, in a second variant, to include fewer neighboring blocks, ctxIdx=IsAffine(A2)+IsAffine(B3). Again, the resulting context value can have the values 0, 1, or 2.
[0310] In a third variant, ctxIdx=IsAffine(current block) to avoid including neighboring blocks. The resulting context value can have the value 0 or 1.
[0311] FIG. 15 is a flowchart of a partial decoding process of some syntax elements related to coding modes according to a third variant. In this diagram, the skip flag (1601), prediction mode (1611), merge flag (1603), merge index (1608), and affine flag (1606) can be decoded. This flowchart is similar to the flowchart of FIG. 12 described above, and therefore a detailed description will be omitted. The difference is that the merge index decoding process takes the affine flag into account so that the affine flag decoded before the merge index can be used when obtaining the context variable for the merge index. This is not the case in VTM2.0. In VTM2.0, the affine flag of the current block always has the same value "0" and therefore cannot be used to obtain the context variable for the merge index.
[0312] Thirteenth embodiment In the tenth embodiment, the context variable for the bit of the merge index of the current block depends on the value of the skip flag for the current block (current coding unit, or CU). In the thirteenth embodiment, instead of directly using the skip flag value to derive the context variable for the target bit of the merge index, the context value for the target bit is derived from the context variable used to encode the skip flag of the current CU. This is possible because the skip flag itself is CABAC encoded and therefore has a context variable. Preferably, the context variable for the target bit of the merge index of the current CU is set equal to (copied from) the context variable used to encode the skip flag of the current CU. The target bit can be only the first bit. Other bits may be bypass encoded as in the first embodiment.
[0313] The context variable for the skip flag of the current CU is derived in the manner specified in VTM2.0. The advantage of this embodiment compared to the VTM2.0 reference software is that the complexity of merge index decoding and decoder design (and encoder design) is reduced without affecting coding efficiency. Indeed, in this embodiment, only a minimum of one CABAC state is required to encode the merge index, instead of five for the current VTM merge index encoding (encoding / decoding). Furthermore, other bits are CABAC bypass coded, reducing the number of operations compared to coding all bits with CABAC, thereby reducing worst-case complexity.
[0314] Fourteenth embodiment In the thirteenth embodiment, the context variable / value of the target bit was derived from the context variable of the skip flag of the current CU. In the fourteenth embodiment, the context value of the target bit is derived from the context variable of the affine flag of the current CU.
[0315] This is possible because the affine flag itself is CABAC coded and therefore has a context variable. Preferably, the context variable for the target bit of the merge index of the current CU is set equal to (copied from) the context variable for the affine flag of the current CU. The target bit can be only the first bit. The other bits are bypass coded as in the first embodiment.
[0316] The context variable for the current CU's affine flag is derived as specified in VTM2.0.
[0317] The advantage of this embodiment compared to the VTM 2.0 reference software is that the complexity of merge index decoding and decoder design (and encoder design) is reduced without affecting coding efficiency. Indeed, in this embodiment, a minimum of one CABAC state is required for the merge index, instead of five for the current VTM merge index coding (encoding / decoding). Furthermore, other bits are CABAC bypass coded, reducing the number of operations compared to coding all bits with CABAC, thereby reducing worst-case complexity.
[0318] Fifteenth embodiment In some of the above-described embodiments, the context variables had more than two values, e.g., three values: 0, 1, and 2. However, to reduce complexity and the number of states to be processed, it is possible to limit the number of allowed context variable values to two, e.g., 0 and 1. This can be achieved, for example, by changing any initial context variable that has a value of 2 to 1. In practice, this simplification has no or only a limited effect on coding efficiency.
[0319] Combinations of the embodiment and other embodiments Any two or more of the above embodiments may be combined.
[0320] The preceding description has focused on encoding and decoding of merge indexes. For example, a first embodiment includes generating a list of merge candidates including ATMVP candidates (in the case of classical merge mode or classical merge skip mode, i.e., non-affine merge mode or non-affine merge skip mode), selecting one of the merge candidates in the list, and generating a merge index for the selected merge candidate using CABAC coding, where one or more bits of the merge index are bypass CABAC coded. In principle, the present invention can be applied to modes other than merge mode (e.g., affine merge mode), including generating a list of motion information predictor candidates (e.g., a list of affine merge candidates or motion vector predictor (MVP) candidates), selecting one of the motion information predictor candidates (e.g., MVP candidates) in the list, and generating an identifier or index for the selected motion information predictor candidate in the list (e.g., the selected affine merge candidate or the selected MVP candidate for predicting the motion vector of the current block). Therefore, the present invention is not limited to merge modes (i.e., classical merge mode and classical merge skip mode), and the index to be coded or decoded is not limited to the merge index. For example, in the development of VVC, it is contemplated that the techniques of the foregoing embodiments may be applied to (or extended to) modes other than merge mode, such as the AMVP mode of HEVC, or its equivalent in VVC, or an affine merge mode, and the appended claims should be construed accordingly.
[0321] As mentioned above, in the aforementioned embodiments, one or more candidate motion information (e.g., motion vectors) for the affine merge mode (affine merge or affine merge skip mode) and / or one or more affine parameters are obtained from first neighboring blocks that are affine coded between spatially adjacent blocks (e.g., positions A1, B1, B0, A0, B2) or temporally related blocks (e.g., the “center” block with its co-located blocks, or its spatial neighbors, such as “H”). These positions are shown in Figures 6a and 6b. To enable this obtaining (e.g., deriving or sharing or "merging") of one or more motion information and / or affine parameters between the current block (or a group of samples / pixel values currently being coded / decoded, such as the current CU) and an adjacent block (spatially adjacent or temporally related to the current block), one or more affine merge candidates are added to a list of merge candidates (i.e., classical merge mode candidates), so that if the selected merge candidate (signaled using a merge index, e.g., using a syntax element such as "merge_idx" in HEVC or its functionally equivalent syntax element) is an affine merge candidate, the current CU / block is coded / decoded using the affine merge mode together with the affine merge candidate.
[0322] As described above, such one or more affine merge candidates for obtaining (e.g., deriving or sharing) one or more motion information and / or affine parameters for the affine merge mode may also be signaled using a separate list (or set) of affine merge candidates (which may be the same as or different from the list of merge candidates used for the classical merge mode).
[0323] According to one embodiment of the present invention, when the techniques of the previous embodiments are applied to affine merge mode, the list of affine merge candidates can be generated using the same technique as the motion vector derivation process for classical merge mode shown in and described in conjunction with Figure 8, or using the same technique as the merge candidate derivation process shown in and described in conjunction with Figure 13. The advantage of sharing the same technique for generating / compiling this list of affine merge candidates (for affine merge mode or affine merge skip mode) and the list of merge candidates (for classical merge mode or classical merge skip mode) is that the complexity of the encoding / decoding process is reduced compared to having separate techniques.
[0324] It should be understood that, according to other embodiments, similar techniques are applied to other inter-prediction modes that require signaling a selected motion information predictor (from multiple candidates) to achieve similar advantages.
[0325] According to another embodiment, a separate technique can be used to generate / compile the list of affine merge candidates, as described below in connection with FIG.
[0326] FIG. 24 is a flowchart illustrating an affine merge candidate derivation process for affine merge modes (affine merge mode and affine merge skip mode). In the first step of the derivation process, five block locations are considered (2401-2405) to obtain / derive spatial affine merge candidates 2413. These locations are the spatial locations indicated in FIG. 6a (and FIG. 6b) by reference numerals A1, B1, B0, A0, and B2. In the next step, the availability of spatial motion vectors is checked to determine whether each of the inter-mode coded blocks associated with each location A1, B1, B0, A0, and B2 is coded in affine mode (e.g., using one of affine merge, affine merge skip, or affine AMVP mode) (2410). At most five motion vectors (i.e., spatial affine merge candidates) are selected / obtained / derived. A predictor is considered available if a predictor exists (e.g., there is information to obtain / derive a motion vector associated with that position), and if the block is not intra-coded, and if the block is affine (i.e., coded using affine mode).
[0327] Next, for each available block position, affine motion information is derived / obtained (2411) (2410). This derivation is performed for the current block based on the affine model of the block position (and its affine model parameters, e.g., as described in connection with Figures 11(a) and 11(b)). Next, a pruning process (2412) is applied to remove candidates that provide the same affine motion compensation (or have the same affine model parameters) as those previously added to the list.
[0328] At the end of this stage, the list of spatially affine merge candidates contains up to five candidates.
[0329] If the number of candidates (Nb_Cand) is strictly less than the maximum number of candidates (2426) (where Max_Cand is the value signaled in the bitstream slice header, which is equal to 5 for affine merge mode but may differ / variable depending on the implementation).
[0330] Next, constructed affine merge candidates (i.e., additional affine merge candidates generated to provide some diversity in addition to approaching the target number, serving a role similar to, for example, combined bi-predictive merge candidates in HEVC) are generated (2428). These constructed affine merge candidates are based on motion vectors associated with neighboring spatial and temporal positions of the current block. First, control points are defined (2418, 2419, 2420, 2421) to generate motion information for generating the affine model. Two of these control points correspond to, for example, v0 and v1 in Figures 11(a) and 11(b). These four control points correspond to the four corners of the current block.
[0331] The motion information for the top-left control point (2418) is obtained from (e.g., by being equal to) the motion information for the block position at position B2 (2405) if it exists and this block is coded in inter mode (2414). Otherwise, the motion information for the top-left control point (2418) is obtained from (e.g., by being equal to) the motion information for the block position at position B3 (2406) (as shown in Figure 6b) if it exists and this block is coded in inter mode (2414), which is not the case, and the motion information for the top-left control point (2418) is obtained from (e.g., by being equal to) the motion information for the block position at position A2 (2407) (as shown in Figure 6b) if it exists and this block is coded in inter mode (2414). If there is no block available for this control point, it is considered unavailable.
[0332] The motion information for the control point top right (2419) is obtained from (e.g., equal to) the motion information for the block position at position B1 (2402) if it exists and this block is coded in inter mode (2415). Otherwise, the motion information for the control point top right (2419) is obtained from (e.g., equal to) the motion information for the block position at position B0 (2403) if it exists and this block is coded in inter mode (2415). If there is no block available at this control point, it is considered unavailable (unavailable).
[0333] The motion information for the bottom left control point (2420) is obtained from (e.g., equal to) the motion information for the block position at position A1 (2401) if it exists and this block is coded in inter mode (2416). Otherwise, the motion information for the bottom left control point (2420) is obtained from (e.g., equal to) the motion information for the block position at position A0 (2404) if it exists and this block is coded in inter mode (2416). If there is no block available for this control point, it is considered unavailable.
[0334] The motion information of the control point bottom right (2421) is obtained from (e.g., equal to) the motion information of a temporal candidate, e.g., the co-located block position at position H (2408) (as shown in Figure 6a), if it exists and this block is coded in inter mode (2417). If there is no block available at this control point, it is considered unavailable (unavailable).
[0335] Based on these control points, up to ten constructed affine merge candidates can be generated (2428). These candidates are generated based on affine models with four, three, or two control points. For example, the first constructed affine merge candidate may be generated using four control points. Then, the next four constructed affine merge candidates are four possibilities that can be generated using four different sets of three control points (i.e., four different possible combinations of sets that include three of the four available control points). Then, the other constructed affine merge candidates are generated using different sets of two control points (i.e., different possible combinations of sets that include two of the four control points).
[0336] If the number of candidates (Nb_Cand) remains strictly less than the maximum number of candidates (Max_Cand) after adding these additional (constructed) affine merge candidates (2430), other additional virtual motion information candidates, such as zero motion vector candidates (or combined bi-predictive merge candidates, if applicable), are added / generated (2432) until the number of candidates in the list of affine merge candidates reaches a target number (e.g., the maximum number of candidates).
[0337] At the end of this process, a list or set of affine merge mode candidates (i.e., a list or set of affine merge mode candidates that are affine merge mode and affine merge skip mode) is generated / constructed (2434). As shown in FIG. 24, the list or set of affine merge (motion vector predictor) candidates is constructed / generated (2434) from a subset of spatial candidates (2401-2407) and temporal candidates (2408). It should be understood that, according to embodiments of the present invention, other affine merge candidate derivation processes having different orders for checking availability, pruning processes, or the number / type of potential candidates (e.g., ATMVP candidates can also be added in a manner similar to the merge candidate list derivation process of FIG. 13 or FIG. 16) can also be used to generate the list / set of affine merge candidates.
[0338] The following embodiments show how a list (or set) of affine merge candidates can be used to signal (e.g., encode or decode) a selected affine merge candidate (which can be signaled using the merge index used for the merge mode, or a separate affine merge index used specifically for the affine merge mode).
[0339] In the following embodiments, a merge mode (i.e., a merge mode other than the affine merge mode defined later, in other words, a classical non-affine merge mode or a classical non-affine merge skip mode) is a type of merge mode in which motion information of either a spatially neighboring block or a temporally related block is obtained for (or derived for, or shared with) the current block; a merge mode predictor candidate (i.e., a merge candidate) is information about one or more spatially neighboring blocks or temporally related blocks from which the current block can obtain / derive motion information in the merge mode; a merge mode predictor is a selected merge mode predictor candidate, the information of which is used when predicting the motion information of the current block and during signaling in a merge mode (e.g., encoding or decoding) process; and an index (e.g., a merge index) identifying a merge mode predictor from a list (or set) of merge mode predictor candidates is The affine merge mode is a type of merge mode in which motion information of one of spatially adjacent or temporally related blocks is obtained for (derived for or shared with) the current block so that the motion information of the current block and / or affine parameters for affine mode processing (or affine motion model processing) can use this obtained / derived / shared motion information; the affine merge mode predictor candidate (i.e., affine merge candidate) is information about one or more spatially adjacent or temporally related blocks from which the current block can obtain / derive motion information in the affine merge mode; the affine merge mode predictor is a selected affine merge mode predictor candidate, whose information can be used in the affine motion model when predicting the motion information of the current block and during signaling in the affine merge mode (e.g., encoding or decoding) processing;An index (e.g., an affine merge index) is signaled that identifies an affine merge mode predictor from a list (or set) of affine merge mode predictor candidates. In the following embodiments, it is understood that an affine merge mode is a merge mode that has its own affine merge index (an identifier that is a variable) to identify one affine merge mode predictor candidate from a list / set of candidates (also known as an "affine merge list" or "sub-block merge list") and has a single index value associated with it, whereas an affine merge index is signaled to identify that particular affine merge mode predictor candidate.
[0340] In the following embodiments, "merge mode" refers to either the classical merge skip mode in HEVC / JEM / VTM or the classical merge mode or any functionally equivalent mode, provided that motion information acquisition (e.g., derivation or sharing) and merge index signaling as described above are used in said modes. It should be understood that "affine merge mode" also refers to either the affine merge mode or the affine merge skip mode (using such acquisition / derivation, if present), or any other functionally equivalent mode, provided that the same features are used in said modes.
[0341] Sixteenth embodiment In a sixteenth embodiment, a motion information predictor index for identifying an affine merge mode predictor (candidate) from a list of affine merge candidates is signaled using CABAC coding, and one or more bits of the motion information predictor index are bypass CABAC coded.
[0342] According to a first variant of the embodiment, in an encoder, a motion information predictor index for an affine merge mode is encoded by generating a list of motion information predictor candidates, selecting one of the motion information predictor candidates in the list as an affine merge mode predictor, and generating a motion information predictor index for the selected motion information predictor candidate using CABAC coding, where one or more bits of the motion information predictor index are bypass CABAC coded. Next, data indicating an index for the selected motion information predictor candidate is included in the bitstream. Next, a decoder generates a list of motion information predictor candidates from the bitstream including this data, decodes the motion information predictor index using CABAC decoding, where one or more bits of the motion information predictor index are bypass CABAC coded, and, when the affine merge mode is used, decodes the motion information predictor index for the affine merge mode by identifying one of the motion information predictor candidates in the list as an affine merge mode predictor using the decoded motion information predictor index.
[0343] According to a further variation of the first variation, one or more of the motion information predictor candidates in the list are also selectable as merge mode predictors when merge mode is used, such that when merge mode is used, a decoder can use a decoded motion information predictor index (e.g., merge index) to identify one of the motion information predictor candidates in the list as the merge mode predictor. In this further variation, an affine merge index is used to signal the affine merge mode predictor (candidate), and signaling the affine merge index is implemented using merge index signaling according to any one of the first to fifteenth embodiments or index signaling similar to merge index signaling used in current VTM or HEVC.
[0344] In this variant, when a merge mode is used, signaling a merge index can be performed using merge index signaling according to any one of the first to fifteenth embodiments or merge index signaling used in current VTM or HEVC. In this variant, different index signaling schemes can be used for signaling an affine merge index and for signaling a merge index. An advantage of this variant is that better coding efficiency can be achieved by using efficient index coding / signaling for both the affine merge mode and the merge mode. Furthermore, in this variant, separate syntax elements can be used for the merge index (such as "Merge_idx[][]" in HEVC or its functional equivalent) and the affine merge index (such as "A_Merge_idx[][]"). This allows the merge index and the affine merge index to be signaled (encoded / decoded) separately.
[0345] According to yet another further variant, when the merge mode is used and one of the motion information predictor candidates in the list is also selectable as the merge mode predictor, CABAC encoding uses the same context variable for at least one bit of the motion information predictor index (e.g., merge index or affine merge index) of the current block for both modes, i.e., when the affine merge mode is used and when the merge mode is used, so that the affine merge index and at least one bit of the merge index share the same context variable. Then, when the merge mode is used, the decoder uses the decoded motion information predictor index to identify one of the motion information predictor candidates in the list as the merge mode predictor, and CABAC decoding uses the same context variable for at least one bit of the motion information predictor index of the current block for both modes, i.e., when the affine merge mode is used and when the merge mode is used.
[0346] According to a second variant of the embodiment, in an encoder, the motion information predictor index is coded by generating a list of motion information predictor candidates, selecting one of the motion information predictor candidates in the list as an affine merge mode predictor when an affine merge mode is used, selecting one of the motion information predictor candidates in the list as a merge mode predictor when a merge mode is used, and using CABAC coding to generate a motion information predictor index for the selected motion information predictor candidate, wherein one or more bits of the motion information predictor index are bypass CABAC coded. Then, data indicating the index for the selected motion information predictor candidate is included in the bitstream. The decoder then decodes the motion information predictor index from the bitstream by generating a list of motion information predictor candidates, decoding the motion information predictor index using CABAC decoding, one or more bits of the motion information predictor index being bypass CABAC decoded, and if affine merge mode is used, using the decoded motion information predictor index to identify one of the motion information predictor candidates in the list as an affine merge mode predictor, and if merge mode is used, using the decoded motion information predictor index to identify one of the motion information predictor candidates in the list as a merge mode predictor.
[0347] According to a further variation of the second variation, the affine merge index signaling and the merge index signaling use the same index signaling scheme according to any one of the first to fifteenth embodiments, or the merge index signaling used in the current VTM or HEVC. An advantage of this further variation is a simple design during implementation, which may also lead to less complexity. In this variation, when the affine merge mode is used, the CABAC encoding of the encoder includes using a context variable for at least one bit of the motion information predictor index (affine merge index) of the current block, where the context variable is separable from another context variable for at least one bit of the motion information predictor index (merge index) when the merge mode is used, and data indicating the use of the affine merge mode is included in the bitstream so that the context variables for the affine merge mode and the merge mode can be distinguished (clearly identified) for the CABAC decoding process. The decoder then obtains, from the bitstream, data for indicating the use of affine merge mode in the bitstream, and when affine merge mode is used, the CABAC decoding uses this data to distinguish between affine merge indexes and context variables for merge indexes. Furthermore, at the decoder, the data for indicating the use of affine merge mode can also be used to generate a list (or set) of affine merge mode predictor candidates when the obtained data indicates the use of affine merge mode, and to generate a list (or set) of merge mode predictor candidates when the obtained data indicates the use of merge mode.
[0348] This variation allows both merge and affine merge indices to be signaled using the same index signaling scheme, while the merge and affine merge indices are still encoded / decoded independently of each other (e.g., by using separate context variables).
[0349] One way to use the same index signaling scheme is to use the same syntax element for both affine merge indexes and merge indexes, i.e., when affine merge mode and when merge mode are used, the motion information predictor index of the selected motion information predictor candidate is coded using the same syntax element in both cases. Then, at the decoder, the motion information predictor index is decoded by parsing the same syntax element from the bitstream, regardless of whether the current block was coded (and decoded) using affine merge mode or merge mode.
[0350] Figure 22 shows partial decoding processing of some syntax elements related to encoding modes (i.e., the same index signaling scheme) according to this variant of the sixteenth embodiment. This figure shows signaling of an affine merge index (2255—“merge idx affine”) for affine merge mode (2257: Yes) and a merge index (2258—“merge idx”) for merge mode (2257: No) with the same index signaling scheme. It should be understood that in some variants, the affine merge candidate list can include ATMVP candidates, just like the merge candidate list of the current VTM. The encoding of the affine merge index is similar to the encoding of the merge index for merge mode, as shown in Figure 10(a), Figure 10(b), or Figure 14. In some variations, even if no ATMVP merge candidates are defined in the affine merge candidate derivation, if ATMVP is enabled for merge mode with up to five other candidates (i.e., six candidates total) so that the maximum number of candidates in the affine merge candidate list matches the maximum number of candidates in the merge candidate list, the affine merge index is encoded as described in Figure 10(b). Thus, each bit of the affine merge index has its own context. All context variables used for the merge index signaling bits are independent of the context variables used for the affine merge index signaling bits.
[0351] According to a further variation, the same index signaling scheme shared by merge index and affine merge index signaling uses CABAC coding for only the first bin, as in the first embodiment. That is, all bits except the first bit of the motion information predictor index are bypass CABAC coded. In this further variation of the sixteenth embodiment, when ATMVP is included as a candidate in one of the lists of merge candidates or affine merge candidates (e.g., when ATMVP is enabled at the SPS level), the coding of each index (i.e., merge index or affine merge index) is modified so that only the first bit of the index is CABAC coded using a single context variable, as shown in FIG. 14. This single context is set in the same way as the current VTM reference software when ATMVP is not enabled at the SPS level. The other bits (the second through fifth bits or the fourth bit if only five candidates are present in the list) are bypass coded. If the merge candidate list does not include ATMVP as a candidate (e.g., if ATMVP is disabled at the SPS level), there are five merge candidates and five affine merge candidates available. Only the first bit of the merge index for merge mode is encoded by CABAC using the first single context variable. And only the first bit of the affine merge index for affine merge mode is encoded by CABAC using the second single context variable. These first and second context variables are set in the same way as the current VTM reference software when ATMVP is not enabled at the SPS level for both the merge index and the affine merge index. The other bits (the second through fourth bits) are bypass-decoded.
[0352] The decoder generates the same list of merge candidates and the same list of affine merge candidates as the encoder. This can be achieved, for example, by using the method of FIG. 24. The same index signaling scheme is used for both merge mode and affine merge mode, but the affine flag (2256) is used to determine whether the currently decoded data is for a merge index or an affine merge index, so that the first and second context variables are separable (or distinguishable) from each other for the CABAC decoding process. That is, the affine flag (2256) is used during the index decoding process (i.e., in step 2257) to determine whether to decode "merge idx 2258" or "merge idx affine 2255." If ATMVP is not included as a candidate in the list of merge candidates (e.g., if ATMVP is disabled at the SPS level), there are five merge candidates in both lists of candidates (for merge mode and affine merge mode). Only the first bit of the merge index is decoded by CABAC using the first single context variable. Then, only the first bit of the affine merge index is decoded by CABAC using the second single context variable. All other bits (second to fourth bits) are bypass-decoded. In contrast to the current reference software, if ATMVP is included as a candidate in the list of merge candidates (e.g., when ATMVP is enabled at the SPS level), only the first bit of the merge index is decoded by CABAC using the first single context variable in decoding the merge index and the second single context variable in decoding the affine merge index. The other bits (second to fifth bits or fourth bit) are bypass-decoded. The decoded index is then used to identify the candidate selected by the encoder from the corresponding list of candidates (i.e., merge candidate or affine merge candidate).
[0353] The advantage of this variant is that by using the same index signaling scheme for both the merge index and the affine merge index, the complexity of index decoding and decoder design (and encoder design) for implementing these two different modes is reduced without significantly affecting coding efficiency. Indeed, with this variant, only two CABAC states (one for each of the first and second single-context variables) are required for index signaling, instead of nine or ten if all bits of the merge index and all bits of the affine merge index are CABAC encoded / decoded. Furthermore, because all other bits (apart from the first bit) are CABAC bypass coded, the worst-case complexity is reduced, and the number of operations required during the CABAC coding / decoding process is reduced compared to coding all bits using CABAC.
[0354] According to yet another variant, CABAC encoding or decoding uses the same context variable for at least one bit of the motion information predictor index of the current block both when the affine merge mode is used and when the merge mode is used. In this variant, the context variable used for the first bit of the merge index and the first bit of the affine merge index does not depend on which index is being coded or decoded; that is, the first and second single context variables (from the previous variant) are not distinguished / separated but are the same single context variable. Thus, in contrast to the previous variant, the merge index and the affine merge index share one context variable during CABAC processing. As shown in Figure 23, the index signaling scheme is the same for both the merge index and the affine merge index; that is, only one type of index, "merge idx (2308)," is coded or decoded for both modes. As far as the CABAC decoder is concerned, the same syntax elements are used for both the merge index and the affine merge index, and there is no need to distinguish between them when considering context variables. Therefore, there is no need to use the affine flag (2306) to determine whether the current block is coded (decoded) in affine merge mode, as in step (2257) of Figure 22, and there is no branching after step 2306 of Figure 23, since only one index ("merge idx") needs to be decoded. The affine flag is used to perform motion information prediction in affine merge mode, i.e., during the prediction process after the CABAC decoder decodes the index ("merge idx"). Furthermore, only the first bit of this index (i.e., the merge index and the affine merge index) is coded by CABAC using one single context, and the other bits are bypass coded as described for the first embodiment.Therefore, in this further variant, one context variable, the first bit of the merge index and the affine merge index, is shared by both merge index and affine merge index signaling. If the size of the list of candidates differs between merge index and affine merge index, the maximum number of bits for signaling the associated index in each case may also differ, i.e., they are independent of each other. Therefore, the number of bypass coding bits can be adjusted accordingly, as needed, according to the value of the affine flag (2306), for example, to enable parsing of data for the associated index from the bitstream.
[0355] The advantage of this variant is that the complexity of the merge index and affine merge index decoding process and decoder design (and encoder design) is reduced without significantly affecting coding efficiency. Indeed, in this further variant, when signaling both the merge index and the affine merge index, only one CABAC state is required, instead of the previous variant or 9 or 10 CABAC states. Furthermore, because all other bits (apart from the first bit) are CABAC bypass coded, the worst-case complexity is reduced, and the number of operations required during the CABAC coding / decoding process is reduced compared to coding all bits by CABAC.
[0356] In the aforementioned variations of this embodiment, affine merge index signaling and merge index signaling can reduce the number of contents and / or share one or more contexts as described in any of the first to fifteenth embodiments. The advantage of this is reduced complexity due to the reduced number of contexts required to encode or decode these indices.
[0357] In the aforementioned variants of this embodiment, the motion information predictor candidate comprises information for obtaining (or deriving) one or more of the direction, the list ID, the reference frame index, and the motion vector. Preferably, the motion information predictor candidate includes information for obtaining the motion vector predictor candidate. In a preferred variant, a motion information predictor index (e.g., an affine merge index) is used to signal the affine merge mode predictor candidate, and the affine merge index signaling is implemented using index signaling similar to the merge index signaling according to any one of the first to fifteenth embodiments or the merge index signaling used in current VTM or HEVC (with the affine merge mode motion information predictor candidate as the merge candidate).
[0358] In the aforementioned variations of this embodiment, the generated list of motion information predictor candidates includes an ATMVP candidate, as in the first embodiment, or as in some variations of the other aforementioned second through fifteenth embodiments. The ATMVP candidate may be included in one or both of the merge candidate list and the affine merge candidate list. Alternatively, the generated list of motion information predictor candidates does not include an ATMVP candidate.
[0359] In the aforementioned variant of this embodiment, the maximum number of candidates that can be included in the list of candidates for the merge index and the affine merge index is fixed. The maximum number of candidates that can be included in the list of candidates for the merge index and the affine merge index may be the same. Then, data for determining (or indicating) the maximum number (or target number) of motion information predictor candidates that can be included in the generated list of motion information predictor candidates is included in the bitstream by the encoder, and the decoder obtains the data for determining the maximum number (or target number) of motion information predictor candidates that can be included in the generated list of motion information predictor candidates from the bitstream. This allows data for decoding the merge index or the affine merge index to be parsed from the bitstream. This data for determining (or indicating) the maximum number (or target number) may be the maximum number (or target number) itself when decoded, or may enable the decoder to determine this maximum / target number in conjunction with other parameters / syntax elements, such as "five_minus_max_num_merge_cand" or "MaxNumMergeCand-1" used in HEVC, or functionally equivalent parameters.
[0360] Alternatively, if the maximum number (or target number) of candidates in the lists of merge index and affine merge index candidates can vary or be different (such as because the use of ATMVP candidates or any other candidates is enabled or disabled for one list but not the other, or because the lists use different candidate list generation / derivation processes), the maximum number (or target number) of motion information predictor candidates that can be included in the generated list of motion information predictor candidates when the affine merge mode and when the merge mode are used can be determined separately, and the encoder includes data for determining the maximum / target number in the bitstream. The decoder then obtains the data for determining the maximum / target number from the bitstream and uses the obtained data to parse or decode the motion information predictor index. The affine flag can then be used to switch, for example, between parsing or decoding the merge index and the affine merge index.
[0361] As mentioned above, one or more of the additional inter-prediction modes (such as an MHII merge mode, a triangle merge mode, and an MMVD merge mode) can be used in addition to or instead of the merge mode or the affine merge mode, and an index (or flag or information) for one or more of the additional inter-prediction modes can be signaled (encoded or decoded). The following embodiments relate to signaling information (such as an index) for the additional inter-prediction modes.
[0362] Seventeenth embodiment Signaling of all inter prediction modes (including merge mode, affine merge mode, MHII merge mode, triangle merge mode, and MMVD merge mode) These multiple inter-prediction "merge" modes are signaled using data provided in the bitstream along with their associated syntax (elements) according to the seventeenth embodiment. Figure 26 shows the decoding process for the inter-prediction mode for a current CU (image portion or block) according to one embodiment of the present invention. As described in connection with Figure 12 (and the skip flag of 1201), the first CU skip flag is extracted from the bitstream (2601). If the CU is not skip (2602), i.e., if the current CU is not processed in skip mode, the pred mode flag (2603) and / or merge flag (2606) are decoded to determine whether the current CU is a merge CU. If the current CU is processed as a merge skip (2602) or merge CU (2607), the MMVD_Skip_Flag or MMVD_Merge_Flag is decoded (2608). If this flag is equal to 1 (2609), the current CU is decoded using MMVD merge mode (i.e., in MMVD merge mode or in MMVD merge mode), resulting in the MMVD merge index being decoded (2610), followed by the MMVD distance index (2611) and the MMVD direction index (2612). If the CU is not an MMVD merge CU (2609), the merge sub-block flag is decoded (2613). This flag is also referred to as the "affine flag" in the previous description. If the current CU is processed in affine merge mode (also known as "sub-block merge" mode) (2614), the merge sub-block index (i.e., the affine merge index) is decoded (2615). If the current CU is not processed in affine merge mode (2614) or skip mode (2616), the MHII merge flag is decoded (2620). If this block is processed in MHII merge mode (2621), the normal merge index (2619) is decoded with the associated intra prediction mode for MHII merge mode (2622). Note that MHII merge mode is only available for non-skip "merge" mode, not for skip mode.If the MHII merge flag is equal to 0 (2621), or if the current CU is not processed in skip mode (2616) or affine merge mode (2614), the triangle merge flag is decoded (2617). If this CU is processed in triangle merge mode (2618), the triangle merge index is decoded (2623). If the current CU is not processed in triangle merge mode (2618), the current CU is a regular merge mode CU, and the merge index is decoded.
[0363] Signaling each merge candidate MMVD Merge Flag / Index Signaling In a first variant of the seventeenth embodiment, only two initial candidates are available for use / selection in the MMVD merge mode. However, if eight possible values for the distance index and four possible values for the direction index are also signaled with the bitstream, the number of potential candidates for use in the MMVD merge mode at the decoder is 64 (two candidates × eight distance indexes × four direction indexes), and each potential candidate is different from another (i.e., unique) if the initial candidate is different. These 64 potential candidates can be evaluated / compared for the MMVD merge mode at the encoder side, and the MMVD merge index (2610) for the selected initial candidate is then signaled with a unary max code. Because only two initial candidates are used, this MMVD merge index (2610) corresponds to a flag. Figure 27(a) shows the encoding of this flag for CABAC coding using one context variable. It should be understood that in another variant, a different number of initial candidates, distance index values, and / or direction index values may be used instead, with the signaling of the MMVD merge index adapted accordingly (e.g., at least one bit is CABAC coded using one context variable).
[0364] Triangle Merge Index Signaling In a first variant of the seventeenth embodiment, triangle merge indexes are signaled differently compared to index signaling for other inter-prediction modes. In triangle merge mode, 40 possible permutations of candidates are available, corresponding to combinations of five initial candidates and two possible types of triangles (see Figures 25(a) and 25(b) , where two possible first block predictors (2501 or 2511) and second block predictors (2502 or 2512) are for each type of triangle). Figure 27(b) shows the encoding of indices for triangle merge mode, i.e., for signaling these candidates. The first bit (i.e., the first bin) is CABAC-decoded in one context. If this first bit is equal to 0, the second bit (i.e., the second bin) is CABAC-bypass decoded. If this second bit is equal to 0, the index corresponds to the first candidate in the list, i.e., index 0 (Cand0). Otherwise (if the second bit is equal to 1), the index corresponds to the second candidate in the list, i.e., index 1 (Cand1). If the first bit is equal to 1, an Exponential-Golomb code is extracted from the bitstream, and the Exponential-Golomb code represents the index of the selected candidate in the list, i.e., selected from index 2 (Cand2) to index 39 (Cand39).
[0365] It will be appreciated that in another variant, a different number of initial candidates may be used instead, and the signaling of the triangle merge index adapted accordingly (e.g., at least one bit is CABAC coded using one context variable).
[0366] ATMVP of affine merge lists In a second variant of the seventeenth embodiment, the ATMVP is available as a candidate in the affine merge candidate list (i.e., in affine merge mode—also known as “sub-block merge” mode). Figure 28 shows a list of affine merge list derivations with this additional ATMVP candidate (2848). This figure is similar to Figure 24 (described above), but since this additional ATMVP candidate (2848) has been added to the list, a detailed description will not be repeated here. It will be understood that in another variant, a different number of initial candidates may be used instead, and the signaling of the triangle merge index will be adapted accordingly (e.g., at least one bit is CABAC coded using one context variable).
[0367] It will be appreciated that in another variant, an ATMVP candidate may be added to the list of candidates for another inter-prediction mode, and the signaling of its index is adapted accordingly (e.g., at least one bit is CABAC coded using one context variable).
[0368] While Figure 26 provides a complete overview of the signaling of all inter-prediction modes (i.e., merge mode, affine merge mode, MHII merge mode, triangle merge mode, and MMVD merge mode) according to another variant, it will be understood that only a subset of the inter-prediction modes may alternatively be used.
[0369] Eighteenth embodiment According to the 18th embodiment, one or both of the triangle merge mode or the MMVD merge mode are available for use in the encoding or decoding process, and one or both of these inter prediction modes share context variables (used in conjunction with CABAC encoding) with another inter prediction mode when signaling its index / flag.
[0370] In further variations of this or the following embodiments, it is understood that one or more of the inter prediction modes may use more than one context variable when signaling its index / flag (e.g., an affine merge mode may use four or five context variables depending on whether ATMVP candidates can also be included in the list for its affine merge index encoding / decoding process).
[0371] For example, before this embodiment or a variant of the following embodiment is implemented, the total number of context variables for signaling all bits of the indexes / flags for all inter-prediction modes may be 7: (regular) merge = 1 (as shown in FIG. 10(a)); affine merge = 4 (as shown in FIG. 10(b) but with one less candidate, e.g., no ATMVP candidate); triangle = MMVD = 1; and MHII (if available) = 0 (shared with regular merge). Then, by implementing the variant, the total number of context variables for signaling all bits of the indexes / flags for all inter-prediction modes may be reduced to 5: (regular) merge = 1 (as shown in FIG. 10(a)); affine merge = 4 (as shown in FIG. 10(b) but with one less candidate, e.g., no ATMVP candidate); and triangle = MMVD = MHII (if available) = 0 (shared with regular merge).
[0372] In another example, before this variant is implemented, the total number of context variables for signaling all bits of indexes / flags for all inter-prediction modes may be 4: (regular) merge = affine merge = triangle = MMVD = 1 (as shown in FIG. 10(a)); and MHII (if available) = 0 (shared with regular merge). Then, by implementing this variant, the total number of context variables for signaling all bits of indexes / flags for all inter-prediction modes is reduced to 2: (regular) merge = affine merge = 1 (as shown in FIG. 10(a)); and triangle = MMVD = MHII (if available) = 0 (shared with regular merge).
[0373] Note that for simplicity of the following description, we will discuss sharing or not sharing one context variable (e.g., only the first bit). This means that in the following description, we often see the simple case of using a context variable to signal only the first bit for each inter-prediction mode, and this context variable is either 1 (a separate / independent context variable is used) or 0 (this bit is bypass CABAC coded or shares the same context variable with another inter-prediction mode, so there is no separate / independent context variable). Different variations of this embodiment and the following embodiments are not limited thereto, and it is understood that other bits of the context variable, or indeed all bits, may be shared / non-shared / bypass CABAC coded in the same way.
[0374] In a first variant of the eighteenth embodiment, all inter prediction modes available for use in the encoding or decoding process share at least some CABAC context.
[0375] In this variant, the index coding and its related parameters (e.g., the number of (initial) candidates) for the inter prediction mode may be set to be the same or similar as long as possible / compatible. For example, to simplify signaling, the number of candidates for the affine merge mode and merge mode may be set to 5 and 6, respectively, the number of initial candidates for the MMVD merge mode may be set to 2, and the maximum number of candidates for the triangle merge mode may be 40. Also, the triangle merge index is not signaled using a unary max code like other inter prediction modes. For this triangle merge mode, only the first bit of the context variable (for the triangle merge index) can be shared with other inter prediction modes. The advantage of this variant is that the design of the encoder and decoder is simplified.
[0376] In a further variant, the CABAC content for all merged inter-prediction mode indexes is shared. This means that only one CABAC context variable is required for the first bit of every index. In yet another variant, if an index contains two or more bits to be CABAC coded, the coding of the additional bits (all CABAC coded bits apart from the first bit) is treated as a separate part (i.e., as if for a separate syntax element as far as the CABAC coding process is concerned), and if two or more indexes have two or more bits to be CABAC coded, one and the same context variable is shared for these CABAC coded "additional" bits. The advantage of this variant is that the amount of CABAC context is reduced. This reduces the storage requirements for the context state that needs to be stored at the encoder and decoder sides without significantly affecting the coding efficiency for most of the sequences processed by the video codec implementing this variant.
[0377] FIG. 29 shows another variant of the decoding process for inter prediction mode. This figure is similar to FIG. 26 but includes an implementation of this variant. In this figure, when the current CU is processed in MMVD merge mode, its MMVD merge index is decoded as the same index as the merge index in regular merge mode (i.e., the "merge index" (2919)). However, unlike regular merge mode, MMVD merge mode only allows two initial candidates to be selected, not six. Because there are only two possibilities, this "shared" index used in MMVD merge mode is essentially a flag. Because the same index is shared, the CABAC context variable is the same for this flag in MMVD merge mode and the first bit of the merge index in merge mode. Next, if it is determined that the current CU should be processed in MMVD merge mode (2925), the distance index (2911) and direction index (2912) are decoded. If it is determined that the current CU should be processed in affine merge mode (2914), its affine merge index is decoded as the same index as the merge index in regular merge mode (i.e., the "merge index" (2919)). However, unlike regular merge mode, the maximum number of candidates (i.e., the maximum number of indices) in affine merge mode is 5, not 6. If it is determined that the current CU should be processed in triangle merge mode (2918), the first bit is decoded as the shared index (2919), resulting in the same CABAC context variables being shared as in regular merge mode. If this CU is processed in triangle merge mode (2926), the remaining bits related to the triangle merge index are decoded (2923).
[0378] Thus, for example, when processing these indices / flags during the CABAC encoding process, the number of separate (independent) context variables used for the first bit of the index / flag for each inter prediction mode is: (regular)merge=1; MHII=Affine Merge=Triangle=MMVD=0 (shared with Regular Merge) is.
[0379] In a second variant, when one or both of the triangle merge mode and the MMVD merge mode are used (i.e., information about the motion information predictor selection of the current CU is processed / encoded / decoded in the associated inter-prediction mode), its / their index signaling shares context variables with the index signaling of the merge mode. In this variant, the CABAC context of the triangle merge index and / or the CABAC context of the MMVD merge index / flag share the same CABAC context of the merge index (in the (regular) merge mode). This means that only one CABAC state is required for at least these three modes.
[0380] In a further variation of the second variation, the CABAC context of the triangle merge index and / or the CABAC context of the MMVD merge index / flag share the same first CABAC context variable of the merge index, e.g., the same context variable of the first bit of the merge index.
[0381] So, for example, when processing these indices / flags during the CABAC encoding process, the number of separate (independent) context variables used for the first bit of the indices / flags is: (regular)merge=1; MHII (if available) = affine merge (if available) = 0 (shared with regular merge) or 1 depending on implementation; Triangle = MMVD = 0 (shared with regular merge) is.
[0382] In yet another variation of the second variation, if two or more context variables are used for triangle merge index CABAC encoding / decoding or two or more context variables are used for MMVD merge index CABAC encoding / decoding, they may all be shared with two or more CABAC context variables used for merge index CABAC encoding / decoding, or at least partially shared whenever compatible.
[0383] The advantage of this second variant is that it reduces the amount of context that needs to be stored, and consequently the amount of state that needs to be stored at the encoder and decoder side, without significantly affecting the coding efficiency of the majority of sequences processed by video codecs that implement them.
[0384] In a third variant, when one or both of the triangle merge mode and the MMVD merge mode are used (i.e., information regarding the motion information predictor selection of the current CU is processed / encoded / decoded in the associated inter prediction mode), its / their index signaling shares context variables with the index signaling of the affine merge mode. In this variant, the CABAC context of the triangle merge index and / or the CABAC context of the MMVD merge index / flag share the same CABAC context of the affine merge index (for the affine merge mode).
[0385] In a further variation of the third variation, the CABAC context of the triangle merge index and / or the CABAC context of the MMVD merge index / flag share the same first CABAC context variable of the affine merge index, e.g., the same context variable of the first bit of the affine merge index.
[0386] So, for example, when processing these indices / flags during the CABAC encoding process, the number of separate (independent) context variables used for the first bit of the indices / flags is: (Canonical) merge (if available) = 0 (shared with affine merge) or 1 depending on implementation; MHII(if available)=0(regular merge and share); affinemerge=1; Triangle = MMVD = 0 (shared with affine merge) is.
[0387] In yet another variation of the third variation, if two or more context variables are used for triangle merge index CABAC encoding / decoding or two or more context variables are used for MMVD merge index CABAC encoding / decoding, they may all be shared or at least partially shared if they are compatible with two or more CABAC context variables used for affine merge index CABAC encoding / decoding.
[0388] In a fourth variant, when the MMVD merge mode is used (i.e., the information regarding the motion information predictor selection of the current CU is processed / encoded / decoded in the MMVD merge mode), the index signaling shares context variables with the index signaling of the merge mode or the affine merge mode. In this variant, the CABAC context of the MMVD merge index / flag is the same CABAC context of the merge index or the same CABAC context of the affine merge index.
[0389] So, for example, when processing these indices / flags during the CABAC encoding process, the number of separate (independent) context variables used for the first bit of the indices / flags is: (regular)merge=1; MHII(if available)=0(regular merge and share); Affine merge (if available) = 0 (shared with regular merge) or 1 depending on implementation; MMVD=0 (regular merge and share) or (Canonical) merge (if available) = 0 (shared with affine merge) or 1 depending on implementation; MHII(if available)=0(regular merge and share); affinemerge=1; MMVD=0 (shared with affine merge) is.
[0390] In a fifth variant, when the triangle merge mode is used (i.e., the information regarding the motion information predictor selection of the current CU is processed / encoded / decoded in the triangle merge mode), the index signaling shares context variables with the index signaling for the merge mode or the affine merge mode. In this variant, the CABAC context of the triangle merge index is the same CABAC context of the merge index or the same CABAC context of the affine merge index.
[0391] So, for example, when processing these indices / flags during the CABAC encoding process, the number of separate (independent) context variables used for the first bit of the indices / flags is: (regular)merge=1; MHII(if available)=0(regular merge and share); Affine merge (if available) = 0 (shared with regular merge) or 1 depending on implementation; Triangles = 0 (regular merge and share) or (Canonical) merge (if available) = 0 (shared with affine merge) or 1 depending on implementation; MHII(if available)=0(regular merge and share); affinemerge=1; Triangle = 0 (shared with affine merge) is.
[0392] In a sixth variant, when the triangle merge mode is used (i.e., the information about the motion information predictor selection of the current CU is processed / encoded / decoded in the triangle merge mode), its index signaling shares a context variable with the index signaling of the MMVD merge mode. In this variant, the CABAC context of the triangle merge index is the same CABAC context of the MMVD merge index. Thus, for example, during the CABAC encoding process, when processing these indexes / flags, the number of separate (independent) context variables used for the first bit of the index / flag is: MMVD=1; triangle=0(shared with MMVD); (Canonical) merge or MHII or affine merge = depending on implementation and availability is.
[0393] In a seventh variant, when one or both of the triangle merge mode and the MMVD merge mode are used (i.e., information regarding the motion information predictor selection of the current CU is processed / encoded / decoded in the associated inter-prediction mode), the index signaling for that / those modes shares a context variable with the index signaling for the inter-prediction mode that can include the ATMVP predictor candidate in its list of candidates, i.e., the inter-prediction mode can have the ATMVP predictor candidate as one of the available candidates. In this variant, the CABAC context of the triangle merge index and / or the MMVD merge index / flag share the same CABAC context of the index of the inter-prediction mode that can use the ATMVP predictor.
[0394] In a further variation, the CABAC context variables for the triangle merge index and / or MMVD merge index / flag share the same first CABAC context variable for the merge index in merge mode with the includable ATMVP candidate, or share the affine merge index in affine merge mode with the includable ATMVP candidate.
[0395] In yet another variant, if two or more context variables are used for triangle merge index CABAC encoding / decoding or two or more context variables are used for MMVD merge index CABAC encoding / decoding, they can all be shared with two or more CABAC context variables used for affine merge indices for affine merge mode or merge indices for merge mode with includable ATMVP candidates, or shared wherever at least partially compatible.
[0396] The advantage of these variations is improved coding efficiency, since ATMVP (predictor) candidates are the predictors that benefit most from CABAC adaptation when compared to other types of predictors.
[0397] Nineteenth embodiment According to the 19th embodiment, one or both of the triangle merge mode or the MMVD merge mode are available for use in the encoding or decoding process, and the indexes / flags for one or both of these inter prediction modes are CABAC bypass coded when signaling the indexes / flags.
[0398] In a first variant of the nineteenth embodiment, all inter prediction modes available for use in the encoding or decoding process have their indices / flags CABAC bypass coded / decoded to signal the indices / flags. In this variant, all indices of all inter prediction modes are coded (e.g., by the bypass coding engine 1705 of FIG. 17) without using CABAC context variables. This means that all bits of the merge index (2619), affine merge index (2615), MMVD merge index (2610), and triangle merge index (2623) in FIG. 26 are CABAC bypass coded. Figures 30(a) to 30(c) show coding of indices / flags according to this embodiment. Figure 30(a) shows MMVD merge index coding of initial MMVD merge candidates. Figure 30(b) shows triangle merge index coding. Figure 30(c) shows affine merge index coding, which can also be easily used for merge index coding.
[0399] So, for example, when processing these indices / flags during the CABAC encoding process, the number of distinct (independent) context variables used for the indices / flags is: (Regular) Merge (if available) = MHII (if available) = Affine Merge (if available) = Triangle (if available) = MMVD (if available) = 0 (all bypass code). The advantage of this variant is a reduction in the amount of context that needs to be stored, and therefore a reduction in the amount of state that needs to be stored at the encoder and decoder side, with only a small impact on the coding efficiency of the majority of sequences processed by a video codec that implements the variant. Note, however, that this variant may be lossy when used to code screen content. This variant represents another compromise between coding efficiency and complexity when compared to other variants / embodiments. The impact on coding efficiency is often small. In fact, with a large number of available inter-prediction modes, the average amount of data required to signal an index for each inter-prediction mode is smaller than the average amount of data required to signal a merge index when only merge mode is available (this comparison is for the same sequence and the same coding efficiency compromise). This means that the efficiency of CABAC encoding / decoding from context-based bin probability adaptation may be inefficient.
[0400] In a second variant, when one or both of the triangle merge mode and the MMVD merge mode are used (i.e., information regarding motion information predictor selection for the current CU is processed / encoded / decoded in the associated inter prediction mode), the / their indexes / flags are signaled by CABAC bypass encoding / decoding the indexes / flags. In this variant, the MMVD merge index and / or the triangle merge index are CABAC bypass coded. Depending on the implementation, i.e., when the merge mode and the affine merge mode are available, the merge index and the affine merge index have their own context. In yet another variant, the context of the affine merge index and the merge index are shared. So, for example, when processing these indices / flags during the CABAC encoding process, the number of separate (independent) context variables used for the first bit of the indices / flags is: (canonical)merge = affinemerge = 0 or 1 depending on implementation; MHII(if available)=0(shared with regular merge); Triangle = MMVD = 0 (bypass coding) is.
[0401] The advantage of these variants is that they improve coding efficiency compared to previous variants by providing yet another compromise between CABAC context reduction and coding efficiency. In practice, triangle merge mode is not often selected. As a result, when its context is removed, i.e., when triangle merge mode uses CABAC bypass coding, the impact on coding efficiency is small. While MMVD merge mode tends to be selected more frequently than triangle merge mode, the probability of selecting the first and second candidates for MMVD merge mode tends to be equal to or greater than other inter-prediction modes, such as merge mode or affine merge mode. Therefore, the benefit gained from using the CABAC coding context for MMVD merge mode is not as great. Another advantage of these variants is a small impact on coding efficiency for screen content sequences, since merge mode is the most influential inter-prediction mode for screen content.
[0402] In a third variant, when the merge mode, triangle merge mode, or MMVD merge mode is used (i.e., information about the motion information predictor selection of the current CU is processed / encoded / decoded in the associated inter-prediction mode), its / their indexes / flags are signaled by CABAC bypass encoding / decoding the indexes / flags. In this variant, the MMVD merge index, triangle merge index, and merge index are CABAC bypass encoded. Therefore, for example, during the CABAC encoding process, when processing these indexes / flags, the number of separate (independent) context variables used for the first bit of the indexes / flags is: (regular) merge = triangle (if available) = MMVD (if available) = 0 (bypass coding); MHII (if available) = 0 (same as regular merge); and Affine Merge = 1 is.
[0403] This variant offers an alternative compromise compared to the other variants, for example, this variant provides a larger coding efficiency reduction for screen content sequences than the previous variants.
[0404] In a fourth variant, when the affine merge mode, triangle merge mode, or MMVD merge mode is used (i.e., information about the motion information predictor selection of the current CU is processed / encoded / decoded in the associated inter prediction mode), its / their indexes / flags are signaled by CABAC bypass encoding / decoding the indexes / flags. In this variant, the affine merge index, MMVD merge index, and triangle merge index are CABAC bypass encoded, and the merge indexes are coded with one or more CABAC contexts. Thus, for example, during the CABAC coding process, when processing these indexes / flags, the number of separate (independent) context variables used for the first bit of the indexes / flags is: (regular)merge=1; affinemerge = triangle(if available) = MMVD(if available) = 0(bypass coding); MHII (if available) = 0 (regular merge and share). The advantage of this variant over the previous variant is increased coding efficiency for screen content sequences.
[0405] In a fifth variant, an index / flag is CABAC bypass coded / decoded to signal an inter-prediction mode available for use in the encoding or decoding process, except when the inter-prediction mode can include an ATMVP predictor candidate in the list of candidates, i.e., the inter-prediction mode can have an ATMVP predictor candidate as one of the available candidates. In this variant, all indices of all inter-prediction modes are CABAC bypass coded, except when the inter-prediction mode can have an ATMVP predictor candidate. Thus, for example, when processing these indexes / flags during the CABAC coding process, the number of separate (independent) context variables used for the first bit of the index / flag is: (regular)merge with includable ATMVP candidates = 1; AffineMerge(if available) = Triangle(if available) = MMVD(if available) = 0(Bypass coding); MHII (if available) = 1 (shared with regular merge) or 0 depending on implementation or affinemergewithencompassableATMVPcandidates=1; (regular) merge(if available) = triangle(if available) = MMVD(if available) = 0(bypass coding); MHII (if available) = 0 (same as regular MERGE) is.
[0406] This variant also introduces another complexity / coding efficiency compromise for most natural sequences. Note, however, that for screen content sequences, it may be preferable to have ATMVP predictor candidates in the regular merge candidate list.
[0407] In a sixth variant, when an inter prediction mode available for use in the encoding or decoding process is not a skip mode (e.g., not one of a regular merge skip mode, an affine merge skip mode, a triangle merge skip mode, or an MMVD merge skip mode), the index / flag is CABAC bypass coded / decoded to signal the index / flag. In this variant, all indices are not processed in skip mode, i.e., CABAC bypass coded for any CU that is not skipped. The index for the skipped CU (i.e., a CU processed in skip mode) may be processed using any one of the CABAC coding techniques described in connection with the previous embodiments / variations (e.g., only the first bit or two or more bits have a context variable, which may or may not be shared).
[0408] Figure 31 is a flowchart of a decoding process for inter prediction mode illustrating this modification. The process in Figure 31 is similar to that in Figure 29, except that it has an additional "skip mode" decision / check step (3127), after which the index / flag ("Merge_idx") is decoded using the context of either CABAC decoding (3119) or CABAC bypass decoding (3128). According to yet another modification, it is understood that the result of the decision / check made in the previous step is that the CU is skip (2902 / 3102) or MMVD skip (2908 / 3108), and the CU uses the skip (2916 / 3116) in Figure 29 or Figure 31 to make a "skip mode" decision / check instead of the additional "skip mode" decision / check step (3127).
[0409] This variation has a lower impact on coding efficiency because skip mode is generally selected more frequently than non-skip mode (i.e., non-skip inter prediction modes such as regular merge mode, MHII merge mode, affine merge mode, triangle merge mode, or MMVD merge mode), and because the first candidate selection is more likely for skip mode than for non-skip mode. Because skip mode is designed for more predictable motion, their indices must also be more predictable. Therefore, the probability of utilizing CABAC encoding / decoding is more likely to be useful for skip mode. However, non-skip mode is more likely to be used when motion is less predictable, so that a more random selection from predictor candidates is more likely. Therefore, CABAC encoding / decoding is less likely to be efficient in non-skip mode.
[0410] Twentieth embodiment According to a twentieth embodiment, data is provided in a bitstream, the data being for determining whether indexes / flags for one or more inter-prediction modes should be signaled by using CABAC bypass encoding / decoding, CABAC encoding / decoding with separate context variables, or CABAC encoding / decoding with one or more shared context variables. For example, such data may be a flag for enabling or disabling the use of one or more independent contexts for index encoding / decoding of inter-prediction modes. Such data may be used to control the use or non-use of context sharing in CABAC encoding / decoding or CABAC bypass encoding / decoding.
[0411] In a variation of the twentieth embodiment, the CABAC context sharing between two or more indices of two or more inter-prediction modes relies on data transmitted in the bitstream, for example, at a level higher than the CU level (e.g., at the level of an image portion larger than the smallest CU, such as the sequence, frame, slice, tile, or CTU level). For example, this data may indicate that for any CU in a particular image portion, the CABAC context of a merge index of a merge mode is shared (or not shared) with one or more other CABAC contexts of another inter-prediction mode.
[0412] In another variation, one or more indexes are CABAC bypass coded in response to data transmitted in the bitstream, for example, at a level higher than the CU level (e.g., at the slice level). For example, this data may indicate that for any CU in a particular image portion, an index of a particular inter-prediction mode should be CABAC bypass coded.
[0413] In one variant, to further improve coding efficiency, at the encoder side, the value of this data for indicating the context sharing of one or more indexes of one or more inter-prediction modes, or CABAC bypass encoding / decoding, can be selected based on how frequently one or more inter-prediction modes are used in previously coded frames. An alternative may be to select the value of this data based on the type of sequence being processed or the type of application in which the variant is implemented.
[0414] The advantage of this embodiment is a controlled increase in coding efficiency compared to the previous embodiment / variant.
[0415] Implementation of embodiments of the present invention One or more of the aforementioned embodiments may be implemented by the processor 311 of the processing device 300 of FIG. 3 , or a corresponding functional module / unit of the decoder 60 of FIG. 5 , of the CABAC coder of FIG. 17 , of the encoder 400 of FIG. 4 , or its corresponding CABAC decoder, which performs the method steps of one or more of the aforementioned embodiments.
[0416] FIG. 19 is a schematic block diagram of a computing device 2000 for implementing one or more embodiments of the present invention. The computing device 2000 may be a device such as a microcomputer, a workstation, or a light portable device. The computing device 2000 includes: a central processing unit (CPU) 2001, such as a microprocessor; random access memory (RAM) 2002 for storing executable code for methods of embodiments of the present invention and registers for recording variables and parameters necessary to implement methods for encoding or decoding at least a portion of an image according to embodiments of the present invention, the capacity of which may be expanded, for example, by an optional RAM connected to an expansion port; read-only memory (ROM) 2003 for storing computer programs for implementing embodiments of the present invention; and a communication bus connected to a network interface (NET) 2004, which typically connects to a communication network over which digital data to be processed is transmitted or received. The network interface (NET) 2004 may be a single network interface or may consist of a set of different network interfaces (e.g., wired and wireless interfaces, or different types of wired or wireless interfaces). Data packets are written to the network interface for transmission or read from the network interface for reception under the control of software applications running on the CPU 2001. A user interface (UI) 2005 may be used to receive input from a user and display information to the user. A hard disk (HD) 2006 may be provided as mass storage. An input / output module (IO) 2007 may be used to send and receive data to external devices such as video sources and displays. The executable code may be stored in either the ROM 2003, the HD 2006, or a removable digital medium such as a disk.According to a variant, the executable code of the program can be received by the communications network via NET 2004, for storage in one of the storage means of the communications device 2000, such as HD 2006, before being executed. The CPU 2001 is adapted to control and direct the execution of instructions or parts of a program or software code of a program according to an embodiment of the invention, the instructions of which are stored in one of the aforementioned storage means. After power-on, the CPU 2001 can execute instructions for a software application from the main RAM memory 2002, for example, after these instructions have been loaded from the program ROM 2003 or the HD 2006. Such a software application, when executed by the CPU 2001, causes the steps of the method according to the invention to be carried out.
[0417] It will also be appreciated that according to other embodiments of the present invention, a decoder according to the aforementioned embodiments is provided in a user terminal such as a computer, a mobile phone (cell phone), a tablet, or any other type of device (e.g., a display device) that can provide / display content to a user. According to yet another embodiment, an encoder according to the aforementioned embodiments is provided in an image capture device that also comprises a camera, video camera, or network camera (e.g., a closed-circuit television or video surveillance camera) that captures and provides content for the encoder to encode. Two such examples are provided below with reference to Figures 20 and 21.
[0418] FIG. 20 is a diagram illustrating a network camera system 2100 including a network camera 2102 and a client device 2104 .
[0419] The network camera 2102 includes an imaging unit 2106 , an encoding unit 2108 , a communication unit 2110 , and a control unit 2112 .
[0420] The network camera 2102 and the client device 2104 are connected to each other via the network 200 so that they can communicate with each other.
[0421] The image capture unit 2106 includes a lens and an image sensor (e.g., a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS)) to capture an image of an object and generate image data based on the image. The image may be a still image or a video image. The image capture unit may also include zoom and / or pan means adapted to zoom or pan (optically or digitally).
[0422] The encoder 2108 encodes the image data using the encoding method described in one or more of the above-described embodiments. The encoder 2108 uses at least one of the encoding methods described in the above-described embodiments. In other examples, the encoder 2108 can use a combination of the encoding methods described in the above-described embodiments.
[0423] The communication unit 2110 of the network camera 2102 transmits the encoded image data encoded by the encoding unit 2108 to the client device 2104. The communication unit 2110 also receives commands from the client device 2104. The commands include commands for setting parameters for encoding by the encoding unit 2108.
[0424] The control unit 2112 controls the other units in the network camera 2102 according to the commands received by the communication unit 2110 .
[0425] The client device 2104 includes a communication unit 2114, a decoding unit 2116, and a control unit 2118. The communication unit 2114 of the client device 2104 transmits commands to the network camera 2102. The communication unit 2114 of the client device 2104 also receives encoded image data from the network camera 2102.
[0426] The decoder 2116 decodes the encoded image data using the decoding method described in one or more of the above embodiments. In other examples, the decoder 2116 can use a combination of the decoding methods described in the above embodiments.
[0427] The control unit 2118 of the client device 2104 controls other units in the client device 2104 in accordance with user operations and commands received by the communication unit 2114. The control unit 2118 of the client device 2104 controls the display device 2120 to display the image decoded by the decoding unit 2116. The control unit 2118 of the client device 2104 also controls the display device 2120 to display a GUI (Graphical User Interface), and specifies parameter values for the network camera 2102, including parameters for encoding by the encoding unit 2108.
[0428] Furthermore, the control unit 2118 of the client device 2104 controls other units within the client device 2104 in response to a user operation input to the GUI displayed by the display device 2120. The control unit 2118 of the client device 2104 controls the communication unit 2114 of the client device 2104 in response to a user operation input to the GUI displayed by the display device 2120 so as to transmit to the network camera 2102 a command specifying parameter values of the network camera 2102.
[0429] The network camera system 2100 can determine whether the camera 2102 utilizes zoom or pan while recording video, and such information can be used when encoding the video stream, as zooming or panning during recording can benefit from the use of an affine mode, which is well suited to encoding complex movements such as zooming, rotation, and / or stretching (which can be a side effect of panning, especially if the lens is a "fisheye" lens).
[0430] FIG. 21 is a diagram illustrating a smartphone 2200.
[0431] The smartphone 2200 includes a communication unit 2202 , a decoding / encoding unit 2204 , a control unit 2206 , and a display unit 2208 .
[0432] The communication unit 2202 receives the encoded image data via the network 200 .
[0433] The decoding / encoding unit 2204 decodes the encoded image data received by the communication unit 2202. The decoding / encoding unit 2204 decodes the encoded image data using the decoding method described in one or more of the above-mentioned embodiments. The decoding / encoding unit 2204 can use at least one of the decoding methods described in the above-mentioned embodiments. In other examples, the decoding / encoding unit 2204 can use a combination of the decoding or encoding methods described in the above-mentioned embodiments.
[0434] The control unit 2206 controls other units in the smartphone 2200 in response to user operations or commands received by the communication unit 2202 or via the input unit. For example, the control unit 2206 controls the display device 2208 to display the image decoded by the decoding unit 2204.
[0435] The smartphone may further include an image recording device 2210 (e.g., a digital camera and associated circuitry) for recording images or video. Such recorded images or video may be encoded by the decoding / encoding unit 2204 under the direction of the control unit 2206. The smartphone may further include a sensor 2212 configured to sense the orientation of the mobile device. Such a sensor may include an accelerometer, gyroscope, compass, global positioning (GPS) unit, or similar position sensor. Such a sensor 2212 may determine whether the smartphone is changing orientation, and such information may be used when encoding the video stream as orientation changes during capture, benefiting from the use of affine modes, which are well suited to encoding complex movements such as rotations.
[0436] Substitutions and Modifications It will be appreciated that the object of the present invention is to ensure that affine modes are utilized in the most efficient way, and the particular example above relates to signaling the use of affine modes depending on the likelihood that the affine modes will be perceived as useful. A further example of this may be applied to an encoder where it is known that complex motion is being coded, for which affine transformations may be particularly efficient. Examples of such cases are: a) Camera zoom in / out b) Portable cameras (e.g., mobile phones) that change orientation during recording (i.e., rotational movement) c) Panning a "fisheye" lens camera (e.g. stretching / distorting parts of the image) Includes.
[0437] Thus, complex motion indications can be increased during the recording process so that affine mode is more likely to be used for slices, frame sequences, or indeed the entire video stream.
[0438] In a further example, affine mode may be more likely to be used depending on the features or functionality of the device used to record the video. For example, a mobile device is more likely to change orientation than (for example) a fixed security camera, so affine mode may be more suitable for encoding video from the former. Examples of features or functionality include the presence / use of zoom means, the presence / use of a position sensor, the presence / use of panning means, whether the device is handheld, or user selection on the device.
[0439] While the present invention has been described with reference to embodiments, it should be understood that the present invention is not limited to the disclosed embodiments. Those skilled in the art will recognize that various changes and modifications can be made without departing from the scope of the present invention, as defined by the appended claims. All features disclosed in this specification (including any accompanying claims, abstract, and drawings), and / or all steps of any method or process so disclosed, may be combined in any combination, except combinations in which at least some of such features and / or steps are mutually exclusive. Each feature disclosed in this specification (including any accompanying claims, abstract, and drawings), unless otherwise specified, may be replaced by an alternative feature serving the same, equivalent, or similar purpose. Thus, unless otherwise specified, each feature disclosed is merely an example of a generic series of equivalent or similar features.
[0440] It will also be understood that any result of the above-mentioned comparison, determination, evaluation, selection, execution, performance, or consideration, e.g., a selection made during an encoding or filtering process, may be indicated in or determinable / inferable from data in the bitstream, e.g., a flag or data indicating the result, and such indicated or determined / inferred result may be used in processing, e.g., during a decoding process, in lieu of actually performing the comparison, determination, evaluation, selection, execution, performance, or consideration.
[0441] In the claims, the word "comprise" does not exclude other elements or steps, and the indefinite articles "a" or "an" do not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.
[0442] Reference signs appearing in the claims are by way of illustration only and shall have no limiting effect on the scope of the claims.
[0443] In the foregoing embodiments, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted over as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit.
[0444] Computer-readable media may include computer-readable storage media, which correspond to tangible media such as data storage media, or communication media, including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communications protocol. In this manner, computer-readable media generally can correspond to (1) tangible computer-readable storage media that are non-transitory, or (2) communication media such as signals or carrier waves. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include computer-readable media.
[0445] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included within the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead cover non-transitory, tangible storage media. Disk and disk as used herein include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where discs typically reproduce data magnetically and discs reproduce data optically with a laser. Combinations of the above should also be included within the scope of computer-readable media.
[0446] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate / logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the term "processor," as used herein, may refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or may be incorporated into a combined codec. Additionally, the techniques may be implemented entirely in one or more circuit or logic elements.
Claims
1. 1. A method for encoding information relating to a motion information predictor, comprising: selecting one of a plurality of motion information predictor candidates; encoding one of a plurality of indexes including a first index and a second index for identifying the selected motion information predictor candidate using Context-based Adaptive Binary Arithmetic Coding (CABAC) coding; Including, the first index is used for a first merge mode in which a block predictor is obtainable from a first block predictor associated with a first region in a block and a second block predictor associated with a second region in the block that is different from the first region; the second index is used for a second merge mode of inter prediction modes that is different from the first merge mode; the CABAC encoding of the first bit of the first index for the first merge mode uses the same context variables as the CABAC encoding of the first bit of the second index for the second merge mode; all bits of the first index except the first bit of the first index are bypass coded, and all bits of the second index except the first bit of the second index are bypass coded; a block predictor can be obtained based on the second merging mode and using an average of an intra-block predictor and an inter-block predictor; Spatial merge candidates are available for the second merge mode. A method characterized by:
2. The method of claim 1 , wherein each of the first and second merge modes is a merge mode that is independent of a merge mode that uses affine motion information.
3. The method of claim 1 , wherein the first region and the second region each have a shape other than rectangular.
4. the first region in the block does not include the bottom left vertex of the block but includes the top right vertex of the block; 4. The method of claim 3, wherein the second region of the block excludes the upper right vertex of the block and includes the lower left vertex of the block.
5. 2. The method of claim 1, wherein a weighted average is applied to the region between the first region and the second region within the block.
6. 1. A method for decoding information related to a motion information predictor, comprising: decoding one of the plurality of indexes, including the first index and the second index, using Context-based Adaptive Binary Arithmetic Coding (CABAC) decoding to identify one of the plurality of motion information predictor candidates; selecting one of the plurality of motion information predictor candidates using the decoded index; Including, the first index is used for a first merge mode in which a block predictor is obtainable from a first block predictor associated with a first region in a block and a second block predictor associated with a second region in the block that is different from the first region; the second index is used for a second merge mode of inter prediction modes that is different from the first merge mode; CABAC decoding of the first bit of the first index for the first merge mode uses the same context variables as CABAC decoding of the first bit of the second index for the second merge mode of an inter prediction mode; all bits of the first index except the first bit of the first index are bypass decoded, and all bits of the second index except the first bit of the second index are bypass decoded; a block predictor can be obtained based on the second merging mode and using an average of an intra-block predictor and an inter-block predictor; Spatial merge candidates are available for the second merge mode. A method characterized by:
7. The method of claim 6 , wherein each of the first and second merge modes is a merge mode that is independent of a merge mode that uses affine motion information.
8. 7. The method of claim 6, wherein the first region and the second region each have a shape other than rectangular.
9. the first region in the block does not include the bottom left vertex of the block but includes the top right vertex of the block; 9. The method of claim 8, wherein the second region of the block excludes the upper right vertex of the block and includes the lower left vertex of the block.
10. 7. The method of claim 6, wherein a weighted average is applied to the area between the first and second areas within the block.
11. 1. An apparatus for encoding information related to a motion information predictor, comprising: means for selecting one of a plurality of motion information predictor candidates; means for encoding one of a plurality of indexes, including a first index and a second index, for identifying the selected motion information predictor candidate using Context-based Adaptive Binary Arithmetic Coding (CABAC) coding; Including, the first index is used for a first merge mode in which a block predictor is obtainable from a first block predictor associated with a first region in a block and a second block predictor associated with a second region in the block that is different from the first region; the second index is used for a second merge mode of inter prediction modes that is different from the first merge mode; the CABAC encoding of the first bit of the first index for the first merge mode uses the same context variables as the CABAC encoding of the first bit of the second index for the second merge mode; all bits of the first index except the first bit of the first index are bypass coded, and all bits of the second index except the first bit of the second index are bypass coded; a block predictor can be obtained based on the second merging mode and using an average of an intra-block predictor and an inter-block predictor; Spatial merge candidates are available for the second merge mode. An apparatus characterized in that
12. 1. An apparatus for decoding information related to a motion information predictor, comprising: means for decoding one of a plurality of indexes, including a first index and a second index, using Context-based Adaptive Binary Arithmetic Coding (CABAC) decoding to identify one of a plurality of motion information predictor candidates; means for selecting one of the plurality of motion information predictor candidates using the decoded index; Including, the first index is used for a first merge mode in which a block predictor is obtainable from a first block predictor associated with a first region in a block and a second block predictor associated with a second region in the block that is different from the first region; the second index is used for a second merge mode of inter prediction modes that is different from the first merge mode; CABAC decoding of the first bit of the first index for the first merge mode uses the same context variables as CABAC decoding of the first bit of the second index for the second merge mode of an inter prediction mode; all bits of the first index except the first bit of the first index are bypass decoded, and all bits of the second index except the first bit of the second index are bypass decoded; a block predictor can be obtained based on the second merging mode and using an average of an intra-block predictor and an inter-block predictor; Spatial merge candidates are available for the second merge mode. An apparatus characterized in that
13. A computer program product for causing a computer to carry out the method of claim 1.
14. A computer program product for causing a computer to carry out the method according to claim 6.
Citation Information
Patent Citations
Methods and apparatus for context-adaptive binary arithmetic coding of syntactic elements
JP2014531819A
Split Based Motion Vector Operation Reduction
US20180352223A1