Video coding and decoding
Patent Information
- Application Number
- JP2024102219
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-12-20
- Filing Date
- 2024-06-25
- Publication Date
- 2025-06-11
- Estimated Expiration
- 2039-11-19
AI Technical Summary
The existing video coding standards, such as HEVC, face challenges in efficiently encoding and decoding motion vectors due to the complexity and inefficiency associated with affine motion modes and Alternative Temporal Motion Vector Prediction (ATMVP), particularly in high-definition and ultra-high-definition video applications.
Implementing methods to bypass Context-Based Adaptive Binary Arithmetic Coding (CABAC) for encoding and decoding motion vector predictor indices, including strategies where certain bits of the motion vector predictor index are bypass CABAC encoded or decoded, and using shared contexts for multiple bits, or context variables dependent on neighboring blocks or block parameters.
Reduces encoding and decoding complexity while maintaining efficiency in encoding motion vectors, especially for affine motion modes and ATMVP, thereby improving compression performance in high-definition and ultra-high-definition video coding.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to video encoding and decoding. [Background technology]
[0002] Recently, the Joint Video Experts Team (JVET), a joint team formed by MPEG and the VCEG of ITU-T Study Group 16, has begun work on a new video coding standard called Versatile Video Coding (VVC). The goal of VVC is to provide significant improvements in compression performance over the existing HEVC standard (i.e., typically twice as much as before) and to be completed in 2020. Primary target applications and services include, but are not limited to, 360-degree and high dynamic range (HDR) video. In total, JVET evaluated responses from 32 parties using formal subjective testing conducted by an independent testing laboratory. Several proposals demonstrated compression efficiency gains of typically 40% or more compared to using HEVC. They showed particular effectiveness for ultra-high definition (UHD) video test material. Compression efficiency gains are therefore expected to far exceed the 50% targeted for the final standard.
[0003] The JVET Search Model (JEM) uses all the HEVC tools. An additional tool that is not present in HEVC is the use of "affine motion mode" when applying motion compensation. Motion compensation in HEVC is limited to translation, but in practice there are many types of motion, e.g. zoom in / out, rotation, perspective motion, and other irregular motions. When utilizing affine motion mode, more complex transformations are applied to blocks, attempting to more accurately predict the formation of such motion. It is therefore desirable to be able to use affine motion mode with reduced complexity while still achieving good coding efficiency.
[0004] Another tool that does not exist in HEVC is the use of Alternative Temporal Motion Vector Prediction (ATMVP). Alternative Temporal Motion Vector Prediction (ATMVP) is a specific motion compensation. Instead of considering only one motion information for the current block from the temporal reference frame, each motion information of each co-located block is considered. This temporal motion vector prediction thus gives a segmentation of the current block with the associated motion information of each sub-block. In the current VTM (VVC Test Model) reference software, ATMVP is signaled as a merge candidate inserted in the list of merge candidates. When ATMVP is enabled at the SPS level, the maximum number of merge candidates is increased by one. Thus, 6 candidates are considered instead of 5 from when this mode is disabled.
[0005] These and other tools described below raise concerns regarding the coding efficiency and complexity of encoding an index (e.g., a merge index) or flag used to signal which candidate was selected from among a list of candidates (e.g., from a list of merge candidates for use with merge mode encoding). Summary of the Invention
[0006] Therefore, a solution to at least one of the aforementioned problems is desirable.
[0007] According to a first aspect of the present invention, there is provided a method for encoding a motion vector predictor index, comprising the steps of: generating a list of motion vector predictor candidates including ATMVP candidates; selecting one of the motion vector predictor candidates in the list; Generate a motion vector predictor index (merge index) for the selected motion vector predictor candidate using CABAC coding, and one or more bits of the motion vector predictor index are bypass CABAC coded. A method is provided, comprising:
[0008] In one embodiment, all bits except the first bit of the motion vector predictor index are bypass CABAC coded.
[0009] According to a second aspect of the present invention, there is provided a method for decoding a motion vector predictor index, comprising the steps of: generating a list of motion vector predictor candidates including ATMVP candidates; Decode a motion vector predictor index using CABAC decoding, where one or more bits of the motion vector predictor index are bypass CABAC decoded; Using the decoded motion vector predictor index to identify one of the motion vector predictor candidates in the list. A method is provided, comprising:
[0010] In one embodiment, all bits except the first bit of the motion vector predictor index are bypass CABAC decoded.
[0011] According to a third aspect of the present invention, there is provided an apparatus for encoding a motion vector predictor index, comprising: means for generating a list of motion vector predictor candidates, the list including the ATMVP candidates; means for selecting one of the motion vector predictor candidates in the list; means for generating a motion vector predictor index (merge index) for a selected motion vector predictor candidate using CABAC coding, where one or more bits of the motion vector predictor index are bypass CABAC coded; An apparatus is provided comprising:
[0012] According to a fourth aspect of the present invention, there is provided an apparatus for decoding a motion vector predictor index, comprising: means for generating a list of motion vector predictor candidates, the list including the ATMVP candidates; means for decoding a motion vector predictor index using CABAC decoding, where one or more bits of the motion vector predictor index are bypass CABAC decoded; means for identifying one of the motion vector predictor candidates in the list using the decoded motion vector predictor index; An apparatus is provided comprising:
[0013] According to a fifth aspect of the present invention, there is provided a method for encoding a motion vector predictor index, comprising the steps of: generating a list of motion vector predictor candidates; selecting one of the motion vector predictor candidates in the list; Using CABAC coding, a motion vector predictor index is generated for a selected motion vector predictor candidate, and two or more bits of the motion vector predictor index share the same context. A method is provided, comprising:
[0014] In one embodiment, all bits of the motion vector predictor index share the same context.
[0015] According to a sixth aspect of the present invention, there is provided a method for decoding a motion vector predictor index, comprising the steps of: generating a list of motion vector predictor candidates; Using CABAC decoding, decode the motion vector predictor index, and two or more bits of the motion vector predictor index share the same context; Using the decoded motion vector predictor index to identify one of the motion vector predictor candidates in the list. A method is provided, comprising:
[0016] In one embodiment, all bits of the motion vector predictor index share the same context.
[0017] According to a seventh aspect of the present invention, there is provided an apparatus for encoding a motion vector predictor index, comprising: means for generating a list of motion vector predictor candidates; means for selecting one of the motion vector predictor candidates in the list; means for generating a motion vector predictor index for a selected motion vector predictor candidate using CABAC coding, wherein two or more bits of the motion vector predictor index share the same context; An apparatus is provided comprising:
[0018] According to an eighth aspect of the present invention, there is provided an apparatus for decoding a motion vector predictor index, comprising: means for generating a list of motion vector predictor candidates; means for decoding a motion vector predictor index using CABAC decoding, where two or more bits of the motion vector predictor index share the same context; means for identifying one of the motion vector predictor candidates in the list using the decoded motion vector predictor index; An apparatus is provided comprising:
[0019] According to a ninth aspect of the present invention, there is provided a method for encoding a motion vector predictor index, comprising the steps of: generating a list of motion vector predictor candidates; selecting one of the motion vector predictor candidates in the list; A motion vector predictor index of a selected motion vector predictor candidate is generated using CABAC coding, and a context variable of at least one bit of the motion vector predictor index of a current block depends on the motion vector predictor index of at least one block adjacent to the current block. A method is provided, comprising:
[0020] In one embodiment, a context variable for at least one bit of a motion vector predictor index depends on the motion vector predictor indexes of each of the at least two neighboring blocks.
[0021] In another embodiment, a context variable of at least one bit of the motion vector predictor index depends on the motion vector predictor index of a left neighboring block to the left of the current block and the motion vector predictor index of an upper neighboring block above the current block.
[0022] In another embodiment, the left adjacent block is A2 and the above adjacent block is B3.
[0023] In another embodiment, the left adjacent block is A1 and the above adjacent block is B1.
[0024] In another embodiment, the context variable has three different possible values.
[0025] Another embodiment includes comparing a motion vector predictor index of at least one neighboring block with an index value of a motion vector predictor index of a current block, and setting said context variable according to a comparison result.
[0026] Another embodiment includes comparing a motion vector predictor index of at least one adjacent block with a parameter representing a bit position of the or one of the bits in the motion vector predictor index of the current block and setting the context variable depending on the comparison result.
[0027] Yet another embodiment includes performing a first comparison, comparing a motion vector predictor index of a first adjacent block with a parameter representing a bit position of the or one of the bits in the motion vector predictor index of the current block, performing a second comparison, comparing a motion vector predictor index of a second adjacent block with the parameter, and setting the context variable depending on results of the first and second comparisons.
[0028] According to a tenth aspect of the present invention, there is provided a method for decoding a motion vector predictor index, comprising the steps of: generating a list of motion vector predictor candidates; Decode a motion vector predictor index using CABAC decoding, where a context variable of at least one bit of the motion vector predictor index of the current block depends on the motion vector predictor index of at least one block adjacent to the current block; Using the decoded motion vector predictor index to identify one of the motion vector predictor candidates in the list. A method is provided, comprising:
[0029] In one embodiment, a context variable for at least one bit of a motion vector predictor index depends on the motion vector predictor indexes of each of the at least two neighboring blocks.
[0030] In another embodiment, a context variable of at least one bit of the motion vector predictor index depends on the motion vector predictor index of a left neighboring block to the left of the current block and the motion vector predictor index of an upper neighboring block above the current block.
[0031] In another embodiment, the left adjacent block is A2 and the above adjacent block is B3.
[0032] In another embodiment, the left adjacent block is A1 and the above adjacent block is B1.
[0033] In another embodiment, the context variable has three different possible values.
[0034] Another embodiment includes comparing a motion vector predictor index of at least one neighboring block with an index value of a motion vector predictor index of a current block, and setting said context variable according to a comparison result.
[0035] Another embodiment includes comparing a motion vector predictor index of at least one adjacent block with a parameter representing a bit position of the or one of the bits in the motion vector predictor index of the current block and setting the context variable depending on the comparison result.
[0036] Yet another embodiment includes performing a first comparison, comparing a motion vector predictor index of a first adjacent block with a parameter representing a bit position of the or one of the bits in the motion vector predictor index of the current block, performing a second comparison, comparing a motion vector predictor index of a second adjacent block with the parameter, and setting the context variable depending on results of the first and second comparisons.
[0037] According to an eleventh aspect of the present invention, there is provided an apparatus for encoding a motion vector predictor index, comprising: means for generating a list of motion vector predictor candidates; means for selecting one of the motion vector predictor candidates in the list; means for generating a motion vector predictor index for a selected motion vector predictor candidate using CABAC coding, wherein a context variable of at least one bit of the motion vector predictor index for a current block depends on the motion vector predictor index of at least one block adjacent to the current block; An apparatus is provided comprising:
[0038] According to a twelfth aspect of the present invention, there is provided an apparatus for decoding a motion vector predictor index, comprising: means for generating a list of motion vector predictor candidates; means for decoding a motion vector predictor index using CABAC decoding, where a context variable for at least one bit of the motion vector predictor index of a current block depends on a motion vector predictor index of at least one block neighboring the current block; means for identifying one of the motion vector predictor candidates in the list using the decoded motion vector predictor index; An apparatus is provided comprising:
[0039] According to a thirteenth aspect of the present invention, there is provided a method for encoding a motion vector predictor index, comprising the steps of: generating a list of motion vector predictor candidates; selecting one of the motion vector predictor candidates in the list; Generate a motion vector predictor index of a selected motion vector predictor candidate using CABAC coding, where a context variable of at least one bit of the motion vector predictor index of a current block depends on a skip flag of the current block. A method is provided, comprising:
[0040] According to a fourteenth aspect of the present invention, there is provided a method for encoding a motion vector predictor index, comprising the steps of: generating a list of motion vector predictor candidates; selecting one of the motion vector predictor candidates in the list; Generate a motion vector predictor index of a selected motion vector predictor candidate using CABAC coding, where a context variable of at least one bit of the motion vector predictor index of a current block depends on another parameter or syntax element of the current block available before decoding the motion vector predictor index. A method is provided, comprising:
[0041] According to a fifteenth aspect of the present invention, there is provided a method for encoding a motion vector predictor index, comprising the steps of: generating a list of motion vector predictor candidates; selecting one of the motion vector predictor candidates in the list; Generate a motion vector predictor index for a selected motion vector predictor candidate using CABAC coding, where a context variable of at least one bit of the motion vector predictor index for a current block depends on another parameter or syntax element of the current block that is an indicator of the complexity of the motion within the current block. A method is provided, comprising:
[0042] According to a sixteenth aspect of the present invention, there is provided a method for decoding a motion vector predictor index, comprising the steps of: generating a list of motion vector predictor candidates; Decode a motion vector predictor index using CABAC decoding, where a context variable for at least one bit of the motion vector predictor index of a current block depends on a skip flag of the current block; Using the decoded motion vector predictor index to identify one of the motion vector predictor candidates in the list. A method is provided, comprising:
[0043] According to a seventeenth aspect of the present invention, there is provided a method for decoding a motion vector predictor index, comprising the steps of: generating a list of motion vector predictor candidates; Decode a motion vector predictor index using CABAC decoding, where a context variable of at least one bit of the motion vector predictor index of a current block depends on another parameter or syntax element of the current block that is available before decoding of the motion vector predictor index; Using the decoded motion vector predictor index to identify one of the motion vector predictor candidates in the list. A method is provided, comprising:
[0044] According to an eighteenth aspect of the present invention, there is provided a method for decoding a motion vector predictor index, comprising the steps of: generating a list of motion vector predictor candidates; Decode a motion vector predictor index using CABAC decoding, where a context variable of at least one bit of the motion vector predictor index of a current block depends on another parameter or syntax element of the current block that is an indicator of motion complexity in the current block; Using the decoded motion vector predictor index to identify one of the motion vector predictor candidates in the list. A method is provided, comprising:
[0045] According to a nineteenth aspect of the present invention, there is provided an apparatus for encoding a motion vector predictor index, comprising: means for generating a list of motion vector predictor candidates; means for selecting one of the motion vector predictor candidates in the list; means for generating a motion vector predictor index for a selected motion vector predictor candidate using CABAC coding, wherein a context variable for at least one bit of the motion vector predictor index for a current block depends on a skip flag for the current block; An apparatus is provided comprising:
[0046] According to a twentieth aspect of the present invention, there is provided an apparatus for encoding a motion vector predictor index, comprising: means for generating a list of motion vector predictor candidates; means for selecting one of the motion vector predictor candidates in the list; means for generating a motion vector predictor index for a selected motion vector predictor candidate using CABAC coding, wherein a context variable of at least one bit of the motion vector predictor index of a current block depends on another parameter or syntax element of the current block that is available prior to decoding of the motion vector predictor index; An apparatus is provided comprising:
[0047] According to a twenty-first aspect of the present invention, there is provided an apparatus for encoding a motion vector predictor index, comprising: means for generating a list of motion vector predictor candidates; means for selecting one of the motion vector predictor candidates in the list; means for generating a motion vector predictor index for a selected motion vector predictor candidate using CABAC coding, wherein a context variable of at least one bit of the motion vector predictor index for a current block depends on another parameter or syntax element of the current block that is an indicator of motion complexity within the current block; An apparatus is provided comprising:
[0048] According to a twenty-second aspect of the present invention, there is provided an apparatus for decoding a motion vector predictor index, comprising: means for generating a list of motion vector predictor candidates; means for decoding a motion vector predictor index using CABAC decoding, where a context variable for at least one bit of the motion vector predictor index of a current block depends on a skip flag of the current block; means for identifying one of the motion vector predictor candidates in the list using the decoded motion vector predictor index; An apparatus is provided comprising:
[0049] According to a twenty-third aspect of the present invention, there is provided an apparatus for decoding a motion vector predictor index, comprising: means for generating a list of motion vector predictor candidates; means for decoding a motion vector predictor index using CABAC decoding, where a context variable for at least one bit of the motion vector predictor index of a current block depends on another parameter or syntax element of the current block that is available prior to decoding of the motion vector predictor index; means for identifying one of the motion vector predictor candidates in the list using the decoded motion vector predictor index; An apparatus is provided comprising:
[0050] According to a twenty-fourth aspect of the present invention, there is provided an apparatus for decoding a motion vector predictor index, comprising: means for generating a list of motion vector predictor candidates; means for decoding a motion vector predictor index using CABAC decoding, where a context variable for at least one bit of the motion vector predictor index of a current block depends on another parameter or syntax element of the current block that is an indicator of motion complexity in the current block; means for identifying one of the motion vector predictor candidates in the list using the decoded motion vector predictor index; An apparatus is provided comprising:
[0051] According to a twenty-fifth aspect of the present invention, there is provided a method for encoding information relating to a motion information predictor, comprising: selecting one of a plurality of motion information predictor candidates; and encoding information for identifying the selected motion information predictor candidate using CABAC encoding, the CABAC encoding using, for at least one bit of the information, the same context variable used for another inter prediction mode when one or both of a triangle merge mode or a motion vector difference (MMVD) merge mode are used.
[0052] According to a twenty-sixth aspect of the present invention, there is provided a method for decoding information relating to a motion information predictor, comprising: decoding information for identifying one of a plurality of motion information predictor candidates using CABAC decoding; selecting one of the plurality of motion information predictor candidates using the decoded information; and, for at least one bit of the information, the CABAC decoding uses the same context variable used for another inter prediction mode when one or both of a triangle merge mode or a motion vector difference (MMVD) merge mode are used.
[0053] Regarding the twenty-fifth or twenty-sixth aspect of the present invention, the following features may be provided according to embodiments thereof.
[0054] Suitably, all bits except the first bit of the information are bypass CABAC coded or bypass CABAC decoded. Suitably, the first bit is CABAC coded or CABAC decoded. Suitably, the other inter prediction mode includes one or both of a merge mode or an affine merge mode. Suitably, the other inter prediction mode includes a Multi-Hypothesis Intra Inter (MHII) merge mode. Suitably, the multiple motion information predictor candidates for the other inter prediction mode include an ATMVP candidate. Suitably, the CABAC coding or CABAC decoding includes using the same context variable for both when a triangle merge mode is used and when an MMVD merge mode is used. Suitably, at least one bit of the information is CABAC coded or CABAC decoded when a skip mode is used. Suitably, the skip mode includes one or more of a merge skip mode, an affine merge skip mode, a triangle merge skip mode, or a merge of motion vector difference (MMVD) merge skip mode.
[0055] According to a twenty-seventh aspect of the present invention, there is provided a method for encoding information relating to a motion information predictor, comprising: selecting one of a plurality of motion information predictor candidates; encoding information for identifying the selected motion information predictor candidate; and, encoding the information, bypassing CABAC encoding of at least one bit of the information when one or both of a triangle merge mode or a motion vector difference (MMVD) merge mode merging are used.
[0056] According to a twenty-eighth aspect of the present invention, there is provided a method for decoding information relating to a motion information predictor, the method comprising: decoding information to identify one of a plurality of motion information predictor candidates; selecting one of the plurality of motion information predictor candidates using the decoded information; and, decoding the information, bypass CABAC decoding of at least one bit of the information when one or both of a triangle merge mode or a motion vector difference (MMVD) merge mode merging are used.
[0057] Regarding the twenty-seventh or twenty-eighth aspect of the present invention, the following features may be provided according to embodiments thereof.
[0058] Suitably, all bits of information except the first bit are bypass CABAC encoded or bypass CABAC decoded. Suitably, the first bit is CABAC encoded or CABAC decoded. Suitably, all bits of said information are bypass CABAC encoded or bypass CABAC decoded when one or both of the triangle merge mode or the MMVD merge mode are used. Suitably, all bits of said information are bypass CABAC encoded or bypass CABAC decoded.
[0059] Suitably, at least one bit of said information is CABAC encoded or CABAC decoded when affine merge mode is used, and suitably, all bits of said information are bypass CABAC encoded or bypass CABAC decoded except when affine merge mode is used.
[0060] Preferably, at least one bit of the information is CABAC encoded or CABAC decoded when one or both of a merge mode or a Multi-Hypothesis Intra Inter (MHII) merge mode are used. Preferably, all bits of the information are bypass CABAC encoded or bypass CABAC decoded except when one or both of a merge mode or a Multi-Hypothesis Intra Inter (MHII) merge mode are used.
[0061] Preferably, at least one bit of the information is CABAC coded or CABAC decoded if the plurality of motion information predictor candidates includes an ATMVP candidate. Preferably, all bits of the information are bypass CABAC coded or bypass CABAC decoded except if the plurality of motion information predictor candidates includes an ATMVP candidate.
[0062] Preferably, at least one bit of the information is CABAC encoded or CABAC decoded when a skip mode is used. Preferably, all bits of the information are bypass CABAC encoded or bypass CABAC decoded except when a skip mode is used. Preferably, the skip mode includes one or more of a merge skip mode, an affine merge skip mode, a triangle merge skip mode, or a merge of motion vector difference (MMVD) merge skip mode.
[0063] Regarding the twenty-fifth, twenty-sixth, twenty-seventh or twenty-eighth aspect of the present invention, the following features may be provided according to embodiments thereof.
[0064] Preferably, the at least one bit includes a first bit of the information. Preferably, the information includes a motion information prediction index or a flag. Preferably, the motion information predictor candidate includes information for obtaining a motion vector.
[0065] Regarding the twenty-fifth or twenty-seventh aspect of the present invention, the following features may be provided according to embodiments thereof.
[0066] Preferably, the method further includes, in the bitstream, information for indicating the use of one of a triangle merge mode, an MMVD merge mode, a merge mode, an affine merge mode, or a Multi-Hypothesis Intra Inter (MHII) merge mode. Preferably, the method further includes, in the bitstream, information for determining a maximum number of motion information predictor candidates that may be included in the multiple motion information predictor candidates.
[0067] Regarding the twenty-sixth or twenty-eighth aspects of the present invention, the following features may be provided according to embodiments thereof.
[0068] Suitably, the method further includes obtaining information from the bitstream to indicate the use of one of a triangle merge mode, an MMVD merge mode, a merge mode, an affine merge mode, or a Multi-Hypothesis Intra Inter (MHII) merge mode. Preferably, the method further includes obtaining information from the bitstream to determine a maximum number of motion information predictor candidates that may be included in the plurality of motion information predictor candidates.
[0069] According to a 29th aspect of the present invention, there is provided an apparatus for encoding information on a motion information predictor, comprising: means for selecting one of a plurality of motion information predictor candidates; and means for encoding information for identifying the selected motion information predictor candidate using CABAC encoding, the CABAC encoding using the same context variables used for another inter prediction mode when one or both of the merging of the triangle merge mode or the motion vector difference (MMVD) merge mode are used for at least one bit of the information. Suitably, the apparatus comprises means for performing the method for encoding information on a motion information predictor according to the 25th or 27th aspect of the present invention.
[0070] According to a 30th aspect of the present invention, there is provided an apparatus for encoding information related to a motion information predictor, comprising: means for selecting one of a plurality of motion information predictor candidates; and means for encoding information for identifying the selected motion information predictor candidate, wherein encoding the information comprises bypass CABAC encoding of at least one bit of the information when one or both of a triangle merge mode or a motion vector difference (MMVD) merge mode are used. Suitably, the apparatus comprises means for executing the method for encoding information related to a motion information predictor according to the 25th or 27th aspect of the present invention.
[0071] According to a 31st aspect of the present invention, there is provided an apparatus for decoding information on a motion information predictor, comprising: means for decoding information for identifying one of a plurality of motion information predictor candidates using CABAC decoding; and means for selecting one of a plurality of motion information predictor candidates using the decoded information, the CABAC decoding comprising using, for at least one bit of the information, the same context variable used for another inter prediction mode when one or both of the merge of the triangle merge mode or the motion vector difference (MMVD) merge mode are used. Suitably, the apparatus comprises means for executing the method for decoding information on a motion information predictor according to the 26th or 28th aspect of the present invention.
[0072] According to a 32nd aspect of the present invention, there is provided an apparatus for decoding information relating to a motion information predictor, comprising means for decoding information for identifying one of a plurality of motion information predictor candidates and means for selecting one of the plurality of motion information predictor candidates using the decoded information, the decoding of the information comprising bypass CABAC decoding of at least one bit of the information when one or both of a triangle merge mode or a motion vector difference (MMVD) merge mode are used. Suitably, the apparatus comprises means for performing the method for decoding information relating to a motion information predictor according to the 26th or 28th aspect of the present invention.
[0073] According to a thirty-third aspect of the present invention, there is provided a method for encoding a motion vector predictor index, comprising: generating a list of motion vector predictor candidates; selecting one of the motion vector predictor candidates in the list; and generating a motion vector predictor index for the selected motion vector predictor candidate using CABAC encoding, wherein a context variable of at least one bit of the motion vector predictor index of a current block is derived from at least one context variable of a skip flag and an affine flag of the current block.
[0074] According to a thirty-fourth aspect of the present invention, there is provided a method for decoding a motion vector predictor index, comprising: generating a list of motion vector predictor candidates; decoding the motion vector predictor index using CABAC decoding; a context variable of at least one bit of the motion vector predictor index of a current block is derived from at least one context variable of a skip flag and an affine flag of the current block; and using the decoded motion vector predictor index to identify one of the motion vector predictor candidates in the list.
[0075] According to a 35th aspect of the present invention, there is provided an apparatus for encoding a motion vector predictor index, the apparatus comprising: means for generating a list of motion vector predictor candidates; means for selecting one of the motion vector predictor candidates in the list; and means for generating a motion vector predictor index for the selected motion vector predictor candidate using CABAC encoding, wherein a context variable of at least one bit of the motion vector predictor index of a current block is derived from a context variable of at least one of a skip flag and an affine flag of the current block.
[0076] According to a 36th aspect of the present invention, there is provided an apparatus for decoding a motion vector predictor index, comprising: means for generating a list of motion vector predictor candidates; means for decoding the motion vector predictor index using CABAC decoding; and means for identifying one of the motion vector predictor candidates in the list using the decoded motion vector predictor index, wherein a context variable of at least one bit of the motion vector predictor index of a current block is derived from a context variable of at least one of a skip flag and an affine flag of the current block.
[0077] According to a thirty-seventh aspect of the present invention, there is provided a method for encoding a motion vector predictor index, comprising generating a list of motion vector predictor candidates, selecting one of the motion vector predictor candidates in the list, and generating a motion vector predictor index for the selected motion vector predictor candidate using CABAC encoding, wherein a context variable of at least one bit of the motion vector predictor index of a current block has only two different possible values.
[0078] According to a thirty-eighth aspect of the present invention, there is provided a method for decoding a motion vector predictor index, comprising generating a list of motion vector predictor candidates and decoding the motion vector predictor index using CABAC decoding, wherein a context variable for at least one bit of the motion vector predictor index of a current block has only two different possible values, and wherein the decoded motion vector predictor index is used to identify one of the motion vector predictor candidates in the list.
[0079] According to a thirty-ninth aspect of the present invention, there is provided an apparatus for encoding a motion vector predictor index, the apparatus comprising: means for generating a list of motion vector predictor candidates; means for selecting one of the motion vector predictor candidates in the list; and means for generating a motion vector predictor index for the selected motion vector predictor candidate using CABAC encoding, wherein a context variable of at least one bit of the motion vector predictor index of a current block has only two different possible values.
[0080] According to a fortieth aspect of the present invention, there is provided an apparatus for decoding a motion vector predictor index, comprising: means for generating a list of motion vector predictor candidates; means for decoding the motion vector predictor index using CABAC decoding; and means for identifying one of the motion vector predictor candidates in the list using the decoded motion vector predictor index, wherein a context variable for at least one bit of the motion vector predictor index of a current block has only two different possible values.
[0081] According to a forty-first aspect of the present invention, there is provided a method for encoding a motion information predictor index, comprising: generating a list of motion information predictor candidates; if an affine merge mode is used, selecting one of the motion information predictor candidates in the list as an affine merge mode predictor; and if a non-affine merge mode is used, selecting one of the motion information predictor candidates in the list as a non-affine merge mode predictor; and generating a motion information predictor index for the selected motion information predictor candidate using CABAC encoding, wherein one or more bits of the motion information predictor index are bypass CABAC encoded.
[0082] Suitably, the CABAC encoding comprises using the same context variable for at least one bit of a motion information predictor index of the current block when the affine merge mode is used and when the non-affine merge mode is used, or the CABAC encoding comprises using a first context variable when the affine merge mode is used or using a second context variable when the non-affine merge mode is used for at least one bit of a motion information predictor index of the current block, and the method further comprises including data indicating use of the affine merge mode in the bitstream when the affine merge mode is used.
[0083] Preferably, the method further includes data for determining a maximum number of motion information predictor candidates that may be included in the generated list of motion information predictor candidates in the bitstream. Preferably, all bits except the first bit of the motion information predictor index are bypass CABAC coded. Suitably, the first bit is CABAC coded. Suitably, the motion information predictor index of the selected motion information predictor candidate is coded using the same syntax element when the affine merge mode and the non-affine merge mode are used.
[0084] According to a forty-second aspect of the present invention, there is provided a method for decoding a motion information predictor index, comprising: generating a list of motion information predictor candidates; decoding the motion information predictor index using CABAC decoding, one or more bits of the motion information predictor index being bypass CABAC decoded; and, if an affine merge mode is used, using the decoded motion information predictor index to identify one of the motion information predictor candidates in the list as an affine merge mode predictor; and, if a non-affine merge mode is used, using the decoded motion information predictor index to identify one of the motion information predictor candidates in the list as a non-affine merge mode predictor.
[0085] Suitably, the CABAC decoding includes using the same context variable for at least one bit of the motion information predictor index of the current block when the affine merge mode and when the non-affine merge mode are used. Alternatively, the method further includes obtaining data indicating use of the affine merge mode from the bitstream, and the CABAC decoding includes using, for at least one bit of the motion information predictor index of the current block, a first context variable if the obtained data indicates use of the affine merge mode, and using a second context variable if the obtained data indicates use of the non-affine merge mode.
[0086] Suitably, the method further includes obtaining data from the bitstream indicating the use of an affine merge mode, and the generated list of motion information predictor candidates includes an affine merge mode predictor candidate if the obtained data indicates the use of an affine merge mode, and a non-affine merge mode predictor candidate if the obtained data indicates the use of a non-affine merge mode.
[0087] Suitably, the method further comprises obtaining data from the bitstream for determining a maximum number of motion information predictor candidates that can be included in the generated list of motion information predictor candidates. Suitably, all bits except the first bit of the motion information predictor index are bypass CABAC decoded. Suitably, the first bit is CABAC decoded. Suitably, decoding the motion information predictor index comprises parsing the same syntax element from the bitstream when the affine merge mode is used and when the non-affine merge mode is used. Suitably, the motion information predictor candidate includes information for obtaining a motion vector. Suitably, the generated list of motion information predictor candidates includes ATMVP candidates. Suitably, the generated list of motion information predictor candidates has the same maximum number of motion information predictor candidates that can be included therein when the affine merge mode is used and when the non-affine merge mode is used.
[0088] According to a 43rd aspect of the present invention, there is provided an apparatus for encoding a motion information predictor index, comprising: means for generating a list of motion information predictor candidates; means for selecting one of the motion information predictor candidates in the list as an affine merge mode predictor if an affine merge mode is used; means for selecting one of the motion information predictor candidates in the list as a non-affine merge mode predictor if a non-affine merge mode is used; and means for generating a motion information predictor index for the selected motion information predictor candidate using CABAC coding, where one or more bits of the motion information predictor index are bypass CABAC coded. Preferably, the apparatus comprises means for performing the method for encoding a motion information predictor index according to the 41st aspect.
[0089] According to a 44th aspect of the present invention, there is provided an apparatus for decoding a motion information predictor index, the apparatus comprising: means for generating a list of motion information predictor candidates; means for decoding the motion information predictor index using CABAC decoding, where one or more bits of the motion information predictor index are bypass CABAC decoded; means for identifying one of the motion information predictor candidates in the list as an affine merge mode predictor using the decoded motion information predictor index if an affine merge mode is used; and means for identifying one of the motion information predictor candidates in the list as a non-affine merge mode predictor using the decoded motion information predictor index if a non-affine merge mode is used. Preferably, the apparatus comprises means for performing the method for decoding a motion information predictor index according to the 42nd aspect.
[0090] According to a forty-fifth aspect of the present invention, there is provided a method for encoding a motion information predictor index for an affine merge mode, comprising generating a list of motion information predictor candidates, selecting one of the motion information predictor candidates in the list as an affine merge mode predictor, and generating a motion information predictor index for the selected motion information predictor candidate using CABAC encoding, wherein one or more bits of the motion information predictor index are bypass CABAC encoded.
[0091] Suitably, if the non-affine merge mode is used, the method further comprises selecting one of the motion information predictor candidates in the list as a non-affine merge mode predictor. Suitably, the CABAC encoding comprises using a first context variable if the affine merge mode is used or using a second context variable if the non-affine merge mode is used for at least one bit of the motion information predictor index of the current block, and the method further comprises including data indicating the use of the affine merge mode in the bitstream if the affine merge mode is used. Alternatively, the CABAC encoding comprises using the same context variable for at least one bit of the motion information predictor index of the current block if the affine merge mode is used and if the non-affine merge mode is used.
[0092] Preferably, the method further comprises data for determining a maximum number of motion information predictor candidates that may be included in the generated list of motion information predictor candidates in the bitstream.
[0093] Preferably, all bits except the first bit of the motion information predictor index are bypass CABAC coded. Suitably, the first bit is CABAC coded. Suitably, the motion information predictor index of the selected motion information predictor candidate is coded using the same syntax element when the affine merge mode and the non-affine merge mode are used.
[0094] According to a forty-sixth aspect of the present invention, there is provided a method for decoding a motion information predictor index for an affine merge mode, comprising: generating a list of motion information predictor candidates; decoding the motion information predictor index using CABAC decoding; and, if one or more bits of the motion information predictor index are bypass CABAC decoded and an affine merge mode is used, using the decoded motion information predictor index to identify one of the motion information predictor candidates in the list as an affine merge mode predictor.
[0095] Suitably, if the non-affine merge mode is used, the method further includes identifying one of the motion information predictor candidates in the list as a non-affine merge mode predictor using the decoded motion information predictor index. Suitably, the method further includes obtaining data indicating the use of the affine merge mode from the bitstream, and the CABAC decoding includes using a first context variable for at least one bit of the motion information predictor index of the current block if the obtained data indicates the use of the affine merge mode, and using a second context variable if the obtained data indicates the use of the non-affine merge mode. Alternatively, the CABAC decoding includes using the same context variable for at least one bit of the motion information predictor index of the current block if the affine merge mode is used and if the non-affine merge mode is used.
[0096] Suitably, the method further includes obtaining data from the bitstream indicating the use of an affine merge mode, and the generated list of motion information predictor candidates includes an affine merge mode predictor candidate if the obtained data indicates the use of an affine merge mode, and a non-affine merge mode predictor candidate if the obtained data indicates the use of a non-affine merge mode.
[0097] Suitably, decoding the motion information predictor index includes parsing the same syntax element from the bitstream when the affine merge mode and when the non-affine merge mode are used. Suitably, the method further includes obtaining data from the bitstream for determining a maximum number of motion information predictor candidates that can be included in the generated list of motion information predictor candidates. Suitably, all bits except the first bit of the motion information predictor index are bypass CABAC decoded. Suitably, the first bit is CABAC decoded. Suitably, the motion information predictor candidate includes information for obtaining a motion vector. Suitably, the generated list of motion information predictor candidates includes ATMVP candidates. Suitably, the generated list of motion information predictor candidates has the same maximum number of motion information predictor candidates that can be included therein when the affine merge mode and when the non-affine merge mode are used.
[0098] According to a 47th aspect of the present invention, there is provided an apparatus for encoding a motion information predictor index for an affine merge mode, comprising: means for generating a list of motion information predictor candidates; means for selecting one of the motion information predictor candidates in the list as an affine merge mode predictor; and means for generating a motion information predictor index for the selected motion information predictor candidate using CABAC coding, wherein one or more bits of the motion information predictor index are bypass CABAC coded. Preferably, the apparatus comprises means for performing the method for encoding a motion information predictor index according to the 45th aspect.
[0099] According to a 48th aspect of the present invention, there is provided an apparatus for decoding a motion information predictor index for an affine merge mode, comprising: means for generating a list of motion information predictor candidates; means for decoding the motion information predictor index using CABAC decoding, where one or more bits of the motion information predictor index are bypass CABAC decoded; and means for identifying one of the motion information predictor candidates in the list as an affine merge mode predictor using the decoded motion information predictor index if the affine merge mode is used. Preferably, the apparatus comprises means for performing the method for decoding a motion information predictor index according to the 46th aspect.
[0100] Yet another aspect of the invention relates to a program which, when executed by a computer or a processor, causes the computer or processor to perform any of the methods of the previous aspects. The program may be provided by itself or may be carried on, by or in a carrier medium. The carrier medium may be non-transitory, for example a storage medium, in particular a computer-readable storage medium. The carrier medium may also be transitory, for example a signal or other transmission medium. The signal may be transmitted over any suitable network, including the Internet.
[0101] Yet another aspect of the invention relates to a camera comprising an apparatus according to any of the apparatus aspects above. In one embodiment, the camera further comprises zooming means. In one embodiment, the camera indicates when said zooming means is operable and is adapted to signal an inter-prediction mode in response to said indication that the zooming means is operable. In another embodiment, the camera further comprises panning means. In another embodiment, the camera indicates when said panning means is operable and is adapted to signal an inter-prediction mode in response to said indication that the panning means is operable.
[0102] According to yet another aspect of the present invention, there is provided a mobile device comprising a camera embodying any of the above camera aspects. In one embodiment, the mobile device further comprises at least one position sensor adapted to sense a change in orientation of the mobile device. In one embodiment, the mobile device is adapted to signal an inter-prediction mode dependent on said sensing a change in orientation of the mobile device.
[0103] Further features of the invention are characterized by the other independent and dependent claims.
[0104] Any feature in one aspect of the invention may be applied to other aspects of the invention in any suitable combination. In particular, method aspects may be applied to apparatus aspects and vice versa. Furthermore, features implemented in hardware may be implemented in software and vice versa. Any references herein to software and hardware features should be interpreted accordingly. Any apparatus features as described herein may also be provided as method features and vice versa. As used herein, means-plus-function features may alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and associated memory.
[0105] It is also to be understood that specific combinations of the various features described and defined in any embodiment of the present invention can be implemented and / or provided and / or used independently. [Brief description of the drawings]
[0106] Reference will now be made, by way of example, to the accompanying drawings in which: [Figure 1] FIG. 1 is a diagram used to explain the coding structure used in HEVC. [Diagram 2] FIG. 2 is a block diagram that illustrates generally a data communications system in which one or more embodiments of the present invention may be implemented. [Diagram 3] FIG. 3 is a block diagram illustrating components of a processing device capable of implementing one or more embodiments of the present invention. [Figure 4] FIG. 4 is a flow chart illustrating steps of an encoding method according to an embodiment of the invention. [Diagram 5] FIG. 5 is a flow chart illustrating steps of a decoding method according to an embodiment of the invention. [Figure 6a] FIG. 6a shows spatial and temporal blocks that can be used to generate a motion vector predictor. [Figure 6b] FIG. 6b shows spatial and temporal blocks that can be used to generate a motion vector predictor. [Figure 7] FIG. 7 shows simplified steps in the process of AMVP predictor set derivation. [Figure 8] FIG. 8 is a schematic diagram of a motion vector derivation process in the merge mode. [Figure 9] FIG. 9 shows the segmentation and temporal motion vector prediction of the current block. [Figure 10] Figure 10(a) shows the encoding of the merge index for HEVC or when ATMVP is not enabled at the SPS level, and Figure 10(b) shows the encoding of the merge index when ATMVP is enabled at the SPS level. [Figure 11] Figure 11(a) shows a simple affine motion field, and Figure 11(b) shows a more complex affine motion field. [Figure 12] FIG. 12 is a flow chart of a partial decoding process for some syntax elements related to coding modes. [Figure 13] FIG. 13 is a flowchart showing merging candidate derivation. [Figure 14] FIG. 14 illustrates the encoding of the merge index according to the first embodiment of the present invention. [Figure 15]FIG. 15 is a flowchart of a partial decoding process of some syntax elements related to the encoding mode in the twelfth embodiment of the present invention. [Figure 16] FIG. 16 is a flowchart showing generation of a list of merging candidates in the twelfth embodiment of the present invention. [Figure 17] FIG. 17 is a block diagram for use in describing a CABAC encoder suitable for use in embodiments of the present invention. [Figure 18] FIG. 18 is a schematic block diagram of a communication system for implementation of one or more embodiments of the present invention. [Figure 19] FIG. 19 is a schematic block diagram of a computing device. [Figure 20] FIG. 20 is a diagram showing a network camera system. [Figure 21] FIG. 21 is a diagram showing a smartphone. [Figure 22] FIG. 22 is a flowchart of a partial decoding process of some syntax elements related to encoding modes according to the sixteenth embodiment. [Figure 23] FIG. 23 is a flow chart illustrating the use of a single index signaling scheme for both merge and affine merge modes according to an embodiment. [Figure 24] FIG. 24 is a flowchart showing the affine merge candidate derivation process in the affine merge mode according to this embodiment. [Diagram 25] 25(a) and 25(b) illustrate the predictor derivation process for triangle merging mode according to one embodiment. [Figure 26] FIG. 26 is a flowchart of a decoding process for an inter prediction mode for a current coding unit according to one embodiment. [Figure 27] Figure 27(a) shows the encoding of flags for merging for the motion vector differential (MMVD) merge mode according to one embodiment, and Figure 27(b) shows the encoding of indices for the triangle merge mode according to one embodiment. [Figure 28] FIG. 28 is a flow diagram illustrating an affine merge candidate derivation process for the affine merge mode using ATMVP candidates, according to one embodiment. [Figure 29] FIG. 29 is a flowchart of a decoding process in the inter prediction mode according to the eighteenth embodiment. [Diagram 30] Figure 30(a) shows the encoding of flags for merging for the motion vector difference (MMVD) merge mode according to the nineteenth embodiment, Figure 30(b) shows the encoding of indexes for the triangle merge mode according to the nineteenth embodiment, and Figure 30(c) shows the encoding of indexes for the affine merge mode or merge mode according to the nineteenth embodiment. [Diagram 31] FIG. 31 is a flowchart of a decoding process in the inter prediction mode according to the nineteenth embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0107] The embodiments of the present invention described below relate to improving the encoding and decoding of index / flag / information / data using CABAC. It should be understood that alternative embodiments of the present invention can also be implemented to improve other context-based arithmetic coding schemes similar in functionality to CABAC. Before describing the embodiments, video encoding and decoding techniques and related encoders and decoders are described.
[0108] In this specification, "signaling" may refer to inserting (providing / including / during encoding) or extracting / obtaining (decoding) bitstream information regarding one or more syntax elements that represent the use, discontinuance, enabling or disabling of a mode (e.g., inter-prediction mode) or other information (such as information regarding a selection).
[0109] Figure 1 relates to the coding structure used in the High Efficiency Video Coding (HEVC) video standard. A video sequence 1 is composed of a sequence of digital images i. Each such digital image is represented by one or more matrices. The matrix coefficients represent pixels.
[0110] The image 2 of the sequence is divided into slices 3, which in some cases constitute the entire image. These slices are divided into non-overlapping coding tree units (CTUs). A coding tree unit (CTU) is the basic processing unit of the High Efficiency Video Coding (HEVC) video standard and conceptually corresponds in structure to the macroblock unit used in some previous video standards. A CTU is sometimes also called a maximum coding unit (LCU). A CTU has luma and chroma component parts, each of which is called a coding tree block (CTB). These different color components are not shown in FIG. 1.
[0111] A CTU is typically of size 64 pixels by 64 pixels for HEVC, but for VVC the size may be 128 pixels by 128 pixels. Each CTU, in turn, may be iteratively divided into smaller variable-size coding units (CUs) 5 using a quadtree decomposition.
[0112] A coding unit is a basic coding element and consists of two types of subunits called prediction units (PUs) and transform units (TUs). The maximum size of a PU or TU is equal to the CU size. A prediction unit corresponds to a partition of a CU for the prediction of pixel values. Various different partitions of a CU into PUs are possible, including a partition into four square PUs and two different partitions into two rectangular PUs, as shown in 6. A transform unit is a basic unit that performs spatial transformation using DCT. A CU can be partitioned into TUs based on a quadtree representation 7. Thus, a slice, a tile, a CTU / LCU, a CTB, a CU, a PU, a TU, or a block of pixels / samples may be referred to as an image part, i.e., a part of an image 2 of a sequence.
[0113] Each slice is embedded in one network abstraction layer (NAL) unit. Furthermore, the coding parameters of a video sequence are stored in a dedicated NAL unit called a parameter set. In HEVC and H.264 / AVC, two types of parameter set NAL units are used: first, the sequence parameter set (SPS) NAL unit, which collects all parameters that do not change during the entire video sequence. Typically, it handles the coding profile, the size of the video frames, and other parameters. Second, the picture parameter set (PPS) NAL unit contains parameters that can change from one picture (or frame) of the sequence to another picture (or frame). HEVC also includes the video parameter set (VPS) NAL unit, which contains parameters that describe the overall structure of the bitstream. VPS is a new type of parameter set defined in HEVC that applies to all layers of the bitstream. A layer can contain multiple temporal sublayers, and all version 1 bitstreams are limited to one layer. HEVC has specific layer extensions for scalability and multiview, which allow multiple layers with a backward-compatible version 1 base layer.
[0114] 2 and 18 illustrate a data communication system in which one or more embodiments of the present invention may be implemented. The data communication system includes a transmitting device, e.g., server 201 in FIG. 2 or content provider 150 in FIG. 18, operable to transmit data packets of a data stream 204 (or bitstream 101 in FIG. 18) to a receiving device, e.g., client terminal 202 in FIG. 2 or content consumer 100 in FIG. 18, via a data communication network 200. The data communication network 200 may be a wide area network (WAN) or a local area network (LAN). Such a network may be, for example, a wireless network (Wifi / 802.11a or b or g), an Ethernet network, an Internet network, or a mixed network made up of several different networks. In a particular embodiment of the present invention, the data communication system may be a digital television broadcasting system in which the server 201 (or content provider 150 in FIG. 18) transmits the same data content to multiple clients (or content consumers).
[0115] The data stream 204 (or bitstream 101) provided by the server 201 (or content provider 150) may consist of multimedia data representing video and audio data. The audio and video data streams may, in some embodiments of the present invention, be captured by the server 201 (or content provider 150) using a microphone and a camera, respectively. In some embodiments, the data streams may be stored on the server 201 (or content provider 150) or may be received by the server 201 (or content provider 150) from another data provider, or may be generated at the server 201 (or content provider 150). The server 201 (or content provider 150) particularly comprises an encoder for encoding the video and audio streams (e.g., the original sequence 151 of the images in FIG. 18) to provide a compressed bitstream 204,101 for transmission, which is a more compact representation of the data presented as input to the encoder.
[0116] In order to obtain a better ratio of the quality of the transmitted data to the amount of the transmitted data, the compression of the video data may for example be according to the HEVC format, or the H.264 / AVC format, or the VVC format.
[0117] The client 202 (or content consumer 100) receives the transmitted bitstream and decodes the reconstructed bitstream to play video images (e.g., video signal 109 in FIG. 18) on a display device and audio data through speakers.
[0118] In the examples of Figure 2 or Figure 18 a streaming scenario is considered, but it will be understood that in some embodiments of the invention data communication between the encoder and the decoder may be performed using a media storage device, such as an optical disc, for example.
[0119] In one or more embodiments of the present invention, a video image may be transmitted along with data representing a compensation offset to be applied to reconstructed pixels of the image to provide filtered pixels in the final image.
[0120] 3 shows a schematic diagram of a processing device 300 configured to implement at least one embodiment of the present invention. The processing device 300 may be a device such as a microcomputer, a workstation, or a light portable device. The device 300 may include: - a central processing unit 311 such as a microprocessor represented by CPU - a read-only memory 307, denoted ROM, for storing a computer program for implementing the invention; a random access memory 312, denoted RAM, for storing the executable code of the method of the embodiment of the invention, as well as registers arranged to record variables and parameters necessary for implementing the method of encoding a sequence of digital images and / or the method of decoding a bitstream according to an embodiment of the invention; a communications interface 302 connected to a communications network 303 over which the digital data to be processed is sent and received; The communication bus 313 is connected to the
[0121] Optionally, the device 300 may also include the following components:
[0122] - data storage means 304, such as a hard disk, for storing computer programs for implementing the methods of one or more embodiments of the invention and data used or generated during the implementation of one or more embodiments of the invention; a disk drive 305 for a disk 306, the disk drive configured to read data from or write data to the disk 306; A screen 309 that displays data and / or acts as a graphical interface with the user using a keyboard 310 or any other pointing / input means. The device 300 may be connected to various peripheral devices, such as a digital camera 320 or a microphone 308, each connected to an input / output card (not shown) for providing multimedia data to the device 300.
[0123] The communication bus provides communication and interoperability between the various elements included in or connected to the device 300. The representation of the bus is not limiting, in particular the central processing unit is operable to communicate instructions to any element of the device 300, either directly or by means of another element of the device 300.
[0124] The disk 306 may be replaced by any information medium, such as, for example, a compact disk (CD-ROM), a ZIP disk or a memory card, rewritable or not, and generally speaking, by any information storage means readable by a microcomputer or microprocessor, integrated or not integrated into the device, possibly removable, and configured to store one or more programs whose execution enables the method of encoding a sequence of digital images and / or the method of decoding a bitstream according to the invention to be performed.
[0125] The executable code may be stored either in the read-only memory 307, on the hard disk 304 or on a removable digital medium such as, for example, the disk 306 as previously explained. According to a variant, the executable code of the program may be received by the communication network 303, via the interface 302, to be stored in one of the storage means of the device 300 before being executed, such as the hard disk 304.
[0126] The central processing unit 311 is arranged to control and direct the execution of instructions or parts of the program or software code of the program according to the invention, with instructions stored in one of the aforementioned storage means. On power-up, the program or programs stored in a non-volatile memory, for example on the hard disk 304 or on the disk 306 or in the read-only memory 307, are transferred to the random access memory 312, which contains registers for storing the program or the executable code of the program, as well as variables and parameters necessary to implement the invention.
[0127] In this embodiment, the device is a programmable device that uses software to implement the invention, however, alternatively, the invention may be implemented in hardware (e.g. in the form of an application specific integrated circuit or ASIC).
[0128] 4 shows a block diagram of an encoder according to at least one embodiment of the present invention. The encoder is represented by connected modules, each module adapted to perform, for example in the form of program instructions to be executed by the CPU 311 of the device 300, at least one corresponding step of a method for implementing at least one embodiment of encoding images of a sequence of images according to one or more embodiments of the present invention.
[0129] An original sequence of digital images i0 to in401 is received as input by the encoder 400. Each digital image is represented by a set of samples, sometimes also called picture elements (hereafter referred to as pixels).
[0130] A bitstream 410 is output by the encoder 400 after the encoding process is performed. The bitstream 410 comprises a number of coding units or slices, each slice comprising a slice header for transmitting coded values of coding parameters used to code the slice, and a slice body comprising coded video data.
[0131] An input digital image i0 to in401 is divided by module 402 into blocks of pixels. The blocks correspond to image portions and may be of variable size (for example 4x4, 8x8, 16x16, 32x32, 64x64, 128x128 pixels, and also several rectangular block sizes can be considered). A coding mode is selected for each input block. Two families of coding modes are provided: coding modes based on spatial predictive coding (intra prediction) and coding modes based on temporal prediction (inter coding, merge, SKIP). Possible coding modes are tested.
[0132] The module 403 implements an intra prediction process in which a given block to be coded is predicted by a predictor calculated from neighbouring pixels of said block to be coded. An indication of the selected intra predictor and the difference between the given block and its predictor are coded to provide a residual when intra coding is selected.
[0133] Temporal prediction is performed by the motion estimation module 404 and the motion compensation module 405. First, a reference image is selected from the set of reference images 416, and the portion of the reference image, also called the reference region or image portion, which is the closest region (closest in terms of pixel value similarity) to the given block to be coded, is selected by the motion estimation module 404. The motion compensation module 405 then uses the selected region to predict the block to be coded. The difference between the selected reference region and the given block, also called the residual block, is calculated by the motion compensation module 405. The selected reference region is indicated by a motion vector.
[0134] Therefore, in both cases (spatial and temporal prediction), the residual is calculated by subtracting the predictor from the original block when the original block is not in SKIP mode.
[0135] In the INTRA prediction implemented by module 403, the prediction direction is coded. In the INTER prediction implemented by modules 404, 405, 416, 418, 417, at least one motion vector or data for identifying such a motion vector is coded for temporal prediction.
[0136] If inter prediction is selected, the motion vector and information related to the residual block are coded. To further reduce the bit rate, assuming that the motion is uniform, the motion vector is coded by the difference to the motion vector predictor. The motion vector predictor from a set of motion information predictor candidates is obtained from the motion vector field 418 by the motion vector predictive coding module 417.
[0137] The encoder 400 further comprises a selection module 406 for selecting a coding mode by applying a coding cost criterion, such as a rate-distortion criterion. To further reduce redundancy, a transform (such as a DCT) is applied to the residual block by a transform module 407, and the resulting transformed data is quantized by a quantization module 408 and entropy coded by an entropy coding module 409. Finally, the coded residual block of the current block being coded is inserted into the bitstream 410 when not in the SKIP mode, which requires the residual block to be coded in the bitstream.
[0138] The encoder 400 also performs decoding of the encoded images to generate reference images (e.g. those in the reference images / pictures 416) for motion estimation of subsequent images. This allows the encoder and decoder receiving the bitstream to have the same reference frames (reconstructed images or image parts are used). The inverse quantization ("dequantization") module 411 performs inverse quantization ("dequantization") of the quantized data, followed by inverse transformation by the inverse transformation module 412. The intra prediction module 413 uses the prediction information to decide which predictor to use for a given block, and the motion compensation module 414 actually adds the residual obtained by module 412 to a reference region obtained from the set of reference images 416.
[0139] Then, post-filtering is applied by module 415 to filter the reconstructed frame of pixels (image or image portion). In an embodiment of the present invention, an SAO loop filter is used in which a compensation offset is added to the pixel values of the reconstructed pixels of the reconstructed image. It is understood that post-filtering does not necessarily have to be performed. Also, any other type of post-filtering can be performed in addition to or instead of SAO loop filtering.
[0140] 5 shows a block diagram of a decoder 60 that may be used to receive data from an encoder according to one embodiment of the present invention. The decoder is represented by connected modules, each module configured to implement a corresponding step of a method implemented by the decoder 60, for example in the form of program instructions executed by the CPU 311 of the device 300.
[0141] The decoder 60 receives a bitstream 61 containing coding units (e.g. data corresponding to image portions, blocks or coding units), each coding unit consisting of a header containing information about coding parameters and a body containing the coded video data. As explained with respect to Fig. 4, the coded video data is entropy coded and the index of the motion vector predictor is coded with a predefined number of bits for a given image portion (e.g. block or CU). The received coded video data is entropy decoded by module 62. The residual data is then inverse quantized by module 63 and then an inverse transform is applied by module 64 to obtain pixel values.
[0142] Mode data indicating the coding mode is also entropy decoded, and based on this mode, INTRA type decoding or INTER type decoding is performed on the coded block (unit / set / group) of image data.
[0143] For INTRA modes, the INTRA predictor is determined by the intra prediction module 65 based on the intra prediction mode specified in the bitstream.
[0144] If the mode is INTER, motion prediction information is extracted from the bitstream to find (identify) the reference region used by the encoder. The motion prediction information includes a reference frame index and a motion vector residual. The motion vector predictor is added to the motion vector residual by the motion vector decoding module 70 to obtain a motion vector.
[0145] A motion vector decoding module 70 applies motion vector decoding for each image portion (e.g., current block or CU) coded by motion prediction. Once the motion vector predictor index for the current block is obtained, the actual value of the motion vector associated with the image portion (e.g., current block or CU) may be decoded and used to apply motion compensation by module 66. The reference image portion indicated by the decoded motion vector is extracted from reference image 68 and motion compensation 66 is applied. Motion vector field data 71 is updated with the decoded motion vector for use in predicting subsequent decoded motion vectors.
[0146] Finally, decoded blocks are obtained. If appropriate, post-filtering is applied by a post-filtering module 67. A decoded video signal 69 is finally obtained and provided by the decoder 60.
[0147] CABAC HEVC uses several types of entropy coding, such as CABAC (Context based Adaptive Binary Arithmetic Coding), Golomb-rice Code, or a simple binary representation called Fixed Length Coding. In most cases, a binary coding process is performed to represent the different syntax elements. This binary coding process is also very specific and depends on the different syntax elements. Arithmetic coding represents the syntax elements according to their current probability. CABAC is an extension of arithmetic coding that separates the probability of syntax elements depending on a "context" defined by a context variable. This corresponds to a conditional probability. The context variables can be derived from the values of the current syntax of the top-left block (A2 in Fig. 6b, explained in detail below) and the top-left block (B3 in Fig. 6b), which have already been decoded.
[0148] CABAC has been adopted as a reference part of the H.264 / AVC and H.265 / HEVC standards. In H.264 / AVC, it is one of two alternative methods of entropy coding. The other method specified in H.264 / AVC is a low-complexity entropy coding technique based on the use of a context-adaptively switched set of variable-length codes, the so-called Context-Adaptive Variable-Length Coding (CAVLC). Compared to CABAC, CAVLC offers reduced implementation costs at the expense of lower compression efficiency. For TV signals with standard or high definition resolution, CABAC typically offers 10-20% bitrate savings over CAVLC at the same objective video quality. In HEVC, CABAC is one of the entropy coding methods used. Many bits are also bypass CABAC coded (also denoted as CABAC bypass coding). In addition, some syntax elements are coded with unary codes or Golomb codes, which are other types of entropy codes.
[0149] FIG. 17 shows the main blocks of the CABAC encoder.
[0150] Input syntax elements that are non-binary values are binarized by a binarizer 1701. The coding strategy of CABAC is based on the discovery that highly efficient coding of syntax element values in hybrid block-based video coders, such as components of motion vector differences or transform coefficient level values, can be achieved by using a binarization scheme as a kind of pre-processing unit for the subsequent stages of context modeling and binary arithmetic coding. In general, a binarization scheme defines a unique mapping of syntax element values to a sequence of binary decisions, so-called bins, which can be "bits", and thus can also be interpreted in terms of a binary code tree. The design of the binarization scheme in CABAC is based on a small number of basic prototypes whose structure allows for simple online computation and which are applied to some suitable model-probability distributions.
[0151] Each bin can be processed in one of two basic ways, depending on the setting of switch 1702. When the switch is in the "regular" setting, the bin is fed to a context modeler 1703 and a regular encoding engine 1704. When the switch is in the "bypass" setting, the context modeler is bypassed and the bin is fed to a bypass encoding engine 1705. Another switch 1706 has similar "regular" and "bypass" settings as switch 1702, so that the bins encoded by the applicable one of the encoding engines 1704 and 1705 can form a bitstream as the output of the CABAC encoder.
[0152] It will be appreciated that the other switch 1706 may be used in conjunction with the storage to group some of the bins (e.g., bins for encoding image portions such as blocks or coding units) encoded by the coding engine 1705 to provide blocks of bypass encoded data in the bitstream, and to group some of the bins (e.g., bins for encoding blocks or coding units) encoded by the coding engine 1704 to provide another block of “regular” (or arithmetically) encoded data in the bitstream. This separate grouping of bypass encoded data and regular encoded data may result in improved throughput during the decoding process (since the bypass encoded data can be processed first / in parallel with the regular CABAC encoded data).
[0153] By decomposing each syntax element value into a sequence of bins, the further processing of each bin value in CABAC depends on an associated coding mode decision, which can be selected as either normal mode or bypass mode. The latter is selected for bins related to code information, or for lower significant bins that are assumed to be uniformly distributed, so that the entire normal binary arithmetic coding process is simply bypassed. In the normal coding mode, each bin value is coded by using a normal binary arithmetic coding engine, and the associated probability model is either determined by a fixed selection without context modeling, or is adaptively selected depending on the associated context model. As an important design decision, the latter case is generally applied only to the most frequently observed bins, while other, usually less frequently observed bins are processed using a joint, typically zero-order probability model. In this way, CABAC allows selective context modeling at the sub-symbol level, thus providing an efficient means to exploit inter-symbol redundancy with a significantly reduced overall modeling or learning cost. For a particular choice of context model, four basic design types were adopted in CABAC, two of which were applied to coding only at the transform coefficient level: The design of these four prototypes is based on a priori knowledge of the typical characteristics of the source data to be modeled and reflects the aim of finding a good compromise between the conflicting objectives of avoiding unnecessary modeling cost overhead and exploiting statistical dependencies to a large extent.
[0154] At the lowest level of processing in CABAC, each bin value enters a binary arithmetic coder in either normal or bypass coding mode. In the latter case, a fast branch of the coding engine with significantly reduced complexity is used, while in the former coding mode, the coding of a given bin value depends on the actual state of an associated adaptive probability model that is passed along with the bin value to the M-coder, the term chosen for the table-based adaptive binary arithmetic coding engine in CABAC.
[0155] A corresponding CABAC decoder then receives the bitstream output from the CABAC encoder and processes the bypass encoded data and the regular CABAC encoded data accordingly. As the CABAC decoder processes the regular CABAC encoded data, the context modeler (and its probability model) is updated so that it can correctly decode / process (e.g., debinarize) the bins that form the bitstream to obtain the syntax elements.
[0156] Inter-coding HEVC uses three different inter modes, namely, inter mode (Advanced Motion Vector Prediction (AMVP) with signaling of motion information difference), "classical" merge mode (i.e., also known as "non-affine merge mode" or "regular" merge mode with no signaling of motion information difference), and "classical" merge skip mode (i.e., also known as "non-affine merge skip" mode or "regular" merge skip mode with no signaling of motion information difference and no signaling of residual data of sample values). The main difference between these modes is the data signaling in the bitstream. For motion vector coding, the current HEVC standard includes a contention-based scheme for motion vector prediction that was not present in previous versions of the standard. This means that for each inter coding mode (AMVP) or merge mode (i.e., "classical / regular" merge mode or "classical / regular" merge skip mode), several candidates are competing with a rate-distortion criterion on the encoder side to find the best motion vector predictor or the best motion information. Then, an index or flag corresponding to the best candidate for the best predictor or motion information is inserted into the bitstream. The decoder can derive the same set of predictors or candidates and use the best one according to the decoded index / flag. In the screen content extension of HEVC, a new coding tool called Intra Block Copy (IBC) is signaled as any of these three inter modes, and the difference between IBC and the equivalent inter mode is done by checking if the reference frame is the current one. IBC is also known as Current Picture Reference (CPR). This can be implemented, for example, by checking the reference index of list L0 and inferring that if this is the last frame in the list, this is an intra block copy. Another way is to compare the picture order counts of the current frame and the reference frame, and if they are equal, this is an intra block copy.
[0157] The design of predictor and candidate derivation is important to achieve the best coding efficiency without disproportionately impacting the complexity. In HEVC, two motion vector derivations are used: one for inter-mode (Advanced Motion Vector Prediction (AMVP)) and one for merge mode (Merge derivation process - for the classical Merge mode and the classical Merge Skip mode). These processes are described below.
[0158] Figures 6a and 6b show spatial and temporal blocks that can be used to generate motion vector predictors, for example, in Advanced Motion Vector Prediction (AMVP) and merge modes of an HEVC encoding and decoding system, and Figure 7 shows simplified steps of the process of AMVP predictor set derivation.
[0159] Two spatial predictors, i.e., two spatial motion vectors for AMVP mode, are selected from among the motion vectors of the top block (indicated by the letter "B") and the left block (indicated by the letter "A"), including the top corner block (block B2) and the left corner block (block A0), and one temporal predictor is selected from among the motion vectors of the bottom right block (H) and the center block (Center) of the collocated blocks, as represented in FIG. 6a.
[0160] Table 1 below outlines the nomenclature used when referencing blocks relative to the current block as shown in Figures 6a and 6b. This nomenclature is used for simplicity, but it will be understood that other labeling systems may be used, especially in future versions of the standard.
[0161] [Table 1]
[0162] It should be noted that the "current block" can be variable in size, such as 4x4, 16x16, 32x32, 64x64, 128x128, or any size in between. The dimensions of the block are preferably a multiple of 2 (i.e., 2^n x 2^m, where n and m are positive integers), which results in a more efficient use of bits when using binary encoding. The current block does not have to be square, although this is often the preferred embodiment due to encoding complexity.
[0163] 7, the first step aims to select the first spatial predictor (Cand1, 706) among the bottom left blocks A0 and A1, whose spatial location is shown in Fig. 6a. To that end, these blocks are selected one after the other in a given (i.e., predefined / preset) order (700, 702), and for each selected block, the following conditions are evaluated in the given order (704), and the first block for which the conditions are satisfied is set as the predictor:
[0164] - Motion vectors from the same reference images and from the same reference list -Motion vectors from the same reference image and other reference lists - scaled motion vectors from different reference images and the same reference list - Scaled motion vectors from different reference pictures and other reference lists If the value is missing, the left predictor is considered to be unavailable, indicating that the associated blocks are either intra-coded or do not exist.
[0165] The next step aims to select a second spatial predictor (Cand2, 716) among the top right block B0, the top block B1 and the top left (top left) block B2, whose spatial locations are shown in Fig. 6a. To that end, these blocks are selected one after the other in a given order (708, 710, 712), and for each selected block, the above mentioned conditions are evaluated in a given order (714), and the first block for which the above mentioned conditions are satisfied is set as the predictor.
[0166] Again, if a value is not found, the above predictor is considered unavailable, indicating that the associated blocks are either intra-coded or do not exist.
[0167] In the next step (718), the two predictors, if both are available, are compared to each other in order to eliminate one of them if they are equal (i.e., same motion vector value, same reference list, same reference index, and same orientation type). If only one spatial predictor is available, the algorithm looks for a temporal predictor in a later step.
[0168] The temporal motion predictor (Cand3, 726) is derived as follows: the bottom right (H, 720) location of the co-located block in the previous / reference frame is first considered in the availability check module 722. If it does not exist or if no motion vector predictor is available, the center (Center, 724) of the co-located block is selected to be checked. These temporal locations (Center and H) are shown in Figure 6a. In any case, scaling 723 is applied to these candidates to match the temporal distance between the current frame and the first frame in the reference list.
[0169] The motion predictor value is then added to the set of predictors. The number of predictors (Nb_Cand) is then compared to the maximum number of predictors (Max_Cand) (728). As mentioned above, the maximum number of predictors (Max_Cand) for the motion vector predictors that the AMVP derivation process needs to generate is 2 in the current version of the HEVC standard.
[0170] If this maximum number is reached, then a final list or set of AMVP predictors (732) is constructed. Otherwise, a zero predictor is added to the list (730). A zero predictor is a motion vector equal to (0,0).
[0171] As shown in FIG. 7, a final list or set of AMVP predictors (732) is constructed from a subset of candidate spatial motion predictors (700-712) and a subset of candidate temporal motion predictors (720, 724).
[0172] As mentioned above, a motion predictor candidate for classical merge mode or classical merge skip mode can represent all the necessary motion information: direction, list, reference frame index, and motion vector (or any subset thereof to perform prediction). An indexed list of several candidates is generated by the merge derivation process. In the current HEVC design, the maximum number of candidates for both merge modes (i.e., classical merge mode and classical merge skip mode) is equal to 5 (4 spatial candidates and 1 temporal candidate).
[0173] FIG. 8 is a schematic diagram of the motion vector derivation process for the merge modes (classical merge mode and classical merge skip mode). In the first step of the derivation process, five block positions are considered (800-808). These positions are the spatial positions shown in FIG. 6a with reference numbers A1, B1, B0, A0, and B2. In a later step, the availability of spatial motion vectors is checked and at most five motion vectors are selected / obtained for consideration (810). If a predictor is present and the block is not intra-coded, the predictor is considered to be available. Thus, the selection of motion vectors corresponding to the five blocks as candidates is done according to the following conditions:
[0174] If a "left" A1 motion vector (800) is available (810), i.e., if it is present and this block is not intra-coded, then the motion vector of the "left" block is selected and used as the first candidate in the candidate list (814).
[0175] If a "top" B1 motion vector (802) is available (810), the candidate "top" block motion vector is compared to the "left" A1 motion vector, if present (812). If the B1 motion vector is equal to the A1 motion vector, B1 is not added to the list of spatial candidates (814). Conversely, if the B1 motion vector is not equal to the A1 motion vector, B1 is added to the list of spatial candidates (814).
[0176] If a "top right" B0 motion vector (804) is available (810), the "top right" motion vector is compared (812) to the B1 motion vector. If the B0 motion vector is equal to the B1 motion vector, the B0 motion vector is not added to the list of spatial candidates (814). Conversely, if the B0 motion vector is not equal to the B1 motion vector, the B0 motion vector is added to the list of spatial candidates (814).
[0177] If a "bottom left" A0 motion vector (806) is available (810), the "bottom left" motion vector is compared (812) to the A1 motion vector. If the A0 motion vector is equal to the A1 motion vector, the A0 motion vector is not added to the list of spatial candidates (814). Conversely, if the A0 motion vector is not equal to the A1 motion vector, the A0 motion vector is added to the list of spatial candidates (814).
[0178] If the list of spatial candidates does not contain four candidates, the availability of a "top left" B2 motion vector (808) is checked (810). If available, it is compared with the A1 and B1 motion vectors. If the B2 motion vector is equal to the A1 or B1 motion vector, the B2 motion vector is not added to the list of spatial candidates (814). Conversely, if the B2 motion vector is not equal to the A1 or B1 motion vector, the B2 motion vector is added to the list of spatial candidates (814).
[0179] At the end of this stage, the list of spatial candidates contains up to four candidates.
[0180] For the temporal candidates, two positions can be used: the bottom right position of the co-located block (816, H shown in FIG. 6a) and the center of the co-located block (818). These positions are shown in FIG. 6a.
[0181] As explained in relation to FIG. 7 for the temporal motion predictor of the AMVP motion vector derivation process, the first step aims to check the availability of a block in the H position (820). Then, if it is not available, the availability of a block in the center position is checked (820). If at least one motion vector of these positions is available, the temporal motion vector may be scaled (822) to the reference frame with index 0, if necessary, for both lists L0 and L1, to create a temporal candidate (824) that is added to the list of merge motion vector predictor candidates. This is placed after the spatial candidate in the list. Lists L0 and L1 are two reference frame lists that contain zero, one or more reference frames.
[0182] If the number of candidates (Nb_Cand) is strictly less than the maximum number of candidates (826) (the Max_Cand information for determining the value is signaled in the bitstream slice header and is equal to 5 in the current HEVC design) and if the current frame is of type B, a combined candidate is generated (828). The combined candidate is generated based on the available candidates of the list of merge motion vector predictor candidates. This mainly consists of combining (pairing) the motion information of one candidate of list L0 with the motion information of one candidate of list L1.
[0183] If the number of candidates (Nb_Cand) is strictly less than the maximum number of candidates (Max_Cand) (830), zero motion candidates are generated (832) until the number of candidates in the list of merge motion vector predictor candidates reaches the maximum number of candidates.
[0184] At the end of this process, a list or set of merge motion vector predictor candidates (i.e., a list or set of candidates for the merge modes, which are the classical merge mode and the classical merge skip mode) is constructed (834). As shown in Figure 8, the list or set of merge motion vector predictor candidates is constructed (834) from a subset of spatial candidates (800-808) and a subset of temporal candidates (816, 818).
[0185] Alternative Temporal Motion Vector Prediction (ATMVP) Alternative temporal motion vector prediction (ATMVP) is a special type of motion compensation. Instead of considering only one motion information for the current block from the temporal reference frame, each motion information of each collocated block is considered. Thus, this temporal motion vector prediction gives a segmentation of the current block with the associated motion information of each subblock, as shown in Figure 9.
[0186] In the VTM reference software, the ATMVP is signaled as a merge candidate inserted into a list of merge candidates (i.e., a list or set of candidates for merge modes that are classical merge mode and classical merge skip mode). When ATMVP is enabled at the SPS level, the maximum number of merge candidates is increased by one. Thus, six candidates are considered instead of five, which was the case when this ATMVP mode was disabled. It is understood that according to one embodiment of the present invention, the ATMVP can be signaled as an affine merge candidate (e.g., an ATMVP candidate) inserted into a list of affine merge candidates (i.e., a separate list or set of candidates for the affine merge mode, described in more detail below).
[0187] Furthermore, when this prediction is enabled at the SPS level, all bins of the merge index (i.e., identifiers or indexes or information for identifying a candidate from a list of merge candidates) are context coded by CABAC. While in HEVC or when ATMVP is not enabled at the SPS level in JEM, only the first bin is context coded and the remaining bins are context bypass coded (i.e., bypass CABAC coded).
[0188] Figure 10(a) shows the encoding of the merge index when ATMVP is not enabled at the SPS level of HEVC or JEM. This corresponds to a unary maximal code. Furthermore, in this Figure 10(a), the first bit is CABAC encoded and the other bits are bypass CABAC encoded.
[0189] Figure 10(b) shows the encoding of the merge index when ATMVP is enabled at the SPS level. All bits are CABAC encoded (from the 1st to the 5th bits). Note that each bit for encoding the index has its own context, in other words, their probabilities of being used for CABAC encoding are separated.
[0190] Affine Mode In HEVC, we only applied the translational motion model for motion compensated prediction (MCP), whereas in the real world, there are many kinds of motions, such as zoom in / out, rotation, perspective motion, and other irregular motions.
[0191] In JEM we apply a simplified affine transformation motion compensation prediction and the general principles of the affine mode are described below based on an excerpt from the JVET-G1001 document presented at the JVET conference in Torino, July 13-21, 2017. This document is incorporated herein in its entirety by reference insofar as it describes other algorithms used in JEM.
[0192] As shown in FIG. 11(a), the affine motion field of a block in this document is described by two control point motion vectors (it should be understood that other affine models, such as those with more control point motion vectors, may also be used according to embodiments of the present invention).
[0193] The motion vector field (MVF) of a block is described by the following equation:
[0194]
number
[0195] where (v0x, v0y) is the motion vector of the control point in the top-left corner, (v1x, v1y) is the motion vector of the control point in the top-right corner, and w is the width of block Cur (the current block).
[0196] To further simplify the motion compensation prediction, we apply subblock-based affine transformation prediction. The subblock size MxN is derived as Equation 2, where MvPre is the motion vector fractional precision (1 / 16 in JEM), and (v2x, v2y) is the motion vector of the top-left control point calculated according to Equation 1.
[0197]
number
[0198] After being derived by Equation 2, M and N may be adjusted downwards if necessary to be divisors of w and h, respectively, where h is the height of the current block Cur.
[0199] To derive a motion vector for each M × N subblock, the motion vector of the center sample of each subblock is calculated according to Equation 1 and rounded to 1 / 16 fractional precision, as shown in Fig. 11(b). Then, a motion compensated interpolation filter is applied to generate a prediction for each subblock with the derived motion vector.
[0200] Affine mode is a motion compensation mode like inter mode (AMVP, "classical" merge, or "classical" merge skip). Its principle is to generate one motion information per pixel according to two or three neighboring motion information. In JEM, affine mode derives one motion information for each 4x4 block as shown in FIG. 11(b) (each square is a 4x4 block, and the whole block in FIG. 11(b) is a 16x16 block divided into 16 such square blocks of 4x4 size, and each 4x4 square block has a motion vector associated with it). It should be understood that in the embodiment of the present invention, affine mode can drive one motion information for blocks of different sizes or shapes as long as one motion information can be derived.
[0201] According to one embodiment, this mode is made available for AMVP and merge modes (i.e., classical merge mode, also called "non-affine merge mode", and classical merge skip mode, also called "non-affine merge skip mode") by enabling the affine mode with a flag. This flag is CABAC coded. In one embodiment, the context depends on the sum of the affine flags of the left block (position A2 in Fig. 6b) and the top-left block (position B3 in Fig. 6b).
[0202] Thus, in JEM, three context variables (0, 1, or 2) are possible for the affine flag given in the later formula.
[0203] Ctx = IsAffine(A2) + IsAffine(B3) Here, IsAffine(block) is a function that returns 0 if the block is not an affine block and returns 1 if the block is affine.
[0204] Affine Merge Candidate Deriving In JEM, the affine merge mode (or affine merge skip mode), also known as the subblock (merge) mode, derives motion information for the current block from the first neighbouring block that is affine among the blocks at positions A1, B1, B0, A0, B2 (i.e., the first neighbouring block coded using the affine mode). These positions are shown in Fig. 6a and 6b. However, how the affine parameters are derived is not fully defined, and the present invention aims to improve at least this aspect, for example by defining affine parameters for the affine merge mode, thereby allowing the selection of a wider selection for affine merge candidates (i.e., not only the first neighbouring block that is affine, but at least one other candidate is available for selection using an identifier such as an index).
[0205] For example, according to some embodiments of the present invention, an affine merge mode having its own list of affine merge candidates (candidates for deriving / obtaining motion information for the affine mode) and an affine merge index (for identifying one affine merge candidate from the list of affine merge candidates) are used to encode or decode a block.
[0206] Affine merge signaling Figure 12 is a flow chart of a partial decoding process for some syntax elements related to coding modes for signaling the use of affine merge mode. In this figure, the skip flag (1201), prediction mode (1211), merge flag (1203), merge index (1208), and affine flag (1206) can be decoded.
[0207] For all CUs in the inter slice, the skip flag is decoded (1201). If the CU is not a skip (1202), the prediction mode (PREDICTION_MODE) is decoded (1211). This syntax element indicates whether the current CU is coded (decoded) in inter or intra mode. Note that if the CU is a skip (1202), its current mode is inter mode. If the CU is not a skip (1202: No), the CU is coded in AMVP or merge mode. If the CU is inter (1212), the merge flag is decoded (1203). If the CU is a merge (1204) or if the CU is a skip (1202: Yes), it is verified / checked (1205) whether the affine flag (1206) needs to be decoded, i.e., a determination is made in (1205) whether the current CU could have been coded in affine mode. This flag is decoded if the current CU is a 2Nx2N CU, which means that in the current VVC, the height and width of the CU are equal. Furthermore, at least one neighboring CU A1 or B1 or B0 or A0 or B2 must be coded in affine mode (either affine merge mode or AMVP mode with affine mode enabled). Finally, the current CU is not a 4x4 CU, and by default CU 4x4 is disabled in the VTM reference software. If this condition (1205) is false, it is certain that the current CU is coded in classical merge mode (or classical merge skip mode) as specified in HEVC, and the merge index is decoded (1208). If the affine flag (1206) is set equal to 1 (1207), then the CU is a merge affine CU (i.e., a CU coded in affine merge mode) or a merge skip affine CU (i.e., a CU coded in affine merge skip mode) and the merge index (1208) does not need to be decoded (because affine merge mode is used, i.e., the CU is decoded using affine mode with the first neighboring block being affine).Otherwise, the current CU is a classical (basic) merge or merge skip CU (i.e., a CU encoded in classical merge or merge skip mode) and the merge candidate index (1208) is decoded.
[0208] Merge candidate derivation FIG. 13 is a flow chart illustrating a merge candidate (i.e., a candidate for classical merge mode or classical merge skip mode) derivation according to one embodiment. This derivation builds on the merge mode motion vector derivation process (i.e., the merge candidate list derivation for HEVC) shown in FIG. 8. The main changes compared to HEVC are the addition of ATMVP candidates (1319, 1321, 1323), the full overlap check of candidates (1325), and the new order of candidates. ATMVP prediction is set as a dedicated candidate since it represents some motion information of the current CU. The value of the first sub-block (top left) is compared with the temporal candidate, and the temporal candidate is not added to the list of merge candidates if they are equal (1320). The ATMVP candidate is not compared with other spatial candidates. This is in contrast to the temporal candidate, which is compared with each spatial candidate already in the list (1325) and is not added to the merge candidate list if it is an overlap candidate.
[0209] When a spatial candidate is added to the list, it is compared (1312) with other spatial candidates in the list, which is not the case in the final version of HEVC.
[0210] In the current VTM version, the list of merge candidates is set in the following order, as determined to provide the best results for the coding test criteria: A1 B1 B0 A0 ATMVP B2 · Temporal · combination · Zero_MV It is important to note that spatial candidate B2 is set after the ATMVP candidate.
[0211] Furthermore, when ATMVP is enabled at the slice level, the maximum number in the list of candidates is 6 instead of 5 in HEVC.
[0212] Other Inter Prediction Modes In the first few embodiments (up to the 16th embodiment) described below, the description describes encoding or decoding of indexes for (regular) merge mode and affine merge mode. In addition to (regular) merge mode and affine merge mode, recent versions of the VVC standard under development also consider additional inter prediction modes. Such additional inter prediction modes currently considered are the Multi-Hypothesis Intra Inter (MHII) merge mode, the triangle merge mode, and the merge of motion vector difference (MMVD) merge mode, which are described below.
[0213] According to variations of these first several embodiments, it will be appreciated that one or more of the additional inter prediction modes may be used in addition to or instead of the merge mode or the affine merge mode, and an index (or flag or information) for one or more of the additional inter prediction modes may be signaled (encoded or decoded) using the same techniques as any one of the merge mode or the affine merge mode.
[0214] MHII (Multi-Hypothesis Intra Inter) merge mode The Multi-Hypothesis Intra Inter (MHII) merge mode is a hybrid that combines the regular merge mode and the intra mode. The block predictor for this mode is obtained as the average between the (regular) merge mode block predictor and the intra mode block predictor. The obtained block predictor is added to the residual of the current block to obtain a reconstructed block. To obtain this merge mode block predictor, the MHII merge mode uses the same number of candidates and the same merge candidate derivation process as the merge mode. Therefore, the index signaling for the MHII merge mode can use the same technique as the index signaling for the merge mode. Also, this mode is only enabled for blocks that are coded / decoded in non-skip mode. Therefore, when the current CU is coded / decoded in skip mode, the MHII is not available in the coding / decoding process.
[0215] Triangle Merge Mode The triangle merge mode is a type of bi-predictive mode that uses motion compensation based on triangle shape. Figures 25(a) and 25(b) show different partition configurations used for its block predictor generation. The block predictor is obtained from the first triangle (first block predictor 2501 or 2511) and the second triangle (second block predictor 2502 or 2512) in the block. Two different configurations are used for this block predictor generation. For the first one, the partition / split between the triangular parts / regions (with which the two block predictor candidates are associated) is from the top left corner to the bottom right corner, as shown in Figure 25(a). For the second one, the partition / split between the triangular regions (with which the two block predictor candidates are associated) is from the top right corner to the bottom left corner, as shown in Figure 25(b). Furthermore, samples around the boundary between the triangular regions are filtered with a weighted average, where the weight depends on the sample position (e.g., distance from the boundary). A separate triangle merge candidate list is generated, and index signaling for the triangle merge mode may use techniques modified accordingly relative to the techniques for index signaling in the merge mode or affine merge mode.
[0216] Merge Motion Vector Difference (MMVD) Merge Mode The MMVD merge mode is a special type of regular merge mode candidate derivation that generates an independent MMVD merge candidate list. The selected MMVD merge candidate for the current CU is obtained by adding an offset value to the motion vector component (mvx or mvy) of one of the MMVD merge candidates. The offset value is added to the component of the motion vector from the first list L0 or the second list L1 depending on the configuration of these reference frames (e.g., both backwards, forwards or both forwards and backwards). The selected MMVD merge candidate is signaled using an index. The offset value is signaled using a distance index between eight possible preset distances (1 / 4-pel, 1 / 2-pel, 1-pel, 2-pel, 4-pel, 8-pel, 16-pel, 32-pel) and a direction index giving the x-axis or y-axis and the sign of the offset. Thus, the index signaling for the MMVD merge mode can use the same technique as the index signaling for the merge mode or the affine merge mode.
[0217] Embodiment Embodiments of the present invention will now be described with reference to the remaining figures, it being noted that embodiments may be combined unless otherwise stated, e.g., certain combinations of embodiments may improve coding efficiency at the expense of increased complexity, which may be acceptable in certain use cases.
[0218] First embodiment As mentioned above, in the VTM reference software, the ATMVP is signaled as a merge candidate inserted into the list of merge candidates. The ATMVP can be enabled or disabled for the entire sequence (at the SPS level). When the ATMVP is disabled, the maximum number of merge candidates is 5. When the ATMVP is enabled, the maximum number of merge candidates is increased by 1, from 5 to 6.
[0219] At the encoder, a list of merging candidates is generated using the method of Figure 13. One merging candidate is selected from the list of merging candidates based on, for example, a rate-distortion criterion. The selected merging candidate is signaled to the decoder in the bitstream using a syntax element called a merge index.
[0220] The current VTM reference software encodes the merge index differently depending on whether ATMVP is enabled or disabled.
[0221] Figure 10(a) shows the encoding of merge indexes when ATMVP is not enabled at the SPS level. The five merge candidates Cand0, Cand1, Cand2, Cand3, and Cand4 are encoded as 0, 10, 110, 1110, and 1111, respectively. This corresponds to unary maximal encoding. Furthermore, the 1st bit is encoded by CABAC using a single context, and the other bits are bypass encoded.
[0222] Figure 10(b) shows the encoding of the merge index when ATMVP is enabled. Six merge candidates Cand0, Cand1, Cand2, Cand3, Cand4, and Cand5 are encoded as 0, 10, 110, 1110, 11110, and 11111, respectively. In this case, all bits of the merge index (from the 1st to the 5th bit) are context coded by CABAC. Each bit has its own context, and there are separate probability models for different bits.
[0223] In the first embodiment of the present invention, as shown in FIG. 14, if the list of merge candidates includes ATMVP as a merge candidate (e.g., ATMVP is enabled at the SPS level), the encoding of the merge index is modified so that only the first bit of the merge index is encoded by CABAC using a single context. The context is set in the same way as the current VTM reference software when ATMVP is not enabled at the SPS level, i.e., the other bits (2nd to 5th) are bypass encoded. If the list of merge candidates does not include ATMVP as a merge candidate (e.g., ATMVP is disabled at the SPS level), there are five merge candidates. Only the first bit of the merge index is encoded by CABAC using a single context. The context is set in the same way as the current VTM reference software when ATMVP is not enabled at the SPS level. The other bits (2nd to 4th) are bypass decoded.
[0224] The decoder generates the same list of merge candidates as the encoder. This can be achieved by using the method of Figure 13. If ATMVP is not included as a merge candidate in the list of merge candidates (e.g., ATMVP is disabled at the SPS level), there are five merge candidates. Only the first bit of the merge index is decoded by CABAC using a single context. The other bits (2nd to 4th bits) are bypass decoded. In contrast to the current reference software, if ATMVP is included as a merge candidate in the list of merge candidates (e.g., ATMVP is enabled at the SPS level), only the first bit of the merge index is decoded by CABAC using a single context in the decoding of the merge index. The other bits (2nd to 5th bits) are bypass decoded. The decoded merge index is used to identify the merge candidate selected by the encoder from the list of merge candidates.
[0225] The advantage of this embodiment compared to the VTM2.0 reference software is that the complexity of the merge index decoding and decoder design (and encoder design) is reduced without affecting the coding efficiency. Indeed, in this embodiment, only one CABAC state is needed for the merge index instead of five for the current VTM merge index coding / decoding. Furthermore, the other bits are CABAC bypass coded, reducing the number of operations compared to coding all bits with CABAC, thus reducing the worst-case complexity.
[0226] Second embodiment In the second embodiment, all bits of the merge index are CABAC coded, but they all share the same context. In this case, there may be a single context like in the first embodiment, shared between the bits. As a result, if the list of merge candidates includes ATMVP as a merge candidate (e.g., ATMVP is enabled at the SPS level), only one context is used, compared to five in the VTM2.0 reference software. The advantage of this embodiment compared to the VTM2.0 reference software is that the complexity of the merge index decoding and the decoder design (and the encoder design) is reduced, without affecting the coding efficiency.
[0227] Alternatively, as described below in connection with the third to sixteenth embodiments, context variables may be shared between bits such that more than one context is available but the current context is shared by the bits.
[0228] When ATMVP is disabled, the same context is still used for all bits.
[0229] This and all subsequent embodiments are applicable even if ATMVP is not an available mode or is disabled.
[0230] In a variation of the second embodiment, any two or more bits of the merge index are CABAC coded and share the same context. Other bits of the merge index are bypass coded. For example, the first N bits of the merge index may be CABAC coded, where N is 2 or more.
[0231] Third embodiment In the first embodiment, the first bit of the merge index was CABAC coded using a single context.
[0232] In a third embodiment, the context variable of a bit of a merge index depends on the value of the merge index of the neighboring block, which allows multiple contexts for a target bit, where each context corresponds to a different value of the context variable.
[0233] A neighboring block may be any block that has already been decoded so that its merge index is available to the decoder by the time the current block is decoded. For example, a neighboring block may be any of blocks A0, A1, A2, B0, B1, B2, and B3 shown in Figure 6b.
[0234] In a first variant, only the first bit is CABAC coded using this context variable.
[0235] In a second variant, the first N bits of the merge index, where N is 2 or more, are CABAC coded and the context variables are shared among these N bits.
[0236] In a third variant, any N bits of the merge index, where N is 2 or more, are CABAC coded and the context variables are shared among these N bits.
[0237] In a fourth variant, the first N bits of the merge index, where N is 2 or more, are CABAC coded, and N context variables are used for these N bits. Assuming that the context variable has K values, KxN CABAC states are used. For example, in this embodiment, with one adjacent block, the context variable can conveniently have two values, e.g., 0 and 1. In other words, 2N CABAC states are used.
[0238] In a fifth variant, any N bits of the merge index, where N is 2 or more, are adaptive PM coded and N context variables are used for these N bits.
[0239] Similar modifications are applicable to the fourth to sixteenth embodiments described below.
[0240] Fourth embodiment In the fourth embodiment, the context variable of a bit of a merge index depends on the respective values of the merge index of two or more neighboring blocks. For example, the first neighboring block is a left block A0, A1 or A2, and the second neighboring block is a top block B0, B1, B2 or B3. The manner of combining two or more merge index values is not particularly limited. An example is shown below.
[0241] The context variables can conveniently have three different values in this case, for example 0, 1 and 2, since there are two adjacent blocks. Therefore, if the fourth variant described in relation to the third embodiment is applied to this embodiment with three different values, K is 3 instead of 2. In other words, 3N CABAC states are used.
[0242] Fifth embodiment In the fifth embodiment, the context variable of a bit of a merge index depends on the respective values of the merge indexes of neighboring blocks A2 and B3.
[0243] Sixth embodiment In the sixth embodiment, the context variables of the bits of the merge index depend on the respective values of the merge indexes of the neighboring blocks A1 and B1. The advantage of this variant is the alignment with the merge candidate derivation. As a result, some decoder and encoder implementations can achieve a reduction in memory accesses.
[0244] Seventh embodiment In the seventh embodiment, the context variable of the bit having the bit position idx_num in the merge index of the current block is obtained according to the following formula:
[0245] ctxIdx=(Merge_index_left==idx_num)+(Merge_index_up==idx_num) Here, Merge_index_left is the merge index of the left block, Merge_index_up is the merge index of the top block, and the symbol == is the equality symbol.
[0246] For example, if there are 6 merge candidates, then 0<=idx_num<=5.
[0247] The left block is block A1 and the top block is block B1 (as in the sixth embodiment), or the left block may be block A2 and the top block may be block B3 (as in the fifth embodiment).
[0248] If the merge index of the left block is equal to idx_num, then the formula (Merge_index_left==idx_num) is equal to 1. The following table shows the result of this formula (Merge_index_left==idx_num).
[0249] [Table 2]
[0250] Of course, the formula table (Merge_index_up==idx_num) is the same.
[0251] The following table shows the unary maximum code for each merge index value and the relative bit position of each bit: This table corresponds to Figure 10(b).
[0252] [Table 3]
[0253] If the left block is not a merge block or an affine merge block (i.e., it is coded using the affine merge mode), then the left block is considered unavailable. Similar conditions apply for the top block.
[0254] For example, if only the first bit is CABAC coded, the context variable ctxIdx is 0 if the top-left block does not have a merge index, or if the left block merge index is not the first index (i.e., not 0), and if the top block merge index is not the first index (i.e., not 0) 1 if one of the left and top blocks but not the other has a merge index equal to the first index 2 if the merge index is equal to the first index for each of the left and top blocks is set equal to
[0255] More generally, for a target bit at CABAC encoded position idx_num, the context variable ctxIdx is 0 if the top-left block has no merge index, or if the left block merge index is not the ith index (if i=idx_num), and if the top block merge index is not the ith index 1 if one of the left and top blocks but not the other has a merge index equal to the i-th index 2 if the merge index is equal to the i-th index for each of the left and top blocks where the i-th index means the first index when i=0, the second index when i=1, and so on.
[0256] Eighth embodiment In the eighth embodiment, the context variable of the bit having the bit position idx_num in the merge index of the current block is obtained according to the following formula:
[0257] Ctx=(Merge_index_left>idx_num)+(Merge_index_up>idx_num), where Merge_index_left is the merge index of the left block, Merge_index_up is the merge index of the top block, and the symbol > means "greater than".
[0258] For example, if there are 6 merge candidates, then 0<=idx_num<=5.
[0259] The left block is block A1 and the top block is block B1 (as in the sixth embodiment), or the left block may be block A2 and the top block may be block B3 (as in the fifth embodiment).
[0260] If the merge index of the left block is greater than idx_num, then the formula (Merge_index_left>idx_num) is equal to 1. If the left block is not a merge block or an affine merge block (i.e., it is coded using the affine merge mode), then the left block is considered unavailable. Similar conditions apply for the top block.
[0261] The following table shows the result of this formula (Merge_index_left>idx_num).
[0262] [Table 4]
[0263] For example, if only the first bit is CABAC coded, the context variable ctxIdx is 0 if the top-left block has no merge index, or if the left block merge index is less than or equal to the first index (i.e., not 0), and if the top block merge index is less than or equal to the first index (i.e., not 0). 1 if one of the left and top blocks but not the other has a merge index greater than the first index 2 if the merge index is greater than the first index for each of the left and top blocks is set equal to
[0264] More generally, for a target bit at CABAC encoded position idx_num, the context variable ctxIdx is 0 if the top-left block has no merge index, or if the left block merge index is less than the i-th index (for i=idx_num), and if the top block merge index is less than or equal to the i-th index 1 if one of the left and top blocks but not the other has a merge index greater than the i-th index 2 if the merge index is greater than the i-th index for each of the left and top blocks is set equal to
[0265] The eighth embodiment further improves the coding efficiency compared to the seventh embodiment.
[0266] Ninth embodiment In the fourth to eighth embodiments, the context variables of the bits of the merge index of the current block depend on the respective values of the merge indexes of two or more adjacent blocks.
[0267] In the ninth embodiment, the context variable of the bit of the merge index of the current block depends on the merge flags of each of two or more neighboring blocks. For example, the first neighboring block is the left block A0, A1 or A2, and the second neighboring block is the top block B0, B1, B2 or B3.
[0268] The merge flag is set to 1 if the block is coded using the merge mode, and to 0 if other modes such as skip mode or affine merge mode are used. Note that in VMT2.0, affine merge is a separate mode from the basic or "classical" merge modes. Affine merge mode can be signaled using a dedicated affine flag. Alternatively, the list of merge candidates may contain affine merge candidates, in which case the affine merge mode may be selected and signaled using the merge index.
[0269] The context variable is then 0 if neither the left adjacent block nor the top adjacent block has its merge flag set to 1 1 if one of the left and top neighbors but not the other has its merge flag set to 1 2 if each of the left and top neighboring blocks has its merge flag set to 1 is set to.
[0270] This simple evaluation achieves improved coding efficiency with respect to VTM 2.0. Another advantage is the lower complexity compared to the seventh and eighth embodiments, since only the merge flags need to be checked, and not the merge indices of the neighboring blocks.
[0271] In a variant, the context variable of the merge index bit of the current block depends on the merge flag of a single neighboring block.
[0272] Tenth embodiment In the third to ninth embodiments, the context variable of the merge index bit of the current block depended on the merge index values or merge flags of one or more neighboring blocks.
[0273] In a tenth embodiment, the context variable of the merge index bit of the current block depends on the value of the skip flag for the current block (current coding unit, or CU). The skip flag is equal to 1 if the current block uses merge skip mode, and is equal to 0 otherwise.
[0274] The skip flag is the first example of another variable or syntax element that has already been decoded or parsed for the current block. This other variable or syntax element is preferably an indicator of the complexity of the motion information in the current block. Since the occurrence of the merge index value depends on the complexity of the motion information, a variable or syntax element such as the skip flag generally correlates with the merge index value.
[0275] More specifically, the merge skip mode is generally selected for still scenes or scenes with constant motion. As a result, the merge index value is generally lower in the merge skip mode than in the classical merge mode used to code the inter prediction including block residuals. This generally occurs for more complex motion. However, the choice between these modes is often also related to the quantization and / or RD criteria.
[0276] This simple evaluation improves coding efficiency over VTM2.0, and is very easy to implement since it does not involve checking neighboring blocks or merge index values.
[0277] In a first variant, the context variable of the bit of the merge index of the current block is simply set equal to the skip flag of the current block. The bit can be only the first bit. The other bits are bypass coded as in the first embodiment.
[0278] In the second variant, all bits of the merge index are CABAC coded, each of them having its own context variable depending on the merge flag, which requires 10 probability states when there are 5 CABAC coded bits in the merge index (corresponding to 6 merge candidates).
[0279] In a third variant, to limit the number of states, only N bits of the merge index are CABAC coded, where N is 2 or more, e.g., the first N bits. This requires 2N states. For example, if the first two bits are CABAC coded, four states are required.
[0280] In general, instead of the skip flag, it is possible to use any other variable or syntax element that has already been decoded or parsed for the current block and that is an indicator of the complexity of the motion information in the current block.
[0281] Eleventh embodiment The eleventh embodiment relates to affine merge signaling as described above with reference to FIGS. 11(a), 11(b) and 12.
[0282] In an eleventh embodiment, the context variable of the CABAC coded bit of the merge index of the current block (current CU), if any, depends on the affine merge candidate in the list of merge candidates. This bit can be only the first bit of the merge index, or the first N bits, where N is 2 or more, or any N bits. The other bits are bypass coded.
[0283] Affine prediction is designed to compensate for complex motion. Therefore, for complex motion, the merge index generally has a higher value than for less complex motion. As a result, if the first Affine Merge candidate is far down the list, or if there is no Affine Merge candidate at all, the merge index of the current CU may have a small value.
[0284] Therefore, the context variables may usefully depend on the presence and / or position of at least one affine merge candidate in the list.
[0285] For example, the context variable is set equal to 1 if A1 is affine, 2 if B1 is affine, 3 if B0 is affine, 4 if A0 is affine, 5 if B2 is affine, and 0 if the neighboring block is not affine.
[0286] When the merge index of the current block is decoded or parsed, the affine flags of the merge candidates at these locations are already checked, so no further memory accesses are required to derive the context of the merge index of the current block.
[0287] This embodiment improves coding efficiency over VTM 2.0: since step 1205 already involves checking the neighboring CU affine modes, no additional memory accesses are required.
[0288] In a first variant, in order to limit the number of states, the context variables are: It is set equal to 0 if the neighboring block is not affine or if A1 or B1 is affine, and 1 if B0, A0, or B2 are affine.
[0289] In a second variant, to limit the number of states, the context variables are: It is set equal to 0 if the neighboring block is not affine, 1 if A1 or B1 are affine, and 2 if B0, A0, or B2 are affine.
[0290] In a third variant, the context variables are: It is set equal to 1 if A1 is affine, 2 if B1 is affine, 3 if B0 is affine, 4 if A0 or B2 are affine, and 0 if the neighboring block is not affine.
[0291] Note that these locations are already checked when the merge index is decoded or parsed, since the affine flag decoding depends on these locations, so no additional memory accesses are required to derive the merge index context that is coded after the affine flag.
[0292] Twelfth embodiment In a twelfth embodiment, signaling the affine mode includes inserting the affine mode as a candidate motion predictor.
[0293] In one example of the twelfth embodiment, affine merge (and affine merge skip) is signaled as a merge candidate (i.e., as one of the merge candidates for use with classical merge mode or classical merge skip mode). In this case, modules 1205, 1206, and 1207 of FIG. 12 are removed. Furthermore, the maximum possible number of merge candidates is incremented so as not to affect the coding efficiency of the merge mode. For example, in the current VTM version, this value is set equal to 6, and therefore, when applying this embodiment to the current version of VTM, the value becomes 7.
[0294] The advantage is a simplified design of syntax elements for merge mode since fewer syntax elements need to be decoded. In some situations, an improvement / change in coding efficiency may be observed.
[0295] Two possibilities for implementing this example are described below.
[0296] The merge index of an affine merge candidate always has the same position in the list, whatever the values of the other merge MVs. The position of a candidate motion predictor indicates its likelihood of being selected, so the higher it is placed in the list (lower index value), the more likely that motion vector predictor is to be selected.
[0297] In the first example, the merge index of an affine merge candidate always has the same position in the list of merge candidates. This means that it has a fixed "merge idx" value. For example, this value can be set equal to 5, since the affine merge mode should represent complex motion, which is not the most probable content. An additional advantage of this embodiment is that the current block can be set as an affine block when it is parsed (not just decoding / reading syntax elements, but decoding the data itself). As a result, this value can be used to determine the CABAC context of the affine flag used for AMVP. Thus, the conditional probability should be improved for this affine flag, and the coding efficiency should be better.
[0298] In a second example, an affine merge candidate is derived together with other merge candidates. In this example, a new affine merge candidate is added to the list of merge candidates (for classical merge mode or classical merge skip mode). Figure 16 shows this example. Compared to Figure 13, the affine merge candidate is the first affine neighbouring block from A1, B1, B0, A0, and B2 (1917). If the same condition as 1205 in Figure 12 is valid (1927), a motion vector field generated with affine parameters is generated and an affine merge candidate is obtained (1929). The list of initial merge candidates can have 4, 5, 6, or 7 candidates according to the use of ATMVP, temporal and affine merge candidates.
[0299] The order among all these candidates is important as the more likely candidates should be processed first to ensure that they are more likely to make the cut in the motion vector candidates, the preferred order is as follows:
[0300] A1, B1, B0, A0, Affine merge, ATMVP, B2, Temporal, combination, Zero_MV It is important to note that the affine merge candidate is placed before the ATMVP candidate but after the four major neighboring blocks. The advantage of placing the affine merge candidate before the ATMVP candidate is that it increases the coding efficiency compared to placing it after the ATMVP and temporal predictor candidates. This coding efficiency improvement depends on the GOP (group of pictures) structure and the QP (Quantization Parameter) settings of each picture in the GOP. However, for most used GOP and QP settings, this order results in an increase in coding efficiency.
[0301] A further advantage of this solution is the clean design of classical merge and classical merge skip modes (i.e., merge modes with additional candidates such as ATMVP or affine merge candidates) for both syntax and derivation processing. Moreover, the merge index of an affine merge candidate can be modified according to the availability or value (duplicate check) of previous candidates in the list of merge candidates. As a result, an efficient signaling can be obtained.
[0302] In a further example, the merge index for the affine merge candidates is variable according to one or several conditions.
[0303] For example, the merge index or position in the list associated with an affine merge candidate varies according to a criterion: the principle is to set a low value for the merge index corresponding to an affine merge candidate if the probability of the affine merge candidate being selected is high (and a higher value if the probability of selection is low).
[0304] In a twelfth embodiment, the affine merge candidates have a merge index value. To improve the coding efficiency of the merge index, it is effective to make the context variables of the bits of the merge index dependent on the affine flags of the neighboring blocks and / or the current block.
[0305] For example, the context variables may be determined using the following formula:
[0306] ctxIdx=IsAffine(A1)+IsAffine(B1)+IsAffine(B0)+IsAffine(A0)+IsAffine(B2) The resulting context value can have the values 0, 1, 2, 3, 4, or 5.
[0307] The affine flag increases coding efficiency.
[0308] In a first variant, to include fewer neighboring blocks, ctxIdx=IsAffine(A1)+IsAffine(B1). The resulting context value can have the values 0, 1 or 2.
[0309] And in the second variant, to include fewer neighboring blocks, ctxIdx=IsAffine(A2)+IsAffine(B3). Again, the resulting context value can have the values 0, 1 or 2.
[0310] In a third variant, ctxIdx=IsAffine(current block) to avoid including adjacent blocks. The resulting context value can have the value 0 or 1.
[0311] FIG. 15 is a flowchart of a partial decoding process of some syntax elements related to the coding mode according to the third variant. In this figure, the skip flag (1601), the prediction mode (1611), the merge flag (1603), the merge index (1608), and the affine flag (1606) can be decoded. This flowchart is similar to the flowchart of FIG. 12 described above, so a detailed description is omitted. The difference is that the merge index decoding process takes into account the affine flag so that the affine flag decoded before the merge index can be used when obtaining the context variable for the merge index. This is not the case in VTM2.0. In VTM2.0, the affine flag of the current block always has the same value "0", so it cannot be used to obtain the context variable for the merge index.
[0312] Thirteenth embodiment In the tenth embodiment, the context variable for the bit of the merge index of the current block depends on the value of the skip flag for the current block (current coding unit, i.e., CU). In the thirteenth embodiment, instead of directly using the skip flag value to derive the context variable for the target bit of the merge index, the context value for the target bit is derived from the context variable used to code the skip flag of the current CU. This is possible because the skip flag itself is CABAC coded and therefore has a context variable. Preferably, the context variable for the target bit of the merge index of the current CU is set equal to (copied from) the context variable used to code the skip flag of the current CU. The target bit can be the first bit only. The other bits may be bypass coded as in the first embodiment.
[0313] The context variable of the skip flag of the current CU is derived in the manner specified in VTM2.0. The advantage of this embodiment compared to the VTM2.0 reference software is that the complexity of the merge index decoding and decoder design (and encoder design) is reduced without affecting the coding efficiency. Indeed, in this embodiment, only one CABAC state is required at minimum to code the merge index, instead of five for the current VTM merge index coding (encoding / decoding). Furthermore, other bits are CABAC bypass coded, reducing the number of operations compared to coding all bits with CABAC, thus reducing the worst-case complexity.
[0314] Fourteenth embodiment In the thirteenth embodiment, the context variable / value of the target bit is derived from the context variable of the skip flag of the current CU. In the fourteenth embodiment, the context value of the target bit is derived from the context variable of the affine flag of the current CU.
[0315] This is possible because the affine flag itself is CABAC coded and therefore has a context variable. Preferably, the context variable for the target bit of the merge index of the current CU is set equal to (copied from) the context variable for the affine flag of the current CU. The target bit can be the first bit only. The other bits are bypass coded as in the first embodiment.
[0316] The affine flag context variable for the current CU is derived as specified in VTM2.0.
[0317] The advantage of this embodiment compared to the VTM2.0 reference software is that the complexity of the merge index decoding and decoder design (and encoder design) is reduced without affecting the coding efficiency. Indeed, in this embodiment, a minimum of only one CABAC state is required for the merge index, instead of five for the current VTM merge index coding (encoding / decoding). Furthermore, other bits are CABAC bypass coded, reducing the number of operations compared to coding all bits with CABAC, thus reducing the worst-case complexity.
[0318] Fifteenth embodiment In some of the above embodiments, the context variables had more than two values, e.g., three values 0, 1, and 2. However, to reduce complexity and the number of states to be processed, it is possible to limit the number of allowed context variable values to two, e.g., 0 and 1. This can be achieved, for example, by changing any initial context variables that have a value of 2 to 1. In practice, this simplification has no or only limited impact on the coding efficiency.
[0319] Combinations of the embodiment and other embodiments Any two or more of the above embodiments may be combined.
[0320] The above description has focused on encoding and decoding of the merge index. For example, the first embodiment includes generating a list of merge candidates including ATMVP candidates (in the case of classical merge mode or classical merge skip mode, i.e., in the case of non-affine merge mode or non-affine merge skip mode), selecting one of the merge candidates in the list, and generating a merge index of the selected merge candidate using CABAC encoding, where one or more bits of the merge index are bypass CABAC encoded. In principle, the present invention can be applied to modes other than the merge mode (e.g., affine merge mode), including generating a list of motion information predictor candidates (e.g., a list of affine merge candidates or a list of motion vector predictor (MVP) candidates), selecting one of the motion information predictor candidates (e.g., MVP candidates) in the list, and generating an identifier or index of the selected motion information predictor candidate (e.g., the selected affine merge candidate or the selected MVP candidate for predicting the motion vector of the current block) in the list. Thus, the present invention is not limited to the merge mode (i.e., the classical merge mode and the classical merge skip mode), and the index to be encoded or decoded is not limited to the merge index. For example, in the development of VVC, it is contemplated that the techniques of the above-described embodiments may be applied (or extended) to modes other than merge mode, such as the AMVP mode of HEVC, or its equivalent in VVC, or an affine merge mode, etc. The appended claims should be construed accordingly.
[0321] As mentioned above, in the aforementioned embodiments, one or more candidate motion information (e.g., motion vectors) for the affine merge mode (affine merge or affine merge skip mode) and / or one or more affine parameters are obtained from a first neighboring block that is affine coded between spatially adjacent blocks (e.g., positions A1, B1, B0, A0, B2) or temporally related blocks (e.g., the "center" block with its collocated blocks, or its spatial neighbors such as "H"). These positions are illustrated in Figures 6a and 6b. To enable this obtaining (e.g., deriving or sharing or "merging") one or more motion information and / or affine parameters between the current block (or a group of samples / pixel values currently being encoded / decoded, such as the current CU) and an adjacent block (spatially adjacent or temporally related to the current block), one or more affine merge candidates are added to a list of merge candidates (i.e., classical merge mode candidates), so that if a selected merge candidate (signaled using a merge index, e.g., using a syntax element such as "merge_idx" or a functionally equivalent syntax element in HEVC) is an affine merge candidate, the current CU / block is encoded / decoded using the affine merge mode together with the affine merge candidate.
[0322] As described above, such one or more affine merge candidates for obtaining (e.g., deriving or sharing) one or more motion information and / or affine parameters for the affine merge mode may also be signaled using a separate list (or set) of affine merge candidates (which may be the same as or different from the list of merge candidates used for the classical merge mode).
[0323] According to one embodiment of the present invention, when the techniques of the previous embodiments are applied to the affine merge mode, the list of affine merge candidates can be generated using the same technique as the motion vector derivation process for the classical merge mode shown in and described in conjunction with Figure 8, or using the same technique as the merge candidate derivation process shown in and described in conjunction with Figure 13. The advantage of sharing the same technique for generating / compiling this list of affine merge candidates (for the affine merge mode or the affine merge skip mode) and the list of merge candidates (for the classical merge mode or the classical merge skip mode) is that the complexity of the encoding / decoding process is reduced compared to having separate techniques.
[0324] It should be understood that, according to other embodiments, similar techniques are applied to other inter-prediction modes that require signaling a selected motion information predictor (from multiple candidates) to achieve similar advantages.
[0325] According to another embodiment, a separate technique can be used to generate / compile the list of affine merge candidates, as described below in conjunction with FIG.
[0326] FIG. 24 is a flow chart showing an affine merge candidate derivation process for affine merge modes (affine merge mode and affine merge skip mode). In the first step of the derivation process, five block positions are considered (2401-2405) to obtain / derive spatial affine merge candidates 2413. These positions are the spatial positions shown in FIG. 6a (and FIG. 6b) with reference numbers A1, B1, B0, A0, and B2. In the next step, the availability of spatial motion vectors is checked to determine whether each of the inter mode coded blocks associated with each position A1, B1, B0, A0, and B2 is coded in affine mode (e.g., using any one of affine merge, affine merge skip, or affine AMVP modes) (2410). At most five motion vectors (i.e., spatial affine merge candidates) are selected / obtained / derived. A predictor is considered available if a predictor exists (e.g., there is information to obtain / derive a motion vector associated with that position), and if the block is not intra-coded, and if the block is affine (i.e., coded using affine mode).
[0327] Then, for each available block position, affine motion information is derived / obtained (2411) (2410). This derivation is performed for the current block based on the affine model of the block position (and its affine model parameters, e.g., as described in connection with Figures 11(a) and 11(b)). Then, a pruning process (2412) is applied to remove candidates that give the same affine motion compensation (or have the same affine model parameters) as each other that were previously added to the list.
[0328] At the end of this stage, the list of spatially affine merge candidates contains up to five candidates.
[0329] If the number of candidates (Nb_Cand) is strictly less than the maximum number of candidates (2426) (where Max_Cand is the value signaled in the bitstream slice header, which is equal to 5 for affine merge mode but may differ / variable depending on the implementation).
[0330] Next, constructed affine merge candidates (i.e., additional affine merge candidates generated to provide some diversity as well as to get closer to the target number, e.g., playing a role similar to the combined bi-predictive merge candidates in HEVC) are generated (2428). These constructed affine merge candidates are based on motion vectors related to neighboring spatial and temporal positions of the current block. First, control points are defined (2418, 2419, 2420, 2421) to generate motion information for generating the affine model. Two of these control points correspond to v0 and v1 in, e.g., Figures 11(a) and 11(b). These four control points correspond to the four corners of the current block.
[0331] The motion information for the top left of the control point (2418) is obtained from (e.g., by making it equal to) the motion information for the block position at position B2 (2405) if it exists and this block is coded in inter mode (2414). Otherwise, the motion information for the top left of the control point (2418) is obtained from (e.g., by making it equal to) the motion information for the block position at position B3 (2406) if it exists and this block is coded in inter mode (2414) and if this is not the case (as shown in Fig. 6b), the motion information for the top left of the control point (2418) is obtained from (e.g., by making it equal to) the motion information for the block position at position A2 (2407) if it exists and this block is coded in inter mode (2414). If there is no block available for this control point, it is considered unavailable (unavailable).
[0332] The motion information for the control point top right (2419) is obtained from (e.g. is equal to) the motion information for the block location at position B1 (2402) if it exists and this block is coded in inter mode (2415). Otherwise, the motion information for the control point top right (2419) is taken from (e.g. is equal to) the motion information for the block location at position B0 (2403) if it exists and this block is coded in inter mode (2415). If there is no block available for this control point, it is considered unavailable (not available).
[0333] The motion information for the control point bottom left (2420) is obtained (e.g., equal to) the motion information for the block position at position A1 (2401) if it exists and this block is coded in inter mode (2416). Otherwise, the motion information for the control point bottom left (2420) is obtained (e.g., equal to) the motion information for the block position at position A0 (2404) if it exists and this block is coded in inter mode (2416). If there is no block available for this control point, it is considered unavailable (not available).
[0334] The motion information of the control point bottom right (2421) is obtained (e.g., equal) from the motion information of a temporal candidate, e.g., the co-located block position (as shown in FIG. 6a) at position H (2408), if it exists and this block is coded in inter mode (2417). If there is no block available at this control point, it is considered unavailable (unavailable).
[0335] Based on these control points, up to ten constructed affine merge candidates may be generated (2428). These candidates are generated based on affine models having four, three, or two control points. For example, a first constructed affine merge candidate may be generated using four control points. Then, the following four constructed affine merge candidates are four possibilities that can be generated using four different sets of three control points (i.e., four different possible combinations of sets that include three of the four available control points). Then, the other constructed affine merge candidates are those that are generated using different sets of two control points (i.e., different possible combinations of sets that include two of the four control points).
[0336] If the number of candidates (Nb_Cand) remains strictly less than the maximum number of candidates (Max_Cand) after adding these additional (constructed) affine merge candidates (2430), other additional virtual motion information candidates, such as zero motion vector candidates (or combined bi-predictive merge candidates, if applicable), are added / generated (2432) until the number of candidates in the list of affine merge candidates reaches a target number (e.g., the maximum number of candidates).
[0337] At the end of this process, a list or set of affine merge mode candidates (i.e., a list or set of affine merge mode candidates that are affine merge mode and affine merge skip mode) is generated / constructed (2434). As shown in FIG. 24, a list or set of affine merge (motion vector predictor) candidates is constructed / generated (2434) from a subset of spatial candidates (2401-2407) and temporal candidates (2408). It should be understood that, according to an embodiment of the present invention, other affine merge candidate derivation processes having different orders for checking availability, pruning processes, or number / type of potential candidates (e.g., ATMVP candidates can also be added in a manner similar to the merge candidate list derivation process of FIG. 13 or FIG. 16) can also be used to generate the list / set of affine merge candidates.
[0338] The following embodiments show how a list (or set) of affine merge candidates can be used to signal (e.g., encode or decode) a selected affine merge candidate (which can be signaled using the merge index used for the merge mode, or a separate affine merge index used specifically for the affine merge mode).
[0339] In the following embodiments, a merge mode (i.e., a merge mode other than the affine merge mode defined later, in other words, a classical non-affine merge mode or a classical non-affine merge skip mode) is a type of merge mode where motion information of either a spatially neighboring block or a temporally related block is obtained for (or derived for or shared with) the current block; a merge mode predictor candidate (i.e., merge candidate) is information about one or more spatially neighboring or temporally related blocks from which the current block can obtain / derive motion information in the merge mode; a merge mode predictor is a selected merge mode predictor candidate, the information being used when predicting the motion information of the current block and during signaling in a merge mode (e.g., encoding or decoding) process; and an index (e.g., merge index) identifying a merge mode predictor from a list (or set) of merge mode predictor candidates is signaled, where affine merge mode is a type of merge mode in which motion information of either spatially adjacent or temporally related blocks is obtained for the current block (derived for the current block or shared with the current block) such that the motion information of the current block and / or affine parameters for affine mode processing (or affine motion model processing) can use this obtained / derived / shared motion information, affine merge mode predictor candidate (i.e., affine merge candidate) is information about one or more spatially adjacent or temporally related blocks from which the current block can obtain / derive motion information in the affine merge mode, affine merge mode predictor is a selected affine merge mode predictor candidate, which information is usable in the affine motion model when predicting the motion information of the current block and during signaling in the affine merge mode (e.g., encoding or decoding) processing,An index (e.g., an affine merge index) is signaled that identifies an affine merge mode predictor from a list (or set) of affine merge mode predictor candidates. In the following embodiments, it will be understood that an affine merge mode is a merge mode that has its own affine merge index (an identifier that is a variable) to identify one affine merge mode predictor candidate from a list / set of candidates (also known as an "affine merge list" or "sub-block merge list"), and has a single index value associated with it, whereas an affine merge index is signaled to identify that particular affine merge mode predictor candidate.
[0340] In the following embodiments, "merge mode" refers to either the classical merge skip mode in HEVC / JEM / VTM or any one of the classical merge modes or any functionally equivalent modes, provided that the motion information acquisition (e.g., derivation or sharing) and merge index signaling as described above are used in said modes. It should be understood that "affine merge mode" also refers to either the affine merge mode or the affine merge skip mode (using such acquisition / derivation, if present), or any other functionally equivalent modes, provided that the same features are used in said modes.
[0341] Sixteenth embodiment In a sixteenth embodiment, a motion information predictor index for identifying an affine merge mode predictor (candidate) from a list of affine merge candidates is signaled using CABAC encoding, and one or more bits of the motion information predictor index are bypass CABAC encoded.
[0342] According to a first variant of the embodiment, in an encoder, a motion information predictor index for an affine merge mode is encoded by generating a list of motion information predictor candidates, selecting one of the motion information predictor candidates in the list as an affine merge mode predictor, and generating a motion information predictor index for the selected motion information predictor candidate using CABAC encoding, where one or more bits of the motion information predictor index are bypass CABAC encoded. Then, data indicating an index for the selected motion information predictor candidate is included in a bitstream. Then, a decoder generates a list of motion information predictor candidates from the bitstream including this data, decodes the motion information predictor index using CABAC decoding, where one or more bits of the motion information predictor index are bypass CABAC decoded, and decodes a motion information predictor index for an affine merge mode by using the decoded motion information predictor index to identify one of the motion information predictor candidates in the list as an affine merge mode predictor when the affine merge mode is used.
[0343] According to a further variant of the first variant, one or more of the motion information predictor candidates in the list are also selectable as merge mode predictors when the merge mode is used, such that the decoder can use the decoded motion information predictor index (e.g., merge index) to identify one of the motion information predictor candidates in the list as a merge mode predictor when the merge mode is used. In this further variant, an affine merge index is used to signal the affine merge mode predictor (candidate), and signaling the affine merge index is implemented using merge index signaling according to any one of the first to fifteenth embodiments, or index signaling similar to merge index signaling used in current VTM or HEVC.
[0344] In this variant, when a merge mode is used, signaling the merge index can be implemented using the merge index signaling according to any one of the first to fifteenth embodiments, or the merge index signaling used in current VTM or HEVC. In this variant, different index signaling schemes can be used for signaling the affine merge index and for signaling the merge index. The advantage of this variant is to achieve better coding efficiency by using efficient index coding / signaling for both affine merge mode and merge mode. Furthermore, in this variant, separate syntax elements can be used for the merge index (such as "Merge_idx[][]" in HEVC or its functional equivalent) and the affine merge index (such as "A_Merge_idx[][]"). This allows the merge index and the affine merge index to be signaled (encoded / decoded) separately.
[0345] According to yet another further variant, when the merge mode is used and one of the motion information predictor candidates in the list is also selectable as a merge mode predictor, the CABAC encoding uses the same context variable for at least one bit of the motion information predictor index (e.g., merge index or affine merge index) of the current block for both modes, i.e., when the affine merge mode is used and when the merge mode is used, so that the affine merge index and at least one bit of the merge index share the same context variable. Then, when the merge mode is used, the decoder uses the decoded motion information predictor index to identify one of the motion information predictor candidates in the list as a merge mode predictor, and the CABAC decoding uses the same context variable for at least one bit of the motion information predictor index of the current block for both modes, i.e., when the affine merge mode is used and when the merge mode is used.
[0346] According to a second variant of the embodiment, in an encoder, the motion information predictor index is encoded by generating a list of motion information predictor candidates, selecting one of the motion information predictor candidates in the list as an affine merge mode predictor when an affine merge mode is used, selecting one of the motion information predictor candidates in the list as a merge mode predictor when a merge mode is used, and generating a motion information predictor index for the selected motion information predictor candidate using CABAC encoding, and one or more bits of the motion information predictor index are bypass CABAC encoded. Then, data indicating an index for this selected motion information predictor candidate is included in the bitstream. The decoder then decodes the motion information predictor index from the bitstream by generating a list of motion information predictor candidates, decoding the motion information predictor index using CABAC decoding, where one or more bits of the motion information predictor index are bypass CABAC decoded, and if an affine merge mode is used, using the decoded motion information predictor index to identify one of the motion information predictor candidates in the list as an affine merge mode predictor, and if a merge mode is used, using the decoded motion information predictor index to identify one of the motion information predictor candidates in the list as a merge mode predictor.
[0347] According to a further variant of the second variant, the affine merge index signaling and the merge index signaling use the same index signaling scheme according to any one of the first to fifteenth embodiments, or the merge index signaling used in the current VTM or HEVC. An advantage of this further variant is a simple design in implementation, which may also lead to less complexity. In this variant, when the affine merge mode is used, the CABAC encoding of the encoder includes using a context variable for at least one bit of the motion information predictor index (affine merge index) of the current block, the context variable being separable from another context variable for at least one bit of the motion information predictor index (merge index) when the merge mode is used, and data indicating the use of the affine merge mode is included in the bitstream such that the context variables for the affine merge mode and the merge mode can be distinguished (clearly identified) for the CABAC decoding process. The decoder then obtains data from the bitstream to indicate the use of affine merge mode in the bitstream, and the CABAC decoding uses this data to distinguish between the affine merge index and the context variables for the merge index when the affine merge mode is used. Furthermore, at the decoder, the data to indicate the use of affine merge mode can also be used to generate a list (or set) of affine merge mode predictor candidates when the obtained data indicates the use of affine merge mode, and to generate a list (or set) of merge mode predictor candidates when the obtained data indicates the use of merge mode.
[0348] This variant allows both the merge index and the affine merge index to be signaled using the same index signaling scheme, while the merge index and the affine merge index are still encoded / decoded independently of each other (e.g., by using separate context variables).
[0349] One way to use the same index signaling scheme is to use the same syntax element for both affine merge index and merge index, i.e., when the affine merge mode and when the merge mode are used, the motion information predictor index of the selected motion information predictor candidate is coded using the same syntax element in both cases. Then, at the decoder, the motion information predictor index is decoded by parsing the same syntax element from the bitstream, regardless of whether the current block was coded (and decoded) using the affine merge mode or the merge mode.
[0350] Figure 22 shows partial decoding processing of some syntax elements related to coding modes (i.e., the same index signaling scheme) according to this variant of the sixteenth embodiment. This figure shows signaling of affine merge index (2255-"merge idx affine") for affine merge mode (2257: Yes) and merge index (2258-"merge idx") for merge mode (2257: No) with the same index signaling scheme. It should be understood that in some variants, the affine merge candidate list can include ATMVP candidates as the merge candidate list of the current VTM. The encoding of the affine merge index is similar to the encoding of the merge index of the merge mode as shown in Figure 10(a), Figure 10(b), or Figure 14. In some variations, even if no ATMVP merge candidates are defined in the affine merge candidate derivation, if ATMVP is enabled for merge mode with up to five other candidates (i.e., six candidates total) such that the maximum number of candidates in the affine merge candidate list matches the maximum number of candidates in the merge candidate list, the affine merge index is encoded as described in Figure 10(b). Thus, each bit of the affine merge index has its own context. All context variables used for the bits of the merge index signaling are independent of the context variables used for the bits of the affine merge index signaling.
[0351] According to a further variant, this same index signaling scheme shared by merge index and affine merge index signaling uses CABAC coding for only the first bin, as in the first embodiment. That is, all bits except the first bit of the motion information predictor index are bypass CABAC coded. In this further variant of the sixteenth embodiment, if ATMVP is included as a candidate in one of the lists of merge candidates or affine merge candidates (e.g., if ATMVP is enabled at the SPS level), the coding of each index (i.e., merge index or affine merge index) is modified so that only the first bit of the index is coded by CABAC using a single context variable as shown in FIG. 14. This single context is set in the same way as the current VTM reference software when ATMVP is not enabled at the SPS level. The other bits (the second to fifth bits or the fourth bit if there are only five candidates in the list) are bypass coded. If the merge candidate list does not include ATMVP as a candidate (e.g., ATMVP is disabled at the SPS level), there are five merge candidates and five affine merge candidates available. Only the first bit of the merge index for the merge mode is encoded by CABAC using the first single context variable. And only the first bit of the affine merge index for the affine merge mode is encoded by CABAC using the second single context variable. These first and second context variables are set in the same way as the current VTM reference software does when ATMVP is not enabled at the SPS level for both merge index and affine merge index. The other bits (second through fourth bits) are bypass decoded.
[0352] The decoder generates the same list of merge candidates and the same list of affine merge candidates as the encoder. This is accomplished, for example, by using the method of FIG. 24. The same index signaling scheme is used for both merge and affine merge modes, but the affine flag (2256) is used to determine whether the currently decoded data is for a merge index or an affine merge index, so that the first and second context variables are separable (or distinguishable) from each other for the CABAC decoding process. That is, the affine flag (2256) is used during the index decoding process (i.e., used in step 2257) to determine whether to decode "merge idx 2258" or "merge idx affine 2255". If ATMVP is not included as a candidate in the list of merge candidates (e.g., if ATMVP is disabled at the SPS level), there are five merge candidates in both lists of candidates (for merge and affine merge modes). Only the first bit of the merge index is decoded by CABAC using the first single context variable. Then, only the first bit of the affine merge index is decoded by CABAC using the second single context variable. All other bits (from the second to the fourth bit) are bypass decoded. In contrast to the current reference software, if ATMVP is included as a candidate in the list of merge candidates (e.g., when ATMVP is enabled at the SPS level), only the first bit of the merge index is decoded by CABAC using the first single context variable in the decoding of the merge index and the second single context variable in the decoding of the affine merge index. The other bits (from the second to the fifth bit or the fourth bit) are bypass decoded. The decoded index is then used to identify the candidate selected by the encoder from the corresponding list of candidates (i.e., merge candidate or affine merge candidate).
[0353] The advantage of this variant is that by using the same index signaling scheme for both the merge index and the affine merge index, the complexity of index decoding and decoder design (and encoder design) for implementing these two different modes is reduced without significantly affecting the coding efficiency. Indeed, with this variant, only two CABAC states (one for each of the first and second single-context variables) are required for index signaling, instead of 9 or 10 if all bits of the merge index and all bits of the affine merge index are CABAC coded / decoded. Furthermore, since all other bits (apart from the first bit) are CABAC bypass coded, the worst-case complexity is reduced, and the number of operations required during the CABAC coding / decoding process is reduced compared to coding all bits by CABAC.
[0354] According to yet another further variant, the CABAC encoding or decoding uses the same context variable for at least one bit of the motion information predictor index of the current block for both when the affine merge mode is used and when the merge mode is used. In this further variant, the context variable used for the first bit of the merge index and the first bit of the affine merge index does not depend on which index is being encoded or decoded, i.e. the first and second single context variables (from the previous variant) are not differentiated / separated but are one and the same single context variable. Thus, in contrast to the previous variant, the merge index and the affine merge index share one context variable during CABAC processing. As shown in Figure 23, the index signaling scheme is the same for both the merge index and the affine merge index, i.e. only one type of index "merge idx (2308)" is encoded or decoded for both modes. As far as the CABAC decoder is concerned, the same syntax elements are used for both the merge index and the affine merge index and there is no need to distinguish between them when considering the context variables. Therefore, there is no need to use the affine flag (2306) to determine whether the current block is coded (decoded) in affine merge mode as in step (2257) of FIG. 22, and there is no branching after step 2306 of FIG. 23, since only one index ("merge idx") needs to be decoded. The affine flag is used to perform motion information prediction in affine merge mode, i.e., during the prediction process after the CABAC decoder decodes the index ("merge idx"). Furthermore, only the first bit of this index (i.e., the merge index and the affine merge index) is coded by CABAC using one single context, and the other bits are bypass coded as described for the first embodiment.Thus, in this further variant, one context variable, the first bit of the merge index and the affine merge index, is shared by both merge index and affine merge index signaling. If the size of the list of candidates is different for merge index and affine merge index, the maximum number of bits for signaling the associated index in each case may also be different, i.e., they are independent of each other. Thus, the number of bypass coding bits can be adjusted accordingly, as necessary, according to the value of the affine flag (2306), for example to enable parsing of data for the associated index from the bitstream.
[0355] The advantage of this variant is that the complexity of the merge index and affine merge index decoding process and the decoder design (and encoder design) is reduced without significantly affecting the coding efficiency. Indeed, in this further variant, only one CABAC state is needed when signaling both the merge index and the affine merge index, instead of the previous variant or 9 or 10 CABAC states. Furthermore, since all other bits (apart from the 1st bit) are CABAC bypass coded, the worst-case complexity is reduced and the number of operations required during the CABAC coding / decoding process is reduced compared to coding all bits by CABAC.
[0356] In the above-mentioned variations of this embodiment, affine merge index signaling and merge index signaling can reduce the number of content and / or share one or more contexts as described in any of the first to fifteenth embodiments. The advantage of this is reduced complexity due to the reduced number of contexts required to encode or decode these indexes.
[0357] In the aforementioned variants of this embodiment, the motion information predictor candidate comprises information for obtaining (or deriving) one or more of the following: direction, ID of the list, reference frame index, and motion vector. Preferably, the motion information predictor candidate includes information for obtaining a motion vector predictor candidate. In a preferred variant, a motion information predictor index (e.g., an affine merge index) is used to signal the affine merge mode predictor candidate, and the affine merge index signaling is implemented using an index signaling similar to the merge index signaling according to any one of the first to fifteenth embodiments or the merge index signaling used in current VTM or HEVC (with the motion information predictor candidate of the affine merge mode as a merge candidate).
[0358] In the above-mentioned variants of this embodiment, the generated list of motion information predictor candidates includes an ATMVP candidate, as in the first embodiment, or as in some variants of the other above-mentioned second to fifteenth embodiments. The ATMVP candidate may be included in one or both of the merge candidate list and the affine merge candidate list. Alternatively, the generated list of motion information predictor candidates does not include an ATMVP candidate.
[0359] In the aforementioned variant of this embodiment, the maximum number of candidates that can be included in the list of candidates for the merge index and the affine merge index is fixed. The maximum number of candidates that can be included in the list of candidates for the merge index and the affine merge index may be the same. Then, data for determining (or indicating) the maximum number (or target number) of motion information predictor candidates that can be included in the generated list of motion information predictor candidates is included in the bitstream by the encoder, and the decoder obtains data for determining the maximum number (or target number) of motion information predictor candidates that can be included in the generated list of motion information predictor candidates from the bitstream. This allows data for decoding the merge index or the affine merge index to be parsed from the bitstream. This data for determining (or indicating) the maximum number (or target number) may be the maximum number (or target number) itself when decoded, or may enable the decoder to determine this maximum / target number in conjunction with other parameters / syntax elements, for example, "five_minus_max_num_merge_cand" or "MaxNumMergeCand-1" or functionally equivalent parameters used in HEVC.
[0360] Alternatively, if the maximum number (or target number) of candidates in the list of candidates for merge index and affine merge index may change or be different (such as because the use of ATMVP candidates or any other candidates is enabled or disabled for one list but not the other list, or because the lists use different candidate list generation / derivation processes), the maximum number (or target number) of motion information predictor candidates that may be included in the generated list of motion information predictor candidates when the affine merge mode is used and when the merge mode is used can be determined separately, and the encoder includes data for determining the maximum number / target number in the bitstream. The decoder then obtains the data for determining the maximum / target number from the bitstream and uses the obtained data to analyze or decode the motion information predictor index. The affine flag can then be used to switch, for example, between analyzing or decoding the merge index and the affine merge index.
[0361] As mentioned above, one or more of the additional inter prediction modes (such as MHII merge mode, triangle merge mode, and MMVD merge mode) may be used in addition to or instead of the merge mode or the affine merge mode, and an index (or flag or information) for one or more of the additional inter prediction modes may be signaled (encoded or decoded). The following embodiments relate to signaling information (such as an index) for the additional inter prediction modes.
[0362] Seventeenth embodiment Signaling for all inter prediction modes (including merge mode, affine merge mode, MHII merge mode, triangle merge mode, and MMVD merge mode) These multiple inter-prediction "merge" modes are signaled using data provided in the bitstream along with their associated syntax (elements) according to the seventeenth embodiment. Figure 26 shows a decoding process for an inter-prediction mode for a current CU (image portion or block) according to an embodiment of the present invention. As described in relation to Figure 12 (and the skip flag in 1201), the first CU skip flag is extracted from the bitstream (2601). If the CU is not skip (2602), i.e., the current CU is not processed in skip mode, the pred mode flag (2603) and / or the merge flag (2606) are decoded to determine whether the current CU is a merge CU. If the current CU is processed in merge skip (2602) or merge CU (2607), the MMVD_Skip_Flag or MMVD_Merge_Flag is decoded (2608). If this flag is equal to 1 (2609), the current CU is decoded using the MMVD merge mode (i.e., in or out of the MMVD merge mode), resulting in the MMVD merge index being decoded (2610), followed by the MMVD distance index (2611) and the MMVD direction index (2612). If the CU is not an MMVD merge CU (2609), the merge subblock flag is decoded (2613). This flag is also indicated as the "affine flag" in the previous description. If the current CU is processed in the affine merge mode (also known as the "subblock merge" mode) (2614), the merge subblock index (i.e., the affine merge index) is decoded (2615). If the current CU is not processed in the affine merge mode (2614) and is not processed in the skip mode (2616), the MHII merge flag is decoded (2620). If this block is processed in MHII merge mode (2621), then the regular merge index (2619) is decoded with the associated intra prediction mode for MHII merge mode (2622). Note that MHII merge mode is only available for non-skip "merge" modes, not for skip modes.If the MHII merge flag is equal to 0 (2621), or if the current CU is not processed in skip mode (2616) & affine merge mode (2614), the triangle merge flag is decoded (2617). If this CU is processed in triangle merge mode (2618), the triangle merge index is decoded (2623). If the current CU is not processed in triangle merge mode (2618), then the current CU is a regular merge mode CU and the merge index is decoded.
[0363] Signaling each merge candidate MMVD Merge Flag / Index Signaling In a first variant of the seventeenth embodiment, only two initial candidates are available for use / selection in the MMVD merge mode. However, if eight possible values for the distance index and four possible values for the direction index are also signaled with the bitstream, the number of potential candidates for use in the MMVD merge mode at the decoder is 64 (2 candidates x 8 distance index x 4 direction index), with each potential candidate being different from another (i.e., unique) if the initial candidate is different. These 64 potential candidates can be evaluated / compared for the MMVD merge mode at the encoder side, and then the MMVD merge index (2610) for the selected initial candidate is signaled with a unary maximum code. Since only two initial candidates are used, this MMVD merge index (2610) corresponds to a flag. Figure 27(a) shows the encoding of this flag, which is CABAC coded using one context variable. It should be appreciated that in another variant, a different number of initial candidates, distance index values, and / or direction index values may be used instead, with the signaling of the MMVD merge index adapted accordingly (e.g., at least one bit is CABAC coded using one context variable).
[0364] Triangle Merge Index Signaling In a first variant of the seventeenth embodiment, the triangle merge index is signaled differently when compared to the index signaling for other inter-prediction modes. In the triangle merge mode, 40 possible permutations of candidates are available, corresponding to the combination of five initial candidates and two possible types of triangles (see Fig. 25(a) and Fig. 25(b), two possible first block predictors (2501 or 2511) and second block predictors (2502 or 2512) for each type of triangle). Fig. 27(b) shows the encoding of the index for the triangle merge mode, i.e., the signaling of these candidates. The first bit (i.e., the first bin) is CABAC decoded in one context. If this first bit is equal to 0, the second bit (i.e., the second bin) is CABAC bypass decoded. If this second bit is equal to 0, the index corresponds to the first candidate in the list, i.e., index 0 (Cand0). Otherwise (if the second bit is equal to 1), the index corresponds to the second candidate in the list, i.e. index 1 (Cand1). If the first bit is equal to 1, an Exponential-Golomb code is extracted from the bit stream and the Exponential-Golomb code represents the index of the selected candidate in the list, i.e. from index 2 (Cand2) to index 39 (Cand39).
[0365] It will be appreciated that in another variant, a different number of initial candidates may be used instead, and the signaling of the triangle merge index adapted accordingly (e.g., at least one bit is CABAC coded using one context variable).
[0366] ATMVP of affine merge lists In a second variant of the seventeenth embodiment, the ATMVP is available as a candidate in the affine merge candidate list (i.e., in the affine merge mode - also known as the "sub-block merge" mode). Figure 28 shows a list of the affine merge list derivation with this additional ATMVP candidate (2848). This figure is similar to Figure 24 (described above), but since this additional ATMVP candidate (2848) has been added to the list, a detailed description will not be repeated here. It will be understood that in another variant, a different number of initial candidates may be used instead, and the signaling of the triangle merge index is adapted accordingly (e.g., at least one bit is CABAC coded using one context variable).
[0367] It will be appreciated that in another variant, the ATMVP candidate may be added to the list of candidates for another inter-prediction mode, and the signaling of its index is adapted accordingly (e.g., at least one bit is CABAC coded using one context variable).
[0368] While Figure 26 provides a complete overview of the signaling for all inter-prediction modes (i.e., merge mode, affine merge mode, MHII merge mode, triangle merge mode, and MMVD merge mode) according to another variant, it will be understood that only a subset of the inter-prediction modes may be used instead.
[0369] Eighteenth embodiment According to an 18th embodiment, one or both of the triangle merge mode or the MMVD merge mode are available for use in the encoding or decoding process, and one or both of these inter prediction modes share context variables (used in conjunction with CABAC encoding) with another inter prediction mode when signaling its index / flag.
[0370] In further variations of this or the following embodiments, it is understood that one or more of the inter prediction modes may use more than one context variable when signaling its index / flag (e.g., an affine merge mode may use four or five context variables, depending on whether ATMVP candidates can also be included in the list for its affine merge index encoding / decoding process).
[0371] For example, before this embodiment or a variant of the following embodiment is implemented, the total number of context variables for signaling all bits of index / flags for all inter-prediction modes may be 7: (regular) merge = 1 (as shown in FIG. 10(a)); affine merge = 4 (as shown in FIG. 10(b) but with one less candidate, e.g., no ATMVP candidate); triangle = MMVD = 1; and MHII (if available) = 0 (shared with regular merge). Then, by implementing the variant, the total number of context variables for signaling all bits of index / flags for all inter-prediction modes may be reduced to 5: (regular) merge = 1 (as shown in FIG. 10(a)); affine merge = 4 (as shown in FIG. 10(b) but with one less candidate, e.g., no ATMVP candidate); and triangle = MMVD = MHII (if available) = 0 (shared with regular merge).
[0372] In another example, before this variant is implemented, the total number of context variables for signaling all bits of index / flags for all inter-prediction modes may be 4: (regular) merge = affine merge = triangle = MMVD = 1 (as shown in FIG. 10(a)); and MHII (if available) = 0 (shared with regular merge). Then, by implementing this variant, the total number of context variables for signaling all bits of index / flags for all inter-prediction modes is reduced to 2: (regular) merge = affine merge = 1 (as shown in FIG. 10(a)); and triangle = MMVD = MHII (if available) = 0 (shared with regular merge).
[0373] Please note that for simplicity of the following description, we will describe sharing or not sharing one context variable (e.g., only the first bit). This means that in the following description, we often see the simple case of using a context variable to signal only the first bit for each inter-prediction mode, and this context variable is either 1 (a separate / independent context variable is used) or 0 (this bit is bypass CABAC coded or shares the same context variable with another inter-prediction mode, so there is no separate / independent context variable). It is understood that different variants of this embodiment and the following embodiments are not limited thereto, and other bits of context variables, or indeed all bits, may be shared / non-shared / bypass CABAC coded in the same way.
[0374] In a first variant of the eighteenth embodiment, all inter prediction modes available for use in the encoding or decoding process share at least some CABAC context.
[0375] In this variant, the index coding and its related parameters (e.g., the number of (initial) candidates) for the inter prediction mode may be set to be the same or similar as long as possible / compatible. For example, to simplify signaling, the number of candidates for the affine merge mode and the merge mode are set to 5 and 6, respectively, the number of initial candidates for the MMVD merge mode is set to 2, and the maximum number of candidates for the triangle merge mode is 40. Also, the triangle merge index is not signaled using a unary max code like other inter prediction modes. For this triangle merge mode, only the first bit of context variables (for the triangle merge index) can be shared with other inter prediction modes. The advantage of this variant is that the design of the encoder and the decoder is simplified.
[0376] In a further variant, the CABAC content for all merged inter prediction mode indexes is shared. This means that only one CABAC context variable is needed for the 1st bit of every index. In yet another variant, if an index contains more than one bit to be CABAC coded, the coding of the additional bits (all CABAC coded bits apart from the 1st bit) is treated as a separate part (i.e., as for another syntax element as far as the CABAC coding process is concerned), and if more than one index has more than one bit to be CABAC coded, one and the same context variable is shared for these CABAC coded "additional" bits. The advantage of this variant is that the amount of CABAC context is reduced. This reduces the storage requirements for the context state that needs to be stored at the encoder and decoder side, without significantly affecting the coding efficiency for the majority of sequences processed by the video codec implementing the variant.
[0377] FIG. 29 shows another further modified decoding process for inter prediction mode. This figure is similar to FIG. 26, but includes the implementation of this modified example. In this figure, when the current CU is processed in MMVD merge mode, its MMVD merge index is decoded as the same index (i.e., "merge index" (2919)) as the merge index in the regular merge mode. However, unlike the regular merge mode, only two initial candidates can be selected in MMVD merge mode, not six. Since there are only two possibilities, this "shared" index used in MMVD merge mode is essentially a flag. Since the same index is shared, the CABAC context variable is the same for this flag in MMVD merge mode and the first bit of the merge index in merge mode. Next, if it is determined that the current CU should be processed in MMVD merge mode (2925), the distance index (2911) and the direction index (2912) are decoded. If it is determined that the current CU is to be processed in the affine merge mode (2914), its affine merge index is decoded as the same index (i.e., the "merge index" (2919)) as the merge index in the regular merge mode. However, in the affine merge mode, unlike in the regular merge mode, the maximum number of candidates (i.e., the maximum number of indices) is 5, not 6. If it is determined that the current CU should be processed in the triangle merge mode (2918), the first bit is decoded (2919) as the shared index, so that the same CABAC context variables are shared as in the regular merge mode. If the CU is to be processed in the triangle merge mode (2926), the remaining bits related to the triangle merge index are decoded (2923).
[0378] Thus, for example, during the CABAC encoding process, when processing these indices / flags, the number of separate (independent) context variables used for the first bit of the index / flag for each inter prediction mode is: (regular)merge=1; MHII=Affine merge=Triangle=MMVD=0 (shared with regular merge) It is.
[0379] In a second variant, when one or both of the triangle merge mode or the MMVD merge mode are used (i.e., information about the motion information predictor selection of the current CU is processed / encoded / decoded in the relevant inter prediction mode), its / their index signaling shares context variables with the index signaling of the merge mode. In this variant, the CABAC context of the triangle merge index and / or the CABAC context of the MMVD merge index / flag share the same CABAC context of the merge index (for the (regular) merge mode). This means that only one CABAC state is needed for at least these three modes.
[0380] In a further variation of the second variation, the CABAC context of the triangle merge index and / or the CABAC context of the MMVD merge index / flag share the same first CABAC context variable of the merge index, e.g., the same context variable of the first bit of the merge index.
[0381] So, for example, when processing these indices / flags during the CABAC encoding process, the number of separate (independent) context variables used for the first bit of the index / flag is: (regular)merge=1; MHII (if available) = affine merge (if available) = 0 (shared with regular merge) or 1 depending on implementation; Triangle=MMVD=0 (shared with regular merge) It is.
[0382] In yet another variant of the second variant, if two or more context variables are used for triangle merge index CABAC encoding / decoding or two or more context variables are used for MMVD merge index CABAC encoding / decoding, they may all be shared or at least partially shared whenever compatible with the two or more CABAC context variables used for merge index CABAC encoding / decoding.
[0383] The advantage of this second variant is that it reduces the amount of context that needs to be stored, and therefore the amount of state that needs to be stored at the encoder and decoder side, without significantly affecting the coding efficiency of the majority of sequences processed by the video codecs that implement them.
[0384] In a third variant, when one or both of the triangle merge mode or the MMVD merge mode are used (i.e., information regarding the motion information predictor selection of the current CU is processed / encoded / decoded in the relevant inter prediction mode), its / their index signaling shares context variables with the index signaling of the affine merge mode. In this variant, the CABAC context of the triangle merge index and / or the CABAC context of the MMVD merge index / flag share the same CABAC context of the affine merge index (for the affine merge mode).
[0385] In a further variation of the third variation, the CABAC context of the triangle merge index and / or the CABAC context of the MMVD merge index / flag share the same first CABAC context variable of the affine merge index, e.g., the same context variable of the first bit of the affine merge index.
[0386] So, for example, when processing these indices / flags during the CABAC encoding process, the number of separate (independent) context variables used for the first bit of the index / flag is: (Canonical)Merge (if available) = 0 (shared with Affine Merge) or 1 depending on implementation; MHII(if available)=0(share with regular merge); affinemerge=1; Triangle = MMVD = 0 (shared with affine merge) It is.
[0387] In yet another variation of the third variation, if two or more context variables are used for triangle merge index CABAC encoding / decoding or two or more context variables are used for MMVD merge index CABAC encoding / decoding, they may all be shared or at least partially shared if compatible with two or more CABAC context variables used for affine merge index CABAC encoding / decoding.
[0388] In a fourth variant, when the MMVD merge mode is used (i.e., the information regarding the motion information predictor selection of the current CU is processed / encoded / decoded in the MMVD merge mode), the index signaling shares context variables with the index signaling of the merge mode or the affine merge mode. In this variant, the CABAC context of the MMVD merge index / flag is the same CABAC context of the merge index, or the same CABAC context of the affine merge index.
[0389] So, for example, when processing these indices / flags during the CABAC encoding process, the number of separate (independent) context variables used for the first bit of the index / flag is: (regular)merge=1; MHII(if available)=0(share with regular merge); affine merge (if available) = 0 (shared with regular merge) or 1 depending on implementation; MMVD=0 (regular merge and share) or (Canonical)Merge (if available) = 0 (shared with Affine Merge) or 1 depending on implementation; MHII(if available)=0(share with regular merge); affinemerge=1; MMVD=0 (shared with affine merging) It is.
[0390] In a fifth variant, when the triangle merge mode is used (i.e., the information regarding the motion information predictor selection of the current CU is processed / encoded / decoded in the triangle merge mode), the index signaling shares context variables with the index signaling for the merge mode or the affine merge mode. In this variant, the CABAC context of the triangle merge index is the same CABAC context of the merge index or the same CABAC context of the affine merge index.
[0391] So, for example, when processing these indices / flags during the CABAC encoding process, the number of separate (independent) context variables used for the first bit of the index / flag is: (regular)merge=1; MHII(if available)=0(share with regular merge); affine merge (if available) = 0 (shared with regular merge) or 1 depending on implementation; Triangles = 0 (regular merge and share) or (Canonical)Merge (if available) = 0 (shared with Affine Merge) or 1 depending on implementation; MHII(if available)=0(share with regular merge); affinemerge=1; Triangle = 0 (shared with affine merge) It is.
[0392] In a sixth variant, when the triangle merge mode is used (i.e., the information on the motion information predictor selection of the current CU is processed / encoded / decoded in the triangle merge mode), its index signaling shares context variables with the index signaling of the MMVD merge mode. In this variant, the CABAC context of the triangle merge index is the same CABAC context of the MMVD merge index. Thus, for example, during the CABAC encoding process, when processing these indexes / flags, the number of separate (independent) context variables used for the first bit of the index / flag is: MMVD=1; triangle=0(shared with MMVD); (Canonical) Merge or MHII or Affine Merge = depending on implementation and availability It is.
[0393] In a seventh variant, when one or both of the triangle merge mode or the MMVD merge mode are used (i.e., information regarding the motion information predictor selection of the current CU is processed / encoded / decoded in the relevant inter prediction mode), its / their index signaling shares a context variable with the index signaling for the inter prediction mode that can include the ATMVP predictor candidate in the list of candidates, i.e., the inter prediction mode can have the ATMVP predictor candidate as one of the available candidates. In this variant, the CABAC context of the triangle merge index and / or the MMVD merge index / flag share the same CABAC context of the index of the inter prediction mode that can use the ATMVP predictor.
[0394] In a further variation, the CABAC context variables for the triangle merge index and / or the MMVD merge index / flag share the same first CABAC context variable for the merge index in merge mode with the includeable ATMVP candidate, or share the affine merge index in affine merge mode with the includeable ATMVP candidate.
[0395] In yet another variant, if two or more context variables are used for triangle merge index CABAC encoding / decoding or two or more context variables are used for MMVD merge index CABAC encoding / decoding, they can all be shared with two or more CABAC context variables used for affine merge indexes of affine merge mode or merge indexes of merge mode with includeable ATMVP candidates, or shared wherever at least partially compatible.
[0396] The advantage of these variations is improved coding efficiency, since the ATMVP (predictor) candidates are the predictors that benefit most from CABAC adaptation when compared to other types of predictors.
[0397] Nineteenth embodiment According to a 19th embodiment, one or both of triangle merge mode or MMVD merge mode are available for use in the encoding or decoding process, and indexes / flags for one or both of these inter prediction modes are CABAC bypass coded when signaling the indexes / flags.
[0398] In a first variant of the nineteenth embodiment, all inter prediction modes available for use in the encoding or decoding process have their indexes / flags CABAC bypass coded / decoded to signal the indexes / flags. In this variant, all indexes of all inter prediction modes are coded (e.g., by the bypass coding engine 1705 of FIG. 17) without using CABAC context variables. This means that all bits of the merge index (2619), affine merge index (2615), MMVD merge index (2610), and triangle merge index (2623) of FIG. 26 are CABAC bypass coded. Figures 30(a)-30(c) show coding of indexes / flags according to this embodiment. Figure 30(a) shows MMVD merge index coding of initial MMVD merge candidates. Figure 30(b) shows triangle merge index coding. Figure 30(c) shows affine merge index coding, which can be easily used for merge index coding as well.
[0399] So, for example, when processing these indices / flags during the CABAC encoding process, the number of separate (independent) context variables used for the indices / flags is: (Regular) Merge (if available) = MHII (if available) = Affine Merge (if available) = Triangle (if available) = MMVD (if available) = 0 (all bypass codes). The advantage of this variant is the reduction in the amount of context that needs to be stored and, therefore, the reduction in the amount of state that needs to be stored at the encoder and decoder side, with only a small impact on the coding efficiency of the majority of sequences processed by the video codec implementing the variant. However, it should be noted that it may be lossy when used to code screen content. This variant represents another compromise between coding efficiency and complexity when compared to other variants / embodiments. The impact on coding efficiency is often small. Indeed, with a large number of inter prediction modes available, the average amount of data required to signal an index for each inter prediction mode is smaller than the average amount of data required to signal a merge index when only a merge mode is available / available (this comparison is for the same sequences and the same coding efficiency compromise). This means that the efficiency of CABAC coding / decoding from the adaptation of bin probabilities based on the context may not be efficient.
[0400] In a second variant, when one or both of the triangle merge mode or the MMVD merge mode are used (i.e., information regarding the motion information predictor selection of the current CU is processed / encoded / decoded in the relevant inter prediction mode), its / their indexes / flags are signaled by CABAC bypass encoding / decoding the indexes / flags. In this variant, the MMVD merge index and / or the triangle merge index are CABAC bypass encoded. Depending on the implementation, i.e., when the merge mode and the affine merge mode are available, the merge index and the affine merge index have their own context. In yet another variant, the context of the affine merge index and the merge index are shared. So, for example, when processing these indices / flags during the CABAC encoding process, the number of separate (independent) context variables used for the first bit of the index / flag is: (canonical)merge = affine merge = 0 or 1 depending on implementation; MHII(if available)=0(shared with regular merge); Triangle = MMVD = 0 (bypass coding) It is.
[0401] The advantage of these variants is that the coding efficiency is improved compared to the previous variants, since there is yet another compromise between the reduction of the CABAC context and the coding efficiency. In fact, the triangle merge mode is not often selected. As a result, the impact on coding efficiency is small when its context is removed, i.e., when the triangle merge mode uses CABAC bypass coding. Although the MMVD merge mode tends to be selected more frequently than the triangle merge mode, the probability of selecting the first and second candidates for the MMVD merge mode tends to be equal to or greater than other inter prediction modes such as the merge mode or the affine merge mode, and therefore the benefit from using the context of CABAC coding is not as great for the MMVD merge mode. Another advantage of these variants is the small coding efficiency impact for screen content sequences, since the most influential inter prediction mode for screen content is the merge mode.
[0402] In a third variant, when merge mode, triangle merge mode, or MMVD merge mode is used (i.e., information regarding the motion information predictor selection of the current CU is processed / encoded / decoded in the relevant inter prediction mode), its / their indexes / flags are signaled by CABAC bypass encoding / decoding the indexes / flags. In this variant, the MMVD merge index, triangle merge index, and merge index are CABAC bypass encoded. Thus, for example, during the CABAC encoding process, when processing these indexes / flags, the number of separate (independent) context variables used for the first bit of the index / flag is: (regular) merge = triangle (if available) = MMVD (if available) = 0 (bypass coding); MHII (if available) = 0 (same as regular merge); and Affine Merge = 1 It is.
[0403] This variant offers an alternative compromise compared to the other variants, for example this variant provides a larger coding efficiency reduction for screen content sequences than the previous variant.
[0404] In a fourth variant, when the affine merge mode, triangle merge mode, or MMVD merge mode is used (i.e., information regarding the motion information predictor selection of the current CU is processed / encoded / decoded in the relevant inter prediction mode), its / their indexes / flags are signaled by CABAC bypass encoding / decoding the indexes / flags. In this variant, the affine merge index, MMVD merge index, and triangle merge index are CABAC bypass encoded, and the merge index is encoded in one or more CABAC contexts. Thus, for example, during the CABAC encoding process, when processing these indexes / flags, the number of separate (independent) context variables used for the first bit of the index / flag is: (regular)merge=1; affine merge = triangle (if available) = MMVD (if available) = 0 (bypass coding); MHII (if available) = 0 (regular merging and sharing). The advantage of this variant over the previous one is increased coding efficiency for screen content sequences.
[0405] In a fifth variant, the inter-prediction modes available for use in the encoding or decoding process are CABAC bypass coded / decoded to signal the index / flag, except when the inter-prediction modes can include an ATMVP predictor candidate in the list of candidates, i.e., the inter-prediction modes can have an ATMVP predictor candidate as one of the available candidates. In this variant, all indices of all inter-prediction modes are CABAC bypass coded, except when the inter-prediction modes can have an ATMVP predictor candidate. Thus, for example, during the CABAC coding process, when processing these indexes / flags, the number of separate (independent) context variables used for the first bit of the index / flag is: (regular) merge with includable ATMVP candidates = 1; AffineMerge(if available) = Triangle(if available) = MMVD(if available) = 0(BypassEncoding); MHII (if available) = 1 (shared with regular merge) or 0 depending on implementation or affinemergewithincludeableATMVPcandidates=1; (regular) merge(if available) = triangle(if available) = MMVD(if available) = 0(bypass coding); MHII (if available) = 0 (same as regular MERGE) It is.
[0406] This variant also introduces other complexity / coding efficiency compromises for most of the natural sequences. However, note that for screen content sequences it may be preferable to have ATMVP predictor candidates in the canonical merge candidate list.
[0407] In a sixth variant, if an inter prediction mode available for use in the encoding or decoding process is not a skip mode (e.g., not one of a regular merge skip mode, an affine merge skip mode, a triangle merge skip mode, or an MMVD merge skip mode), the index / flag is CABAC bypass coded / decoded to signal the index / flag. In this variant, all indexes are CABAC bypass coded for any CU that is not processed in skip mode, i.e., is not skipped. The index for the skipped CU (i.e., the CU processed in skip mode) may be processed using any one of the CABAC coding techniques described in connection with the previous embodiments / variations (e.g., only the first bit, or more than one bit, has a context variable, which may or may not be shared).
[0408] Figure 31 is a flowchart of the decoding process of the inter prediction mode showing this modification. The process of Figure 31 is similar to Figure 29, except that it has an additional "skip mode" decision / check step (3127), after which the index / flag ("Merge_idx") is decoded using the context of either CABAC decoding (3119) or CABAC bypass decoding (3128). According to yet another modification, it is understood that the result of the decision / check made in the previous step is that the CU is skip (2902 / 3102), MMVD skip (2908 / 3108), and the CU uses the skip (2916 / 3116) of Figure 29 or Figure 31 to make the "skip mode" decision / check instead of the additional "skip mode" decision / check step (3127).
[0409] This variant has a lower impact on coding efficiency because skip modes are generally selected more frequently than non-skip modes (i.e., non-skip inter prediction modes such as regular merge mode, MHII merge mode, affine merge mode, triangle merge mode, or MMVD merge mode), and because the selection of the first candidate is more likely for skip modes than for non-skip modes. Since SKIP modes are designed for more predictable motion, their indices must also be more predictable. Thus, the probability of utilizing CABAC coding / decoding is more likely to be useful for skip modes. However, non-skip modes are more likely to be used when the motion is less predictable, so that a more random selection from the predictor candidates is more likely to occur. Thus, in non-skip modes, CABAC coding / decoding is less likely to be efficient.
[0410] Twentieth embodiment According to a twentieth embodiment, data is provided in a bitstream, the data being for determining whether an index / flag for one or more of the inter-prediction modes should be signaled by using CABAC bypass encoding / decoding, CABAC encoding / decoding with separate context variables, or CABAC encoding / decoding with one or more shared context variables. For example, such data may be a flag for enabling or disabling the use of one or more independent contexts for index encoding / decoding of the inter-prediction modes. Such data may be used to control the use or non-use of context sharing in CABAC encoding / decoding or CABAC bypass encoding / decoding.
[0411] In a variation of the twentieth embodiment, the CABAC context sharing between two or more indexes of two or more inter-prediction modes relies on data transmitted in the bitstream, for example, at a level higher than the CU level (e.g., at the level of an image portion larger than the smallest CU, such as the sequence, frame, slice, tile, or CTU level). For example, this data may indicate that for any CU in a particular image portion, the CABAC context of a merge index of a merge mode is shared (or not shared) with one or more other CABAC contexts of another inter-prediction mode.
[0412] In another variation, one or more indexes are CABAC bypass coded in response to data transmitted in the bitstream, for example, at a level higher than the CU level (e.g., at the slice level). For example, this data may indicate that for any CU in a particular image portion, an index of a particular inter-prediction mode should be CABAC bypass coded.
[0413] In one variant, to further improve the coding efficiency, at the encoder side, the value of this data for indicating the context sharing of one or more indices of one or more inter-prediction modes, or CABAC bypass coding / decoding, can be selected based on how frequently one or more inter-prediction modes are used in previously coded frames. An alternative may be to select the value of this data based on the type of sequence to be processed, or the type of application in which the variant is implemented.
[0414] An advantage of this embodiment is a controlled increase in coding efficiency compared to the previous embodiment / variant.
[0415] Implementation of the embodiments of the present invention One or more of the above-mentioned embodiments may be implemented by the processor 311 of the processing device 300 of FIG. 3, or a corresponding functional module / unit of the decoder 60 of FIG. 5, of the CABAC coder of FIG. 17, of the encoder 400 of FIG. 4, or its corresponding CABAC decoder, which performs the method steps of one or more of the above-mentioned embodiments.
[0416] Fig. 19 is a schematic block diagram of a computing device 2000 for the implementation of one or more embodiments of the present invention. The computing device 2000 may be a device such as a microcomputer, a workstation, or a light portable device. The computing device 2000 comprises: - a central processing unit (CPU) 2001, such as a microprocessor; - a random access memory (RAM) 2002 for storing executable code of the method of the present invention and registers for recording variables and parameters necessary for implementing the method for encoding or decoding at least a part of an image according to the embodiments of the present invention, the capacity of which may be expanded, for example, by an optional RAM connected to an expansion port; - a read-only memory (ROM) 2003 for storing computer programs for implementing the embodiments of the present invention; - a communication bus connected to a network interface (NET) 2004, which is typically connected to a communication network over which the digital data to be processed is transmitted or received. The network interface (NET) 2004 may be a single network interface or may consist of a set of different network interfaces (e.g. wired and wireless interfaces, or different types of wired or wireless interfaces). Data packets are written to the network interface for transmission or read from the network interface for reception under the control of software applications running on the CPU 2001. - A user interface (UI) 2005 may be used to receive input from a user and display information to the user. - A hard disk (HD) 2006 may be provided as mass storage device. - An input / output module (IO) 2007 may be used to send and receive data to external devices such as video sources and displays. The executable code may be stored either in the ROM 2003, the HD 2006 or on a removable digital medium such as a disk.According to a variant, the executable code of the program can be received by the communication network, via NET 2004, in order to be stored in one of the storage means of the communication device 2000, such as HD 2006, before being executed. The CPU 2001 is adapted to control and direct the execution of instructions or parts of a program or software code of a program according to an embodiment of the invention, the instructions of which are stored in one of the aforementioned storage means. After power-on, the CPU 2001 can execute instructions for a software application, for example from the main RAM memory 2002, after these instructions have been loaded from the program ROM 2003 or from the HD 2006. Such a software application, when executed by the CPU 2001, causes the steps of the method according to the invention to be carried out.
[0417] It is also understood that according to other embodiments of the present invention, a decoder according to the aforementioned embodiment is provided in a user terminal such as a computer, a mobile phone (cell phone), a tablet, or any other type of device (e.g., a display device) that can provide / display content to a user. According to yet another embodiment, an encoder according to the aforementioned embodiment is provided in an image capture device that also comprises a camera, video camera, or network camera (e.g., a closed circuit television or video surveillance camera) that captures and provides content for the encoder to encode. Two such examples are provided below with reference to Figures 20 and 21.
[0418] FIG. 20 is a diagram illustrating a network camera system 2100 including a network camera 2102 and a client device 2104 .
[0419] The network camera 2102 includes an imaging unit 2106 , an encoding unit 2108 , a communication unit 2110 , and a control unit 2112 .
[0420] A network camera 2102 and a client device 2104 are connected to each other via a network 200 so as to be able to communicate with each other.
[0421] The imaging unit 2106 includes a lens and an imager (e.g., a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS)) to capture an image of an object and generate image data based on the image. The image may be a still image or a video image. The imaging unit may also include zoom means and / or pan means adapted to zoom or pan (optically or digitally).
[0422] The encoder 2108 encodes the image data using the encoding methods described in one or more of the above embodiments. The encoder 2108 uses at least one of the encoding methods described in the above embodiments. In other examples, the encoder 2108 can use a combination of the encoding methods described in the above embodiments.
[0423] The communication unit 2110 of the network camera 2102 transmits the encoded image data encoded by the encoding unit 2108 to the client device 2104. The communication unit 2110 also receives a command from the client device 2104. The command includes a command for setting parameters for encoding by the encoding unit 2108.
[0424] The control unit 2112 controls other units within the network camera 2102 according to commands received by the communication unit 2110 .
[0425] The client device 2104 includes a communication unit 2114, a decoding unit 2116, and a control unit 2118. The communication unit 2114 of the client device 2104 transmits commands to the network camera 2102. In addition, the communication unit 2114 of the client device 2104 receives encoded image data from the network camera 2102.
[0426] The decoder 2116 decodes the encoded image data using the decoding methods described in one or more of the above embodiments. In other examples, the decoder 2116 may use a combination of the decoding methods described in the above embodiments.
[0427] The control unit 2118 of the client device 2104 controls other units in the client device 2104 in accordance with user operations and commands received by the communication unit 2114. The control unit 2118 of the client device 2104 controls the display device 2120 to display the image decoded by the decoding unit 2116. The control unit 2118 of the client device 2104 also controls the display device 2120 to display a GUI (Graphical User Interface), and specifies parameter values of the network camera 2102 including parameters for encoding by the encoding unit 2108.
[0428] Furthermore, the control unit 2118 of the client device 2104 controls other units in the client device 2104 in response to a user operation input to a GUI displayed by the display device 2120. The control unit 2118 of the client device 2104 controls the communication unit 2114 of the client device 2104 to transmit a command specifying parameter values of the network camera 2102 to the network camera 2102 in response to a user operation input to a GUI displayed by the display device 2120.
[0429] The network camera system 2100 can determine whether the camera 2102 utilizes zoom or pan while recording video, and such information can be used when encoding the video stream, as zooming or panning during recording can benefit from the use of an affine mode, which is well suited for encoding complex movements such as zooming, rotation, and / or stretching (which can be side effects of panning, especially if the lens is a "fisheye" lens).
[0430] FIG. 21 is a diagram showing a smartphone 2200.
[0431] The smartphone 2200 includes a communication unit 2202, a decoding / encoding unit 2204, a control unit 2206, and a display unit 2208.
[0432] The communication unit 2202 receives the encoded image data via the network 200 .
[0433] The decoding / encoding unit 2204 decodes the encoded image data received by the communication unit 2202. The decoding / encoding unit 2204 decodes the encoded image data using the decoding method described in one or more of the above-mentioned embodiments. The decoding / encoding unit 2204 may use at least one of the decoding methods described in the above-mentioned embodiments. In other examples, the decoding / encoding unit 2204 may use a combination of the decoding or encoding methods described in the above-mentioned embodiments.
[0434] The control unit 2206 controls other units in the smartphone 2200 in response to a user operation or a command received by the communication unit 2202 or via the input unit. For example, the control unit 2206 controls the display device 2208 to display an image decoded by the decoding unit 2204.
[0435] The smartphone may further include an image recording device 2210 (e.g., a digital camera and associated circuitry) for recording images or videos. Such recorded images or videos may be encoded by the decoding / encoding unit 2204 under the direction of the control unit 2206. The smartphone may further include a sensor 2212 configured to sense the orientation of the mobile device. Such sensors may include an accelerometer, gyroscope, compass, global positioning (GPS) unit or similar position sensor. Such a sensor 2212 may determine whether the smartphone is changing orientation, and such information may be used when encoding the video stream as orientation changes during capture and may benefit from the use of an affine mode, which is well suited to encoding complex movements such as rotations.
[0436] Substitutions and Modifications It will be appreciated that the aim of the present invention is to ensure that affine modes are utilized in the most efficient way, and the particular example given above relates to signalling the use of affine modes depending on the likelihood that the affine modes are perceived to be useful. A further example of this may be applied to an encoder where it is known that complex motion is being coded, where affine transformations may be particularly efficient. Examples of such cases are: a) Camera zoom in / out b) Portable cameras (e.g., mobile phones) that change orientation during capture (i.e., rotational movement) c) Panning a "fish-eye" lens camera (e.g. stretching / distorting parts of the image) Includes.
[0437] Thus, complex motion indications can be increased during the recording process such that affine mode is more likely to be used for slices, frame sequences, or indeed the entire video stream.
[0438] In a further example, the affine mode is more likely to be used depending on the features or functionality of the device used to record the video. For example, a mobile device is more likely to change orientation than a fixed security camera (for example), so the affine mode may be more suitable for encoding video from the former. Examples of features or functionality include the presence / use of a zoom mechanism, the presence / use of a position sensor, the presence / use of a pan mechanism, whether the device is portable, or a user selection on the device.
[0439] Although the present invention has been described with reference to embodiments, it should be understood that the present invention is not limited to the disclosed embodiments. It will be understood by those skilled in the art that various changes and modifications can be made without departing from the scope of the present invention, as defined in the appended claims. All of the features disclosed in this specification (including any accompanying claims, abstract, and drawings), and / or all of the steps of any method or process so disclosed, can be combined in any combination, except combinations in which at least some of such features and / or steps are mutually exclusive. Each feature disclosed in this specification (including any accompanying claims, abstract, and drawings) can be replaced by an alternative feature serving the same, equivalent, or similar purpose, unless otherwise specified. Thus, unless otherwise specified, each feature disclosed is merely one example of a generic series of equivalent or similar features.
[0440] It will also be understood that any result of the above-mentioned comparisons, determinations, evaluations, selections, executions, performances, or considerations, e.g., selections made during an encoding or filtering process, may be indicated in or determinable / inferable from data in the bitstream, e.g., flags or data indicating the result, and such that the indicated or determined / inferred result may be used in processing, e.g., during a decoding process, in lieu of actually performing the comparisons, determinations, evaluations, selections, executions, performances, or considerations.
[0441] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite articles "a" or "an" do not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.
[0442] Reference signs appearing in the claims are by way of illustration only and shall have no limiting effect on the scope of the claims.
[0443] In the above-described embodiments, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted over a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit.
[0444] A computer-readable medium may include a computer-readable storage medium, which corresponds to a tangible medium, such as a data storage medium, or a communication medium, which includes any medium that facilitates transfer of a computer program from one place to another, for example according to a communication protocol. In this manner, a computer-readable medium may generally correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium, such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0445] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of media. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead cover non-transitory tangible storage media. Disk and disk as used herein include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically and disks reproduce data optically with a laser. Combinations of the above should also be included within the scope of computer readable media.
[0446] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate / logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, the term "processor" as used herein may refer to any of the foregoing structures, or any other structure suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or may be incorporated into a combined codec. Also, the techniques may be implemented entirely in one or more circuit or logic elements.
Claims
1. 1. A method for encoding information relating to a motion information predictor, comprising the steps of: selecting one of a plurality of motion information predictor candidates; encoding one of a plurality of indexes including a first index and a second index for identifying the selected motion information predictor candidate using a context-based adaptive binary arithmetic coding (CABAC) coding; Including, the first index is used for a first merge mode in which a block predictor is obtainable from a first block predictor associated with a first region in a block and a second block predictor associated with a second region in the block that is different from the first region; the second index is used for a second merge mode of an inter prediction mode different from the first merge mode; the CABAC encoding of the first bit of the first index for the first merge mode uses the same context variables as the CABAC encoding of the first bit of the second index for the second merge mode; all bits of the first index except the first bit of the first index are bypass coded, and all bits of the second index except the first bit of the second index are bypass coded; Based on the second merging mode and using an average of the intra-block predictor and the inter-block predictor, a block predictor can be obtained. A method comprising:
2. The method according to claim 1, characterized in that each of the first merge mode and the second merge mode is a merge mode independent of a merge mode that uses affine motion information.
3. The method of claim 1, wherein each of the first region and the second region has a shape other than rectangular.
4. The first region in the block does not include a lower left vertex of the block and includes an upper right vertex of the block; 4. The method of claim 3, wherein the second region of the block excludes the upper right vertex of the block and includes the lower left vertex of the block.
5. The method of claim 1, wherein a weighted average is applied to the area between the first region and the second region within the block.
6. 1. A method for decoding information relating to a motion information predictor, comprising the steps of: decoding one of the plurality of indexes, including the first index and the second index, using a context-based adaptive binary arithmetic coding (CABAC) decoding to identify one of the plurality of motion information predictor candidates; selecting one of the plurality of motion information predictor candidates using the decoded index; and Including, the first index is used for a first merge mode in which a block predictor is obtainable from a first block predictor associated with a first region in a block and a second block predictor associated with a second region in the block that is different from the first region; the second index is used for a second merge mode of an inter prediction mode different from the first merge mode; CABAC decoding of the first bit of the first index for the first merge mode uses the same context variables as CABAC decoding of the first bit of the second index for the second merge mode of an inter prediction mode; all bits of the first index except the first bit of the first index are bypass decoded, and all bits of the second index except the first bit of the second index are bypass decoded; Based on the second merging mode and using an average of the intra-block predictor and the inter-block predictor, a block predictor can be obtained. A method comprising:
7. The method of claim 6, wherein each of the first merge mode and the second merge mode is a merge mode independent of a merge mode that uses affine motion information.
8. The method of claim 6, wherein each of the first region and the second region has a shape other than rectangular.
9. The first region in the block does not include a lower left vertex of the block and includes an upper right vertex of the block; 9. The method of claim 8, wherein the second region of the block excludes the top right vertex of the block and includes the bottom left vertex of the block.
10. The method of claim 6, wherein a weighted average is applied to the area between the first region and the second region within the block.
11. 1. An apparatus for encoding information relating to a motion information predictor, comprising: means for selecting one of a plurality of motion information predictor candidates; means for encoding one of a plurality of indexes including a first index and a second index for identifying the selected motion information predictor candidate using a Context-based Adaptive Binary Arithmetic Coding (CABAC) coding; Including, the first index is used for a first merge mode in which a block predictor is obtainable from a first block predictor associated with a first region in a block and a second block predictor associated with a second region in the block that is different from the first region; the second index is used for a second merge mode of an inter prediction mode different from the first merge mode; the CABAC encoding of the first bit of the first index for the first merge mode uses the same context variables as the CABAC encoding of the first bit of the second index for the second merge mode; all bits of the first index except the first bit of the first index are bypass coded, and all bits of the second index except the first bit of the second index are bypass coded; Based on the second merging mode and using an average of the intra-block predictor and the inter-block predictor, a block predictor can be obtained. An apparatus comprising:
12. 1. An apparatus for decoding information relating to a motion information predictor, comprising: means for decoding one of the plurality of indexes, including the first index and the second index, using a context-based adaptive binary arithmetic coding (CABAC) decoding to identify one of the plurality of motion information predictor candidates; means for selecting one of the plurality of motion information predictor candidates using the decoded index; Including, the first index is used for a first merge mode in which a block predictor is obtainable from a first block predictor associated with a first region in a block and a second block predictor associated with a second region in the block that is different from the first region; the second index is used for a second merge mode of an inter prediction mode different from the first merge mode; CABAC decoding of the first bit of the first index for the first merge mode uses the same context variables as CABAC decoding of the first bit of the second index for the second merge mode of an inter prediction mode; all bits of the first index except the first bit of the first index are bypass decoded, and all bits of the second index except the first bit of the second index are bypass decoded; Based on the second merging mode and using an average of the intra-block predictor and the inter-block predictor, a block predictor can be obtained. An apparatus comprising:
13. A computer program product for causing a computer to carry out the method according to claim 1.
14. A computer program product for causing a computer to carry out the method according to claim 6.