Method and apparatus for deriving interpolation filter index for current block
By constructing a history-based motion information candidate list to inherit interpolation filter indexes, the method improves coding efficiency and compression performance in video coding, addressing the overhead issues in VVC.
Patent Information
- Application Number
- JP2025109437
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-10-02
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The challenge in video coding is the significant amount of data required for video transmission and storage, which is addressed by improving compression techniques, particularly in Versatile Video Coding (VVC) where switchable interpolation filters increase signaling overhead.
A method for constructing a history-based motion information candidate list to inherit interpolation filter indexes, allowing selection of appropriate filters for improved coding efficiency and quality by propagating interpolation filter indexes through the list.
This approach enhances coding efficiency and compression performance by reducing computational complexity and ensuring the quality of the coded signal, particularly in merge or skip modes.
Smart Images

Figure 2025160189000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This patent application claims priority to U.S. Provisional Patent Application No. 62 / 836,072, filed April 19, 2019, U.S. Provisional Patent Application No. 62 / 845,938, filed May 10, 2019, U.S. Provisional Patent Application No. 62 / 909,761, filed October 2, 2019, and U.S. Provisional Patent Application No. 62 / 909,763, filed October 2, 2019. The disclosures of the aforementioned patent applications are incorporated herein by reference in their entireties.
[0002] Embodiments of the present disclosure relate generally to the field of picture processing, and more particularly to inter-prediction, and in particular to methods and apparatus for deriving interpolation filter indexes for a current block, such as merging procedures for parameters of switchable interpolation filters. [Background technology]
[0003] Video coding (video encoding and decoding) is used in a wide range of digital video applications, such as broadcast digital TV, video transmission over the Internet and mobile networks, real-time conversation applications such as video chat, video conferencing, DVD and Blu-ray discs, video content acquisition and editing systems, and camcorders in security applications.
[0004] The amount of video data required to render even a relatively short video can be significant, which can pose challenges when the data is to be streamed or otherwise transmitted over communication networks with limited bandwidth capacity. Therefore, video data is generally compressed before being transmitted over modern communication networks. Because memory resources may be limited, video size can also be an issue when the video is stored on a storage device. Often, video compression devices use software and / or hardware at the source to code the video data before transmission or storage, thereby reducing the amount of data needed to represent a digital video image. The compressed data is then received at the destination by a video decompression device, which decodes the video data. With limited network resources and an ever-increasing demand for higher video quality, improved compression and decompression techniques that increase compression ratios with little or no sacrifice in picture quality are desirable.
[0005] Recently, switchable interpolation filters for half-pixel (half-pel) positions have been introduced in Versatile Video Coding (VVC). Switching between half-pel luma interpolation filters is done depending on the precision of the motion vector. For half-pel motion vector precision, an alternative half-pel interpolation filter can be used, and the alternative half-pel interpolation filter is indicated by an additional syntax element indicating which interpolation filter is used, thus increasing the signaling overhead. Summary of the Invention [Means for solving the problem]
[0006] Embodiments of the present application can achieve inheritance of half-pixel interpolation filter indexes when a history-based motion information candidate list is used. Therefore, a dedicated interpolation filter is selected instead of the default interpolation filter, and it aims to provide an apparatus and method for constructing a history-based motion information candidate list so as to improve the quality of the prediction signal and the coding efficiency.
[0007] Embodiments of the present application can achieve inheritance of half-pixel interpolation filter indexes when a history-based motion information candidate list is used. Therefore, it aims to provide an apparatus and method for inter prediction for a current block coded in skip / merge mode so that the quality of the video signal can be improved.
[0008] The above and other objects are achieved by the subject matter of the independent claims. Further implementation forms are apparent from the dependent claims, the description, and the drawings.
[0009] According to a first aspect of the present invention, a method for constructing a history-based motion information (HMI) candidate list is provided. The method can be executed by an encoding device or a decoding device, and the method includes a step of obtaining a history-based motion information candidate list, where the HMI list is an ordered list of N history-based motion information candidates H related to (or including) the motion information of a plurality of preceding blocks (for example, N preceding blocks) preceding the block, k = 0,..., N - 1, N is an integer greater than 0 (for example, N is an integer greater than 0 and less than or equal to a predefined number (0 < N <= 5)), and each history-based motion information candidate is an element, that is, k and the like, and i) one or more motion vectors MV of the corresponding preceding blocks (such as luma motion vectors mvL0 and / or mvL1 with 1 / 16 fractional sample accuracy, where mvL0 and mvL1 correspond to the L0 and L1 reference picture lists); ii) one or more reference picture indices corresponding to the MV of the corresponding preceding block (such as reference picture indices refIdxL0 and / or refIdxL1, where refIdxL0 and refIdxL1 correspond to the L0 and L1 reference picture lists); and iii) Interpolation filter (IF) index (e.g., IF index of the corresponding preceding block or IF index related to the corresponding preceding block); and including motion information of the corresponding preceding block, updating the HMI list based on the block motion information, wherein the block motion information is based on the elements, i.e. i) one or more motion vectors MV for the block (such as luma motion vectors mvL0 and / or mvL1 with 1 / 16 fractional sample precision); ii) one or more reference picture indices (such as reference indices refIdxL0 and / or refIdxL1) corresponding to the MV of the block; and iii) Interpolation filter index (e.g., IF index of the block or IF index relative to the block) and Includes.
[0010] In an example, an interpolation filter (IF) index may refer to a fractional sample interpolation filter (IF) index, and in particular, an IF index refers to a half-pixel (half-pel) interpolation filter index or a half-sample interpolation filter index (hpelIfIdx). The terms “half-pixel interpolation filter” and “half-sample interpolation filter” may be used interchangeably in this disclosure. A half-sample interpolation filter index indicates a half-pixel interpolation filter used to interpolate half-pixel values when at least one of the motion vectors of the corresponding block points to a half-pixel position. For example, if one or more motion vectors (MVs) (element i) of a history-based motion information candidate have at least one MV pointing to a half-pixel position, the interpolation filter (IF) index (element iii) of the history-based motion information candidate indicates a half-pixel interpolation filter used to interpolate half-pixel values (i.e., the interpolation filter index (element iii) only has an effect for HMI candidates that include half-pel MVs). If one or more motion vectors (MVs) (element i) of the history-based motion information candidate do not have any MVs pointing to half-pel positions, the interpolation filter (IF) index (element iii) of the history-based motion information candidate becomes meaningless (i.e., this IF index does not have any effect on non-half-pel MVs, and the value of the IF index for non-half-pel MVs does not have any meaning. Its value can be set to any value, for example, 0 / FALSE). In this case, the interpolation filter (IF) index (element iii) is assigned a default value that will not be used in later steps. The same applies to the motion information of a block. In other words, the interpolation filter index (element iii) of the motion information of a block becomes meaningful when at least one of the MVs of the block points to a half-pel position. If none of the MVs point to a half-pel position, the interpolation filter index (element iii) of the motion information is assigned a default value that will not be used in later steps. In one exemplary implementation, the IF index is always stored in the HMI list regardless of the fractional part of the MV, even though the IF index is meaningless in some cases.The HMI list can be implemented in this way for design simplicity. It can be seen that if both MVs of the corresponding block do not point to half-pixel positions, the value assigned to the IF index does not have any effect on the decoding result.
[0011] In another exemplary implementation, the interpolation filter (IF) index may be replaced by an index of an interpolation filter (IF) set, and the index of the IF set indicates a switchable IF set among the multiple IF sets. In an example, each IF set includes an interpolation filter for each fractional position. Meanwhile, the IFs for the same fractional position may be equal among some IF sets. For example, among the multiple IF sets, the same filter may exist for some fractional positions and different filters may exist for some fractional positions, and in particular, the IFs for each fractional position may be switched according to the index of the IF set. In some cases, switching between two sets of interpolation filters may be understood as switching between two interpolation filters.
[0012] In one exemplary implementation, the half-pixel interpolation filter index indicates a half-pixel interpolation filter among a set of half-pixel interpolation filters, and the half-pixel interpolation filter is used to interpolate half-pixel values only when at least one of the one or more motion vectors points to a half-pixel position. If the L0 and / or L1 motion vector points to a half-pixel (half-pel) position, an interpolation filter is selected according to the half-pixel interpolation filter index and used for interpolation of samples during motion compensation for the corresponding prediction list (prediction direction) (L0 and / or L1).
[0013] It may be noted that the block and N previous blocks may be within a slice of a frame or within a frame. In an example, a history-based motion information candidate list (table) is emptied when a new slice is encountered. When a new slice is encountered, a construction process is invoked. In another example, the HMI list / table may be reset at the row of each new CTU within the slice.
[0014] It may be understood that the N preceding blocks may be one or more preceding blocks. A preceding block refers to an already coded or decoded block before the current block in coding or decoding order. In an example, block P may use an HMVP table that includes one or more coded / decoded blocks before block P. The HMVP table is updated after deriving motion information for block P. After the HMVP table is updated, block Q following block P may use the updated HMVP table. Block Q is coded or decoded after block P in decoding or coding order.
[0015] It can be understood that after updating the HMVP list, there may be M history-based motion information candidates in the updated HMVP list, where M is less than or equal to a predefined number (such as 5), and M>=N.
[0016] When the index of the HMI list starts from 1, the HMI list includes N history-based motion information candidates H related to the motion information of multiple preceding blocks preceding the block. k It can further be seen that k is an ordered list of k=1, . . . , N.
[0017] Therefore, an improved method is provided that enables the inheritance of interpolation filter indexes within a history-based motion information candidate list. In particular, the interpolation filter (IF) index of a preceding block is stored in the corresponding history-based motion information candidate in the history-based motion information candidate list. When the history-based motion information candidate list is used directly or indirectly for inter-prediction of a block coded in merge or skip mode, the interpolation filter (IF) index can be borrowed from the corresponding motion information candidate without using a separate syntax element. Propagating the IF index through the history-based motion information candidate list allows an appropriate interpolation filter to be used for the block (instead of using a predefined interpolation filter), which ensures the quality of the coded signal. As a result, the technology presented herein provides the advantage of improving coding efficiency and, consequently, improving the overall compression performance of the video coding method.
[0018] It is noted that the terms “block,” “coding block,” or “image block” used in this disclosure may include a transform unit (TU), a prediction unit (PU), a coding unit (CU), etc. In multi-objective video coding (VVC), the transform unit and the coding unit are mostly aligned, except in some scenarios when TU tiling or subblock transform (SBT) is used. It may be understood that the terms “block,” “image block,” “coding block,” and “picture block” may be used interchangeably herein. The terms “sample” and “pixel” may also be used interchangeably in this disclosure. The terms “predicted sample value” and “predicted pixel value” may be used interchangeably in this disclosure. The terms “sample location” and “pixel location” may be used interchangeably in this disclosure.
[0019] It should further be understood that the terms "history-based motion information candidate list," "HMI list," "HMVP list," "HMVP table," and "HMVP LUT" may be used interchangeably in this disclosure.
[0020] It should be understood that the HMVP list is constructed using motion information of one or more coded / decoded previous blocks. The HMVP list is used to store motion information from nearby blocks (but not necessarily from adjacent blocks like normal spatial merge candidates). The idea of HMVP is to use motion information from previous blocks (blocks from some spatial neighborhood) that are spatially close to a block but not necessarily adjacent to it.
[0021] In a possible implementation form of the method according to the first aspect itself, the step of updating the HMI list comprises updating the following elements of the history-based motion information candidates of each of the HMI lists: i) one or more motion vectors MV, and ii) One or more reference picture indices corresponding to the MV If at least one of the elements is different from the corresponding element of the block's motion information, the block's motion information is added to the HMI list as a candidate of history-based motion information H k where k = N.
[0022] If the index of the HMI list starts from 1, adding a history-based motion information candidate H that contains the motion information of the block in the HMI list. k It can be understood that this may refer to adding k = N+1.
[0023] It is enabled to add the motion information of a block as a motion information candidate based on the history of its last position in the HMI list.
[0024] In a possible implementation form of the method according to the first aspect itself, the step of updating the HMI list comprises: As a result of the comparison, the following elements of the candidate motion information based on the history of the HMI list, namely: i) one or more motion vectors MV, and ii) One or more reference picture indices corresponding to the MV is the same as the corresponding element of the block's motion information, delete the history-based motion information candidate from the HMI list, and add the block's motion information to the HMI list as the history-based motion information candidate H k where k = N-1.
[0025] If the index of the HMI list starts from 1, adding the block's motion information to the HMI list is done by adding the block's motion information to the HMI list. k It can be understood that this can refer to adding as k = N, where k = N.
[0026] The motion information of the block may be added as motion information candidate based on the history of its last position in the HMI list.
[0027] In any of the above-described implementations of the first aspect or a possible implementation of the method according to the first aspect itself, the step of updating the HMI list comprises: If N is equal to a predefined number, select the history-based motion information candidate H with k = 0 from the HMI list. k and list the block's motion information in the HMI list as the history-based motion information candidate H k The method includes adding the following steps:
[0028] If the index of the HMI list starts from 1, deleting the candidate H k Deleting can refer to deleting a block, where k = 1, and adding it to the HMI list based on the block's motion information history. kwhere k=N.
[0029] It is enabled to remove the history-based motion information candidate of the first position in the HMI list and add the motion information of the block as the history-based motion information candidate of the last position in the HMI list.
[0030] In any of the above-described implementations of the first aspect or a possible implementation of a method according to the first aspect itself, the method comprises: comparing whether the motion vector of any history-based motion information candidate is the same as the corresponding motion vector of the block; and comparing whether the reference picture index of any history-based motion information candidate is the same as the corresponding reference picture index of the block.
[0031] In an alternative design, the method comprises: comparing whether at least one of the motion vectors of each history-based motion information candidate (i.e., HMVP candidate) is different from the corresponding motion vector of the block; and comparing whether at least one of the reference picture indexes of each HMVP candidate is different from the corresponding reference picture index of the block.
[0032] Therefore, it is possible to use only MVs and reference picture indices in the pruning process while updating the HMVP table without comparing interpolation filter indices. Therefore, a good tradeoff between complexity and diversity of HMVP candidates can be achieved. In particular, enabling comparisons based only on MVs and reference picture indices can avoid additional computational operations and reduce computational complexity. Each comparison operation incurs additional computation during the HMVP table update process and the merge candidate construction process. Therefore, if comparison operations can be reduced or eliminated, computational complexity can be reduced, thereby improving coding efficiency. Furthermore, enabling comparisons based only on MVs and reference picture indices can preserve the diversity of HMVP records. Having two HMVP records that have the same MVs and reference indices and differ only in their IF indices is inefficient because these two records are not sufficiently different. Therefore, it is reasonable to consider these two HMVP records to be the same during the HMVP table update process. In this case, a new record that differs from an existing record only in its IF index is not added to the HMVP table. As a result, "old" / existing records that are "sufficiently different" (have different MV or reference indexes) from the other records are maintained. In other words, in order for a new record to be added to an HMVP table, this new record should not just be bitwise different from an existing record; this new record must be "significantly different." From a coding efficiency perspective, it is more efficient to have two records in an HMVP table that have different MV or reference indexes than two records that differ only by their IF index.
[0033] In any of the above-mentioned implementations of the first aspect or in a possible implementation of the method according to the first aspect itself, the predefined number is five or six.
[0034] In any of the above-mentioned implementations of the first aspect or a possible implementation form of a method according to the first aspect itself, a half-sample interpolation filter index included in the history-based motion information candidate indicates a half-sample interpolation filter from a set of half-sample interpolation filters, and the half-sample interpolation filter is applied to interpolate a half-sample value only when at least one of the one or more MVs of the history-based motion information candidate points to a half-sample position.
[0035] In the prior art, a default IF index (corresponding to a default IF) was always used for a merge candidate obtained from an HMVP table. According to the present invention, an IF index is propagated through the HMVP table, so that one of a set of interpolation filters can be used according to the IF index, and in the example, one of two interpolation filters (a default interpolation filter and an alternative interpolation filter) can be used according to the IF index. Thus, a dedicated interpolation filter is selected instead of the default interpolation filter, which in turn increases the reliability of the reference and thus improves the quality of the predicted signal and the coding efficiency.
[0036] It is noted that the terms "alternative half-pixel interpolation filter," "switchable interpolation filter (SIF)," or "half-pixel interpolation filter" may be used interchangeably in this disclosure.
[0037] An appropriate interpolation filter (IF) can be selected depending on the content. For regions with sharp edges, a regular DCT-based IF can be used. For smooth regions (or when preserving sharp edges is not required), an alternative 6-tap IF (Gaussian filter) can be used. For merge mode, this IF index is borrowed from the corresponding motion information candidate. For blocks coded in merge mode, an alternative IF can be used when the motion information candidate is obtained from the HMVP table. Propagating the IF index through the HMVP table enables the use of an appropriate IF for a block, which has the advantage of increasing coding efficiency. Without the proposed mechanism, a default IF index (corresponding to an 8-tap DCT-based IF) would always be used for HMVP merge candidates, and the content details of the current block (whether sharp edges need to be preserved or not) cannot be taken into account.
[0038] According to a second aspect of the present invention, there is provided a method for inter prediction for a block in a frame of a video signal, the method comprising: constructing a historical motion information candidate (HMI) list, the HMI list being a list of N historical motion information candidates H related to (or including) motion information of a plurality of preceding blocks (e.g., N preceding blocks) preceding the block; k where k=0, ..., N-1, N is an integer greater than 0, and each history-based motion information candidate corresponds to a previous block, with elements, i.e., i) one or more motion vectors MV of the preceding block; ii) one or more reference picture indices corresponding to the MVs of the preceding blocks; and iii) an interpolation filter index (e.g., an interpolation filter index of the preceding block or an interpolation filter index related to the preceding block); and adding one or more history-based motion information candidates from the HMI list to a motion information candidate list for the block; deriving motion information for the block based on the motion information candidate list; Includes.
[0039] The motion information candidate list may be a merge candidate list.
[0040] In an alternative or additional design, according to a second aspect of the present invention, there is provided a method for inter prediction for a block in a frame of a video signal, the method comprising: constructing a history-based motion information candidate list, the HMI list including N history-based motion information candidates H associated with (or including) motion information of a plurality of preceding blocks (e.g., N preceding blocks) preceding the block; k where k=0, ..., N-1, N is an integer greater than 0, and at least one history-based motion information candidate is i) one or more motion vectors (MVs), where at least one of the MVs points to a half-pixel position; ii) one or more reference picture indices corresponding to one or more MVs; and iii) Interpolation filter index of the preceding block a step of including an element for a corresponding predecessor block, adding one or more history-based motion information candidates from the HMI list to a motion information candidate list for the block; and deriving motion information for the block based on the motion information candidate list.
[0041] The motion information candidate list may be a merge candidate list.
[0042] It can be appreciated that the history-based motion information candidates are added to the merge candidate list as history-based merge candidates.
[0043] In the example, the HMI list has a length of N, where N is 5 or 6.
[0044] Therefore, an improved method is provided that enables the inheritance of interpolation filter indexes within a history-based motion information candidate list. In particular, the interpolation filter (IF) index of a preceding block is stored in the corresponding history-based motion information candidate in the history-based motion information candidate list. When the history-based motion information candidate list is used directly or indirectly for inter-prediction of a block coded in merge or skip mode, the interpolation filter (IF) index can be borrowed from the corresponding motion information candidate without using a separate syntax element. Propagating the IF index through the history-based motion information candidate list allows an appropriate interpolation filter to be used for the block (instead of using a predefined interpolation filter), which ensures the quality of the coded signal. As a result, the technology presented herein provides the advantage of improving coding efficiency and, consequently, improving the overall compression performance of the video coding method.
[0045] In a possible implementation form of the method according to the second aspect itself, the half-sample interpolation filter is applied only when at least one of the one or more MVs of the derived motion information points to a half-sample position, and the half-sample interpolation filter is indicated by a half-sample interpolation filter index included in the derived motion information.
[0046] In any of the above-mentioned implementations of the second aspect or a possible implementation form of a method according to the second aspect itself, a half-sample interpolation filter index included in the history-based motion information candidate indicates a half-sample interpolation filter from a set of half-sample interpolation filters, and the half-sample interpolation filter is applied to interpolate a half-sample value only when at least one of the one or more MVs of the history-based motion information candidate points to a half-sample position.
[0047] In any of the above-described implementations of the second aspect or in a possible implementation form of the method according to the second aspect itself, the history-based motion information candidate further includes one or more bi-prediction weight indexes. The term bi-prediction weight index bcw_idx is also referred to as a generalized bi-prediction weight index GBIdx and / or a bi-prediction with CU-level weights (BCW) index. Alternatively, this index may be abbreviated as BWI, which simply refers to a bi-prediction weight index.
[0048] In any of the above-mentioned implementations of the second aspect or possible implementations of the method according to the second aspect itself, The following elements of the candidate motion information based on each history in the HMI list: i) one or more motion vectors MV, and ii) One or more reference picture indices corresponding to the MV If at least one of the elements is different from the corresponding element of the block's motion information, the block's motion information is added to the HMI list as a candidate of history-based motion information H k where k=N.
[0049] In any of the above-mentioned implementations of the second aspect or possible implementations of the method according to the second aspect itself, The following elements of candidate motion information based on the history of the HMI list: i) one or more motion vectors MV, and ii) One or more reference picture indices corresponding to the MV is the same as the corresponding element of the block's motion information, delete the history-based motion information candidate from the HMI list, and add the block's motion information to the HMI list as the history-based motion information candidate H k where k=N−1.
[0050] In any of the above-mentioned implementations of the second aspect or possible implementations of the method according to the second aspect itself, If N is equal to a predefined number, select the history-based motion information candidate H with k = 0 from the HMI list. k and list the block's motion information in the HMI list as the history-based motion information candidate H k In an example, the predefined number is 5.
[0051] In any of the above-mentioned implementations of the second aspect or a possible implementation of a method according to the second aspect itself, the method comprises: comparing whether the corresponding motion vector of any history-based motion information candidate is the same as the motion vector of the block; and comparing whether the corresponding reference picture index of any history-based motion information candidate is the same as the reference picture index of the block.
[0052] In an alternative design, the comparing step may include: comparing whether at least one of the motion vectors of each history-based motion information candidate is different from the corresponding motion vector of the block; and comparing whether at least one of the reference picture indexes of each HMVP candidate is different from the corresponding reference picture index of the block.
[0053] In one of the above-described implementations of the second aspect or in a possible implementation form of the method according to the second aspect itself, the motion information candidate list is used for merge mode or skip mode, in other words, the current block is coded in merge mode or skip mode.
[0054] In any of the above-mentioned implementations of the second aspect or a possible implementation form of the method according to the second aspect itself, the step of deriving motion information for the block based on the motion information candidate list comprises: The method includes deriving motion information referenced by a candidate index from a motion information candidate list as motion information for the current block, where the candidate index is parsed or derived from the bitstream.
[0055] In any of the above-mentioned implementations of the second aspect or possible implementations of the method according to the second aspect itself, When at least one of one or more motion vectors MV included in the derived motion information points to a half-pixel position, obtaining a predicted sample value of the block by applying a half-pixel interpolation filter to pixel values of a reference picture pointed to by the MV, wherein the half-pixel interpolation filter is indicated by an interpolation filter index included in the derived motion information; When the motion vector MV included in the derived motion information does not point to a half-pixel position, the method further includes a step of obtaining a predicted sample value of the block by applying a default interpolation filter to pixel values of the reference picture pointed to by the MV.
[0056] The encoding and decoding methods defined in the claims, the description and the figures can be performed by an encoding device and a decoding device, respectively.
[0057] According to a third aspect of the present invention, there is provided an apparatus for constructing a history-based motion information candidate list, the apparatus comprising: a history-based motion information candidate list obtaining unit configured to obtain a history-based motion information candidate list, wherein the HMI list includes N history-based motion information candidates H related to motion information of a plurality of blocks preceding the block; k where k=0, ..., N-1, N is an integer greater than 0, and each history-based motion information candidate is an element, i.e., i) one or more motion vectors MV; ii) one or more reference picture indices corresponding to the MV; and iii) Interpolation filter index a history-based motion information candidate list obtaining unit, a history-based motion information candidate list updating unit configured to update an HMI list based on motion information of a block, wherein the motion information of the block includes elements, namely: i) one or more motion vectors MV; ii) one or more reference picture indices corresponding to the MV; and iii) Interpolation filter index and a history-based motion information candidate list update unit, including:
[0058] The method according to the first aspect of the invention may be performed by an apparatus according to the third aspect of the invention, the further features and modes of implementation of which correspond to those of the apparatus according to the first aspect of the invention.
[0059] According to a fourth aspect of the present invention, there is provided an apparatus for inter prediction for a block, the apparatus comprising: a list management unit configured to construct a history-based motion information candidate list, the HMI list comprising N history-based motion information candidates H associated with motion information of a plurality of blocks preceding the block; kwhere k=0, ..., N-1, N is an integer greater than 0, and each history-based motion information candidate is an element, i.e., i) one or more motion vectors MV; ii) one or more reference picture indices corresponding to the MV; and iii) Interpolation filter index a list management unit, further configured to add one or more history-based motion information candidates from the HMI list to a motion information candidate list for the block; a motion information derivation unit configured to derive motion information for the block based on the motion information candidate list; Includes.
[0060] The method according to the second aspect of the invention may be performed by an apparatus according to the fourth aspect of the invention. Further features and modes of implementation of the apparatus according to the fourth aspect of the invention correspond to the features and modes of implementation of the apparatus according to the second aspect of the invention.
[0061] According to a fifth aspect, the invention relates to an encoder (20) including processing circuitry for carrying out the method according to the first or second aspect per se or in the form of an implementation thereof.
[0062] According to a sixth aspect, the present invention relates to a decoder (30) including processing circuitry for carrying out the method according to the first or second aspect itself or in the form of an implementation thereof.
[0063] According to a seventh aspect, the present invention relates to a decoder, comprising: one or more processors; and a non-transitory computer-readable storage medium coupled to the processor and storing programming for execution by the processor, the programming, when executed by the processor, configuring the decoder to perform a method according to the first or second aspect itself or an implementation thereof.
[0064] According to an eighth aspect, the present invention relates to an encoder, comprising: one or more processors; and a non-transitory computer-readable storage medium coupled to the processor and storing programming for execution by the processor, the programming, when executed by the processor, configuring the encoder to perform a method according to the first or second aspect itself or an implementation thereof.
[0065] According to a ninth aspect, the present invention relates to a non-transitory storage medium comprising a bitstream encoded / decoded by the method of any one of the above aspects.
[0066] An apparatus for encoding or decoding a video stream may include a processor and a memory, the memory storing instructions for causing the processor to perform a method according to any one of the above aspects.
[0067] For each of the encoding or decoding methods disclosed herein, a computer-readable storage medium is proposed, the storage medium storing thereon instructions that, when executed, cause one or more processors to encode or decode video data, the instructions causing the one or more processors to perform a method according to any one of the above aspects.
[0068] Furthermore, for each of the encoding or decoding methods disclosed herein, a computer program product is proposed, the computer program product comprising a program code for performing the method according to any one of the above aspects.
[0069] The details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will become apparent from the description, drawings, and claims.
[0070] In the following, embodiments of the invention will be explained in more detail with reference to the accompanying figures and drawings. [Brief explanation of the drawings]
[0071] [Figure 1A] 1 is a block diagram illustrating an example of a video coding system configured to implement embodiments of the present invention. [Figure 1B] FIG. 2 is a block diagram illustrating another example of a video coding system configured to implement embodiments of the present invention. [Figure 2] 1 is a block diagram illustrating an example of a video encoder configured to implement embodiments of the present invention. [Figure 3] 1 is a block diagram illustrating an exemplary structure of a video decoder configured to implement embodiments of the present invention. [Figure 4] FIG. 1 is a block diagram illustrating an example of an encoding device or a decoding device. [Figure 5] FIG. 10 is a block diagram showing another example of an encoding device or a decoding device. [Figure 6] FIG. 2 is a diagram illustrating an example of a current block and its spatial neighboring blocks; [Figure 7] FIG. 2 is a diagram illustrating a current block and an upper neighboring block; [Figure 8] 1 is a flow diagram of a method according to an embodiment of the present disclosure. [Figure 9]1 is a block diagram illustrating a method for deriving an interpolation filter index for a current block (e.g., a coding unit or coding block) within a coding tree block (CTB) or coding tree unit (CTU). [Figure 10] FIG. 10 is a diagram illustrating an example of building an HMVP list according to an embodiment of the present disclosure. [Figure 11] FIG. 10 is a diagram illustrating another example of building an HMVP list according to an embodiment of the present disclosure. [Figure 12] FIG. 10 illustrates an example of an HMVP list and its traversal order according to an embodiment of the present disclosure. [Figure 13A] 1 is a flow diagram illustrating an example of a method for constructing an HMVP list. [Figure 13B] 10 is a flow diagram illustrating another example of a method for constructing an HMVP list. [Figure 14] 1 is a flow diagram illustrating an example of a method for inter prediction for blocks within a frame of a video signal. [Figure 15] 1 is a flow diagram illustrating an example of an HMI list update method. [Figure 16] FIG. 1 is a block diagram of an apparatus according to an embodiment of the present disclosure. [Figure 17] FIG. 2 is a block diagram of another apparatus according to an embodiment of the present disclosure. [Figure 18] 1 is a block diagram illustrating an exemplary structure of a content supply system for implementing a content distribution service. [Figure 19] FIG. 2 is a block diagram illustrating the structure of an example terminal device. DETAILED DESCRIPTION OF THE INVENTION
[0072] In the following, the same reference signs, unless otherwise specified, refer to identical or at least functionally equivalent features.
[0073] In the following description, reference is made to the accompanying drawings which form a part of this disclosure and which show, by way of illustration, certain aspects of embodiments of the invention or in which embodiments of the invention may be practiced. It is understood that embodiments of the invention may be practiced in other ways and may include structural or logical changes not shown in the drawings. Therefore, the following detailed description is not to be taken in a limiting sense, and the scope of the present invention is defined by the appended claims.
[0074] For example, it is understood that disclosure related to a described method can also apply to a corresponding device or system configured to perform the method, and vice versa. For example, when one or more particular method steps are described, a corresponding device may include one or more units, e.g., functional units, for performing the described one or more method steps (e.g., one unit that performs one or more steps, or multiple units that each perform one or more of the multiple steps), even if such one or more units are not explicitly described or shown in a figure. On the other hand, for example, when a particular apparatus is described based on one or more units, e.g., functional units, a corresponding method may include one step for performing the function of the one or more units (e.g., one step that performs the function of one or more units, or multiple steps that each perform one or more functions of multiple units), even if such one or more steps are not explicitly described or shown in a figure. Furthermore, it is understood that features of various exemplary embodiments and / or aspects described herein can be combined with each other unless expressly stated otherwise.
[0075] Video coding generally refers to the processing of a sequence of pictures that form a video or a video sequence. Instead of the term "picture," the terms "frame" or "image" may be used synonymously in the field of video coding. Video coding (or coding in general) includes two parts: video encoding and video decoding. Video encoding is performed on the source side and generally involves processing the original video picture (e.g., by compression) to reduce the amount of data needed to represent the video picture (for more efficient storage and / or transmission). Video decoding is performed on the destination side and generally involves the reverse processing compared to the encoder to reconstruct the video picture. Embodiments referring to "coding" a video picture (or pictures in general) are understood to relate to "encoding" or "decoding" the video picture or respective video sequence. The combination of the encoding and decoding parts is also called a codec (coding and decoding).
[0076] In the case of lossless video coding, the original video picture can be reconstructed (assuming there is no transmission loss or other data loss during storage or transmission), i.e., the reconstructed video picture has the same quality as the original video picture. In the case of lossy video coding, further compression, for example by quantization, is performed to reduce the amount of data representing the video picture, which cannot be perfectly reconstructed at the decoder, i.e., the quality of the reconstructed video picture is lower or worse than the quality of the original video picture.
[0077] Some video coding standards belong to the group of "lossy hybrid video codecs" (i.e., combine spatial and temporal prediction in the sample domain with 2D transform coding for applying quantization in the transform domain). Each picture of a video sequence is generally partitioned into a set of non-overlapping blocks, and coding is generally performed at the block level. In other words, at an encoder, video is generally processed, i.e., encoded, at the block (video block) level, for example, by generating a prediction block using spatial (intra-picture) prediction and / or temporal (inter-picture) prediction, subtracting the prediction block from a current block (the block currently being / to be processed) to obtain a residual block, transforming the residual block, and quantizing the residual block in the transform domain to reduce the amount of data to be transmitted (compression); whereas at a decoder, an inverse process is applied to the coded or compressed block compared to the encoder to reconstruct the current block for representation. Furthermore, the encoder replicates the decoder's processing loop so that both generate the same prediction (eg, intra and inter prediction) and / or reconstruction for processing, i.e., coding, subsequent blocks.
[0078] In the following, embodiments of a video coding system 10, a video encoder 20 and a video decoder 30 are described based on FIGS.
[0079] 1A is a schematic block diagram illustrating an example coding system 10 that may utilize the techniques of the present application, e.g., video coding system 10 (or coding system 10 for short). A video encoder 20 (or encoder 20 for short) and a video decoder 30 (or decoder 30 for short) of video coding system 10 illustrate examples of devices that may be configured to perform techniques according to various examples described in the present application.
[0080] As shown in FIG. 1A, coding system 10 includes a source device 12 configured to provide encoded picture data 21 to, for example, a destination device 14 for decoding the encoded picture data 13.
[0081] The source device 12 includes an encoder 20 and may additionally, i.e., optionally, include a picture source 16, a preprocessor (or preprocessing unit) 18, for example, a picture preprocessor 18, and a communication interface or communication unit 22.
[0082] Picture source 16 may include or be any kind of picture capturing device, e.g., a camera for capturing real-world pictures, and / or any kind of picture generating device, e.g., a computer graphics processor for generating computer-animated pictures, or any kind of other device for acquiring and / or providing real-world pictures, computer-generated pictures (e.g., screen content, virtual reality (VR) pictures), and / or any combination thereof (e.g., augmented reality (AR) pictures). Picture source may be any kind of memory or storage device for storing any of the above-mentioned pictures.
[0083] To distinguish from the preprocessor 18 and the processing performed by the preprocessing unit 18, the picture or picture data 17 may also be referred to as a raw picture or raw picture data 17.
[0084] The pre-processor 18 is configured to receive (raw) picture data 17 and perform pre-processing on the picture data 17 to obtain a pre-processed picture 19 or pre-processed picture data 19. The pre-processing performed by the pre-processor 18 may include, for example, cropping, color format conversion (e.g., from RGB to YCbCr), color correction, or noise removal. It may be understood that the pre-processing unit 18 may be any component.
[0085] Video encoder 20 is configured to receive pre-processed picture data 19 and provide encoded picture data 21 (further details are described below, eg, with reference to FIG. 2).
[0086] The communication interface 22 of the source device 12 may be configured to receive the encoded picture data 21 and transmit the encoded picture data 21 (or any further processed version thereof) via a communication channel 13 to another device, e.g., a destination device 14 or any other device, for storage or direct reconstruction.
[0087] The destination device 14 includes a decoder 30 (e.g., a video decoder 30) and may additionally, i.e., optionally, include a communication interface or communication unit 28, a post-processor 32 (or post-processing unit 32), and a display device 34.
[0088] The communications interface 28 of the destination device 14 is configured to receive the encoded picture data 21 (or any further processed version thereof), for example, directly from the source device 12 or from any other source, for example, a storage device, for example, a storage device for encoded picture data, and to provide the encoded picture data 21 to the decoder 30.
[0089] The communication interface 22 and the communication interface 28 may be configured to transmit or receive the encoded picture data 21 or the encoded data 13 via a direct communication link between the source device 12 and the destination device 14, e.g., a direct wired or wireless connection, or via any type of network, e.g., a wired or wireless network or any combination thereof, or any type of private and public network, or any type of combination thereof.
[0090] The communications interface 22 may be configured, for example, to package the encoded picture data 21 into a suitable format, e.g., packets, and / or process the encoded picture data using any type of transmission encoding or processing for transmission over a communications link or network.
[0091] The communications interface 28 forming the counterpart of the communications interface 22 may be configured, for example, to receive the transmitted data and process the transmitted data using any type of corresponding transmission decoding or processing and / or depackaging to obtain the encoded picture data 21.
[0092] Both communication interface 22 and communication interface 28 may be configured as unidirectional communication interfaces, as indicated by the arrows for communication channel 13 in FIG. 1A pointing from source device 12 toward destination device 14, or as bidirectional communication interfaces, and may be configured, for example, to send and receive messages, for example, to set up connections and to confirm and exchange communication links and / or any other information related to data transmission, e.g., transmission of encoded picture data.
[0093] The decoder 30 is configured to receive encoded picture data 21 and provide decoded picture data 31 or decoded pictures 31 (further details are described below, for example, based on Figure 3 or Figure 5).
[0094] Post-processor 32 of destination device 14 is configured to post-process decoded picture data 31 (also referred to as reconstructed picture data), e.g., decoded picture 31, to obtain post-processed picture data 33, e.g., post-processed picture 33. The post-processing performed by post-processing unit 32 may include, e.g., color format conversion (e.g., from YCbCr to RGB), color correction, cropping, or resampling, or any other processing to prepare, e.g., decoded picture data 31, for display by, e.g., display device 34.
[0095] Display device 34 of destination device 14 is configured to receive post-processed picture data 33, for example, to display the picture to a user or viewer. Display device 34 may be or include any type of display for showing the reconstructed picture, e.g., an integrated or external display or monitor. The display may include, for example, a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a microLED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display.
[0096] 1A depicts source device 12 and destination device 14 as separate devices, an embodiment of the devices may include both or both functionality: source device 12 or corresponding functionality and destination device 14 or corresponding functionality. In such an embodiment, source device 12 or corresponding functionality and destination device 14 or corresponding functionality may be implemented using the same hardware and / or software or by separate hardware and / or software or any combination thereof.
[0097] As will be apparent to those skilled in the art based on the description, the functions of different units or the presence and (exact) division of functions within source device 12 and / or destination device 14 shown in FIG. 1A may vary depending on the actual device and application.
[0098] Encoder 20 (e.g., video encoder 20) or decoder 30 (e.g., video decoder 30), or both encoder 20 and decoder 30, may be implemented by processing circuitry shown in FIG. 1B, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, hardware, dedicated to video coding, or any combination thereof. Encoder 20 may be implemented by processing circuitry 46 to embody various modules discussed in connection with encoder 20 of FIG. 2 and / or any other encoder system or subsystem described herein. Decoder 30 may be implemented by processing circuitry 46 to embody various modules discussed in connection with decoder 30 of FIG. 3 and / or any other decoder system or subsystem described herein. The processing circuitry may be configured to perform various operations discussed later. If the techniques are implemented partially in software, as shown in FIG. 5, a device may store instructions for the software on a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Either video encoder 20 and video decoder 30 may be incorporated as part of a combined encoder / decoder (codec) within a single device, for example, as shown in FIG. 1B.
[0099] Source device 12 and destination device 14 may include any of a wide range of devices, including any type of handheld or fixed device, e.g., a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (such as a content service server or content distribution server), a broadcast receiver device, a broadcast transmitter device, etc., and may use no operating system or any type of operating system. In some cases, source device 12 and destination device 14 may be capable of wireless communication. Thus, source device 12 and destination device 14 may be wireless communication devices.
[0100] In some cases, the video coding system 10 shown in FIG. 1A is merely an example, and the techniques of this disclosure may be applied to video coding situations (e.g., video encoding or video decoding) that do not necessarily involve any data communication between an encoding device and a decoding device. In other examples, data may be retrieved from local memory, streamed over a network, etc. A video encoding device may encode data and store it in memory, and / or a video decoding device may retrieve and decode data from memory. In some examples, encoding and decoding are performed by devices that do not communicate with each other but simply encode data to and / or retrieve and decode data from memory.
[0101] For ease of explanation, embodiments of the present invention are described herein with reference to reference software, e.g., High-Efficiency Video Coding (HEVC), or Versatile Video Coding (VVC), the next-generation video coding standard developed by the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Joint Collaboration Team on Video Coding (JCT-VC) of the Motion Picture Experts Group (MPEG). Those skilled in the art will understand that embodiments of the present invention are not limited to HEVC or VVC.
[0102] Encoder and encoding method FIG. 2 shows a schematic block diagram of an exemplary video encoder 20 configured to implement the techniques of the present application. In the example of FIG. 2, the video encoder 20 includes an input 201 (or input interface 201), a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a loop filter unit 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy coding unit 270, and an output 272 (or output interface 272). The mode selection unit 260 may include an inter prediction unit 244, an intra prediction unit 254, and a partitioning unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 shown in FIG. 2 may also be referred to as a hybrid video encoder or a video encoder using a hybrid video codec.
[0103] The residual calculation unit 204, the transform processing unit 206, the quantization unit 208, and the mode selection unit 260 may be considered to form a forward signal path of the encoder 20, while the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 may be considered to form a backward signal path of the video encoder 20, which corresponds to the signal path of a decoder (see video decoder 30 in FIG. 3 ). The inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 may also be considered to form a “built-in decoder” of the video encoder 20.
[0104] Picture & Picture Division (Picture & Block) Encoder 20 may be configured, for example, to receive via input 201 picture 17 (or picture data 17), e.g., a picture of a sequence of pictures forming a video or a video sequence. The received picture or picture data may also be preprocessed picture 19 (or preprocessed picture data 19). For simplicity, the following description refers to picture 17. Picture 17 may also be called a current picture or a picture to be coded (particularly in video coding, to distinguish the current picture from other pictures, e.g., already coded and / or decoded pictures of the same video sequence, i.e., the video sequence that also includes the current picture).
[0105] A (digital) picture is or can be considered as a two-dimensional array or matrix of samples having intensity values. The samples of the array may also be called pixels (short for picture element) or pels. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the picture. For color representation, three color components are generally used, i.e., a picture may be represented or include three sample arrays. In an RBG format or color space, a picture includes corresponding red, green, and blue sample arrays. However, in video coding, each pixel is generally represented in a luminance and chrominance format or color space, e.g., YCbCr, which includes a luminance component denoted by Y (although L may be used instead) and two chrominance components denoted by Cb and Cr. The luminance (or luma for short) component Y represents brightness or gray level intensity (e.g., similar to a grayscale picture), while the two chrominance (or chroma for short) components Cb and Cr represent chromaticity or color information components. Thus, a picture in YCbCr format includes a luminance sample array of luminance sample values (Y) and two chrominance sample arrays of chrominance values (Cb and Cr). A picture in RGB format can be converted or transformed to YCbCr format, and vice versa, a process also known as color transformation or conversion. If a picture is monochrome, the picture may include only a luminance sample array. Thus, a picture may be, for example, an array of luma samples in a monochrome format, or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0106] Embodiments of video encoder 20 may include a picture partitioning unit (not shown in FIG. 2) configured to partition picture 17 into multiple (usually non-overlapping) picture blocks 203. These blocks may also be called root blocks, macroblocks (H.264 / AVC), or coding tree blocks (CTBs) or coding tree units (CTUs) (H.265 / HEVC and VVC). The picture partitioning unit may be configured to use the same block size for all pictures of a video sequence and a corresponding grid that defines the block size, or to vary the block size among pictures or subsets or groups of pictures, and to partition each picture into corresponding blocks.
[0107] In further embodiments, the video encoder may be configured to directly receive blocks 203 of picture 17, e.g., one, some, or all of the blocks that form picture 17. Picture blocks 203 may also be referred to as current picture blocks or picture blocks to be coded.
[0108] Similar to picture 17, picture block 203, although smaller in dimensions than picture 17, is also considered or can be considered a two-dimensional array or matrix of samples having intensity values (sample values). In other words, block 203 may include, for example, one sample array (e.g., a luma array for a monochrome picture 17, or a luma or chroma array for a color picture), or three sample arrays (e.g., a luma and two chroma arrays for a color picture 17), or any other number and / or type of array, depending on the applied color format. The number of samples in the horizontal and vertical directions (or axes) of block 203 defines the size of block 203. Thus, a block may be, for example, an MxN (M columns by N rows) array of samples or an MxN array of transform coefficients.
[0109] The embodiment of video encoder 20 shown in FIG. 2 may be configured to encode picture 17 block by block, eg, encoding and prediction is performed for each block 203.
[0110] Calculating residuals The residual calculation unit 204 may be configured to calculate the residual block 205 (also referred to as the residual 205) based on the picture block 203 and the predictive block 265 (further details about the predictive block 265 are provided later), for example, by subtracting the sample values of the predictive block 265 from the sample values of the picture block 203 on a sample-by-sample (pixel-by-pixel) basis to obtain the residual block 205 in the sample domain.
[0111] conversion The transform processing unit 206 may be configured to apply a transform, such as a discrete cosine transform (DCT) or a discrete sine transform (DST), to the sample values of the residual block 205 to obtain transform coefficients 207 in a transform domain. The transform coefficients 207, also referred to as transform residual coefficients, may represent the residual block 205 in the transform domain.
[0112] The transform processing unit 206 may be configured to apply an integer approximation of a DCT / DST, such as the transform specified for H.265 / HEVC. Compared to an orthogonal DCT transform, such an integer approximation is generally scaled by a particular factor. To maintain the norm of the residual blocks processed by the forward and inverse transforms, an additional scaling factor is applied as part of the transform process. The scaling factor is generally selected based on particular constraints, such as the scaling factor being a power of two for shift operations, the bit depth of the transform coefficients, a trade-off between accuracy and implementation cost, etc. For example, a particular scaling factor may be specified for the inverse transform, e.g., by the inverse transform processing unit 212 (and the corresponding inverse transform, e.g., by the inverse transform processing unit 312 in the video decoder 30), and a corresponding scaling factor for the forward transform, e.g., by the transform processing unit 206 of the encoder 20, may be specified accordingly.
[0113] An embodiment of the video encoder 20 (respectively, the transform processing unit 206) may be configured to output transform parameters, e.g., a certain transform or transforms, either as is or encoded or compressed by the entropy coding unit 270, so that the video decoder 30 may receive the transform parameters and use them for decoding.
[0114] Quantization The quantization unit 208 may be configured to quantize the transform coefficients 207, for example, by applying scalar quantization or vector quantization, to obtain quantized coefficients 209. The quantized coefficients 209 may also be referred to as quantized transform coefficients 209 or quantized residual coefficients 209.
[0115] The quantization process may reduce the bit depth associated with some or all of the transform coefficients 207. For example, an n-bit transform coefficient may be truncated to an m-bit transform coefficient during quantization, where n is greater than m. The degree of quantization may be modified by adjusting a quantization parameter (QP). For example, for scalar quantization, different scaling may be applied to achieve finer or coarser quantization. A smaller quantization step size corresponds to finer quantization, while a larger quantization step size corresponds to coarser quantization. An applicable quantization step size may be indicated by the quantization parameter (QP). The quantization parameter may, for example, be an index into a predefined set of applicable quantization step sizes. For example, a small quantization parameter may correspond to fine quantization (small quantization step size) and a large quantization parameter may correspond to coarse quantization (large quantization step size), or vice versa. Quantization may include division by a quantization step size, and corresponding and / or inverse dequantization by the inverse quantization unit 210 may include multiplication by the quantization step size. Some standards, e.g., HEVC, may be configured to determine the quantization step size using a quantization parameter. Generally, the quantization step size may be calculated based on the quantization parameter using a fixed-point approximation of an equation involving division. Additional scaling factors may be introduced for quantization and dequantization to restore the norm of the residual block, which may be modified due to scaling used in the fixed-point approximation of the equation for the quantization step size and the quantization parameter. In one exemplary implementation, the scaling of the inverse transform and dequantization may be combined. Alternatively, customized quantization tables may be used, e.g., signaled from the encoder to the decoder in the bitstream. Quantization is a lossy operation, and loss increases as the quantization step size increases.
[0116] Embodiments of video encoder 20 (respectively, quantization unit 208) may be configured to output a quantization parameter (QP), e.g., as is or encoded by entropy encoding unit 270, such that video decoder 30 may receive the quantization parameter and apply it for decoding.
[0117] inverse quantization Inverse quantization unit 210 is configured to apply the inverse quantization of quantization unit 208 to the quantized coefficients to obtain dequantized coefficients 211, e.g., by applying the inverse of the quantization scheme applied by quantization unit 208, based on or using the same quantization step size as quantization unit 208. The dequantized coefficients 211, also referred to as dequantized residual coefficients 211, may correspond to transform coefficients 207—although they are generally not identical to the transform coefficients due to loss due to quantization.
[0118] Inverse transformation The inverse transform processing unit 212 is configured to apply an inverse transform of the transform applied by the transform processing unit 206, for example, an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST) or other inverse transform, to obtain a reconstructed residual block 213 (or corresponding dequantized coefficients 213) in the sample domain. The reconstructed residual block 213 may also be referred to as a transform block 213.
[0119] Rebuild The reconstruction unit 214 (e.g., an adder or summator 214) is configured to add the transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265, for example, by adding the sample values of the reconstructed residual block 213 and the sample values of the prediction block 265—sample by sample—to obtain a reconstructed block 215 in the sample domain.
[0120] filtering The loop filter unit 220 (or "loop filter" 220 for short) is configured to filter the reconstructed block 215 to obtain a filtered block 221, or generally, to filter reconstructed samples to obtain filtered samples. The loop filter unit is configured, for example, to smooth pixel transitions or otherwise improve video quality. The loop filter unit 220 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, e.g., a bilateral filter, an adaptive loop filter (ALF), a sharpening filter, a smoothing filter, or a collaborative filter, or any combination thereof. Although the loop filter unit 220 is shown in FIG. 2 as being an in-loop filter, in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may also be referred to as a filtered reconstructed block 221.
[0121] Embodiments of video encoder 20 (respectively, loop filter unit 220) may be configured to output loop filter parameters (such as sample adaptive offset information) either as is or encoded by entropy coding unit 270, e.g., so that decoder 30 may receive and apply the same or respective loop filter parameters for decoding.
[0122] Decoded Picture Buffer The decoded picture buffer (DPB) 230 may be a memory that stores reference pictures or generally reference picture data for encoding video data by video encoder 20. The DPB 230 may be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The decoded picture buffer (DPB) 230 may be configured to store one or more filtered blocks 221. The decoded picture buffer 230 may be further configured to store other already-filtered blocks, e.g., already-reconstructed filtered blocks 221, of the same current picture or a different picture, e.g., an already-reconstructed picture, and may provide a complete already-reconstructed, i.e., decoded, picture (and corresponding reference blocks and samples) and / or a partially-reconstructed current picture (and corresponding reference blocks and samples), e.g., for inter-prediction. The decoded picture buffer (DPB) 230 may also be configured to store one or more unfiltered reconstructed blocks 215 or generally unfiltered reconstructed samples, for example, if the reconstructed blocks 215 are not filtered by the loop filter unit 220, or to store any other further processed version of the reconstructed blocks or samples.
[0123] Mode selection (classification & prediction) The mode selection unit 260 includes a partitioning unit 262, an inter prediction unit 244, and an intra prediction unit 254, and is configured to receive or obtain original picture data, e.g., original block 203 (current block 203 of current picture 17), and reconstructed picture data, e.g., filtered and / or unfiltered reconstructed samples or blocks of the same (current) picture and / or from one or more already decoded pictures, e.g., from the decoded picture buffer 230 or other buffer (e.g., a line buffer, not shown). The reconstructed picture data is used as reference picture data for prediction, e.g., inter prediction or intra prediction, to obtain a prediction block 265 or predictor 265.
[0124] The mode selection unit 260 may be configured to determine or select a partitioning and prediction mode (e.g., intra or inter prediction mode) for the prediction mode of the current block (which does not include partitioning) and generate a corresponding prediction block 265 used for calculating the residual block 205 and reconstructing the reconstructed block 215.
[0125] Embodiments of the mode selection unit 260 may be configured to select a partitioning and prediction mode (e.g., from partitioning and prediction modes supported by or available to the mode selection unit 260) that provides the best match, or in other words, the smallest residual (smallest residual means better compression for transmission or storage), or the smallest signaling overhead (smallest signaling overhead means better compression for transmission or storage), or that considers or balances both. The mode selection unit 260 may be configured to determine the partitioning and prediction mode based on rate-distortion optimization (RDO), i.e., to select the prediction mode that provides the smallest rate-distortion. Terms such as “best,” “minimum,” “optimal,” etc. in this context do not necessarily refer to the overall “best,” “minimum,” “optimal,” etc., but may also refer to satisfying termination or selection criteria such as values above or below a threshold, or other constraints that potentially lead to a “suboptimal selection,” but that reduce complexity and processing time.
[0126] In other words, the partitioning unit 262 may be configured to partition the block 203 into smaller block partitions or sub-blocks (which also form blocks) using, for example, quadtree partitioning (QT), binary partitioning (BT), or ternary tree partitioning (TT), or any combination thereof, iteratively, and to perform prediction on, for example, each of the block partitions or sub-blocks, wherein the mode selection includes selecting a tree structure for the partitioned block 203, and a prediction mode is applied to each of the block partitions or sub-blocks.
[0127] Below, the partitioning (eg, by partitioning unit 260) and prediction processes (by inter-prediction unit 244 and intra-prediction unit 254) performed by exemplary video encoder 20 are described in more detail.
[0128] Division The partitioning unit 262 can partition (or divide) the current block 203 into smaller segments, e.g., square or rectangular sized smaller blocks. These smaller blocks (which may also be called subblocks) can be further partitioned into even smaller segments. This is also called tree partitioning or hierarchical tree partitioning, where, for example, a root block at root tree level 0 (hierarchical level 0, depth 0) can be recursively partitioned, e.g., into two or more blocks at the next lower tree level, e.g., a node at tree level 1 (hierarchical level 1, depth 1), and these blocks can be partitioned again into two or more blocks at the next lower level, e.g., tree level 2 (hierarchical level 2, depth 2), and so on, until, e.g., a termination criterion is met, e.g., a maximum tree depth or a minimum block size is reached and the partitioning is terminated. Blocks that are not further partitioned are also called leaf blocks or leaf nodes of the tree. A tree that uses a partition into two partitions is called a binary tree (BT), a tree that uses a partition into three partitions is called a ternary tree (TT), and a tree that uses a partition into four partitions is called a quad tree (QT).
[0129] As mentioned above, the term "block" as used herein may be a portion of a picture, in particular a square or rectangular portion. For example, in the context of HEVC and VVC, a block may be or correspond to a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), and a transform unit (TU), and / or a corresponding block, such as a coding tree block (CTB), a coding block (CB), a transform block (TB), or a prediction block (PB).
[0130] For example, a coding tree unit (CTU) may be or include a CTB of luma samples, two corresponding CTBs of chroma samples for a picture with a three-sample arrangement, or a CTB of samples for a picture coded using three separate color planes and a syntax structure used to code a monochrome picture or sample. Correspondingly, a coding tree block (CTB) may be an NxN block of samples for some value of N such that the division of the components into CTBs is a partition. A coding unit (CU) may be or include a coding block of luma samples, two corresponding coding blocks of chroma samples for a picture with a three-sample arrangement, or a coding block of samples for a picture coded using three separate color planes and a syntax structure used to code a monochrome picture or sample. Correspondingly, a coding block (CB) may be an MxN block of samples for some values of M and N such that the division of the CTB into coding blocks is a partition.
[0131] For example, in an HEVC embodiment, a coding tree unit (CTU) may be divided into CUs by using a quadtree structure represented as a coding tree. The decision of whether to code a picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction is made at the CU level. Each CU may be further divided into one, two, or four PUs according to a PU partition type. Within one PU, the same prediction process is applied, and related information is transmitted to the decoder based on the PU. After obtaining residual blocks by applying a prediction process based on the PU partition type, the CU may be partitioned into transform units (TUs) by another quadtree structure similar to the coding tree for the CU.
[0132] For example, in an embodiment according to the latest video coding standard currently under development, called Versatile Video Coding (VVC), quadtree and binary tree (QTBT) partitioning is used to partition coding blocks. In the QTBT block structure, CUs may have either square or rectangular shapes. For example, coding tree units (CTUs) are first partitioned by a quadtree structure. The leaf nodes of the quadtree are further partitioned by a binary tree or ternary (or triple) tree structure. The leaf nodes of the partitioning tree are called coding units (CUs), and their segmentation is used for prediction and transform processing without any further partitioning. This means that CUs, PUs, and TUs have the same block size in the QTBT coding block structure. In parallel, multi-partitioning, such as ternary tree partitioning, has also been proposed to be used with the QTBT block structure.
[0133] In one example, mode select unit 260 of video encoder 20 may be configured to perform any combination of the partitioning techniques described herein.
[0134] As mentioned above, video encoder 20 is configured to determine or select a best or optimal prediction mode from a (predetermined) set of prediction modes, which may include, for example, intra-prediction modes and / or inter-prediction modes.
[0135] Intra prediction The set of intra prediction modes may include, for example, the 35 different intra prediction modes defined in HEVC, e.g., non-directional modes such as DC (or average) mode and planar mode, or directional modes, or may include, for example, the 67 different intra prediction modes defined for VVC, e.g., non-directional modes such as DC (or average) mode and planar mode, or directional modes.
[0136] The intra prediction unit 254 is configured to generate the intra prediction block 265 using reconstructed samples of neighboring blocks of the same current picture according to an intra prediction mode from a set of intra prediction modes.
[0137] The intra prediction unit 254 (or generally the mode selection unit 260) is further configured to output the intra prediction parameters (or generally information indicating the selected intra prediction mode for the block) to the entropy encoding unit 270 in the form of a syntax element 266 for inclusion in the encoded picture data 21, for example, so that the video decoder 30 may receive the prediction parameters and use them for decoding.
[0138] Inter Prediction The set (or possible) inter prediction modes depends on the available reference pictures (i.e., for example, previous at least partially decoded pictures stored in DBP230) as well as other inter prediction parameters, such as whether the entire reference picture is used to search for the best matching reference block or only a portion of the reference picture, for example, a search window area around the area of the current block, and / or whether pixel interpolation, for example, half / semi-pel and / or quarter-pel interpolation, is applied.
[0139] In addition to the prediction modes mentioned above, skip mode and / or direct mode may be applied.
[0140] The inter prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (neither shown in FIG. 2). The motion estimation unit may be configured to receive or obtain, for motion estimation, the picture block 203 (current picture block 203 of current picture 17) and the decoded picture 231, or at least one or more already reconstructed blocks, e.g., reconstructed blocks of one or more other / different already decoded pictures 231. For example, a video sequence may include the current picture and the already decoded picture 231, or in other words, the current picture and the already decoded picture 231 may be part of or form a sequence of pictures that form a video sequence.
[0141] The encoder 20 may be configured to, for example, select a reference block from multiple reference blocks of the same or different pictures among multiple other pictures, and provide the reference picture (or reference picture index) and / or an offset (spatial offset) between the position (x, y coordinates) of the reference block and the position of the current block to the motion estimation unit as an inter-prediction parameter. This offset is also called a motion vector (MV).
[0142] The motion compensation unit is configured to obtain, e.g., receive, inter prediction parameters and perform inter prediction based on or using the inter prediction parameters to obtain inter prediction block 265. Motion compensation performed by the motion compensation unit may include fetching or generating a predictive block based on motion / block vectors determined by motion estimation, possibly performing sub-pixel precision interpolation. Interpolation filtering can generate additional pixel samples from known pixel samples, thus potentially increasing the number of candidate predictive blocks that can be used to code the picture block. As described in more detail below, interpolation filtering may be performed using one or more alternative interpolation filters depending on the precision of the motion vector. Upon receiving a motion vector for the PU of the current picture block, the motion compensation unit may find the predictive block to which the motion vector points in one of the reference picture lists.
[0143] The motion compensation unit may also generate syntax elements associated with the blocks and video slices for use by video decoder 30 in decoding picture blocks of the video slices.
[0144] Entropy Coding The entropy coding unit 270 is configured to apply, for example, an entropy coding algorithm or scheme (e.g., a variable length coding (VLC) scheme, a context adaptive VLC scheme (CAVLC), an arithmetic coding scheme, binarization, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding method or technique) or bypass (uncompressed) to the quantized coefficients 209, the inter-prediction parameters, the intra-prediction parameters, the loop filter parameters, and / or other syntax elements to obtain coded picture data 21 that may be output via an output 272, for example, in the form of coded bitstream 21, such that, for example, video decoder 30 may receive the parameters and use them for decoding. The encoded bitstream 21 may be transmitted to video decoder 30 or stored in memory for later transmission or retrieval by video decoder 30 .
[0145] Other structural variations of the video encoder 20 may be used to encode the video stream. For example, a non-transform-based encoder 20 may directly quantize the residual signal for a particular block or frame without a transform processing unit 206. In another implementation, the encoder 20 may have the quantization unit 208 and the inverse quantization unit 210 combined into a single unit.
[0146] Decoder and decoding method 3 shows an example of a video decoder 30 configured to implement the techniques of the present application. The video decoder 30 is configured to receive coded picture data 21 (e.g., coded bitstream 21), e.g., coded by encoder 20, to obtain a decoded picture 331. The coded picture data or bitstream includes information for decoding the coded picture data, e.g., picture blocks of a coded video slice, as well as data representing associated syntax elements.
[0147] 3, decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., summer 314), a loop filter 320, a decoded picture buffer (DBP) 330, an inter prediction unit 344, and an intra prediction unit 354. Inter prediction unit 344 may be or include a motion compensation unit. Video decoder 30 may, in some examples, perform a decoding path that is generally the reverse of the encoding path described in connection with video encoder 100 of FIG. 2.
[0148] As described in connection with encoder 20, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, loop filter 220, decoded picture buffer (DPB) 230, inter prediction unit 344, and intra prediction unit 354 are also considered to form a “built-in decoder” of video encoder 20. Accordingly, inverse quantization unit 310 may be functionally identical to inverse quantization unit 110, inverse transform processing unit 312 may be functionally identical to inverse transform processing unit 212, reconstruction unit 314 may be functionally identical to reconstruction unit 214, loop filter 320 may be functionally identical to loop filter 220, and decoded picture buffer 330 may be functionally identical to decoded picture buffer 230. Accordingly, the descriptions given with respect to the respective units and functions of video encoder 20 apply mutatis mutandis to the respective units and functions of video decoder 30.
[0149] Entropy Decoding The entropy decoding unit 304 is configured to parse the bitstream 21 (or the coded picture data 21 generally), e.g., to perform entropy decoding on the coded picture data 21 to obtain, e.g., quantized coefficients 309 and / or decoded coding parameters (not shown in FIG. 3 ), e.g., any or all of inter-prediction parameters (e.g., reference picture indices and motion vectors), intra-prediction parameters (e.g., intra-prediction modes or indices), transform parameters, quantization parameters, loop filter parameters, and / or other syntax elements. The entropy decoding unit 304 may be configured to apply a decoding algorithm or scheme corresponding to the encoding scheme described in connection with the entropy coding unit 270 of the encoder 20. The entropy decoding unit 304 may be further configured to provide the inter-prediction parameters, intra-prediction parameters, and / or other syntax elements to the mode selection unit 360 and to provide other parameters to other units of the decoder 30. The video decoder 30 may receive syntax elements at the video slice level and / or the video block level.
[0150] inverse quantization Inverse quantization unit 310 may be configured to receive a quantization parameter (QP) (or information generally related to inverse quantization) and quantized coefficients from encoded picture data 21 (e.g., by parsing and / or decoding by entropy decoding unit 304), and apply inverse quantization to the decoded quantized coefficients 309 based on the quantization parameter to obtain dequantized coefficients 311, which may also be referred to as transform coefficients 311. The inverse quantization process may include using the quantization parameter determined by video encoder 20 for each video block within a video slice to determine the degree of quantization and, similarly, the degree of inverse quantization to be applied.
[0151] Inverse transformation The inverse transform processing unit 312 may be configured to receive the dequantized coefficients 311, also referred to as transform coefficients 311, and apply a transform to the dequantized coefficients 311 to obtain reconstructed residual blocks 213 in the sample domain. The reconstructed residual blocks 213 may also be referred to as transform blocks 313. The transform may be an inverse transform, e.g., an inverse DCT, an inverse DST, an inverse integer transform, or a conceptually similar inverse transform process. The inverse transform processing unit 312 may further be configured to receive transform parameters or corresponding information from the coded picture data 21 (e.g., by parsing and / or decoding by the entropy decoding unit 304) to determine the transform to apply to the dequantized coefficients 311.
[0152] Rebuild The reconstruction unit 314 (e.g., an adder or summator 314) may be configured to add the reconstructed residual block 313 to the prediction block 365, for example, by adding the sample values of the reconstructed residual block 313 and the sample values of the prediction block 365, to obtain a reconstructed block 315 in the sample domain.
[0153] filtering Loop filter unit 320 (either in the coding loop or after the coding loop) is configured to filter reconstructed block 315 to, for example, smooth pixel transitions or otherwise improve video quality, to obtain filtered block 321. Loop filter unit 320 may include one or more loop filters, such as a deblocking filter, a sample adaptive offset (SAO) filter, or one or more other filters, e.g., a bilateral filter, an adaptive loop filter (ALF), a sharpening filter, a smoothing filter, or a collaborative filter, or any combination thereof. Although loop filter unit 320 is shown in FIG. 3 as being an in-loop filter, in other configurations, loop filter unit 320 may be implemented as a post-loop filter.
[0154] Decoded Picture Buffer The decoded video blocks 321 of the picture are then stored in a decoded picture buffer 330, which stores the decoded picture 331 as a reference picture for subsequent motion compensation with respect to other pictures and / or for output on a display, respectively.
[0155] The decoder 30 is configured to output the decoded pictures 311 for presentation or viewing to a user, for example via an output 312.
[0156] prediction The inter prediction unit 344 may be identical to the inter prediction unit 244 (especially the motion compensation unit), and the intra prediction unit 354 may be functionally identical to the inter prediction unit 254, performing the partitioning or partitioning decision and prediction based on the partitioning and / or prediction parameters or respective information received from the encoded picture data 21 (e.g., by analyzing and / or decoding by the entropy decoding unit 304). The mode selection unit 360 may be configured to perform prediction (intra or inter prediction) for each block based on the (filtered or unfiltered) reconstructed picture, block, or respective sample to obtain a prediction block 365.
[0157] When a video slice is coded as an intra-coded (I) slice, intra prediction unit 354 of mode select unit 360 is configured to generate predictive block 365 for a picture block of the current video slice based on the signaled intra prediction mode and data from already decoded blocks of the current picture. When a video picture is coded as an inter-coded (i.e., B or P) slice, inter prediction unit 344 (e.g., a motion compensation unit) of mode select unit 360 is configured to generate predictive block 365 for a video block of the current video slice based on the motion vector and other syntax elements received from entropy decoding unit 304. For inter prediction, the predictive block may be generated from one of the reference pictures in one of the reference picture lists. Video decoder 30 may construct the reference frame lists, List 0 and List 1, using a default construction technique based on the reference pictures stored in DPB 330.
[0158] Mode select unit 360 is configured to determine prediction information for video blocks of the current video slice by parsing motion vectors and other syntax elements, and to use the prediction information to generate predictive blocks for the current video block being decoded. For example, mode select unit 360 uses some of the received syntax elements to determine the prediction mode (e.g., intra or inter prediction) used to code the video blocks of the video slice, the slice type for inter prediction (e.g., B slice, P slice, or GPB slice), construction information for one or more of the reference picture lists for the slice, motion vectors for each inter-coded video block of the slice, the status of inter prediction for each inter-coded video block of the slice, and other information for decoding video blocks in the current video slice.
[0159] Other variations of video decoder 30 may be used to decode encoded picture data 21. For example, decoder 30 may generate an output video stream without loop filtering unit 320. For example, a non-transform-based decoder 30 may directly inverse quantize the residual signal for a particular block or frame without inverse transform processing unit 312. In another implementation, video decoder 30 may have inverse quantization unit 310 and inverse transform processing unit 312 combined into a single unit.
[0160] It should be understood that in the encoder 20 and the decoder 30, the processing result of the current step may be further processed and then output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, further operations such as clip or shift may be performed on the processing result of the interpolation filtering, motion vector derivation, or loop filtering.
[0161] It should be noted that further operations may be applied to the derived motion vector of the current block (including, but not limited to, control point motion vectors in affine mode, lower-block motion vectors in affine, planar, and ATMVP modes, temporal motion vectors, etc.). For example, the value of a motion vector is constrained to a predetermined range according to its representation bits. If the representation bits of a motion vector are bitDepth, the range is -2^(bitDepth-1) to 2^(bitDepth-1)-1, where "^" means exponentiation. For example, if bitDepth is set equal to 16, the range is -32768 to 32767, and if bitDepth is set equal to 18, the range is -131072 to 131071.
[0162] 4 is a schematic diagram of a video coding device 400 according to an embodiment of the present disclosure. The video coding device 400 is suitable for implementing the disclosed embodiments as described herein. In an embodiment, the video coding device 400 may be a decoder, such as the video decoder 30 of FIG. 1A, or an encoder, such as the video encoder 20 of FIG. 1A.
[0163] Video coding device 400 includes an incoming port 410 (or input port 410) and a receiver unit (Rx) 420 for receiving data, a processor, logic unit, or central processing unit (CPU) 430 for processing data, a transmitter unit (Tx) 440 and an outgoing port 450 (or output port 450) for transmitting data, and a memory 460 for storing data. Video coding device 400 may also include optical-electrical (OE) and electrical-optical (EO) components coupled to the incoming port 410, receiver unit 420, transmitter unit 440, and outgoing port 450 for emitting or receiving optical or electrical signals.
[0164] The processor 430 is implemented by hardware and software. The processor 430 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), FPGA, ASIC, and DSP. The processor 430 communicates with the incoming port 410, the receiver unit 420, the transmitter unit 440, the outgoing port 450, and the memory 460. The processor 430 includes a coding module 470. The coding module 470 implements the disclosed embodiments described above. For example, the coding module 470 implements, processes, prepares, or provides various coding operations. Thus, the inclusion of the coding module 470 significantly improves the functionality of the video coding device 400 and results in the transition of the video coding device 400 to different states. Alternatively, the coding module 470 is implemented as instructions stored in the memory 460 and executed by the processor 430.
[0165] Memory 460 may include one or more disks, tape drives, and solid-state drives, and may be used to store programs when such programs are selected for execution, as well as an overflow data storage device for storing instructions and data read during execution of the programs. Memory 460 may be, for example, volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random access memory (SRAM).
[0166] FIG. 5 is a simplified block diagram of an apparatus 500 that may be used as either or both of source device 12 and destination device 14 of FIG. 1, according to an example embodiment.
[0167] Processor 502 of apparatus 500 may be a central processing unit. Alternatively, processor 502 may be any other type of device or devices, existing or later developed, that are capable of manipulating or processing information. While the disclosed implementations may be performed by a single processor, e.g., processor 502, as shown, speed and efficiency advantages may be realized by using two or more processors.
[0168] The memory 504 of the apparatus 500 may, in implementation, be a read-only memory (ROM) device or a random-access memory (RAM) device. Any other suitable type of storage device may be used as the memory 504. The memory 504 may include code and data 506 accessed by the processor 502 using a bus 512. The memory 504 may further include an operating system 508 and application programs 510, which include at least one program that enables the processor 502 to perform the methods described herein. For example, the application programs 510 may include applications 1 through N, which further include a video coding application that performs the methods described herein.
[0169] The apparatus 500 may also include one or more output devices, such as a display 518. The display 518, in one example, may be a touch-sensitive display that combines a display with touch-sensing elements operable to sense touch input. The display 518 may be coupled to the processor 502 via the bus 512.
[0170] Although shown here as a single bus, bus 512 of device 500 may be comprised of multiple buses. Additionally, secondary storage 514 may be directly coupled to other components of device 500 or may be accessed over a network, and may include a single integrated unit such as a memory card or multiple units such as multiple memory cards. Accordingly, device 500 may be implemented in a wide variety of configurations.
[0171] The concepts presented herein are explained in more detail below.
[0172] Motion Vector Prediction In the current VVC design, spatial motion vector prediction is used. Spatial motion vector prediction means that during inter prediction, motion information from spatial neighboring blocks is used to predict the motion vector of the current inter block. In particular, in merge and skip modes, motion vectors from adjacent spatial neighboring blocks of the current block are used. In merge and skip modes, so-called HMVP candidates may be used. HMVP candidates include motion information from history-based spatial neighboring blocks. "History-based" means that motion information from blocks earlier than the current block in decoding order is used. Such preceding blocks are from the same frame as the current block and are in some spatial neighborhood around the current block, but are not necessarily adjacent blocks like regular spatial merge candidates.
[0173] Building a list of merge candidates The merge candidate list is constructed based on the following candidates: · Up to four spatial merge candidates derived from five spatially neighboring blocks as shown in Figure 6; One temporal merge candidate derived from two temporally co-located blocks, Additional merge candidates including combined bi-predictive candidates and zero motion vector candidates. The construction of the merge candidate list is explained in more detail below by referring to FIG.
[0174] spatial candidate The first set of candidates in the merge candidate list are spatially neighboring blocks, as shown in Figure 6. For inter-prediction block merging, up to four candidates are inserted into the merge list by sequentially examining A1, B1, B0, A0, and B2 in this order. Instead of simply examining whether neighboring blocks are available and contain motion information, further redundancy checks are performed before considering all motion data of neighboring blocks as merge candidates. These redundancy checks can be divided into two categories with two different purposes: Avoid having candidates with redundant motion data in the HMI list, · Prevents merging two segments that can be expressed by other means, generating redundant syntax.
[0175] History-based motion vector prediction To further improve motion vector prediction, techniques have been proposed that use motion information from non-adjacent CUs (the motion information includes one reference picture index / multiple reference picture indexes and one motion vector / multiple motion vectors). One such technique is history-based motion vector prediction (HMVP). HMVP uses a look-up table (LUT) that contains motion information from CUs that have already been coded. Essentially, the HMVP method consists of two main parts: 1. The HMVP LUT construction and update method shown in Figures 10 and 11; 2. Use of HMVP LUT to construct the merge candidate list (or AMVP candidate list) shown in Fig. 12.
[0176] HMVP LUT construction and update method The LUT is maintained during the encoding and / or decoding process. The LUT is emptied when a new slice is encountered. Whenever the current CU is inter-coded, the associated motion information is added to the last entry of the table as a new HMVP candidate. The size of the LUT (denoted as N) is a parameter of the HMVP method. If the number of HMVP candidates from already coded CUs is more than the size of this LUT, a table update method is applied, so that this LUT always contains no more than N already coded motion candidates. Two table update methods have been proposed: 1. First-in, first-out (FIFO) LUT update method shown in Fig. 10; 2. The constrained FIFO LUT update method shown in Fig. 11.
[0177] FIFO LUT update method According to the FIFO LUT update method, the oldest candidate (the 0th table entry) is deleted from the table before inserting a new candidate. This process is illustrated in Figure 10. In this figure, H0 is the oldest (0th) HMVP candidate and X is the new candidate.
[0178] Although this updating method is relatively low in complexity, when this method is applied, some of the elements of the LUT may be the same (may contain the same motion information), so the data in the LUT may be redundant, and the diversity of the motion information in the LUT is lower than the method in which duplicate candidates are removed.
[0179] Constraint FIFO LUT update method To further improve the coding efficiency, a constrained FIFO LUT update method is introduced. According to this method, a redundancy check is applied before inserting a new HMVP candidate into the table. The redundancy check is performed by checking whether the motion information from the new candidate X is already in the LUT of the candidate H. m This means determining whether the motion information contained in such a candidate H mIf not found, a simple FIFO method is used, otherwise the following steps are performed: 1. H m All LUT entries after H are shifted one position to the left (towards the beginning of the table), resulting in candidate H m is removed from the table, freeing up one position at the end of the LUT. 2. A new candidate, X, is added to the first empty position in the table.
[0180] An example of the use of the constraint FIFO LUT update method is shown in FIG.
[0181] Using HMVP LUT for motion vector coding HMVP candidates may be used in the merge candidate list building process and / or the AMVP candidate list building process.
[0182] Using the HMVP LUT in constructing the merge candidate list In some cases, the HMVP candidate is merged from the last entry to the first entry after the temporal merge candidate (e.g., H N-1 , H N-2 , ..., H0) are inserted into the merge list. The traversal order of the LUT is shown in Figure 12. If an HMVP candidate is equal to one of the candidates already given in the merge list, such an HMVP candidate is not added to the HMVP list. Due to the limited size of the merge list, some HMVP candidates, especially those at the beginning of the LUT, may not be used in the merge list construction process for the current CU.
[0183] Use of HMVP LUT in the AMVP candidate list construction process The HMVP LUT constructed for merge mode can also be used for AMVP. The difference is that only some entries from this LUT are used for constructing the AMVP candidate list. More specifically, only the last M entries of the HMVP LUT are used (e.g., M equals 4). During the AMVP candidate list construction process, HMVP candidates are listed from the end to the (NK)th entry after the TMVP candidates, i.e., H N-1 , H N-2 , ..., H N-K is inserted into the AMVP candidate list. The LUT traversal order is shown in Figure 12.
[0184] Only HMVP candidates that have the same reference picture as the AMVP target reference picture are used. If an HMVP candidate is equal to one of the candidates already given in the HMI list, this HMVP candidate is not used for building the AMVP candidate list. Due to the limited size of the AMVP candidate list size, some HMVP candidates may not be used in the AMVP list building process for the current CU.
[0185] Switchable Interpolation Filters The motion vector difference of a translationally inter-predicted block can be coded with three different precisions (i.e., quarter-pel, full-pel, and four-pel). The interpolation filter (IF) used for each fractional position is fixed. In this disclosure, a switchable interpolation filter (SIF) technique enables the use of one or two alternative luma interpolation filters for half-pel positions. Switching between available luma interpolation filters can be performed at the CU level. To reduce signaling overhead, the switching depends on the precision of the motion vectors used. For this purpose, the Adaptive Motion Vector Resolution (AMVR) scheme is extended to also support half-pel luma motion vector precision. Only in this half-pel motion vector precision mode can an alternative half-pel interpolation filter be used, and the alternative half-pel interpolation filter is indicated by an additional syntax element indicating which interpolation filter is used. In skip or merge modes using spatial merge candidates, the value of this syntax element can be inherited from neighboring blocks.
[0186] Half-pel AMVR mode An additional AMVR mode for non-affine, non-merged inter-coded CUs is introduced, which enables signaling of motion vector differences with half-pel precision. The existing AMVR scheme of the current VVC draft is simply extended as follows: If amvr_flag == 1, then immediately after the syntax element amvr_flag, there is a new context-modeled binary syntax element hpel_amvr_flag, which indicates the use of the new half-pel AMVR mode if hpel_amvr_flag == 1. Otherwise, that is, if hpel_amvr_flag == 0, the choice between full-pel and 4-pel AMVR modes is indicated by the syntax element amvr_precision_flag, as in the current VVC draft.
[0187] Alternative Luma Half-Pel Interpolation Filter For non-affine, non-merged inter-coded CUs that use half-pel motion vector precision (i.e., half-pel AMVR mode), switching between the HEVC / VVC half-pel luma interpolation filter and one or more alternative half-pel interpolators may be done based on the value of a new syntax element if_idx (interpolation filter index). The syntax element if_idx is signaled only for half-pel AMVR mode. For skip / merge modes that use spatial merge candidates, the value of the interpolation filter index is inherited from neighboring blocks.
[0188] It may be understood that the fractional position of a motion vector may be represented, for example, by the luma position (xFracL, yFracL) in units of fractional samples. The motion vectors of the selected merging candidates may be represented by refMvLX and refMvLX, where mvLX = mvL0 or mvL1.
[0189] In the example, xFrac L = refMvLX
[0000] & 15 (8-738) yFrac L = refMvLX
[0001] & 15 (8-739) is.
[0190] xFrac L (or yFrac L ) is equal to 0 (i.e., the MV points to an integer position), no interpolation is used. L is in the range [1, 15]), then f L [ xFrac L ] is used. The luma interpolation filter coefficients f for each fractional sample position p (p is in the range [1, 15]) are L [ p ] is specified in Table 8-8.
[0191] This Table 8-8 (Table 1) is an example of a set of interpolation filters, and one interpolation filter is selected depending on the fractional position. One interpolation filter (interpolation filter coefficient) may be one line in Table 8-8 (Table 1). In an example, the set of interpolation filters of the present disclosure may have the same interpolation filter for all positions except for half sample positions (fractional position: 1 / 2).
[0192] Table 8-8 below shows the HEVC / VVC interpolation filter coefficients f for each fractional sample position p (p is in the range [1, 15], with a precision of 1 / 16 pel (pixel)). L In this table, when p=8, the interpolation filter coefficient f L [ p ] is a half-pel interpolation filter coefficient. As discussed above, additional interpolation filters can be added as alternatives to this half-pel interpolation filter to allow switching between these half-pel interpolation filters. Some examples of alternative half-pel interpolation filters are shown below.
[0193] [Table 1]
[0194] An alternative implementation using a 6-tap half-pel interpolation filter In an example, a 6-tap interpolation filter may be used as an alternative to the normal HEVC / VVC half-pel interpolation filter shown in Table 8-8. Table 1 below shows the mapping between the value of the syntax element if_idx (or derived IF index) and the selected half-pel luma interpolation filter.
[0195] [Table 2]
[0196] Implementation using two alternative 8-tap half-pel interpolation filters In another example, two 8-tap interpolation filters can be used as an alternative to the regular HEVC / VVC half-pel interpolation filters shown in Table 8-8. Table 2 below shows the mapping between the value of the syntax element if_idx and the selected half-pel luma interpolation filters.
[0197] [Table 3]
[0198] Implementation using two alternative 6-tap half-pel interpolation filters In another example, two 6-tap interpolation filters can be used as an alternative to the regular HEVC / VVC half-pel interpolation filters shown in Table 8-8. Table 3 below shows the mapping between the value of the syntax element if_idx and the selected half-pel luma interpolation filters.
[0199] [Table 4]
[0200] As shown in Table 4 of the interpolation filters, the interpolation filters for half-pel positions (see line "8" in this Table 4) can be switched in the present disclosure, where alternative or switchable half-sample interpolation filters are used to interpolate half-sample values when the corresponding MV points to a half-sample position.
[0201] [Table 5]
[0202] More specifically, the following aspects are described: 1. Modification of the history-based motion information candidate list (i.e., HMI list) construction / update method. In addition to the motion information of one or more coded / decoded blocks preceding a block, the interpolation filter (IF) index (e.g., half-pel interpolation filter index (hpelIfIdx)) of the preceding block is stored in the HMI list. In particular, the IF index is also stored in the HMI candidate or record in the HMI list. In this way, the IF index can be propagated through the HMI list, achieving coding consistency and higher coding efficiency. 2. Interpolation filter index (half-pel interpolation filter index) derivation procedure for merge mode: If a block has a merge candidate index corresponding to a history-based candidate, the IF index (half-pel interpolation filter index) of this history-based candidate is used for the current block. 3. Propagation of SIF index across CTU boundaries. Based on the current SIF design, when the SIF technology is applied in a mode that inherits motion information from upper spatial neighboring blocks, if the current block is at the upper boundary of the CTU, the line memory is increased. In the description presented herein, the position of the current block is examined. If the current block is at the upper boundary of the CTU, when inheriting motion information from the upper left (B0), upper (B1), and upper right (B2) neighboring blocks, the IF index is not inherited and instead uses a default value to reduce the cost of line memory.
[0203] An example of SIF index propagation across CTU boundaries is shown in FIG. 7. In this example, motion information is inherited from a neighboring block in B1 (above), which belongs to a CTU that is not the same as the CTU containing the current block 700. In this case, the SIF index of the block in B1 must be stored in a line buffer in the prior art. The present invention prevents SIF index propagation in such cases, thus reducing line buffer size requirements. During the construction of the merge list, the position of the current block is examined. As shown in FIG. 7, if the current block is at the top boundary of the CTU, while inheriting motion information from the neighboring blocks in the upper left (B0), above (B1), and upper right (B2), the IF index is not inherited and uses a default value to reduce line memory costs. Details are described below with reference to FIGS. 8 and 9.
[0204] FIG. 13A shows a flow chart of a construction method 1300 for a history-based motion information candidate list (ie, HMI list), which includes the following steps. In step 1301, a history-based motion information candidate list is obtained, and the HMI list includes N history-based motion information candidates H related to (including) the motion information of N preceding blocks preceding the block. k where k=0, ..., N-1, and N is an integer greater than 0, and each history-based motion information candidate includes the following elements: i) one or more motion vectors MV of the preceding block; ii) one or more reference picture indices corresponding to the MVs of the preceding blocks; and iii) The interpolation filter index of the preceding block. In step 1303, update the HMI list according to the motion information of the block, where the motion information of the block includes the following elements: i) one or more motion vectors MV of the block; ii) one or more reference picture indices corresponding to the MV of the block; and iii) The interpolation filter index for the block.
[0205] It may be noted that one or more MVs of a block refer to MVs corresponding to the L0 and L1 reference picture lists. The same is true for the reference picture indexes.
[0206] As shown in FIG. 13B, step 1301 can be step 1311, which includes loading a history-based motion information candidate list (HMI table), and step 1303 can be step 1313, which includes updating the history-based motion information candidate list (table) using motion information of a decoded block. The HMI table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is emptied when a new slice is encountered. When an inter-coded block of a slice exists, the block is decoded based on the motion information candidate list including the history-based motion information candidate (step 1302), and the associated motion information of the block is added to the last entry of the table as a new HMVP candidate (step 1303).
[0207] 14 shows a flow chart of a method for inter-prediction of a block in a frame of a video signal, which includes the following steps: In step 1401, construct a history-based motion information candidate list (i.e., HMVP list), where the HMVP list includes N history-based motion information candidates H related to the motion information of multiple blocks preceding the block. k where k=0, ..., N-1, and N is an integer greater than 0, and each history-based motion information candidate includes the following elements: i) one or more motion vectors MV; ii) one or more reference picture indices corresponding to the MV; and iii) Interpolation filter index. In step 1402, adding one or more history-based motion information candidates from the HMI list to a motion information candidate list for the block; In step 1403, motion information for the block is derived based on the motion information candidate list.
[0208] It can be understood that the motion information candidate list refers to the merge candidate list as follows:
[0209] It can be appreciated that the history-based merge candidates are included in the motion information candidate list in step 1403 .
[0210] FIG. 15 shows a flow diagram of a method for constructing and updating a history-based motion information candidate list (i.e., an HMI list). In step 1501, an HMI list is constructed. In step 1502, at least one of elements i) and ii) of each history-based motion information candidate in the HMVP list is compared with a corresponding element of the current block. Step 1502 includes comparing whether the motion vectors of the history-based motion information candidates in the history-based motion information candidate list are the same as the corresponding motion vector of the block, and comparing whether the reference picture indexes of the history-based motion information candidates are the same as the corresponding reference picture index of the block. In an alternative design, step 1502 includes comparing whether at least one of the motion vectors of each history-based motion information candidate is different from the corresponding motion vector of the block, and comparing whether at least one of the reference picture indexes of each HMVP candidate is different from the corresponding reference picture index of the block. The result of the element-based comparison is referred to as the comparison result in FIG. 15.
[0211] If the comparison result is that at least one of the following elements i) and ii) of each history-based motion information candidate in the history-based motion information candidate list is different from the corresponding element of the motion information of the block, the motion information of the current block is added to the last position of the HMVP list (step 1503). Otherwise, if the following elements i) and ii) of the history-based motion information candidate in the history-based motion information candidate list are the same as the corresponding element of the motion information of the block, the history-based motion information candidate is removed from the history-based motion information candidate list, and the history-based motion information candidate HMVP containing the motion information of the block is added to the last position of the history-based motion information candidate list. k is added, where k=N-1 (step 1504).
[0212] And the above comparison is only performed looking at the differences in terms of MV and reference picture indexes, without comparing the IF indexes.
[0213] Further embodiments are summarized in the following aspects.
[0214] A method of deriving an interpolation filter index (or an index of a set of interpolation filters) for a current block according to a first aspect of the present invention, comprising the steps of: Record H of N movements relative to N preceding blocks of a frame k where k=0, ..., N-1, N is 1 or greater, and each motion record includes one or more motion vectors, one or more reference picture indices corresponding to the one or more motion vectors, and interpolation filter indices (or interpolation filter set indices) corresponding to the one or more motion vectors (e.g., the same filter index or the same filter set index for both of the two MVs); determining a historical motion information candidate (such as an HMVP candidate) for the current block based on the historical motion information list (such as determining an HMVP candidate for the current block from an HMVP list or an HMVP table).
[0215] In a possible implementation form of the device according to the first aspect itself, wherein the step of determining a history-based motion information candidate for the current block based on the history-based motion information list includes: Deriving, inferring, or determining an interpolation filter index (or an index of a set of interpolation filters) of record Hk as an interpolation filter index (or an index of a set of interpolation filters) for the current block, wherein a determined or selected history-based motion information candidate (such as an HMVP candidate) corresponds to record Hk.
[0216] In any of the above-mentioned implementations of the first aspect or a possible implementation form of a device according to the first aspect itself, the motion records in the history-based motion information list are ordered in the order in which the motion records of the preceding blocks are obtained from the bitstream.
[0217] In any of the above-mentioned implementations of the first aspect or a possible implementation of a device according to the first aspect itself, the history-based motion information list has a length of N, where N is five.
[0218] In any of the above-described implementations of the first aspect or a possible implementation of a device according to the first aspect itself, the step of constructing a history-based motion information list (HMVL) comprises: Before adding the motion information of the current block to the HMVL, checking whether each element of the HMVL is different from the motion information of the current block; and adding the motion information of the current block to the HMVL only if each element of the HMVL differs from the motion information of the current block.
[0219] In any of the above-described implementations of the first aspect or a possible implementation form of a device according to the first aspect itself, checking whether each element of the HMVL is different from the motion information of the current block includes: Comparing the corresponding motion vectors, and This involves comparing the corresponding reference picture indexes.
[0220] In any of the above-described implementations of the first aspect or a possible implementation form of a device according to the first aspect itself, checking whether each element of the HMVL is different from the motion information of the current block includes: Includes comparison of interpolation filter indices.
[0221] In any of the above-mentioned implementations of the first aspect or a possible implementation form of a device according to the first aspect itself, the device further includes a step of deriving motion information from motion information of a first block, the first block having a predetermined spatial or temporal positional relationship with the current block.
[0222] In any of the above-described implementations of the first aspect or possible implementations of a device according to the first aspect itself, The method further includes deriving motion information from motion information of a second block, the second block being reconstructed before the current block.
[0223] In any of the above-mentioned implementations of the first aspect or a possible implementation form of a device according to the first aspect itself, wherein the history-based motion information list (HMIL or HMVP table) is a subset of the candidate motion information list of the current block when the current block is in merge mode, or is a subset of the candidate predicted motion information list of the current block when the current block is in AMVP mode.
[0224] In any of the above-described implementations of the first aspect or in a possible implementation form of a device according to the first aspect itself, only one interpolation filter set index corresponds to one or more motion vectors of the HMVP candidate (such as the same filter set index for both of the two MVs), or One or more interpolation filter set indices correspond to one or more motion vectors of the HMVP candidate, respectively.
[0225] A method of inter prediction for a current block according to a second aspect of the present invention, comprising: inter-predicting a current block, comprising deriving an interpolation filter index (or an index of a set of interpolation filters) for the current block; Deriving an interpolation filter index (or an index of a set of interpolation filters) for the current block determining an HMVP candidate for the current block from an HMVP list (e.g., an HMVP table), where the HMVP candidate includes at least one motion vector, at least one reference picture index corresponding to the at least one motion vector, and at least one interpolation filter index (or interpolation filter set index) corresponding to the at least one motion vector (e.g., only one interpolation filter index or only one interpolation filter set index for all candidates); deriving, inferring, or determining an interpolation filter index (or an index of a set of interpolation filters) of the determined or selected HMVP candidate as an interpolation filter index (or an index of a set of interpolation filters) for the current block; A method wherein one or more candidates (e.g., each candidate) of the HMVP list includes at least one motion vector and an interpolation filter index (or an index of a set of at least one interpolation filter) corresponding to the at least one motion vector.
[0226] In a possible implementation of the method according to the second aspect itself, the interpolation filter index (or the index of the set of interpolation filters) corresponds to one or more motion vectors of the HMVP candidate, or One or more interpolation filter indices (or indices of a set of one or more interpolation filters) correspond to one or more motion vectors of the HMVP candidate.
[0227] A method for deriving an interpolation filter for a coding unit coded in merge mode based on a position of a current coding unit within a CTU according to a third aspect of the present invention, comprising: Parsing or deriving a first merge index from the bitstream; selecting a merge candidate from the merge candidate list according to a first merge index; determining whether the current coding unit overlaps the top or left boundary of the CTU; If the current coding unit overlaps the top or left boundary of the CTU, setting an interpolation filter index (or an index of a set of interpolation filters) for the current coding unit to a predefined value; otherwise, setting an interpolation filter index (or an index of the set of interpolation filters) for the current coding unit equal to the interpolation filter index (or an index of the set of interpolation filters) of the selected merge candidate; selecting a first set of interpolation filters from a set of N interpolation filters (e.g., a set of N predefined interpolation filters) based on an interpolation filter index (or an interpolation filter set index), where N is an integer greater than or equal to 2; For each motion vector of the selected merge candidate, selecting an interpolation filter from the first set of interpolation filters based on the fractional position (e.g., luma position (xFracL, yFracL) in fractional samples) of the motion vector.
[0228] In a possible implementation of the method according to the third aspect itself, The method further includes constructing a merge candidate list, each candidate including one or more motion vectors and an interpolation filter index (or an interpolation filter set index specifying one of N interpolation filter sets (e.g., N predefined interpolation filter sets)).
[0229] In any of the above-mentioned implementations of the third aspect or a possible implementation form of a method according to the third aspect itself, wherein the step of determining whether the current coding unit overlaps with the top or left boundary of a CTB or CTU includes determining whether the top left corner of the current block (such as a luma position (xCb, yCb) specifying the top left sample of the current coding block relative to the top left luma sample of the current picture) overlaps with the top boundary of the CTU that includes the current coding unit.
[0230] In any of the above-described implementations of the third aspect or a possible implementation form of the method according to the third aspect itself, determining whether an upper left corner of the current block overlaps with an upper boundary of the CTU that includes the current coding unit includes: Get the vertical position (y coordinate) of the top left corner of the current block, Calculating the remainder after dividing the obtained vertical position (y coordinate) by the height of the CTU; If the calculated remainder is equal to 0, infer that the top left corner of the current block overlaps with the top boundary of the current CTU; otherwise, infer that the top left corner of the current coding unit does not overlap with the top boundary of the current CTU.
[0231] In any of the above-described implementations of the third aspect or a possible implementation form of the method according to the third aspect itself, determining whether a top left corner of the current coding unit overlaps with a top boundary of a CTU that includes the current coding unit includes: Calculating a first value as the floor value of the current block's top-left vertical coordinate (coordinate y) divided by the height of the CTU; calculating a second value as the floor value of the top left corner of the vertical coordinator of the inherited neighboring block divided by the height of the CTU; If the second value is equal to the first value, inferring that the top left corner of the current coding unit overlaps the top boundary of the current CTU.
[0232] In any of the above-described implementations of the third aspect or a possible implementation form of the method according to the third aspect itself, determining whether a top left corner of the current coding unit overlaps with a top boundary of a CTU that includes the current coding unit includes: calculating a third value as (yCb >> CtbLog2SizeY) << CtbLog2SizeY, where yCb is the top-left vertical coordinate (y coordinate) of the current block, ">>" is a logical or arithmetic right bit shift, "<<" is a logical or arithmetic left bit shift, and CtbLog2SizeY is a binary logarithmic scale of the size of the CTU; If (yCb - 1) is less than a third value, inferring that the top left corner of the current coding unit overlaps the top boundary of the current CTU.
[0233] In any of the above-mentioned implementations of the third aspect or a possible implementation form of the method according to the third aspect itself, the second predefined area includes or encompasses the top left corner of the CTU containing the current block.
[0234] In any of the above-described implementations of the third aspect or a possible implementation form of the method according to the third aspect itself, determining whether the top left corner of the current block overlaps with the left boundary of the CTU that includes the current block includes: Obtaining the horizontal position (x coordinate) of the top left corner of the current coding unit; Calculating the remainder after dividing the obtained horizontal position (x coordinate) by the width of the CTU; if the calculated remainder is equal to 0, inferring that the top-left corner of the current coding unit overlaps with the left boundary of the current CTU; If not, it includes inferring that the top left corner of the current coding unit does not overlap with the left boundary of the current CTU.
[0235] In any of the above-described implementations of the third aspect or a possible implementation form of the method according to the third aspect itself, determining whether the top left corner of the current coding unit overlaps with the left boundary of the CTU that includes the current coding unit includes: Calculating a fourth value as a floor value, where the top-left horizontal coordinate of the current block (coordinate x) is divided by the width of the CTU to obtain the floor value; Calculating a fifth value as a floor value, where the top left corner of the vertical coordinator of the inherited neighboring block is divided by the width of the CTU to obtain the floor value; If the fifth value is equal to the fourth value, inferring that the top left corner of the current coding unit overlaps with the left boundary of the current CTU.
[0236] In any of the above-described implementations of the third aspect or a possible implementation form of the method according to the third aspect itself, determining whether the top left corner of the current coding unit overlaps with the left boundary of the CTU that includes the current coding unit includes: Calculating a sixth value as ( xCb >> CtbLog2SizeX ) << CtbLog2SizeX, where xCb is the top-left vertical coordinate (y coordinate) of the current block, ">>" is a logical or arithmetic right bit shift, and "<<" is a logical or arithmetic left bit shift, and CtbLog2SizeX is a binary logarithmic scale of the width of the CTU; If (xCb - 1) is less than a sixth value, inferring that the top-left corner of the current coding unit overlaps the left boundary of the current CTU.
[0237] In a possible implementation of any of the above-mentioned implementations of the third aspect or the method according to the third aspect itself, instead of a combination of left and right shift operations on N bits, a logical AND with a bit mask containing bit 0 in the lowest N positions and bit 1 elsewhere (e.g., ( yCb >> CtbLog2SizeY ) << CtbLog2SizeY can be calculated as a logical AND of yCb with a bit mask containing bit 0 in the lowest CtbLog2SizeY positions and bit 1 elsewhere).
[0238] In any of the above-mentioned implementations of the third aspect or a possible implementation of the method according to the third aspect itself, the selected interpolation filter is applied to the reference samples to generate predicted samples that fall fractionally between the reference samples; or The selected interpolation filter is used to generate predicted samples within the current coding unit (eg, to generate predicted samples for lower-order blocks of the current coding unit).
[0239] A method of inter prediction for a current block according to a fourth aspect of the present invention, comprising: A method wherein, when a condition is met, at least two luma positions have the same half-sample interpolation filter index and the same bi-prediction weight index.
[0240] In a possible implementation of the method according to the fourth aspect itself, When availableA1 is equal to TRUE, luma positions (xNbA1, yNbA1) and (xNbB1, yNbB1), or luma positions (xNbA1, yNbA1) and (xNbB0, yNbB0), or luma positions (xNbA1, yNbA1) and (xNbA0, yNbA0), or luma positions (xNbA1, yNbA1) and (xNbA0, yNbA0), or luma positions (xNbA1, yNbA1) and (xNbA0, yNbA0), or luma positions (xNbA1, yNbA1) and (xNbB2, yNbB2) have the same bi-prediction weight index and the same half-sample interpolation filter index.
[0241] In any of the above-mentioned implementations of the fourth aspect or a possible implementation form of a method according to the fourth aspect itself, when availableB1 is equal to TRUE, luma positions (xNbB1, yNbB1) and (xNbB0, yNbB0), or luma positions (xNbB1, yNbB1) and (xNbA0, yNbA0), or luma positions (xNbB1, yNbB1) and (xNbB2, yNbB2) have the same bi-prediction weight index and the same half-sample interpolation filter index.
[0242] In any of the above-mentioned implementations of the fourth aspect or in a possible implementation form of the method according to the fourth aspect itself, when availableB0 is equal to TRUE, the luma positions (xNbB0, yNbB0) and (xNbA0, yNbA0), or the luma positions (xNbB0, yNbB0) and (xNbB2, yNbB2) have the same bi-prediction weight index and the same half-sample interpolation filter index.
[0243] In a possible implementation form of any of the above-mentioned implementations of the fourth aspect or the method according to the fourth aspect itself, when availableA0 is equal to TRUE, the luma positions (xNbA0, yNbA0) and (xNbB2, yNbB2) have the same bi-prediction weight index and the same half-sample interpolation filter index.
[0244] A method of inter prediction for a current block according to a fifth aspect of the present invention, comprising: The method, wherein when the condition is met, the MVP candidate and the merge candidate have the same motion vector and the same reference index.
[0245] In any of the above-mentioned implementations of the fifth aspect or a possible implementation of the method according to the fifth aspect itself, obtaining half-sample interpolation filter indexes; When the condition is met, the MVP candidate and the merge candidate have the same half-sample interpolation filter index as well as the same motion vector and the same reference index.
[0246] In any of the above-mentioned implementations of the fifth aspect or a possible implementation of the method according to the fifth aspect itself, obtaining half-sample interpolation filter indexes; When the condition involving half-sample interpolation filter index is met, the MVP candidate and merge candidate have the same motion vector and the same reference index.
[0247] Possible implementation details of the proposed method's derivation of motion information including interpolation filter set indices for a block based on the merge candidate list (see step 1403 of method 1400 shown in FIG. 14) are described below in the form of amendments to the VVC Working Draft Specification. The amendments are highlighted.
[0248] 8.5.2 Derivation Process for Motion Vector Components and Reference Indices 8.5.2.1 Overview The inputs to this process are: - Luma position ( xCb, yCb ) of the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture - variable cbWidth that specifies the width of the current coding block in luma samples - Variable cbHeight specifying the height of the current coding block in luma samples.
[0249] The output of this process is: - Luma motion vectors mvL0
[0000]
[0000] and mvL1
[0000]
[0000] with 1 / 16 fractional sample precision - Reference indices refIdxL0 and refIdxL1 - Prediction list usage flags predFlagL0
[0000]
[0000] and predFlagL1
[0000]
[0000] - Half-sample interpolation filter index hpelIfIdx - Bi-prediction weight index bcwIdx.
[0250] Let variable LX be the RefPicList[X] of the current picture, where X is either 0 or 1.
[0251] For the derivation of the variables mvL0
[0000]
[0000] and mvL1
[0000]
[0000] , refIdxL0 and refIdxL1, and predFlagL0
[0000]
[0000] and predFlagL1
[0000]
[0000] , the following applies: - if general_merge_flag[ xCb ][ yCb ] is equal to 1, the derivation process for luma motion vectors for merge mode specified in Section 8.5.2.2 is called with the luma position ( xCb, yCb ), variables cbWidth and cbHeight as input, and the outputs are the luma motion vectors mvL0
[0000]
[0000] , mvL1
[0000]
[0000] , reference indices refIdxL0, refIdxL1, prediction list usage flags predFlagL0
[0000]
[0000] and predFlagL1
[0000]
[0000] , half sample interpolation filter index hpelIfIdx, bi-prediction weight index bcwIdx, and merge candidate list mergeCandList. - Otherwise, the following applies: - For X, which is replaced by either 0 or 1 in the variables predFlagLX
[0000]
[0000] , mvLX
[0000]
[0000] , and refIdxLX, in PRED_LX, and in the syntax elements ref_idx_lX and MvdLX, the following ordered steps are applied: 1. The variables refIdxLX and predFlagLX
[0000]
[0000] are derived as follows: - if inter_pred_idc[ xCb ][ yCb ] is equal to PRED_LX or PRED_BI, refIdxLX = ref_idx_lX[ xCb ][ yCb ] (8-292) predFlagLX
[0000] [0 ] = 1 (8-293) Otherwise, the variables refIdxLX and predFlagLX
[0000] [0] are defined by: refIdxLX = -1 (8-294) predFlagLX
[0000]
[0000] = 0 (8-295) 2. The variable mvdLX is derived as follows: mvdLX
[0000] = MvdLX[ xCb ][ yCb ]
[0000] (8-296) mvdLX
[0001] = MvdLX[ xCb ][ yCb ]
[0001] (8-297) 3. When predFlagLX
[0000]
[0000] is equal to 1, the derivation process for luma motion vector prediction in Section 8.5.2.8 is called with the position of the luma coding block (xCb, yCb), the width of the coding block cbWidth, the height of the coding block cbHeight, and the variable refIdxLX as input, and the output is mvpLX. 4. When predFlagLX
[0000]
[0000] is equal to 1, the luma motion vector mvLX
[0000]
[0000] is derived as follows: uLX
[0000] = ( mvpLX
[0000] + mvdLX
[0000] + 2 18 ) % 2 18 (8-298) mvLX
[0000]
[0000]
[0000] = ( uLX
[0000] >= 2 17 ) ? ( uLX
[0000] - 2 18 ) : uLX
[0000] (8-299) uLX
[0001] = ( mvpLX
[0001] + mvdLX
[0001] + 218 ) % 2 18 (8-300) mvLX
[0000]
[0000]
[0001] = ( uLX
[0001] >= 2 17 ) ? ( uLX
[0001] - 2 18 ) : uLX
[0001] (8-301) NOTE 1 - The resulting value of mvLX
[0000]
[0000]
[0000] and mvLX
[0000]
[0000]
[0001] specified above is always -2 17 and 2 17 -1 includes -2 17 From 2 17 It falls within the range of -1. The half-sample interpolation filter index hpelIfIdx is derived as follows: hpelIfIdx = AmvrShift == 3 ? 1 : 0 (8-302) The bi-prediction weight index bcwIdx is set equal to bcw_idx[ xCb ][ yCb ].
[0252] If all of the following conditions are true, then refIdxL1 is set equal to −1, predFlagL1 is set equal to 0, and bcwIdx is set equal to 0: - predFlagL0
[0000]
[0000] is equal to 1 - predFlagL1
[0000]
[0000] is equal to 1 - (cbWidth + cbHeight) equals 12.
[0253] The update process for the history-based motion vector predictor list specified in Section 8.5.2.16 is invoked with the luma motion vectors mvL0
[0000]
[0000] and mvL1
[0000]
[0000] , the reference indices refIdxL0 and refIdxL1, the prediction list usage flags predFlagL0
[0000]
[0000] and predFlagL1
[0000]
[0000] , the bi-prediction weight index bcwIdx, and the half-sample interpolation filter index hpelIfIdx.
[0254] It can be understood that the method applies for both uni-prediction and bi-prediction. In the VVC working draft specification, two reference indices and two prediction list usage flags are transmitted, but in the case of uni-prediction, predFlagL1 is set equal to 0, which means that L1 prediction is not used, and in this case refIdxL1 is set equal to -1.
[0255] Possible implementation details of the proposed method for deriving merge candidates based on history (see step 1402 of method 1400 shown in FIG. 14) are described below in the form of modifications to the VVC draft specification. The modifications are highlighted.
[0256] 8.5.2.6 History-Based Merge Candidate Derivation Process The inputs to this process are: - Merge candidate list mergeCandList - numCurrMergeCand, the number of available merge candidates in the list.
[0257] The output to this process is: - Fixed merge candidate list mergeCandList - Fixed number of merge candidates in the list numCurrMergeCand.
[0258] The variables isPrunedA1 and isPrunedB1 are both set equal to FALSE. For each candidate in HmvpCandList[hMvpIdx] with index hMvpIdx = 1..NumHmvpCand, the following ordered steps are repeated until numCurrMergeCand is equal to MaxNumMergeCand - 1. 1. The variable sameMotion is derived as follows: If all of the following conditions are true for any merge candidate N, where N is A1 or B1, then sameMotion and isPrunedN are both set equal to TRUE: - hMvpIdx is 2 or less - Candidate HmvpCandList[ NumHmvpCand - hMvpIdx ] and merge candidate N have the same motion vector and the same reference index - isPrunedN equals FALSE - Otherwise, sameMotion is set equal to FALSE. 2. When sameMotion is equal to FALSE, the candidate HmvpCandList[ NumHmvpCand - hMvpIdx ] is added to the merge candidate list as follows: mergeCandList[ numCurrMergeCand++ ] = HmvpCandList[ NumHmvpCand - hMvpIdx ] (8-381)
[0259] The details of a first possible implementation of the proposed method for updating the history-based motion information (HMVP) candidate list (see steps 1303, 1313 shown in Figures 13A and 13B) are described below in the form of amendments to the VVC draft specification. The amendments are highlighted.
[0260] 8.5.2.16 Update Process for History-Based Motion Vector Predictor Candidate List The inputs to this process are: - Luma motion vectors mvL0 and mvL1 with 1 / 16 fractional sample precision - Reference indices refIdxL0 and refIdxL1 - Prediction list usage flags predFlagL0 and predFlagL1 - Biprediction weight index gbiIdx - index of the set of half-sample interpolation filters hpelIfIdx
[0261] The MVP candidate hMvpCand consists of luma motion vectors mvL0 and mvL1, reference indices refIdxL0 and refIdxL1, prediction list usage flags predFlagL0 and predFlagL1, bi-prediction weight index gbiIdx, and half-sample interpolation filter set index hpelIfIdx.
[0262] The candidate list HmvpCandList is modified using the candidate hMvpCands by the following ordered steps: 1. The variable identicalCandExist is set equal to FALSE and the variable removeIdx is set equal to 0. 2. When NumHmvpCand is greater than 0, for each index hMvpIdx where hMvpIdx = 0..NumHmvpCand - 1, the following steps are applied until identicalCandExist is equal to TRUE. When hMvpCand is equal to HmvpCandList[ hMvpIdx ], identicalCandExist is set equal to TRUE and removeIdx is set equal to hMvpIdx. 3. The candidate list HmvpCandList is updated as follows: - If identicalCandExist is equal to TRUE or NumHmvpCand is equal to MaxNumMergeCand - 1, the following applies: - For each index i, where i = ( removeIdx + 1 )..( NumHmvpCand - 1 ), HmvpCandList[ i - 1 ] is set equal to HmvpCandList[ i ]. - HmvpCandList[ NumHmvpCand - 1 ] is set equal to mvCand. - Otherwise (identicalCandExist is equal to FALSE and NumHmvpCand is less than MaxNumMergeCand - 1), the following applies: - HmvpCandList[ NumHmvpCand++ ] is set equal to mvCand.
[0263] The details of a second possible implementation of the proposed method for updating the history-based motion information (HMVP) candidate list (see steps 1303, 1313 shown in Figures 13A and 13B, and steps 1502-1504 shown in Figure 15) are described below in the form of amendments to the VVC draft specification. The amendments are highlighted.
[0264] 8.5.2.16 Update Process for History-Based Motion Vector Predictor Candidate List The inputs to this process are: - Luma motion vectors mvL0 and mvL1 with 1 / 16 fractional sample precision - Reference indices refIdxL0 and refIdxL1 - Prediction list usage flags predFlagL0 and predFlagL1 - Biprediction weight index bcwIdx - Half-sample interpolation filter index hpelIfIdx.
[0265] The MVP candidate hMvpCand consists of luma motion vectors mvL0 and mvL1, reference indices refIdxL0 and refIdxL1, prediction list usage flags predFlagL0 and predFlagL1, bi-prediction weight index bcwIdx, and half-sample interpolation filter index hpelIfIdx.
[0266] The candidate list HmvpCandList is modified using the candidate hMvpCands by the following ordered steps: 4. The variable identicalCandExist is set equal to FALSE and the variable removeIdx is set equal to 0. 5. When NumHmvpCand is greater than 0, for each index hMvpIdx where hMvpIdx = 0..NumHmvpCand - 1, the following steps are applied until identicalCandExist is equal to TRUE. - When hMvpCand and HmvpCandList[hMvpIdx] have the same motion vector and the same reference index, identicalCandExist is set equal to TRUE and removeIdx is set equal to hMvpIdx. 6. The candidate list HmvpCandList is updated as follows: - If identicalCandExist is equal to TRUE or NumHmvpCand is equal to 5, the following applies: - For each index i, where i = ( removeIdx + 1 )..( NumHmvpCand - 1 ), HmvpCandList[ i - 1 ] is set equal to HmvpCandList[ i ]. - HmvpCandList[ NumHmvpCand - 1 ] is set equal to hMvpCand. - Otherwise (identicalCandExist is equal to FALSE and NumHmvpCand is less than 5), the following applies: - HmvpCandList[ NumHmvpCand++ ] is set equal to hMvpCand.
[0267] As can be seen from the above, the second implementation specifies elements i) and ii) of the HMVP candidate to be compared, while the first implementation specifies all elements of the HMVP candidate to be compared (elements i), ii), and iii), etc.
[0268] The embodiments and exemplary embodiments have their respective methods and corresponding apparatuses.
[0269] FIG. 16 shows a diagram of an apparatus 1600 for constructing a history-based motion information candidate list, which includes an HMI list obtaining unit 1601 and an HMI list updating unit 1603 .
[0270] The history-based motion information (HMI) candidate list obtaining unit 1601 is configured to obtain a history-based motion information candidate list, and the HMI list includes N history-based motion information candidates H related to motion information of a plurality of blocks preceding a block. k where k=0, ... , N-1, and N is an integer greater than 0. Each history-based motion information candidate is an element, i.e., iv) one or more motion vectors MV; v) one or more reference picture indices corresponding to the MV; and vi) Interpolation filter index and the history-based motion information candidate list updating unit 1603 is configured to update the HMI list according to the motion information of the block, and the motion information of the block includes elements, namely: iv) one or more motion vectors MV; v) one or more reference picture indices corresponding to the MV; and vi) Interpolation filter index Includes.
[0271] It will be understood that the HMI list obtaining unit 1601 and the HMI list updating unit 1603 (corresponding to the inter-prediction module) of the encoder 20 or the decoder 30 provided in this embodiment of the present application are functional entities for implementing various execution steps included in the corresponding methods described above, that is, they have functional entities for completely implementing the steps of the method of the present application as well as the extensions and variations of these steps. For details, please refer to the above description of the corresponding methods. For the sake of brevity, the details will not be described again in this specification.
[0272] 17 shows a diagram of an inter-prediction apparatus 1700 according to an embodiment of the present disclosure. The apparatus 1700 is provided for determining motion information for a current block of a frame. The apparatus 1700 includes a list management unit 1701 configured to construct an HMVP list, which is an ordered list of N history-based candidates Hk related to motion information of N previous blocks of the frame preceding the current block, where k=0, ..., N-1, and N is 1 or greater, where each or at least one history-based candidate includes motion information including elements: i) one or more motion vectors MV; ii) one or more reference picture indices corresponding to the MVs; and iii) an interpolation filter index or an index of a set of interpolation filters (e.g., half-pel interpolation filter index), where the HMVP list management unit 1701 is further configured to add one or more history-based candidates from the HMVP list to a motion information candidate list for the current block; and an information derivation unit 1703 configured to derive motion information based on the motion information candidate list.
[0273] In the implementation, the list management unit 1701 is configured to compare at least one element of each history-based candidate in the HMI list with a corresponding element of the current block, and the motion information adding unit is configured to add the motion information of the current block to the HMI list if, as a result of the comparison, at least one element of each history-based candidate in the HMI list is different from the corresponding element of the motion information of the current block.
[0274] Correspondingly, in one example, the exemplary structure of the apparatus 1700 may correspond to the encoder 20 of Figure 2. In another example, the exemplary structure of the apparatus 1700 may correspond to the decoder 30 of Figure 3.
[0275] In another example, the exemplary structure of the apparatus 1700 may correspond to the inter prediction unit 244 of Figure 2. In another example, the exemplary structure of the apparatus 1700 may correspond to the inter prediction unit 344 of Figure 3.
[0276] It will be understood that the list management unit 1701 and the information derivation unit 1703 (corresponding to the inter-prediction module) of the encoder 20 or the decoder 30 provided in this embodiment of the present application are functional entities for implementing various execution steps included in the corresponding methods described above, that is, they have functional entities for completely implementing the steps of the methods of the present application as well as the extensions and variations of these steps. For details, please refer to the above description of the corresponding methods. For the sake of brevity, the details will not be described again in this specification.
[0277] More specifically, the following aspects related to the propagation of SIF indices across CTU boundaries are described.
[0278] As mentioned above, based on the current SIF design, when the SIF technology is applied in a mode of inheriting motion information from the spatial neighboring block above, if the current block is at the upper boundary of the CTU / CTB, the line memory is increased. In the description presented in this specification, the position of the current block is checked. If the current block is at the upper boundary of the CTU / CTB, when inheriting motion information from the upper left (B0), upper (B1), and upper right (B2) neighboring blocks, the IF index is not inherited from the neighboring block, but instead uses a default value to reduce the use of line memory.
[0279] In an aspect of the present disclosure, a method of inter prediction for a current block is provided, the method comprising: The method includes inter-predicting the block, which includes deriving an interpolation filter index for the current block based on the position of the current block (e.g., a coding unit or coding block) within a coding tree block (CTB) or coding tree unit (CTU) and an interpolation filter index inherited from a selected merging candidate.
[0280] FIG. 8 shows a flow diagram of a method for deriving an index of a set of interpolation filters for a current block (such as a coding unit or coding block) within a coding tree block (CTB) or coding tree unit (CTU), including:
[0281] In step 803, the method includes determining whether the current block overlaps a predefined area of the CTB or CTU (such as the top or left border of the CTB or CTU).
[0282] In step 804, the method includes setting an index of an interpolation filter set of the current block as an index of an interpolation filter set of the selected candidate if the current block does not overlap with a predefined area of the CTU (e.g., the current block does not overlap with the top or left boundary of the CTU or the CTB). The selected candidate may be, for example, a selected merge candidate or a selected MVP candidate. The selected candidate may also be a neighboring block corresponding to the selected merge candidate.
[0283] In step 805, the method includes setting an index of the set of interpolation filters for the current block to a predefined value if the current block overlaps a predefined area of the CTB or CTU (such as the top or left boundary of the CTB or CTU).
[0284] Further, in steps 801-802, the method includes building a candidate list, the details of which will not be described again here for the sake of brevity.
[0285] To determine whether the current block is on the top boundary of the CTU 900, the vertical coordinate (yCb) of the upper left corner of the current block is checked, as shown in Figure 9. Assuming that the size of the CTU 900 is equal to (1 << CtbLog2SizeY) x (1 << CtbLog2SizeY), if (yCb >> CtbLog2SizeY) << CtbLog2SizeY is not equal to yCb, the current block is not on the top boundary of the CTU 900 (scenario 1); otherwise (if (yCb >> CtbLog2SizeY) << CtbLog2SizeY is equal to yCb), the current block is on the top boundary of the CTU 900 (scenario 2).
[0286] According to embodiments of the present disclosure, the selected candidates (e.g., selected merge candidates) are spatial merge candidates.
[0287] According to an embodiment of the present disclosure, the vertical position associated with the spatial merge candidate is less than the vertical position of the current block, or the vertical position of the neighboring block corresponding to the spatial merge candidate is less than the vertical position of the current block.
[0288] According to an embodiment of the present disclosure, the spatial merge candidate is the top right candidate (B0 shown in FIG. 6), the top candidate (B1 shown in FIG. 6), or the top left candidate (B2 shown in FIG. 6).
[0289] According to an embodiment of the present disclosure, wherein the merge candidate is an affine merge candidate, the affine merge candidate is an inherited affine merge candidate, where "inherited" means that (i) the candidate is derived based on neighboring affine blocks, (ii) the affine model of the current block is inherited from the affine model of the neighboring affine blocks, or (iii) the affine parameters of the current block are derived based on the affine parameters of the neighboring affine blocks.
[0290] According to an embodiment of the present disclosure, the inherited affine merge candidate is derived based on one of the spatially neighboring blocks, which may include the bottom-left block (such as A0 shown in FIG. 6), the left block (such as A1 shown in FIG. 6), the top-right block (such as B0 shown in FIG. 6), the top block (such as B1 shown in FIG. 6), or the top-left block (such as B2 shown in FIG. 6).
[0291] According to an embodiment of the present disclosure, the inherited affine merge candidates are derived based on blocks that have vertical positions less than the vertical position of the current block.
[0292] According to an embodiment of the present disclosure, the inherited affine merge candidate is derived based on the top right block (such as B0 shown in FIG. 6), the block above (such as B1 shown in FIG. 6), or the top left block (such as B2 shown in FIG. 6).
[0293] According to embodiments of the present disclosure, the selected candidate (e.g., the selected merge candidate) is a lower-order block merge candidate.
[0294] According to an embodiment of the present disclosure, herein the predefined area of the CTU coincides with the CTB or the CTU.
[0295] According to an embodiment of the present disclosure, the step of determining whether the current block overlaps with the predefined area is performed based on the position of the top-left corner of the coding unit (e.g., the horizontal and vertical position of the top-left sample of the current block) (e.g., the luma position (xCb, yCb) specifying the top-left sample of the current block relative to the top-left luma sample of the current picture).
[0296] According to an embodiment of the present disclosure, a current block is inferred to overlap a predefined area (such as the top boundary of a CTU) if the top left corner of the current block overlaps a second predefined area (such as the top left corner of a CTU).
[0297] According to an embodiment of the present disclosure, the second predefined area includes or encompasses the top boundary of the CTU that contains the current block (e.g., the top left corner of the CTU includes or encompasses the top boundary or left boundary of the CTU that contains the current block).
[0298] According to an embodiment of the present disclosure, the step of determining whether the current block overlaps with a predefined area of a CTB or CTU includes determining whether the top left corner of the current block (e.g., luma position (xCb, yCb) specifying the top left sample of the current block relative to the top left luma sample of the current picture) overlaps with the top boundary of the CTU that includes the current coding unit.
[0299] According to an embodiment of the present disclosure, determining whether the top left corner of the current block overlaps with the top boundary of the CTU that includes the current coding unit includes: Get the vertical position (y coordinate) of the top left corner of the current block, Calculating the remainder after dividing the obtained vertical position (y coordinate) by the height of the CTU; If the calculated remainder is equal to 0, infer that the top left corner of the current block overlaps with the top boundary of the current CTU; otherwise, infer that the top left corner of the current coding unit does not overlap with the top boundary of the current CTU.
[0300] According to an embodiment of the present disclosure, determining whether the top left corner of the current coding unit overlaps with the top boundary of the CTU that includes the current coding unit includes: Calculating a first value as a floor value, where the top-left vertical coordinate (coordinate y) of the current block is divided by the height of the CTU to obtain the floor value; Calculating a second value as a floor value, where the top left corner of the vertical coordinator of the inherited neighboring block is divided by the height of the CTU to obtain the floor value; If the second value is equal to the first value, inferring that the top left corner of the current coding unit overlaps the top boundary of the current CTU.
[0301] According to an embodiment of the present disclosure, determining whether the top left corner of the current coding unit overlaps with the top boundary of the CTU that includes the current coding unit includes: Calculating a third value as ( yCb >> CtbLog2SizeY ) << CtbLog2SizeY, where yCb is the upper-left vertical coordinate (y coordinate) of the current block, ">>" is a logical or arithmetic right bit shift, "<<" is a logical or arithmetic left bit shift, and CtbLog2SizeY is a binary logarithmic scale of the size of the CTU or CTB; If (yCb - 1) is less than a third value, inferring that the top left corner of the current coding unit overlaps the top boundary of the current CTU.
[0302] According to an embodiment of the present disclosure, the second predefined area includes or encompasses the top left corner of the CTU that includes the current block.
[0303] According to an embodiment of the present disclosure, determining whether the top left corner of the current block overlaps with the left boundary of the CTU that includes the current block includes: Obtaining the horizontal position (x coordinate) of the top left corner of the current coding unit; Calculating the remainder after dividing the obtained horizontal position (x coordinate) by the width of the CTU; if the calculated remainder is equal to 0, inferring that the top-left corner of the current coding unit overlaps with the left boundary of the current CTU; If not, it includes inferring that the top left corner of the current coding unit does not overlap with the left boundary of the current CTU.
[0304] According to an embodiment of the present disclosure, determining whether the top-left corner of the current coding unit overlaps with the left boundary of the CTU that includes the current coding unit includes: Calculating a fourth value as a floor value, where the top-left horizontal coordinate of the current block (coordinate x) is divided by the width of the CTU to obtain the floor value; Calculating a fifth value as a floor value, where the top left corner of the vertical coordinator of the inherited neighboring block is divided by the width of the CTU to obtain the floor value; If the fifth value is equal to the fourth value, inferring that the top left corner of the current coding unit overlaps with the left boundary of the current CTU.
[0305] According to an embodiment of the present disclosure, determining whether the top-left corner of the current coding unit overlaps with the left boundary of the CTU that includes the current coding unit includes: Calculating a sixth value as ( xCb >> CtbLog2SizeX ) << CtbLog2SizeX, where xCb is the top-left vertical coordinate (y coordinate) of the current block, ">>" is a logical or arithmetic right bit shift, and "<<" is a logical or arithmetic left bit shift, and CtbLog2SizeX is a binary logarithmic scale of the width of the CTU; If (xCb - 1) is less than a sixth value, inferring that the top-left corner of the current coding unit overlaps the left boundary of the current CTU.
[0306] According to an embodiment of the present disclosure, instead of a combination of left and right shift operations on N bits, a logical AND with a bit mask containing bits 0 in the lowest N positions and bits 1 elsewhere (e.g., (yCb >> CtbLog2SizeY) << CtbLog2SizeY can be calculated as the logical AND of yCb with a bit mask containing bits 0 in the lowest CtbLog2SizeY positions and bits 1 elsewhere). For example, if yCb is in the range [0, 2 32 - 1] and CtbLog2SizeY is equal to 7, the value of yCb & 0xFFFFFF80 can be calculated instead of ( yCb >> CtbLog2SizeY ) << CtbLog2SizeY, where 0xFFFFFF80 is a bitmask that contains zeros in the last 7 positions and ones everywhere else.
[0307] According to an embodiment of the present disclosure, a logical or arithmetic shift is used to calculate the floor value of the division result. (e.g., a / 2 n An exemplary floor value for a can be calculated as a >> n).
[0308] According to an embodiment of the present disclosure, the second predefined area includes or encompasses only the upper boundary of the CTU that includes the current block.
[0309] According to an embodiment of the present disclosure, the second predefined area includes or encompasses only the left boundary of the CTU that includes the current block.
[0310] According to an embodiment of the present disclosure, the second predefined area includes or encompasses only the top and left boundaries of the CTU that includes the current block.
[0311] According to an embodiment of the present disclosure, the step of setting the index of the set of interpolation filters for the current block to a predefined value comprises: Setting an index of the set of interpolation filters for the current block to a seventh value, the seventh value being determined before construction of the merge list.
[0312] According to an embodiment of the present disclosure, determining the seventh value includes: determining an interpolation filter set index of one of the spatial neighboring blocks of the current block; and setting a seventh value equal to the determined interpolation filter set index.
[0313] According to an embodiment of the present disclosure, "one of the spatially neighboring blocks" means the left neighboring block (this block is referenced as A1 in FIG. 6).
[0314] Possible implementation details of the proposed method of propagating SIF indices across CTU boundaries (the process is shown in Figures 8 and 9) are described below in the form of modifications to the working draft specification of the SIF proposal). The modifications are highlighted.
[0315] 8.5.2.3 Derivation Process for Spatial Merge Candidates The inputs to this process are: - Luma position ( xCb, yCb ) of the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture - variable cbWidth that specifies the width of the current coding block in luma samples - Variable cbHeight specifying the height of the current coding block in luma samples.
[0316] The output of this process, where X is either 0 or 1, is: - Availability flags of neighboring coding units availableFlagA0, availableFlagA1, availableFlagB0, availableFlagB1, and availableFlagB2 - Reference indices of neighboring coding units refIdxLXA0, refIdxLXA1, refIdxLXB0, refIdxLXB1, and refIdxLXB2 - Prediction list usage flags of neighboring coding units predFlagLXA0, predFlagLXA1, predFlagLXB0, predFlagLXB1, and predFlagLXB2 - Motion vectors mvLXA0, mvLXA1, mvLXB0, mvLXB1, and mvLXB2 with 1 / 16 fractional sample precision of neighboring coding units - Half-sample interpolation filter indices hpelIfIdxA0, hpelIfIdxA1, hpelIfIdxB0, hpelIfIdxB1, and hpelIfIdxB2 - Bi-prediction weight indices gbiIdxA0, gbiIdxA1, gbiIdxB0, gbiIdxB1, and gbiIdxB2.
[0317] For the derivation of availableFlagA1, refIdxLXA1, predFlagLXA1, and mvLXA1, the following applies: The luma position ( xNbA1, yNbA1) within the neighboring luma coding block is set equal to ( xCb - 1, yCb + cbHeight - 1 ). - The availability derivation process for the block specified in clause 6.4 is called with the current luma position (xCurr, yCurr) set equal to (xCb, yCb) and the neighboring luma positions (xNbA1, yNbA1) as input, and the output is assigned to the availability flag availableA1 of the block. The variables availableFlagA1, refIdxLXA1, predFlagLXA1, and mvLXA1 are derived as follows: If availableA1 is equal to FALSE, then availableFlagA1 is set equal to 0, both components of mvLXA1 are set equal to 0, refIdxLXA1 is set equal to −1, predFlagLXA1 is set equal to 0, X is 0 or 1, and gbiIdxA1 is set equal to 0. - Otherwise, availableFlagA1 is set equal to 1 and the following allocation is made: mvLXA1= MvLX[ xNbA1][ yNbA1] (8-294) refIdxLXA1= RefIdxLX[ xNbA1][ yNbA1] (8-295) predFlagLXA1= PredFlagLX[ xNbA1][ yNbA1] (8-296) hpelIfIdxA1= HpelIfIdx[ xNbA1][ yNbA1] (8-297) gbiIdxA1= GbiIdx[ xNbA1][ yNbA1] (8-297)
[0318] For the derivation of availableFlagB1, refIdxLXB1, predFlagLXB1, and mvLXB1, the following applies: The luma position ( xNbB1, yNbB1) within the neighboring luma coding block is set equal to ( xCb + cbWidth - 1, yCb - 1 ). - The availability derivation process for the block specified in clause 6.4 is called with the current luma position (xCurr, yCurr) set equal to (xCb, yCb) and the neighboring luma positions (xNbB1, yNbB1) as input, and the output is assigned to the availability flag availableB1 of the block. The variables availableFlagB1, refIdxLXB1, predFlagLXB1, and mvLXB1 are derived as follows: If one or more of the following conditions are true, then availableFlagB1 is set equal to 0, both components of mvLXB1 are set equal to 0, refIdxLXB1 is set equal to −1, predFlagLXB1 is set equal to 0, X is 0 or 1, and gbiIdxB1 is set equal to 0: - availableB1 equals FALSE. availableA1 is equal to TRUE and the luma positions (xNbA1, yNbA1) and (xNbB1, yNbB1) have the same motion vector and the same reference index. - Otherwise, availableFlagB1 is set equal to 1 and the following allocation is made: mvLXB1= MvLX[ xNbB1][ yNbB1] (8-298) refIdxLXB1= RefIdxLX[ xNbB1][ yNbB1] (8-299) predFlagLXB1= PredFlagLX[ xNbB1][ yNbB1] (8-300) If ( yCb - 1 ) < ( ( yCb >> CtbLog2SizeY ) << CtbLog2SizeY ), hpelIfIdxB1= 2 Otherwise, hpelIfIdxB1= HpelIfIdx[ xNbB1][ yNbB1] gbiIdxB1= GbiIdx[ xNbB1][ yNbB1] (8-301)
[0319] For the derivation of availableFlagB0, refIdxLXB0, predFlagLXB0, and mvLXB0, the following applies: The luma position ( xNbB0, yNbB0) within the neighboring luma coding block is set equal to ( xCb + cbWidth, yCb - 1 ). - The availability derivation process for the block specified in clause 6.4 is called with the current luma position (xCurr, yCurr) set equal to (xCb, yCb) and the neighboring luma positions (xNbB0, yNbB0) as input, and the output is assigned to the availability flag availableB0 of the block. The variables availableFlagB0, refIdxLXB0, predFlagLXB0, and mvLXB0 are derived as follows: If one or more of the following conditions are true, then availableFlagB0 is set equal to 0, both components of mvLXB0 are set equal to 0, refIdxLXB0 is set equal to -1, predFlagLXB0 is set equal to 0, X is 0 or 1, and gbiIdxB0 is set equal to 0: - availableB0 equals FALSE. availableB1 is equal to TRUE and the luma positions (xNbB1, yNbB1) and (xNbB0, yNbB0) have the same motion vector and the same reference index. - availableA1 is equal to TRUE, the luma positions (xNbA1, yNbA1) and (xNbB0, yNbB0) have the same motion vector and the same reference index, and merge_triangle_flag[xCb][yCb] is equal to 1. - Otherwise, availableFlagB0 is set equal to 1 and the following allocation is made: mvLXB0= MvLX[ xNbB0][ yNbB0] (8-302) refIdxLXB0= RefIdxLX[ xNbB0][ yNbB0] (8-303) predFlagLXB0= PredFlagLX[ xNbB0][ yNbB0] (8-304) If ( yCb - 1 ) < ( ( yCb >> CtbLog2SizeY ) << CtbLog2SizeY ), hpelIfIdxB0= 2 Otherwise, hpelIfIdxB0= HpelIfIdx[ xNbB0][ yNbB0] (8-305) gbiIdxB0= GbiIdx[ xNbB0][ yNbB0] (8-305)
[0320] For the derivation of availableFlagA0, refIdxLXA0, predFlagLXA0, and mvLXA0, the following applies: The luma position ( xNbA0, yNbA0) within the neighboring luma coding block is set equal to ( xCb - 1, yCb + cbWidth ). - The availability derivation process for the block specified in clause 6.4 is called with the current luma position (xCurr, yCurr) set equal to (xCb, yCb) and the neighboring luma positions (xNbA0, yNbA0) as input, and the output is assigned to the availability flag availableA0 of the block. The variables availableFlagA0, refIdxLXA0, predFlagLXA0, and mvLXA0 are derived as follows: If one or more of the following conditions are true, then availableFlagA0 is set equal to 0, both components of mvLXA0 are set equal to 0, refIdxLXA0 is set equal to −1, predFlagLXA0 is set equal to 0, X is 0 or 1, and gbiIdxA0 is set equal to 0: - availableA0 equals FALSE. availableA1 is equal to TRUE and the luma positions (xNbA1, yNbA1) and (xNbA0, yNbA0) have the same motion vector and the same reference index. - availableB1 is equal to TRUE, the luma positions ( xNbB1, yNbB1) and ( xNbA0, yNbA0) have the same motion vector and the same reference index, and merge_triangle_flag[ xCb ][ yCb ] is equal to 1. - availableB0 is equal to TRUE, the luma positions ( xNbB0, yNbB0) and ( xNbA0, yNbA0) have the same motion vector and the same reference index, and merge_triangle_flag[ xCb ][ yCb ] is equal to 1. - Otherwise, availableFlagA0 is set equal to 1 and the following allocation is made: mvLXA0= MvLX[ xNbA0][ yNbA0] (8-306) refIdxLXA0= RefIdxLX[ xNbA0][ yNbA0] (8-307) predFlagLXA0= PredFlagLX[ xNbA0][ yNbA0] (8-308) hpelIfIdxA0= HpelIfIdx[ xNbA0][ yNbA0] (8-309) gbiIdxA0= GbiIdx[ xNbA0][ yNbA0] (8-309)
[0321] For the derivation of availableFlagB2, refIdxLXB2, predFlagLXB2, and mvLXB2, the following applies: The luma position ( xNbB2, yNbB2) in the neighboring luma coding block is set equal to ( xCb - 1, yCb - 1 ). - The availability derivation process for the block specified in clause 6.4 is called with the current luma position (xCurr, yCurr) set equal to (xCb, yCb) and the neighboring luma positions (xNbB2, yNbB2) as input, and the output is assigned to the availability flag availableB2 of the block. The variables availableFlagB2, refIdxLXB2, predFlagLXB2, and mvLXB2 are derived as follows: If one or more of the following conditions are true, then availableFlagB2 is set equal to 0, both components of mvLXB2 are set equal to 0, refIdxLXB2 is set equal to -1, predFlagLXB2 is set equal to 0, X is 0 or 1, and gbiIdxB2 is set equal to 0: - availableB2 equals FALSE. availableA1 is equal to TRUE and the luma positions (xNbA1, yNbA1) and (xNbB2, yNbB2) have the same motion vector and the same reference index. availableB1 is equal to TRUE and the luma positions (xNbB1, yNbB1) and (xNbB2, yNbB2) have the same motion vector and the same reference index. - availableB0 is equal to TRUE, the luma positions ( xNbB0, yNbB0) and ( xNbB2, yNbB2) have the same motion vector and the same reference index, and merge_triangle_flag[ xCb ][ yCb ] is equal to 1. - availableA0 is equal to TRUE, the luma positions ( xNbA0, yNbA0) and ( xNbB2, yNbB2) have the same motion vector and the same reference index, and merge_triangle_flag[ xCb ][ yCb ] is equal to 1. - availableFlagA0+ availableFlagA1+ availableFlagB0+ availableFlagB1 equals 4 and merge_triangle_flag[ xCb ][ yCb ] equals 0. - Otherwise, availableFlagB2 is set equal to 1 and the following allocation is made: mvLXB2= MvLX[ xNbB2][ yNbB2] (8-310) refIdxLXB2= RefIdxLX[ xNbB2][ yNbB2] (8-311) predFlagLXB2= PredFlagLX[ xNbB2][ yNbB2] (8-312) If ( yCb - 1 ) < ( ( yCb >> CtbLog2SizeY ) << CtbLog2SizeY ), hpelIfIdxB2= 2 Otherwise, hpelIfIdxB2= HpelIfIdx[ xNbB2][ yNbB2] (8-313) gbiIdxB2= GbiIdx[ xNbB2][ yNbB2] (8-313)
[0322] As can be understood from the above, the half-pixel interpolation filter indexes of neighboring blocks of a current block are determined based on whether the current block overlaps the boundary of a CTU. For example, equations (8-298) to (8-301) show steps for determining the half-pixel interpolation filter index of neighboring block B1 (see Figures 8 and 9). Equations (8-302) to (8-305) show steps for determining the half-pixel interpolation filter index of neighboring block B0 (see Figures 8 and 9). Equations (8-310) to (8-313) show steps for determining the half-pixel interpolation filter index of neighboring block B2 (see Figures 8 and 9).
[0323] Based on the above, the present disclosure is directed to storing an SIF index in an HMVP table (or propagating the SIF index through the HMVP table) and using the SIF index for HMVP candidates in the merge list construction process. The SIF method is used to select an appropriate interpolation filter (IF) depending on the content; that is, for areas with sharp edges, a normal DCT-based IF is used, and for smooth areas (or when maintaining sharp edges is not required), an alternative 6-tap IF (Gaussian filter) is used. For normal inter prediction, the IF index is explicitly signaled, while for merge mode, not only are the MV and reference picture indexes borrowed from the corresponding spatial candidate for merge (HMVP merge candidate), but the IF index is also borrowed from the corresponding spatial candidate for merge. This is in contrast to a conventional design in which the IF index was not propagated through the HMVP table. Therefore, in a conventional design, for blocks coded in merge mode and merge candidates obtained from the HMVP table, the alternative IF may not be used. The HMVP table is used to store motion information from neighboring blocks (but not necessarily from adjacent blocks as in normal spatial merge candidates). The idea of HMVP is to use motion information from blocks that are spatially close to the current block but not necessarily adjacent (blocks from some spatial neighborhood). Thus, for example, if the current block contains smooth content and the neighboring blocks contain mostly sharp content, borrowing IF indices from the neighboring blocks is not efficient. However, smooth content may be in some spatial neighborhood of the current block, and motion information of such blocks may be stored in the HMVP table. Propagating IF indices through the HMVP table as shown herein enables the use of an appropriate IF for the current block (e.g., a Gaussian filter may be selected for smooth content or for cases where maintaining sharp edges is not required).This brings the advantage of increasing coding efficiency. Without the present invention, the default IF index (corresponding to the 8-tap DCT-based IF) is always used for HMVP merging candidates, and the content details of the current block (whether sharp edges need to be preserved or not) cannot be taken into consideration.
[0324] Additionally, this disclosure is also directed to using only MV and reference picture indices (without using SIF indices) in the pruning process during updating of the HMVP table.
[0325] When a new element is added to an HMVP record, it must be determined whether this new element is used in record comparison. A simple approach would be to use all elements of the HMVP record in the record comparison (default C-style structure comparison). However, in this disclosure, the IF index is not used in HMVP record comparison. There are two reasons for this design.
[0326] The first reason is to avoid additional computational complexity. Each comparison operation incurs additional computational operations in the HMVP table update process and merge candidate construction process. Therefore, if comparison operations can be reduced or eliminated, computational complexity can be reduced, thereby increasing coding efficiency. From an implementation perspective, if unnecessary comparisons can be avoided, a better implementation can be achieved. Therefore, instead of the default C-style structure comparison of HMVP records, the elements of HMVP records are divided into two subsets: elements used for record comparison and elements not used for record comparison.
[0327] The second reason is to maintain diversity among HMVP records. For example, having two HMVP records with the same MV and reference indexes but differing only in their IF indexes is inefficient because these two records are not "different enough." Instead, it is more efficient to consider these HMVP records the same during the HMVP table update process. In this case, a new record that differs only in its IF index from an existing record is not added to the HMVP table. As a result, the "old" record that is "sufficiently different" from the other records (has a different MV or reference index) is maintained. In other words, for a new record to be added to the HMVP table, this new record should not only be bitwise different from the existing record; the new record must be "significantly different." From a coding efficiency perspective, it is more efficient to have two records in the HMVP table with different MV or reference indexes than two records that differ only in their IF indexes.
[0328] Furthermore, the present disclosure also addresses constraints on merging parameters of switchable interpolation filters (SIFs) to save line memory. Compared with previous designs of SIFs, the presented disclosure introduces a method for applying SIFs with motion information inheritance tools without increased line memory, which saves line memory bandwidth. For high-resolution cases, the line memory savings significantly reduce on-chip memory costs.
[0329] The modified IF index derivation method improves coding efficiency by using more appropriate IF indexes for CUs coded in merge mode and having merge indices corresponding to history-based merge candidates.
[0330] The mathematical operators used in this application are similar to those used in the C programming language and may refer to mathematical operators specified in the HEVC standard. However, the results of integer division and arithmetic shift operations are more strictly defined, and additional operations such as exponentiation and division of real values are defined. The numbering and counting rules generally start from 0, e.g., "first" is equivalent to number 0, "second" is equivalent to number 1, and so on.
[0331] The following is a description of the application of the encoding and decoding methods shown in the above embodiments and the systems that use them.
[0332] 18 is a block diagram showing a content supply system 3100 for realizing a content distribution service. The content supply system 3100 includes a capture device 3102, a terminal device 3106, and optionally a display 3126. The capture device 3102 communicates with the terminal device 3106 via a communication link 3104. The communication link may include the communication channel 13 described above. The communication link 3104 may include, but is not limited to, WIFI, Ethernet, cable, wireless (3G / 4G / 5G), USB, or any combination thereof.
[0333] The capture device 3102 may generate data and encode the data according to the encoding method described in the above embodiment. Alternatively, the capture device 3102 may deliver the data to a streaming server (not shown), which encodes the data and transmits the encoded data to the terminal device 3106. The capture device 3102 may include, but is not limited to, a camera, a smartphone or smart pad, a computer or laptop, a video conferencing system, a PDA, an in-vehicle device, or any combination thereof. For example, the capture device 3102 may include the source device 12 described above. When the data includes video, a video encoder 20 included in the capture device 3102 may actually perform the video encoding process. When the data includes audio (i.e., voice), an audio encoder included in the capture device 3102 may actually perform the audio encoding process. In some practical scenarios, the capture device 3102 delivers the encoded video and audio data by multiplexing them together. In other practical scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. The capture device 3102 delivers the encoded audio data and the encoded video data separately to the terminal device 3106 .
[0334] In the content supply system 3100, the terminal device 3106 receives and plays the encoded data. The terminal device 3106 can be a device having data reception and restoration capabilities, such as a smartphone or smart pad 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a set-top box (STB) 3116, a video conferencing system 3118, a video surveillance system 3120, a personal digital assistant (PDA) 3122, an in-vehicle device 3124, or any combination thereof, capable of decoding the above-mentioned encoded data. For example, the terminal device 3106 can include the above-mentioned destination device 14. When the encoded data includes video, the video decoder 30 included in the terminal device is prioritized to perform video decoding. When the encoded data includes audio, the audio decoder included in the terminal device is prioritized to perform audio decoding processing.
[0335] For terminal devices with a display, such as a smartphone or smart pad 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a personal digital assistant (PDA), or an in-vehicle device 3124, the terminal device can provide the decoded data to its display. For terminal devices without a display, such as an STB 3116, a video conferencing system 3118, or a video surveillance system 3120, an external display 3126 is contacted to receive and show the decoded data.
[0336] When each device in this system performs encoding or decoding, the picture encoding device or picture decoding device shown in the above embodiments may be used.
[0337] 19 is a diagram illustrating an example structure of a terminal device 3106. After the terminal device 3106 receives a stream from the capture device 3102, a protocol progression unit 3202 analyzes the transmission protocol of the stream. The protocol may include, but is not limited to, Real Time Streaming Protocol (RTSP), Hypertext Transfer Protocol (HTTP), HTTP Live Streaming Protocol (HLS), MPEG-DASH, Real Time Transport Protocol (RTP), Real Time Messaging Protocol (RTMP), or any type of combination thereof.
[0338] After the protocol progression unit 3202 processes the stream, a stream file is generated. The file is output to the demultiplexing unit 3204. The demultiplexing unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As mentioned above, in some practical scenarios, for example, in a video conference system, the encoded audio data and encoded video data are not multiplexed. In this situation, the encoded data is sent to the video decoder 3206 and the audio decoder 3208 without passing through the demultiplexing unit 3204.
[0339] The demultiplexing process generates a video elementary stream (ES), an audio ES, and optionally subtitles. A video decoder 3206, which includes the video decoder 30 described in the above embodiment, decodes the video ES according to the decoding method shown in the above embodiment to generate video frames, and supplies this data to a synchronization unit 3212. An audio decoder 3208 decodes the audio ES to generate audio frames, and supplies this data to the synchronization unit 3212. Alternatively, the video frames may be stored in a buffer (not shown in FIG. 19 ) before being supplied to the synchronization unit 3212. Similarly, the audio frames may be stored in a buffer (not shown in FIG. 19 ) before being supplied to the synchronization unit 3212.
[0340] The synchronization unit 3212 synchronizes video and audio frames and provides the video / audio to a video / audio display 3214. For example, the synchronization unit 3212 synchronizes the presentation of video and audio information. The information may be coded in a syntax that uses timestamps for the presentation of the coded audio and visual data as well as timestamps for the delivery of the data stream itself.
[0341] If subtitles are included in the stream, the subtitle decoder 3210 decodes the subtitles, synchronizes the subtitles with the video and audio frames, and provides the video / audio / subtitles to the video / audio / subtitle display 3216.
[0342] The present invention is not limited to the above-mentioned system, and either the picture encoding device or the picture decoding device of the above-mentioned embodiments may be incorporated into other systems, for example, a system of an automobile.
[0343] It should be noted that, although embodiments of the present invention have been described primarily in terms of video coding, embodiments of coding system 10, encoder 20, and decoder 30 (and correspondingly, system 10), as well as other embodiments described herein, may also be configured for processing or coding of still pictures, i.e., processing or coding of individual pictures independent of any preceding or subsequent pictures, similar to video coding. Generally, when picture processing coding is limited to a single picture 17, only inter prediction units 244 (encoder) and 344 (decoder) may not be available. All other functions (also called tools or technologies) of the video encoder 20 and the video decoder 30, such as residual calculation 204 / 304, transform 206, quantization 208, inverse quantization 210 / 310, (inverse) transform 212 / 312, partitioning 262 / 362, intra prediction 254 / 354, and / or loop filter 220, 320, and entropy coding 270, and entropy decoding 304, may be equally used for processing still pictures.
[0344] For example, embodiments of the encoder 20 and decoder 30 and the functionality described herein in connection with the encoder 20 and decoder 30 may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on a computer-readable medium or transmitted over a communication medium as one or more instructions or code and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which correspond to tangible media such as data storage media, or communication media, including any medium that facilitates transfer of a computer program from one place to another, for example via a communication protocol. Thus, generally, computer-readable media may correspond to (1) tangible computer-readable storage media that are non-transitory or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0345] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio wave, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio wave, and microwave are included within the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead cover non-transitory, tangible storage media. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically while discs reproduce data optically using a laser. Combinations of the above should also be included within the scope of computer-readable media.
[0346] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the term "processor," as used herein, may refer to any of the above structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Also, the techniques may be implemented entirely in one or more circuits or logic elements.
[0347] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). Various components, modules, or units are described in this disclosure to highlight aspects of the functionality of a device configured to perform the disclosed techniques, but do not necessarily require realization by different hardware units. Rather, as noted above, the various units may be combined in a codec hardware unit or provided by a collection of interoperable hardware units including one or more processors as described above in conjunction with suitable software and / or firmware. [Explanation of symbols]
[0348] 10 Video coding system, coding system 12 Source Devices 13 Encoded picture data, communication channel 14 Destination Device 16 Picture Source 17 Picture, Picture Data, Raw Picture, Raw Picture Data, Monochrome Picture, Color Picture, Current Picture 18 Preprocessor, preprocessing unit, picture preprocessor 19 Preprocessed Picture, Preprocessed Picture Data 20 Video Encoder, Encoder 21 Encoded picture data, encoded bitstream 22 Communication interface, communication unit 28 Communication interface, communication unit 30 decoder, video decoder 31 Decoded Picture Data, Decoded Picture 32 Post-processor, post-processing unit 33 Post-processed picture data, post-processed picture 34 Display Devices 46 Processing Circuit 100 Video Encoder 201 Input, input interface 203 Picture Block, Original Block, Current Block, Segmented Block, Current Picture Block 204 Residual Calculation Unit, Residual Calculation 205 Residual Block, Residual 206 Conversion Processing Unit, Conversion 207 Conversion Factor 208 Quantization Unit, Quantization 209 Quantized Coefficients, Quantized Transform Coefficients, Quantized Residual Coefficients 210 Inverse quantization unit, inverse quantization 211 Dequantized Coefficients, Dequantized Residual Coefficients 212 Inverse Transform Processing Unit, (Inverse) Transform 213 Reconstructed residual block, dequantized coefficients, transform block 214 Reconstruction Unit, Adder, Summer 215 reconstructed blocks 216 buffers 220 Loop filter unit, loop filter 221 filtered blocks, filtered reconstructed blocks 230 Decoded Picture Buffer (DPB) 231 decoded pictures 244 Inter Prediction Units 254 Intra prediction unit, Inter prediction unit, Intra prediction 260 Mode Selection Unit 262 Division Unit, Division 265 prediction block, predictor 266 Syntax Elements 270 Entropy Coding Unit, Entropy Coding 272 Output, Output Interface 304 Entropy Decoding Unit, Residual Calculation, Entropy Decoding 309 Quantized Coefficients 310 Inverse Quantization Unit, Inverse Quantization 311 Dequantized Coefficients, Transform Coefficients 312 Inverse Transform Processing Unit, (inverse) transformation, output 313 Reconstructed Residual Blocks 314 Reconstruction Unit, Summer, Adder 315 reconstructed blocks 320 Loop filter, loop filter unit, loop filtering unit 321 filtered blocks, decoded video blocks 330 Decoded Picture Buffer (DPB), Decoded Picture Buffer (DBP) 331 decoded pictures 344 Inter Prediction Unit 354 Intra prediction unit, Intra prediction 360 Mode Selection Unit 362 Division 365 predicted blocks 400 Video Coding Device 410 Incoming port, input port 420 Receiver Unit (Rx) 430 Processor, Logic Unit, Central Processing Unit (CPU) 440 Transmitter Unit (Tx) 450 outgoing and outgoing ports 460 memory 470 Coding Module 500 devices 502 processor 504 memory 506 Data 508 Operating Systems 510 Application Program 512 Bus 514 Secondary memory 518 Display 700 current block 900 CTU 1300 Construction Method 1400 methods 1600 equipment 1601 HMI list acquisition unit 1603 HMI List Update Unit 1700 Inter Prediction Device 1701 List Management Unit 1703 Information Derivation Unit 3100 Contents Supply System 3102 Capture Device 3104 Communication Links 3106 Terminal Device 3108 Smartphones, smart pads 3110 Computers, Laptops 3112 Network Video Recorder (NVR) / Digital Video Recorder (DVR) 3114 TV 3116 Set-top box (STB) 3118 Video Conference System 3120 Video Surveillance System 3122 Personal Digital Assistant (PDA) 3124 In-Vehicle Devices 3126 Display 3202 Protocol Progression Unit 3204 Demultiplexing Unit 3206 Video Decoder 3208 Audio Decoder 3210 Subtitle Decoder 3212 Synchronous Unit 3214 Video / Audio Display 3216 Video / Audio / Subtitle Display
Claims
1. 1. A method for constructing a history-based motion information candidate list, comprising: obtaining a history-based motion information candidate list, wherein the history-based motion information candidate list includes N history-based motion information candidates H H including motion information of N preceding blocks preceding the block; k where k=0, ..., N-1, N is an integer greater than 0, and each history-based motion information candidate is an element, i.e., i) one or more motion vectors (MVs); ii) one or more reference picture indexes corresponding to said one or more MVs; and iii) Interpolation filter index and updating the history-based motion information candidate list based on motion information of the block, wherein the motion information of the block includes elements: i) one or more motion vectors (MVs) of said block; ii) one or more reference picture indexes corresponding to the MV of the block; and iii) Interpolation filter index and A method comprising:
2. updating the history-based motion information candidate list, The following elements of each historical motion information candidate in the historical motion information candidate list: i) said one or more motion vectors (MVs); and ii) the one or more reference picture indexes corresponding to the one or more MVs; If at least one of the elements is different from the corresponding element of the motion information of the block, the history-based motion information candidate H k 2. The method of claim 1 , further comprising adding:
3. updating the history-based motion information candidate list, The following elements of the history-based motion information candidates of the history-based motion information candidate list: i) one or more motion vectors (MVs), and ii) one or more reference picture indexes corresponding to said one or more MVs; is the same as the corresponding element of the motion information of the block, remove the history-based motion information candidate from the history-based motion information candidate list, and add the history-based motion information candidate H k 2. The method of claim 1 , further comprising adding k=N−1.
4. updating the history-based motion information candidate list, If N is equal to a predefined number, select a history-based motion information candidate H with k=0 from the history-based motion information candidate list. k and the motion information of the block is deleted from the history-based motion information candidate list, and the motion information of the block is added to the history-based motion information candidate list with k=N-1 as the history-based motion information candidate H k 4. The method of claim 1, further comprising adding:
5. comparing whether the motion vectors of the history-based motion information candidates in the history-based motion information candidate list are the same as the corresponding motion vectors of the blocks; comparing whether the reference picture index of the history-based motion information candidate is the same as the corresponding reference picture index of the block; 4. The method of claim 2 or 3, further comprising:
6. comparing whether at least one of the motion vectors of each history-based motion information candidate is different from the corresponding motion vector of the block; comparing whether at least one of the reference picture indexes of each HMVP candidate is different from the corresponding reference picture index of the block; 4. The method of claim 2 or 3, further comprising:
7. 7. The method of claim 1, wherein the interpolation filter index included in the history-based motion information candidate indicates a half-sample interpolation filter from a set of half-sample interpolation filters, and the half-sample interpolation filter is applied to interpolate half-sample values only when at least one of the one or more MVs of the history-based motion information candidate points to a half-sample position.
8. 1. A method for inter prediction for a block within a frame of a video signal, comprising: constructing a history-based motion information candidate list, the history-based motion information candidate list including N history-based motion information candidates H H including motion information of N preceding blocks preceding the block; k where k=0, ..., N-1, N is an integer greater than 0, and each history-based motion information candidate is an element, i.e., i) one or more motion vectors (MVs); ii) one or more reference picture indexes corresponding to said MV; and iii) Interpolation filter index and adding one or more history-based motion information candidates from the history-based motion information candidate list to a motion information candidate list for the block; deriving motion information for the block based on the motion information candidate list; A method comprising:
9. 9. The method of claim 8, wherein an alternative half-sample interpolation filter is applied only when at least one of the one or more MVs of the derived motion information points to a half-sample position, and the alternative half-sample interpolation filter is indicated by an interpolation filter index included in the derived motion information.
10. 9. The method of claim 8, wherein the interpolation filter index included in the history-based motion information candidate indicates a half-sample interpolation filter from a set of half-sample interpolation filters, and the half-sample interpolation filter is applied to interpolate half-sample values only when at least one of the one or more MVs of the history-based motion information candidate points to a half-sample position.
11. The following elements of each historical motion information candidate in the historical motion information candidate list: i) said one or more motion vectors (MVs); and ii) the one or more reference picture indexes corresponding to the one or more MVs; If at least one of the elements is different from the corresponding element of the motion information of the block, the history-based motion information candidate H k 11. The method of claim 8, further comprising the step of adding:
12. The following elements of the history-based motion information candidates of the history-based motion information candidate list: i) one or more motion vectors (MVs), and ii) one or more reference picture indices corresponding to said MV; is the same as the corresponding element of the motion information of the block, remove the history-based motion information candidate from the history-based motion information candidate list, and add the history-based motion information candidate H k 11. The method of claim 8, further comprising the step of adding:
13. If N is equal to a predefined number, select a history-based motion information candidate H with k=0 from the history-based motion information candidate list. k and the motion information of the block is deleted from the history-based motion information candidate list, and the motion information of the block is added to the history-based motion information candidate list with k=N-1 as the history-based motion information candidate H k 13. The method of claim 8, further comprising adding as
14. comparing whether the motion vectors of the history-based motion information candidates in the history-based motion information candidate list are the same as the corresponding motion vectors of the blocks; comparing whether the reference picture index of the history-based motion information candidate is the same as the corresponding reference picture index of the block; 13. The method of claim 11 or 12, further comprising:
15. comparing whether at least one of the motion vectors of each history-based motion information candidate is different from the corresponding motion vector of the block; comparing whether at least one of the reference picture indexes of each HMVP candidate is different from the corresponding reference picture index of the block; 13. The method of claim 11 or 12, further comprising:
16. The method of any one of claims 8 to 15, wherein the history-based motion information candidate list has a length of N, where N is 5 or 6.
17. The method of claim 8 , wherein the motion information candidate list is used for a merge mode or a skip mode.
18. deriving the motion information for the block based on the motion information candidate list, 18. The method of claim 8, comprising deriving the motion information referenced by a candidate index from the motion information candidate list as the motion information of a current block.
19. 19. The method of claim 8, further comprising the step of: when at least one of one or more motion vectors (MVs) included in the derived motion information points to a half-sample position, obtaining predicted sample values of the block by applying a half-sample interpolation filter to sample values of a reference picture pointed to by the MV, wherein the half-sample interpolation filter is indicated by a half-sample interpolation filter index included in the derived motion information, and the reference picture is indicated by the one or more reference picture indexes included in the derived motion information.
20. An encoder (20) including processing circuitry for carrying out the method of any one of claims 1 to 19.
21. A decoder (30) including processing circuitry for carrying out the method of any one of claims 1 to 19.
22. 20. A computer program product comprising program code for carrying out the method of any one of claims 1 to 19.
23. one or more processors; a non-transitory computer-readable storage medium coupled to the processor and storing programming for execution by the processor, the programming, when executed by the processor, configuring a decoder to perform the method of any one of claims 1 to 19; and A decoder containing
24. one or more processors; a non-transitory computer-readable storage medium coupled to the processor and storing programming for execution by the processor, the programming, when executed by the processor, configuring an encoder to perform the method of any one of claims 1 to 19; and Encoder including.
25. A non-transitory computer readable medium carrying program code that, when executed by a computing device, causes the computing device to perform the method of any one of claims 1 to 19.
Citation Information
Patent Citations
Method and apparatus for deriving an interpolation filter index for a current block
JP7708823B2
Systems and methods of switching interpolation filters
US20180098066A1
Video decoder and methods
WO2020085954A1
Using interpolation filters for history based motion vector prediction
WO2020200236A1