Decoding method, decoder, computer program product and readable storage medium for blocks in a video signal frame
By constructing a history-based motion information candidate list and inheriting the half-pixel interpolation filter index, the problem of improper use of interpolation filter index in inter prediction in the prior art is solved, and the effect of improving video signal quality and decoding efficiency is achieved.
Patent Information
- Application Number
- CN202510137799.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-02
- Filing Date
- 2020-04-20
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-04-20
AI Technical Summary
The existing video encoding technology is difficult to effectively utilize historical motion information in inter-frame prediction, resulting in improper use of interpolation filter indexes, affecting the quality and decoding efficiency of video signals.
By constructing a history-based motion information candidate list, inheriting the half-pixel interpolation filter index, and replacing the default interpolation filter with appropriate interpolation filters in inter prediction, to improve the quality and decoding efficiency of the predicted signal.
It improves the quality and decoding efficiency of video signals, and improves the overall compression performance of video encoding and decoding methods.
Smart Images

Figure CN119766999B_ABST
Abstract
Description
[0001] This application is a divisional application. The application number of the original application is 202080029582.X, and the original application date is April 20, 2020. The entire contents of the original application are incorporated into this application by reference.
[0002] Related Applications Cross-Application
[0003] This application claims priority to U.S. Provisional Patent Application No. 62 / 836,072 filed on April 19, 2019, U.S. Provisional Patent Application No. 62 / 845,938 filed on May 10, 2019, U.S. Provisional Patent Application No. 62 / 909,761 filed on October 2, 2019, and U.S. Provisional Patent Application No. 62 / 909,763 filed on October 2, 2019. The entire contents of the above patent applications are incorporated into this application by reference. Technical Field
[0004] Embodiments of the present invention generally relate to the field of image processing, more specifically to inter-frame prediction, and in particular to a method and apparatus for deriving an interpolation filter index for a current block, such as a fusion process for switchable interpolation filter parameters. Background Art
[0005] Video coding (video encoding and decoding) is widely used in digital video applications such as broadcast digital TV, video transmission over the Internet and mobile networks, real-time conversation applications (such as video chat), video conferencing, DVD and Blu-ray Discs, video content acquisition and editing systems, and cameras for security applications.
[0006] Even in the case of shorter videos, a large amount of video data needs to be described, which can cause difficulties when the data is to be streamed or otherwise sent across a communication network with limited bandwidth capacity. Therefore, video data is often compressed before being sent across modern telecommunications networks. Since memory resources may be limited, the size of the video may also become an issue when storing the video on a storage device. Video compression devices typically use software and / or hardware on the source side to encode the video data before sending or storing, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received on the destination side by a video decompression device that decodes the video data. With limited network resources and a growing demand for higher video quality, there is a need for improved compression and decompression techniques that can increase the compression ratio with little impact on image quality.
[0007] Recently, a switchable interpolation filter for half-pixel positions has been introduced in Versatile Video Coding (VVC). The switching of the half-pixel luminance interpolation filter depends on the precision of the motion vector. In the case of using half-pixel motion vector precision, an alternative half-pixel interpolation filter can be used, and the alternative half-pixel interpolation filter can be represented by an additional syntax element indicating which interpolation filter to use, thus increasing the indication overhead. Summary of the Invention
[0008] The purpose of the embodiments of the present application is to provide an apparatus and method for constructing a history-based motion information candidate list, so that when using the history-based motion information candidate list, the half-pixel (half-pixel / half-pel) interpolation filter index can be inherited, so as to select a suitable interpolation filter to replace the default interpolation filter, and improve the quality of the predicted signal and the decoding efficiency.
[0009] The purpose of the embodiments of the present application is to provide an apparatus and method for encoding a current block in an inter prediction skip / fusion mode, so that when using a history-based motion information candidate list, the half-pixel interpolation filter index can be inherited, thereby improving the quality of the video signal.
[0010] The above and other purposes are achieved by the subject matter claimed in the independent claims. Other implementations are apparent in the dependent claims, the description, and the drawings.
[0011] According to a first aspect of the present invention, there is provided a method for constructing a history-based motion information (HMI) candidate list, which can be executed by an encoding device or a decoding device, and the method includes:
[0012] Obtaining a history-based motion information candidate list, where the HMI list is an ordered list of N history-based motion information candidates H k (k = 0,..., N - 1), and the N history-based motion information candidates H k are associated with a plurality of previous blocks (e.g., N previous blocks) before a block, N is an integer greater than 0 (e.g., N is an integer greater than 0 and less than or equal to a predefined value (such as 0 < N <= 5)), and each history-based motion information candidate includes the motion information of the corresponding previous block, and the motion information includes elements:
[0013] (i) one or more motion vectors MV corresponding to the previous block (e.g., luma motion vectors mvL0 and / or mvL1 with 1 / 16 fractional sample accuracy, where mvL0 and mvL1 correspond to reference picture list L0 and reference picture list L1),
[0014] (ii) one or more reference picture indexes corresponding to the MV of the corresponding previous block (e.g., reference picture indexes refIdxL0 and / or refIdxL1, refIdxL0 and refIdxL1 corresponding to reference picture list L0 and reference picture list L1),
[0015] (iii) an interpolation filter (IF) index (eg, an IF index of the corresponding previous block or an IF index associated with the corresponding previous block);
[0016] The HMI list is updated according to the motion information of the block, where the motion information of the block includes the elements:
[0017] (i) one or more motion vectors MV of the block (e.g. luma motion vectors mvL0 and / or mvL1 with 1 / 16 fractional sample accuracy),
[0018] (ii) one or more reference picture indices corresponding to the MV of the block (e.g., reference indices refIdxL0 and / or refIdxL1),
[0019] (iii) an interpolation filter index (eg, an IF index of the block or an IF index associated with the block).
[0020] In one example, the interpolation filter (IF) index may refer to a fractional sample interpolation filter (IF) index. Specifically, the IF index refers to a half-pixel (half-pixel / half-pel) interpolation filter index or a half-sample interpolation filter index (hpelIfIdx). In the present invention, the terms "half-pixel interpolation filter" and "half-sample interpolation filter" are interchangeable. The half-sample interpolation filter index represents a half-pixel interpolation filter used to interpolate half-pixel values when at least one of the motion vectors of the corresponding block points to a half-pixel position. For example, if one or more motion vectors (motion vector, MV) (element (i)) of the historical motion information candidate include at least one MV pointing to a half-pixel position, then the interpolation filter (IF) index (element (iii)) of the historical motion information candidate indicates a half-pixel interpolation filter used to interpolate half-pixel values (i.e., the interpolation filter index (element (iii)) is only valid for HMI candidates including half-pixel MVs). If one or more motion vectors (motion vector, MV) (element (i)) of the historical motion information candidate do not include an MV pointing to a half-pixel position, then the interpolation filter (IF) index (element (iii)) of the historical motion information candidate does not work (i.e., the IF index is invalid for non-half-pixel MVs, and the value of the IF index for non-half-pixel MVs does not work either. The interpolation filter index can be set to any value, such as 0 or false. In this case, the interpolation filter (IF) can be set to 0. The interpolation filter, IF) index (element (iii)) is assigned a default value, and the default value is no longer used in subsequent steps. The same is true for the motion information of the block. In other words, when at least one of the MVs of the block points to a half-pixel position, the interpolation filter index (element (iii)) in the motion information of the block takes effect. If all MVs of the block do not point to a half-pixel position, a default value is assigned to the interpolation filter index (element (iii)) in the motion information, and the default value is no longer used in subsequent steps. In an exemplary implementation, the IF index is always stored in the HMI list, regardless of the fractional part of the MV, even if the IF index does not work in some cases. In this way, the design of the HMI list can be simplified. It can be understood that if neither of the two MVs of the corresponding block points to a half-pixel position, the value assigned to the IF index will not have any effect on the decoding result.
[0021] In another exemplary implementation, the interpolation filter (IF) index can be replaced by an interpolation filter (IF) set index, and the IF set index represents a switchable IF set in multiple IF sets. In one example, each IF set includes an interpolation filter for each fractional position. At the same time, the IFs for the same fractional position can be equal in a few IF sets. For example, there are the same filters for some fractional positions and different filters for some fractional positions in multiple IF sets, and in particular, the IFs for the corresponding fractional positions can be switched according to the IF set index. In some cases, switching between two groups of interpolation filters can be understood as switching between two interpolation filters.
[0022] In an exemplary implementation, the half-pixel interpolation filter index represents a half-pixel interpolation filter in a set of half-pixel interpolation filters. The half-pixel interpolation filter is used to interpolate half-pixel values only when at least one of the one or more motion vectors points to a half-pixel position. If the motion vector in L0 and / or L1 points to a half-pixel (half-pixel / half-pel) position, an interpolation filter is selected according to the half-pixel interpolation filter index, and the interpolation filter is used for sample interpolation during motion compensation of the corresponding prediction list (prediction direction) (L0 and / or L1).
[0023] It should be noted that the block and the N previous blocks may be located in a slice of a frame or in a frame. In one example, when a new slice exists, the history-based motion information candidate list (table) is vacated. When a new slice exists, the construction process is called. In another example, the HMI list / table may be reset according to each new CTU row in the slice.
[0024] It is understandable that the N previous blocks may be one or more previous blocks. The previous block refers to a block that is encoded or decoded before the current block in the encoding or decoding order. In one example, block P may adopt the HMVP table of one or more encoding / decoding blocks before block P. After the motion information of block P is derived, the HMVP table is updated. After the HMVP table is updated, the block Q after block P may adopt the updated HMVP table. Block Q is encoded or decoded after block P in the decoding or encoding order.
[0025] It can be understood that after the HMVP list is updated, there may be M history-based motion information candidates in the updated HMVP list, where M is less than or equal to a predefined value (such as 5), and M>=N.
[0026] It can also be understood that if the index of the HMI list starts from 1, the HMI list is N candidate HMIs based on historical motion information. k An ordered list of (k=1, ..., N), wherein the N historical motion information candidates H k Associated with motion information of multiple previous blocks before a block.
[0027] Therefore, an improved method is provided, which allows the inheritance of the interpolation filter index in the history-based motion information candidate list. In particular, the interpolation filter (IF) index of the previous block is stored in the corresponding history-based motion information candidate of the history-based motion information candidate list. When the history-based motion information candidate list is directly or indirectly used for inter-frame prediction of a block encoded in a fusion or skip mode, the interpolation filter (IF) index can be borrowed from the corresponding motion information candidate without using a separate syntax element. The IF index is propagated through the history-based motion information candidate list, allowing the use of a suitable interpolation filter for the block (rather than using a predefined interpolation filter), thereby ensuring the quality of the encoded signal. Therefore, the technology provided in this article is conducive to improving decoding efficiency, thereby improving the overall compression performance of the video coding and decoding method.
[0028] It should be noted that the terms "block", "coding block" or "image block" used in the present invention may include transform units (TU), prediction units (PU), coding units (CU), etc. In versatile video coding (VVC), transform units and coding units are aligned in most cases, except for a few scenes using TU tiling or sub-block transform (SBT). It is understandable that the terms "block", "image block", "coding block" and "image block" are interchangeable in the present invention. The terms "sample" and "pixel" are also interchangeable in the present invention. The terms "predicted sample value" and "predicted pixel value" are interchangeable in the present invention. The terms "sample position" and "pixel position" are interchangeable in the present invention.
[0029] It should also be understood that the terms “history-based motion information candidate list”, “HMI list”, “HMVP list”, “HMVP table” and “HMVP look-up table (LUT)” may be interchangeable in the present invention.
[0030] It should be understood that the HMVP list is constructed using motion information of one or more encoded / decoded previous blocks. The HMVP list is used to store motion information of neighboring blocks (but not necessarily adjacent blocks like conventional spatial fusion candidates). The idea of HMVP is to use motion information of previous blocks that are spatially close to a block but not necessarily adjacent to the block (such as blocks in some spatial neighborhood).
[0031] According to the method of the first aspect, in a possible implementation manner, updating the HMI list includes: if at least one of the following elements of each history-based motion information candidate in the HMI list is different from the corresponding element in the motion information of the block, taking the motion information of the block as the history-based motion information candidate HMI list; k (k=N) is added to the HMI list, wherein the at least one element is:
[0032] (i) the one or more motion vectors MV,
[0033] (ii) the one or more reference picture indices corresponding to the MV.
[0034] It can be understood that if the index in the HMI list starts from 1, the adding may refer to adding the history-based motion information candidate Hk (k=N+1) containing the motion information of the block to the HMI list.
[0035] The motion information of the block is allowed to be added as a history-based motion information candidate to the last position of the HMI list.
[0036] According to the method of the first aspect, in a possible implementation manner, updating the HMI list includes:
[0037] If the comparison result of the following elements of the history-based motion information candidate in the HMI list and the corresponding elements in the motion information of the block is the same, the history-based motion information candidate is deleted from the HMI list, and the motion information of the block is used as the history-based motion information candidate HMI list: k (k=N–1) is added to the HMI list, where the element is:
[0038] (i) one or more motion vectors MV,
[0039] (ii) one or more reference picture indexes corresponding to the MV.
[0040] It can be understood that if the index in the HMI list starts from 1, the adding may refer to taking the motion information of the block as the history-based motion information candidate HMI. k (k=N) is added to the HMI list.
[0041] The motion information of the block may be added as a history-based motion information candidate to the last position of the HMI list.
[0042] According to the method in any one of the first aspect or the implementation manner of the first aspect, in a possible implementation manner, updating the HMI list includes:
[0043] If N is equal to a predefined value, the history-based motion information candidate HMI is deleted from the HMI list. k (k=0), and the motion information of the block is used as the history-based motion information candidate H k (k=N–1) is added to the HMI list.
[0044] It can be understood that if the index in the HMI list starts from 1, the deletion may refer to deleting the history-based motion information candidate HMI from the HMI list. k (k=1), the adding means taking the motion information of the block as the motion information candidate H based on history k (k=N) is added to the HMI list.
[0045] It is allowed to delete the history-based motion information candidate located at the first position in the HMI list, and add the motion information of the block as the history-based motion information candidate to the last position of the HMI list.
[0046] According to the method in any one of the first aspect or the implementation manner of the first aspect, in a possible implementation manner, the method further includes:
[0047] comparing whether a motion vector of any history-based motion information candidate is the same as a corresponding motion vector of the block;
[0048] A comparison is made as to whether the reference picture index of any history-based motion information candidate is the same as the corresponding reference picture index of the block.
[0049] In an alternative design, the method further comprises:
[0050] comparing whether at least one of the motion vectors of each history-based motion information candidate (ie, HMVP candidate) is different from a corresponding motion vector of the block;
[0051] At least one of the reference picture indexes of each HMVP candidate is compared to see if it is different from a corresponding reference picture index of the block.
[0052] Therefore, it is allowed to use only MV and reference image index in the pruning process without comparing the interpolation filter index when updating the HMVP table. Thus, a better trade-off can be achieved between the complexity and diversity of HMVP candidates. In particular, it is allowed to compare only based on MV and reference image index, which can avoid additional calculation operations, thereby reducing the calculation complexity. Each comparison operation will generate additional calculations in the update process of the HMVP table and the construction process of the fusion candidate. Therefore, if the comparison operation can be reduced or eliminated, the calculation complexity can be reduced, thereby improving the decoding efficiency. In addition, it is allowed to compare only based on MV and reference image index, and the diversity of HMVP records can also be maintained. If two HMVP records contain the same MV and the same reference index, and only contain different IF indexes, they are invalid because the two records are not completely different. At this time, in the update process of the HMVP table, it can be reasonably considered that the two HMVP records are the same. At this time, if a new record has only a different IF index compared with the existing records in the HMVP table, the new record will not be added to the HMVP table. Therefore, "old" / existing records that are "completely different" from other records (having different MVs or reference indexes) will be retained. In other words, if a new record is to be added to the HMVP table, the new record should not only be bit-wise different from the existing record, but must also be "substantially different". From the perspective of decoding efficiency, it is more efficient for the HMVP table to contain two records with different MVs or different reference indexes than to contain two records with only different IF indexes.
[0053] According to the method in any one of the first aspect or the implementation manners of the first aspect, in a possible implementation manner, the predefined value is 5 or 6.
[0054] According to the method described in any one of the first aspect or the implementation manner of the first aspect, in a possible implementation manner, the half-sample interpolation filter index included in the history-based motion information candidate represents a half-sample interpolation filter in a half-sample interpolation filter set; the half-sample interpolation filter is applied to interpolate half-sample values only when at least one MV among the one or more MVs of the history-based motion information candidate points to a half-sample position.
[0055] In the prior art, a default interpolation filter (IF) index (corresponding to a default interpolation filter) is always used for the fusion candidate obtained from the HMVP table. In the present invention, the IF index is propagated through the HMVP table, so one interpolation filter in a set of interpolation filters can be used according to the IF index. In one example, one of two interpolation filters (a default interpolation filter and an alternative interpolation filter) can also be used according to the IF index. Therefore, a suitable interpolation filter is selected instead of using a default interpolation filter to improve reference reliability, and improve the quality of the prediction signal and decoding efficiency.
[0056] It should be noted that the terms “alternative half-pixel interpolation filter”, “switchable interpolation filter (SIF)” or “half-pixel interpolation filter” are interchangeable in the present invention.
[0057] A suitable interpolation filter (IF) can be selected according to the content. For areas with sharp edges, a conventional DCT-based interpolation filter can be used. For smooth areas (or areas where sharp edges do not need to be retained), an alternative 6-tap interpolation filter (Gaussian filter) can be used. For fusion mode, the IF index can be borrowed from the corresponding motion information candidate. For blocks encoded in fusion mode, an alternative interpolation filter can be used when the motion information candidate is obtained from the HMVP table. Propagating the IF index through the HMVP table allows the use of a suitable interpolation filter for the block, which is beneficial to improving decoding efficiency. If the proposed mechanism is not adopted, the default IF index (corresponding to an 8-tap DCT-based interpolation filter) is usually used for the HMVP fusion candidate, and the specific content of the current block (whether sharp edges need to be retained) will not be considered.
[0058] According to a second aspect of the present invention, there is provided a method for inter-predicting a block of a frame of a video signal, the method comprising:
[0059] Construct a history-based motion information (HMI) candidate list, wherein the HMI list is N history-based motion information candidates H k An ordered list of (k=0, ..., N-1) of N history-based motion information candidates H kis associated with (or includes) motion information of a plurality of previous blocks (eg, N previous blocks) before the block, where N is an integer greater than 0, each history-based motion information candidate corresponds to one previous block, and includes elements:
[0060] (i) one or more motion vectors of the previous block,
[0061] (ii) one or more reference picture indices corresponding to the MV of the previous block,
[0062] (iii) an interpolation filter index (eg, an interpolation filter index of the previous block or an interpolation filter index associated with the previous block);
[0063] adding one or more history-based motion information candidates in the HMI list to a motion information candidate list for the block;
[0064] The motion information of the block is derived according to the motion information candidate list.
[0065] The motion information candidate list may be a fusion candidate list.
[0066] In an alternative or additional design, a second aspect of the present invention provides a method for inter-frame prediction of a block of a video signal frame, the method comprising:
[0067] Construct a history-based motion information candidate list, wherein the HMI list is N history-based motion information candidate H k An ordered list of (k=0, ..., N-1) of N history-based motion information candidates H k Being associated with (or including) motion information of a plurality of previous blocks (e.g., N previous blocks) before the block, where N is an integer greater than 0, at least one history-based motion information candidate includes an element corresponding to the previous block, including:
[0068] (i) one or more motion vectors (MVs), wherein at least one MV points to a half-pixel position;
[0069] (ii) one or more reference picture indices corresponding to the one or more MVs,
[0070] (iii) an interpolation filter index of the previous block;
[0071] adding one or more history-based motion information candidates in the HMI list to a motion information candidate list for the block;
[0072] The motion information of the block is derived according to the motion information candidate list.
[0073] The motion information candidate list may be a fusion candidate list.
[0074] It can be understood that the history-based motion information candidates are added to the fusion candidate list as history-based fusion candidates.
[0075] In one example, the length of the HMI list is N, and N is 5 or 6.
[0076] Therefore, an improved method is provided, which allows the inheritance of the interpolation filter index in the history-based motion information candidate list. In particular, the interpolation filter (IF) index of the previous block is stored in the corresponding history-based motion information candidate of the history-based motion information candidate list. When the history-based motion information candidate list is directly or indirectly used for inter-frame prediction of a block encoded in a fusion or skip mode, the interpolation filter (IF) index can be borrowed from the corresponding motion information candidate without using a separate syntax element. The IF index is propagated through the history-based motion information candidate list, allowing the use of a suitable interpolation filter for the block (rather than using a predefined interpolation filter), thereby ensuring the quality of the encoded signal. Therefore, the technology provided in this article is conducive to improving decoding efficiency, thereby improving the overall compression performance of the video coding and decoding method.
[0077] According to the method described in the second aspect, in a possible implementation manner, a half-sample interpolation filter is applied only when at least one MV of one or more MVs of the derived motion information points to a half-sample position, and the half-sample interpolation filter is represented by a half-sample interpolation filter index contained in the derived motion information.
[0078] According to the method described in any one of the second aspect or the implementation manner of the second aspect, in one possible implementation manner, the half-sample interpolation filter index included in the history-based motion information candidate represents a half-sample interpolation filter in a half-sample interpolation filter set; the half-sample interpolation filter is applied to interpolate half-sample values only when at least one MV among the one or more MVs of the history-based motion information candidate points to a half-sample position.
[0079] According to the method described in any one of the second aspect or the implementation manner of the second aspect, in a possible implementation manner, the history-based motion information candidate further includes one or more bidirectional prediction weight indexes. The term "bidirectional prediction weight index bcw_idx" is also referred to as a generalized bidirectional prediction weight index GBIdx and / or a CU-level bidirectional prediction weight (BCW) index. Alternatively, the index may be referred to as BWI, i.e., a bidirectional prediction weight index.
[0080] According to the method in any one of the second aspect or the implementation manner of the second aspect, in a possible implementation manner, the method further includes:
[0081] If at least one of the following elements of each history-based motion information candidate in the HMI list is different from the corresponding element in the motion information of the block, the motion information of the block is taken as the history-based motion information candidate HMI: k (k=N) is added to the HMI list, wherein the at least one element is:
[0082] (i) one or more motion vectors MV,
[0083] (ii) one or more reference picture indexes corresponding to the MV.
[0084] According to the method in any one of the second aspect or the implementation manner of the second aspect, in a possible implementation manner, the method further includes:
[0085] If the following elements of the history-based motion information candidate in the HMI list are the same as the corresponding elements in the motion information of the block, the history-based motion information candidate is deleted from the HMI list, and the motion information of the block is used as the history-based motion information candidate HMI list. k (k=N–1) is added to the HMI list, where the element is:
[0086] (i) one or more motion vectors MV,
[0087] (ii) one or more reference picture indexes corresponding to the MV.
[0088] According to the method in any one of the second aspect or the implementation manner of the second aspect, in a possible implementation manner, the method further includes:
[0089] If N is equal to a predefined value, the history-based motion information candidate HMI is deleted from the HMI list. k (k=0), and the motion information of the block is used as the history-based motion information candidate H k(k=N-1) is added to the HMI list. In one example, the predefined value is 5.
[0090] According to the method in any one of the second aspect or the implementation manner of the second aspect, in a possible implementation manner, the method further includes:
[0091] comparing whether a corresponding motion vector of any history-based motion information candidate is the same as the motion vector of the block;
[0092] A comparison is made as to whether the corresponding reference picture index of any history-based motion information candidate is the same as the reference picture index of the block.
[0093] In an alternative design, the comparison includes:
[0094] comparing whether at least one of the motion vectors of each history-based motion information candidate is different from a corresponding motion vector of the block;
[0095] At least one of the reference picture indexes of each HMVP candidate is compared to see if it is different from a corresponding reference picture index of the block.
[0096] According to the method of any one of the second aspect or the implementation manner of the second aspect, in a possible implementation manner, the motion information candidate list is used for a merge mode or a skip mode. In other words, the current block is encoded using a merge mode or a skip mode.
[0097] According to the method in any one of the second aspect or the implementation manner of the second aspect, in a possible implementation manner, deriving the motion information of the block according to the motion information candidate list includes:
[0098] The motion information indicated by the candidate index is derived from the motion information candidate list as the motion information of the current block, wherein the candidate index is parsed or derived from a bitstream.
[0099] According to the method in any one of the second aspect or the implementation manner of the second aspect, in a possible implementation manner, the method further includes:
[0100] When at least one MV of the one or more motion vectors MV included in the derived motion information points to a half-pixel position, obtaining a predicted sample value of the block by applying a half-pixel interpolation filter to a pixel value of the reference image pointed to by the MV, wherein the half-pixel interpolation filter is represented by an interpolation filter index included in the derived motion information;
[0101] When none of the motion vectors MV included in the derived motion information points to a half-pixel position, the predicted sample value of the block is obtained by applying a default interpolation filter to the pixel value of the reference image pointed to by the MV.
[0102] The encoding method and the decoding method defined in the claims, the specification and the drawings can be executed by an encoding device and a decoding device, respectively.
[0103] According to a third aspect of the present invention, there is provided an apparatus for constructing a history-based motion information candidate list, the apparatus comprising:
[0104] A history-based motion information candidate list obtaining unit is used to obtain a history-based motion information candidate list, wherein the HMI list is N history-based motion information candidate HMIs. k An ordered list of (k=0, ..., N-1) of N history-based motion information candidates H k Associated with motion information of multiple blocks before a block, N is an integer greater than 0, and each history-based motion information candidate includes the elements:
[0105] (i) one or more motion vectors MV,
[0106] (ii) one or more reference image indices corresponding to the MV,
[0107] (iii) interpolation filter index;
[0108] A history-based motion information candidate list updating unit is configured to update the HMI list according to the motion information of the block, wherein the motion information of the block includes the following elements:
[0109] (i) one or more motion vectors MV,
[0110] (ii) one or more reference image indices corresponding to the MV,
[0111] (iii) Interpolation filter index.
[0112] The method provided in the first aspect of the present invention may be executed by the apparatus provided in the third aspect of the present invention. More features and implementations of the apparatus provided in the third aspect of the present invention correspond to the features and implementations of the method provided in the first aspect of the present invention.
[0113] According to a fourth aspect of the present invention, there is provided a device for inter-frame prediction block, the device comprising
[0114] A list management unit is used to construct a history-based motion information candidate list, wherein the HMI list is N history-based motion information candidate HMIs. kAn ordered list of (k=0, ..., N-1) of N history-based motion information candidates H k Associated with motion information of multiple blocks before the block, N is an integer greater than 0, and each history-based motion information candidate includes the elements:
[0115] (i) one or more motion vectors MV,
[0116] (ii) one or more reference image indices corresponding to the MV,
[0117] (iii) interpolation filter index;
[0118] The list management unit is further configured to add one or more history-based motion information candidates in the HMI list to the motion information candidate list of the block;
[0119] The motion information derivation unit is used to derive the motion information of the block according to the motion information candidate list.
[0120] The method provided in the second aspect of the present invention may be executed by the apparatus provided in the fourth aspect of the present invention. More features and implementations of the apparatus provided in the fourth aspect of the present invention correspond to the features and implementations of the method provided in the second aspect of the present invention.
[0121] According to a fifth aspect of the present invention, there is provided an encoder (20), comprising a processing circuit for executing the method described in the first aspect or the second aspect and the implementation manner of the first aspect or the second aspect.
[0122] According to a sixth aspect of the present invention, a decoder (30) is provided, the decoder comprising a processing circuit for executing the method described in the first aspect or the second aspect and the implementation manner of the first aspect or the second aspect.
[0123] According to a seventh aspect of the present invention, a decoder is provided. The decoder comprises:
[0124] one or more processors;
[0125] A non-transitory computer-readable storage medium, which is coupled to the processor and stores a program executed by the processor, and when the processor executes the program, causes the decoder to execute the method described in the first aspect or the second aspect and the implementation method of the first aspect or the second aspect.
[0126] According to an eighth aspect of the present invention, there is provided an encoder. The encoder comprises:
[0127] one or more processors;
[0128] A non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium is coupled to the processor and stores a program executed by the processor, and when the processor executes the program, the encoder executes the method described in the first aspect or the second aspect and the implementation method of the first aspect or the second aspect.
[0129] According to a ninth aspect of the present invention, there is provided a non-transitory storage medium, the non-transitory storage medium comprising a code stream encoded / decoded using the method described in any one of the above aspects.
[0130] A device for encoding or decoding a video stream may include a processor and a memory, wherein the memory stores instructions so that the processor executes any one of the methods described above.
[0131] For each encoding method or decoding method disclosed herein, a computer-readable storage medium is provided, wherein the storage medium includes instructions stored thereon, and when the instructions are executed, one or more processors are caused to encode or decode video data. The instructions cause the one or more processors to perform any of the methods described above.
[0132] In addition, for each encoding method or decoding method disclosed herein, a computer program product is provided, wherein the computer program product comprises program code for executing any one of the methods described above.
[0133] The following drawings and description set forth one or more embodiments in detail. Other features, objects, and advantages are apparent in the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0134] The embodiments of the present invention are described in more detail below with reference to the accompanying drawings and schematic diagrams, in which:
[0135] Figure 1A is a block diagram of an exemplary video decoding system for implementing an embodiment of the present invention;
[0136] Figure 1B is a block diagram of another exemplary video decoding system for implementing an embodiment of the present invention;
[0137] Figure 2 is a block diagram of an exemplary video encoder for implementing an embodiment of the present invention;
[0138] Figure 3 is a block diagram of an exemplary structure of a video decoder for implementing an embodiment of the present invention;
[0139] Figure 4 is a block diagram of an exemplary encoding device or decoding device;
[0140] Figure 5 is a block diagram of another exemplary encoding device or decoding device;
[0141] Figure 6 An example of a current block and spatially neighboring blocks of the current block is schematically shown;
[0142] Figure 7 The current block and the upper adjacent block are schematically shown;
[0143] Figure 8 A flowchart of a method provided by an embodiment of the present invention;
[0144] Fig. 9 A block diagram of a method for deriving an interpolation filter index for a current block (eg, coding unit or coding block) within a coding tree block (CTB) or coding tree unit (CTU);
[0145] Fig.10 An example of constructing an HMVP list provided by an embodiment of the present invention is schematically shown;
[0146] Fig.11 Another example of constructing an HMVP list provided by an embodiment of the present invention is schematically shown;
[0147] Fig.12 An exemplary HMVP list and its traversal order provided by an embodiment of the present invention are shown;
[0148] Fig.13A is a flowchart of an example method for constructing an HMVP list;
[0149] Fig. 13B is a flowchart of another example of a method for constructing an HMVP list;
[0150] Fig.14 is a flowchart of an example of a method for inter-frame prediction of blocks of a video signal frame;
[0151] Fig.15 A flowchart of an example of an HMI list update method;
[0152] Fig.16 A block diagram of a device provided by an embodiment of the present invention;
[0153] Fig.17 A block diagram of another device provided by an embodiment of the present invention;
[0154] Fig.18 A block diagram of an exemplary structure of a content providing system for implementing a content distribution service;
[0155] Fig.19 A block diagram of an exemplary structure of a terminal device.
[0156] In the following, identical reference numerals denote identical features or at least functionally equivalent features, unless expressly stated otherwise. DETAILED DESCRIPTION
[0157] In the following description, reference is made to the accompanying drawings that form a part of the present invention, which illustrate by way of illustration specific aspects of embodiments of the present invention or specific aspects in which embodiments of the present invention may be used. It should be understood that embodiments of the present invention may be used in other aspects and may include structural changes or logical changes not described in the accompanying drawings. Therefore, the following detailed description should not be understood in a restrictive sense, and the scope of the present invention is defined by the appended claims.
[0158] For example, it should be understood that the disclosure in conjunction with the described method may also be applicable to the corresponding device or system for performing the method, and vice versa. For example, if one or more specific method steps are described, the corresponding device may include one or more units such as functional units to perform the one or more method steps described (e.g., one unit performs one or more steps, or each of the multiple units performs one or more steps of the multiple steps), even if the one or more units are not explicitly described or illustrated in the drawings. On the other hand, for example, if a specific device is described based on one or more units such as functional units, the corresponding method may include a step to perform the function of one or more units (e.g., one step performs the function of one or more units, or each of the multiple steps performs the function of one or more units of the multiple units), even if the one or more steps are not explicitly described or illustrated in the drawings. Further, it should be understood that, unless otherwise explicitly stated, the features of the various exemplary embodiments and / or aspects described herein may be combined with each other.
[0159] Video decoding generally refers to processing a sequence of images that form a video or video sequence. In the field of video decoding, the terms "frame" and "picture / image" can be used as synonyms. Video decoding (or generally referred to as decoding) includes two parts: video encoding and video decoding. Video encoding is performed at the source end and generally includes processing (e.g., compressing) the original video image to reduce the amount of data required to represent the video image (for more efficient storage and / or transmission). Video decoding is performed at the destination end and generally includes an inverse process relative to the encoder to reconstruct the video image. The "decoding" of the video images (or generally referred to as images) involved in the embodiments should be understood as the "encoding" or "decoding" of the video images or respective video sequences. The encoding part and the decoding part are also collectively referred to as codecs (encoding and decoding, codec).
[0160] In the case of lossless video decoding, the original video image can be reconstructed, that is, the reconstructed video image has the same quality as the original video image (assuming no transmission loss or other data loss during storage or transmission). In the case of lossy video decoding, further compression is performed through quantization, etc. to reduce the amount of data required to represent the video image. At this time, the decoder side cannot fully reconstruct the video image, that is, the quality of the reconstructed video image is lower or inferior to the quality of the original video image.
[0161] Several video coding standards belong to the group of "lossy hybrid video codecs" (i.e., combining spatial and temporal prediction in the sample domain with 2D transform coding for quantization in the transform domain). Each picture in a video sequence is usually segmented into a set of non-overlapping blocks, and decoding is usually performed at the block level. In other words, the encoder side usually processes, i.e., encodes the video at the block (video block) level, for example, by generating a prediction block through spatial (intra-frame) prediction and / or temporal (inter-frame) prediction; subtracting the prediction block from the current block (the block currently being processed / to be processed) to obtain a residual block; transforming and quantizing the residual block in the transform domain to reduce the amount of data to be transmitted (compressed), while the decoder side applies the inverse process relative to the encoder to the coded or compressed block to reconstruct the current block for representation. In addition, the processing steps of the encoder and the decoder are the same, so the encoder and the decoder generate the same predictions (e.g., intra-frame predictions and inter-frame predictions) and / or reconstructions for processing, i.e., decoding of subsequent blocks.
[0162] In the following embodiment of the video decoding system 10, the video encoder 20 and the video decoder 30 are combined Figures 1A to 3 Give a description.
[0163] Figure 1A1 is a schematic block diagram of an exemplary decoding system 10, for example, a video decoding system 10 (or simply decoding system 10) that can utilize the techniques of the present application. The video encoder 20 (or simply encoder 20) and the video decoder 30 (or simply decoder 30) in the video decoding system 10 represent examples of devices that can be used to perform various techniques according to various examples described in the present application.
[0164] like Figure 1A As shown, the decoding system 10 includes a source device 12 , which is used to provide encoded image data 21 , for example, to a destination device 14 ; the destination device 14 is used to decode the encoded image data 21 .
[0165] The source device 12 includes an encoder 20 and, in addition or optionally, an image source 16 , a pre-processor (or pre-processing unit) 18 , such as an image pre-processor 18 , and a communication interface or communication unit 22 .
[0166] The image source 16 may include or may be any type of image capture device, such as a camera for capturing real-world images, and / or any type of image generation device or any type of other device, wherein the image capture device is, for example, a camera for capturing real-world images, the image generation device is, for example, a computer graphics processor for generating computer-animated images, and other devices are used to obtain and / or provide real-world images, computer-generated images (e.g., screen content, virtual reality (VR) images), and / or any combination thereof (e.g., augmented reality (AR) images). The image source may be any type of memory (memory / storage) for storing any of the above images.
[0167] In order to distinguish the pre-processor 18 and the processing performed by the pre-processing unit 18 , the image or image data 17 may also be referred to as a raw image or raw image data 17 .
[0168] The preprocessor 18 is used to receive (raw) image data 17 and preprocess the image data 17 to obtain a preprocessed image 19 or preprocessed image data 19. The preprocessing performed by the preprocessor 18 may include, for example, trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or denoising. It is understood that the preprocessing unit 18 may be an optional component.
[0169] The video encoder 20 is used to receive the pre-processed image data 19 and provide the encoded image data 21 (hereinafter referred to as Figure 2 etc. for detailed description).
[0170] The communication interface 22 in the source device 12 may be used to receive the encoded image data 21 and send the encoded image data 21 (or any other processed version) to the destination device 14 or any other device via the communication channel 13 for storage or direct reconstruction.
[0171] The destination device 14 includes a decoder 30 (eg, a video decoder 30 ) and, additionally and optionally, a communication interface or communication unit 28 , a post-processor 32 (or a post-processing unit 32 ), and a display device 34 .
[0172] The communication interface 28 in the destination device 14 is used to receive the encoded image data 21 (or any other processed version) directly from the source device 12 or from any other source device such as a storage device, for example, a storage device that stores the encoded image data, and provide the encoded image data 21 to the decoder 30.
[0173] The communication interface 22 and the communication interface 28 can be used to send or receive the encoded image data 21 or the encoded data 13 through a direct communication link between the source device 12 and the destination device 14, such as a direct wired or wireless connection, etc., or through any type of network, such as a wired network, a wireless network or any combination thereof, any type of private network and public network or any type of combination thereof.
[0174] For example, the communication interface 22 may be used to encapsulate the encoded image data 21 into a suitable format such as a message, and / or process the encoded image data using any type of transmission coding or processing for transmission over a communication link or network.
[0175] The communication interface 28 corresponds to the communication interface 22 , and may be configured to receive transmission data, and process the transmission data using any type of corresponding transmission decoding or processing and / or decapsulation process to obtain the encoded image data 21 .
[0176] The communication interface 22 and the communication interface 28 can be configured as follows Figure 1A The unidirectional communication interface indicated by the arrow corresponding to the communication channel 13 pointing from the source device 12 to the destination device 14, or configured as a bidirectional communication interface, can be used to send and receive messages, etc. to establish a connection, confirm and exchange any other information related to the communication link and / or data transmission such as encoded image data transmission, etc.
[0177] The decoder 30 is used to receive the encoded image data 21 and provide decoded image data 31 or decoded image 31 (hereinafter, based on, for example, Figure 3 or Figure 5 for detailed description).
[0178] The post-processor 32 in the destination device 14 is used to post-process the decoded image data 31 (also referred to as reconstructed image data), for example, the decoded image 31, to obtain post-processed image data 33, for example, the post-processed image 33. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (for example, from YCbCr format to RGB format), color correction, cropping or resampling, or any other processing, for example, for generating the decoded image data 31 for display by the display device 34, etc.
[0179] The display device 34 in the destination device 14 is used to receive the post-processed image data 33 to display the image to a user or viewer, etc. The display device 34 can be or include any type of display for presenting the reconstructed image, such as an integrated or external display screen or display. For example, the display can include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display.
[0180] although Figure 1A The source device 12 and the destination device 14 are shown as independent devices, but the device embodiment may also include the source device 12 and the destination device 14 at the same time or include the functions of the source device 12 and the destination device 14 at the same time, that is, include the source device 12 or the corresponding functions of the source device 12 and the destination device 14 or the corresponding functions of the destination device 14 at the same time. In these embodiments, the source device 12 or the corresponding functions of the source device 12 and the destination device 14 or the corresponding functions of the destination device 14 may be implemented using the same hardware and / or software or by separate hardware and / or software or any combination thereof.
[0181] According to the description, Figure 1A It will be apparent to the skilled person that the presence and (exact) division of different units or functions in the source device 12 and / or the destination device 14 shown may vary depending on the actual device and application.
[0182] The encoder 20 (eg, video encoder 20) or the decoder 30 (eg, video decoder 30), or both the encoder 20 and the decoder 30 may be configured as follows: Figure 1BThe processing circuits shown may be implemented by one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, video decoding dedicated processors, or any combination thereof. The encoder 20 may be implemented by processing circuits 46 to include reference Figure 2 The various modules discussed in connection with the encoder 20 shown and / or any other encoder system or subsystem described herein. The decoder 30 may be implemented by processing circuitry 46 to include reference Figure 3 The various modules discussed in the decoder 30 shown and / or any other decoder system or subsystem described herein. The processing circuits may be used to perform the various operations described below. Figure 5 As shown, if part of the technology is implemented by software, the device can store the instructions of the software in a suitable non-transitory computer-readable storage medium, and execute the instructions in hardware through one or more processors, thereby performing the technology of the present invention. Figure 1B As shown, either of video encoder 20 and video decoder 30 may be integrated into a single device as part of a combined encoder / decoder (CODEC).
[0183] Source device 12 and destination device 14 may include any of a variety of devices, including any type of handheld or fixed device, such as a notebook or laptop computer, a mobile phone, a smart phone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (such as a content service server or a content distribution server), a broadcast receiving device, a broadcast transmitting device, etc., and may or may not use any type of operating system. In some cases, source device 12 and destination device 14 may be equipped with components for wireless communication. Therefore, source device 12 and destination device 14 may be wireless communication devices.
[0184] In some cases, Figure 1AThe video decoding system 10 shown is exemplary only, and the techniques of the present application may be applicable to video decoding settings (e.g., video encoding or video decoding) that do not necessarily include any data communication between an encoding device and a decoding device. In other examples, data is retrieved from a local memory, streamed over a network, and the like. A video encoding device may encode data and store the data in a memory, and / or a video decoding device may retrieve data from a memory and decode the data. In some examples, encoding and decoding are performed by devices that do not communicate with each other but only encode data to a memory and / or retrieve data from a memory and decode data.
[0185] For ease of description, for example, the embodiments of the present invention are described with reference to the High-Efficiency Video Coding (HEVC), Versatile video coding (VVC) reference software, and the next generation video coding standard developed by the Joint Collaboration Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG). A person of ordinary skill in the art should understand that the embodiments of the present invention are not limited to the HEVC or VVC standards.
[0186] Encoders and encoding methods
[0187] Figure 2 FIG. 2 is a schematic block diagram of an exemplary video encoder 20 for implementing the technology of the present application. Figure 2 In the example of , the video encoder 20 includes an input terminal 201 (or an input interface 201), a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a loop filter 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy coding unit 270, and an output terminal 272 (or an output interface 272). The mode selection unit 260 may include an inter-frame prediction unit 244, an intra-frame prediction unit 254, and a segmentation unit 262. The inter-frame prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). Figure 2 The illustrated video encoder 20 may also be referred to as a hybrid video encoder or a hybrid video codec-based video encoder.
[0188] The residual calculation unit 204, the transform processing unit 206, the quantization unit 208 and the mode selection unit 260 constitute the forward signal path of the encoder 20; the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the inter-frame prediction unit 244 and the intra-frame prediction unit 254 constitute the backward signal path of the video encoder 20. Among them, the backward signal path of the video encoder 20 corresponds to the decoder (see Figure 3 The inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter-frame prediction unit 244 and the intra-frame prediction unit 254 also constitute the “built-in decoder” of the video encoder 20.
[0189] Images and Image Segmentation (Images and Blocks)
[0190] The encoder 20 may be configured to receive an image 17 (or image data 17) via an input 201 or the like, for example, an image in an image sequence forming a video or a video sequence. The received image or image data may also be a preprocessed image 19 (or preprocessed image data 19). For simplicity, it is referred to as image 17 in the following description. Image 17 may also be referred to as a current image or an image to be decoded (particularly in video decoding, in order to distinguish the current image from other images, other images may be, for example, previously encoded images and / or previously decoded images in the same video sequence, i.e., a video sequence simultaneously including the current image).
[0191] A (digital) image is or can be considered to be a two-dimensional array or matrix of samples with intensity values. The samples in the array can also be called pixels (pixel / pel) (short for picture elements). The number of samples in the array or image in the horizontal and vertical directions (or axes) determines the size and / or resolution of the image. In order to represent color, three color components are usually used, that is, the image can be represented as or can include three sample arrays. In the RGB format or color space, the image includes corresponding red, green, and blue sample arrays. However, in video decoding, each pixel is usually represented in a brightness and chrominance format or color space, for example, the YCbCr format, which includes a brightness component represented by Y (sometimes also represented by L) and two chrominance components represented by Cb and Cr. The brightness (or luma for short) component Y represents the brightness or grayscale intensity (for example, in a grayscale image), while the two chrominance (or chroma for short) components Cb and Cr represent the chrominance or color information components. Accordingly, an image in YCbCr format includes a luma sample array of luma sample values (Y) and two chroma sample arrays of chroma values (Cb and Cr). An image in RGB format can be converted or transformed into YCbCr format and vice versa. This process is also called a color conversion or conversion process. If the image is black and white, the image may include only a luma sample array. Accordingly, the image may be, for example, a luma sample array in black and white format or a luma sample array and two corresponding chroma sample arrays in 4:2:0, 4:2:2 and 4:4:4 color formats.
[0192] In an embodiment of the video encoder 20, the video encoder 20 may include an image segmentation unit ( Figure 2 ), for dividing the image 17 into a plurality of (usually non-overlapping) image blocks 203. These blocks may also be referred to as root blocks or macroblocks (H.264 / AVC standard) or as coding tree blocks (CTB) or coding tree units (CTU) (H.265 / HEVC and VVC standards). The image segmentation unit may be used to: use the same block size for all images in a video sequence and define the block size using a corresponding grid, or to change the block size between images or image subsets or groups of images and divide each image into corresponding blocks.
[0193] In other embodiments, the video encoder may be configured to directly receive the image block 203 of the image 17, for example, one, several or all image blocks constituting the image 17. The image block 203 may also be referred to as a current image block or an image block to be decoded.
[0194] Like image 17, image block 203 is or can be considered as a two-dimensional array or matrix composed of samples with intensity values (sample values), but the size of image block 203 is smaller than that of image 17. In other words, image block 203 may include one sample array (for example, when image 17 is a black and white image, image block 203 includes a brightness array; when image 17 is a color image, image block 203 includes a brightness array or a chrominance array), or three sample arrays (for example, when image 17 is a color image, image block 203 includes a brightness array and two chrominance arrays), or any other number and / or type of arrays determined by the color format used. The number of samples of image block 203 in the horizontal direction and the vertical direction (or axis) defines the size of image block 203. Accordingly, a certain image block may be, for example, an M×N (M columns×N rows) sample array, or an M×N transform coefficient array, etc.
[0195] exist Figure 2 In the illustrated embodiment of the video encoder 20 , the video encoder 20 may be configured to encode the image 17 block by block, for example, encoding and predicting each image block 203 .
[0196] Residual calculation
[0197] The residual calculation unit 204 can be used to calculate the residual block 205 (also referred to as residual 205) according to the image block 203 and the prediction block 265 (the prediction block 265 will be described in detail below) in the following manner, for example, by subtracting the sample value of the prediction block 265 from the sample value of the image block 203 sample by sample (pixel by pixel) to obtain the residual block 205 in the sample domain.
[0198] Transform
[0199] The transform processing unit 206 may be used to perform a transform such as discrete cosine transform (DCT) or discrete sine transform (DST) on the sample values of the residual block 205 to obtain a transform coefficient 207 in the transform domain. The transform coefficient 207 may also be referred to as a transform residual coefficient, representing the residual block 205 in the transform domain.
[0200] The transform processing unit 206 can be used to perform integer approximation of DCT / DST, for example, for transforms specified by H.265 / HEVC. Compared with the orthogonal DCT transform, the integer approximation is usually scaled based on a certain factor. Other scaling factors are used as part of the transform process to maintain the norm of the residual block processed by the forward transform and the inverse transform. The scaling factor is usually selected based on certain constraints, such as whether the scaling factor is a power of 2 for the shift operation, the bit depth of the transform coefficients, and the trade-off between accuracy and implementation cost. For example, a specific scaling factor is specified for the inverse transform on the encoder 20 side through the inverse transform processing unit 212, etc. (and for the corresponding inverse transform on the video decoder 30 side through the inverse transform processing unit 312, etc.), and accordingly, the corresponding scaling factor can be specified for the forward transform on the encoder 20 side through the transform processing unit 206, etc.
[0201] In an embodiment of the video encoder 20, the video encoder 20 (correspondingly, the transform processing unit 206) can be used to, for example, directly output or output transform parameters of one or more transform types after being encoded or compressed by the entropy coding unit 270, so that the video decoder 30 can receive and use the transform parameters for decoding.
[0202] Quantification
[0203] The quantization unit 208 is configured to quantize the transform coefficient 207 by performing scalar quantization or vector quantization, etc., to obtain a quantized coefficient 209. The quantized coefficient 209 may also be referred to as a quantized transform coefficient 209 or a quantized residual coefficient 209.
[0204] The quantization process may reduce the bit depth associated with some or all of the transform coefficients 207. For example, an n-bit transform coefficient may be rounded down to an m-bit transform coefficient during quantization, where n is greater than m, and the degree of quantization may be modified by adjusting a quantization parameter (QP). For example, for scalar quantization, different scaling may be performed to achieve finer or coarser quantization. The smaller the quantization step size, the finer the quantization; the larger the quantization step size, the coarser the quantization. A quantization parameter (QP) may be used to indicate a suitable quantization step size. For example, a quantization parameter may be an index of a predefined set of suitable quantization step sizes. For example, a small quantization parameter may correspond to fine quantization (small quantization step size), a large quantization parameter may correspond to coarse quantization (large quantization step size), and vice versa. The quantization operation may include dividing by the quantization step size, and the corresponding dequantization or inverse dequantization operation performed by the dequantization unit 210 or the like may include multiplying by the quantization step size. In some embodiments, according to some standards such as HEVC, the quantization step size may be determined using a quantization parameter. Typically, the quantization step size may be calculated using a fixed-point approximation of an equation including a division operation according to the quantization parameter. Other scaling factors may be introduced for quantization and dequantization to recover the norm of the residual block. Since scaling is used in the fixed-point approximation of the equations for the quantization step size and the quantization parameter, the norm of the residual block may be modified. In an exemplary implementation, the scaling in the inverse transform and dequantization may be combined. Alternatively, a custom quantization table may be used and the custom quantization table may be indicated (signaled) from the encoder to the decoder in a bitstream. Quantization is a lossy operation, where the larger the quantization step size, the greater the loss.
[0205] In an embodiment of the video encoder 20 , the video encoder 20 (correspondingly, the quantization unit 208 ) may be configured to, for example, directly output a quantization parameter (QP) or output the quantization parameter after being encoded by the entropy encoding unit 270 , so that the video decoder 30 may receive and use the quantization parameter for decoding.
[0206] Dequantization
[0207] The inverse quantization unit 210 is used to perform an inverse quantization on the quantized coefficients opposite to the quantization performed by the quantization unit 208 to obtain dequantized coefficients 211, for example, performing an inverse quantization scheme opposite to the quantization scheme performed by the quantization unit 208 according to or using the same quantization step size as the quantization unit 208. The dequantized coefficients 211 may also be referred to as dequantized residual coefficients 211, which correspond to the transform coefficients 207, but due to the loss caused by quantization, the dequantized coefficients 211 are usually not completely equal to the transform coefficients.
[0208] Inverse Transform
[0209] The inverse transform processing unit 212 is used to perform an inverse transform opposite to the transform performed by the transform processing unit 206, for example, an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST), to obtain a reconstructed residual block 213 (or a corresponding dequantized coefficient 211) in the sample domain. The reconstructed residual block 213 may also be referred to as a transform block 213.
[0210] reconstruction
[0211] The reconstruction unit 214 (e.g., adder or summer 214) is used to add the transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265, for example, adding the sample values of the reconstructed residual block 213 and the sample values of the prediction block 265 sample by sample to obtain the reconstructed block 215 in the sample domain.
[0212] Filtering
[0213] The loop filter unit 220 (or simply referred to as "loop filter" 220) is used to filter the reconstruction block 215 to obtain the filter block 221, or is generally used to filter the reconstructed sample to obtain the filtered sample value. For example, the loop filter unit is used to smooth the sudden change of pixels or improve the video quality. The loop filter unit 220 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as a bilateral filter, an adaptive loop filter (ALF), a sharpening filter, a smoothing filter, a collaborative filter, or any combination thereof. Although the loop filter unit 220 is Figure 2 2 is shown as an in-loop filter, but in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filter block 221 may also be referred to as a filter reconstruction block 221.
[0214] In an embodiment of the video encoder 20, the video encoder 20 (correspondingly, the loop filter unit 220) can be used to, for example, directly output or output loop filter parameters (such as sample adaptive offset information) after being encoded by the entropy coding unit 270, so that the decoder 30 can receive and use the same loop filter parameters or different loop filters for decoding.
[0215] Decode image buffer
[0216] The decoded picture buffer (DPB) 230 may be a memory for storing reference pictures or generally storing reference picture data for use when the video encoder 20 encodes video data. The DPB 230 may be formed by any of a variety of memory devices, such as a dynamic random access memory (DRAM) including synchronous DRAM (SDRAM), a magnetoresistive RAM (MRAM), a resistive RAM (RRAM), or other types of memory devices. The decoded picture buffer (DPB) 230 may be used to store one or more filter blocks 221. The decoded picture buffer 230 may also be used to store other previous filter blocks of the same current picture or a different picture such as a previously reconstructed picture, such as a previously reconstructed and filtered block 221, and may provide a complete previously reconstructed picture, i.e., a decoded picture (and corresponding reference blocks and samples) and / or a partially reconstructed current picture (and corresponding reference blocks and samples), for example, to perform inter-frame prediction. The decoded image buffer 230 may also be used to store one or more unfiltered reconstructed blocks 215, or generally store unfiltered reconstructed samples, for example, reconstructed blocks 215 that have not been filtered by the loop filter unit 220, or reconstructed blocks or reconstructed samples that have not undergone any other processing.
[0217] Mode selection (segmentation and prediction)
[0218] The mode selection unit 260 includes a segmentation unit 262, an inter-frame prediction unit 244, and an intra-frame prediction unit 254. The mode selection unit 260 is used to receive or obtain original image data such as an original block 203 (current block 203 of the current image 17) and reconstructed image data such as filtered and / or unfiltered reconstructed samples or reconstructed blocks of the same (current) image and / or one or more previously decoded images from the decoded image buffer 230 or other buffers (e.g., line buffers, not shown in the figure). The reconstructed image data is used as reference image data required for prediction such as inter-frame prediction or intra-frame prediction, and is used to obtain a prediction block 265 or a prediction value 265.
[0219] The mode selection unit 260 can be used to determine or select a partition mode for the current block prediction mode (including a non-partition mode) and a prediction mode (such as an intra-frame prediction mode or an inter-frame prediction mode), and generate a corresponding prediction block 265, which is used for the calculation of the residual block 205 and the reconstruction of the reconstruction block 215.
[0220] In an embodiment of the mode selection unit 260, the mode selection unit 260 may be configured to select a segmentation and prediction mode (e.g., selected from modes supported or available by the mode selection unit 260). The segmentation and prediction mode provides the best match, i.e., the minimum residual (the minimum residual means better compression performance for transmission or storage), or provides the minimum indication overhead (the minimum indication overhead means better compression performance for transmission or storage), or considers both of the above or strikes a balance between the above. The mode selection unit 260 may be configured to determine the segmentation and prediction mode according to rate distortion optimization (RDO), i.e., select the prediction mode that provides the minimum rate distortion. In this document, the terms "best", "minimum", "optimal", etc. do not necessarily mean "best", "minimum", "optimal", etc. in general, but may also refer to situations where termination or selection criteria are met, for example, a value exceeds or falls below a threshold or other limit, which may result in a "suboptimal selection" but reduce complexity and processing time.
[0221] The partitioning unit 262 may be used to partition the block 203 into smaller block parts or sub-blocks (sub-blocks again forming blocks), for example, by iteratively using quad-tree partitioning (QT), binary-tree partitioning (BT) or triple-tree partitioning (TT) or any combination thereof, and for example, to predict each block part or sub-block. The mode selection includes selecting a tree structure of the partitioned block 203 and selecting a prediction mode to be applied to each block part or sub-block.
[0222] The segmentation process (eg, performed by segmentation unit 260) and the prediction process (eg, performed by inter-prediction unit 244 and intra-prediction unit 254) performed by exemplary video encoder 20 are described in detail below.
[0223] segmentation
[0224] The segmentation unit 262 can segment (or divide) the current block 203 into smaller parts, such as square or rectangular blocks. These blocks (also referred to as sub-blocks) can be further segmented into smaller parts. This is also called tree segmentation or hierarchical tree segmentation, where a root block at root tree level 0 (layer 0, depth 0) or the like can be recursively segmented into at least two blocks at the next lower tree level, such as nodes at tree level 1 (layer 1, depth 1). These blocks can be segmented into at least two blocks at the next lower level, such as tree level 2 (layer 2, depth 2), etc., until the segmentation is terminated due to meeting the end criteria, such as reaching the maximum tree depth or the minimum block size. Blocks that are not further segmented are also referred to as leaf blocks or leaf nodes of the tree. A tree segmented into two parts is called a binary tree (binary-tree, BT), a tree segmented into three parts is called a ternary tree (ternary-tree, TT), and a tree segmented into four parts is called a quadtree (quad-tree, QT).
[0225] As mentioned above, the term "block" used in this document may be a portion of an image, in particular a square or rectangular portion. For example, with reference to HEVC and VVC, a block may be or may correspond to a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), and a transform unit (TU), and / or to a corresponding block, for example, a coding tree block (CTB), a coding block (CB), a transform block (TB), or a prediction block (PB).
[0226] For example, a coding tree unit (CTU) may be or may include one CTB of luma samples in an image with 3 sample arrays, two corresponding CTBs of chroma samples in the image, or one CTB of samples in a black and white image or in an image coded using 3 independent color planes and syntax structures. These syntax structures are used to decode samples. Accordingly, a coding tree block (CTB) may be an N×N block of samples, where N may be set to a certain value so that a component is divided into CTBs, which is partitioning. A coding unit (CU) may be or may include one coding block of luma samples in an image with 3 sample arrays, two corresponding coding blocks of chroma samples in the image, or one coding block of samples in a black and white image or in an image coded using 3 independent color planes and syntax structures. These syntax structures are used to decode samples. Accordingly, a coding block (CB) may be an M×N block of samples, where M and N may be set to a certain value so that one CTB is divided into coding blocks, which is partitioning.
[0227] In some embodiments, for example according to HEVC, a coding tree unit (CTU) can be divided into multiple CUs by a quadtree structure represented as a coding tree. A decision is made at the CU level whether to use inter-frame (temporal) prediction or intra-frame (spatial) prediction to decode an image area. Each CU can be further divided into 1, 2 or 4 PUs according to the PU partition type. The same prediction process is performed within a PU, and relevant information is sent to the decoder in units of PUs. After the prediction process is performed according to the PU partition type to obtain the residual block, the CU can be divided into transform units (TUs) according to another quadtree structure similar to the coding tree for the CU.
[0228] In some embodiments, for example, according to the latest video coding standard currently under development (called Versatile Video Coding (VVC)), quad-tree and binary-tree (quad-tree and binary-tree, QTBT) segmentation is used to segment the coding block. In the QTBT block structure, a CU can be square or rectangular in shape. For example, the coding tree unit (CTU) is first segmented by a quadtree structure. The quadtree leaf nodes are further segmented by a binary tree or ternary tree structure. The segmented leaf nodes are called coding units (CUs), and such segmentation is used for prediction and transform processing without any further segmentation. This means that in the QTBT coding block structure, the block sizes of CU, PU, and TU are the same. At the same time, multiple segmentations such as ternary tree segmentation can be used with the QTBT block structure.
[0229] In one example, mode select unit 260 in video encoder 20 may be used to perform any combination of the segmentation techniques described herein.
[0230] As described above, the video encoder 20 is used to determine or select the best or optimal prediction mode from a (predetermined) prediction mode set, which may include, for example, an intra-frame prediction mode and / or an inter-frame prediction mode.
[0231] Intra prediction
[0232] The intra prediction mode set may include 35 different intra prediction modes, such as non-directional modes like DC (or mean) mode and planar mode, or directional modes as defined in HEVC, or may include 67 different intra prediction modes, such as non-directional modes like DC (or mean) mode and planar mode, or directional modes as defined in VVC.
[0233] The intra prediction unit 254 is configured to generate an intra prediction block 265 using reconstructed samples of adjacent blocks of the same current image according to an intra prediction mode in the intra prediction mode set.
[0234] The intra-frame prediction unit 254 (or generally referred to as the mode selection unit 260) is also used to output the intra-frame prediction parameters (or generally referred to as information representing the selected intra-frame prediction mode of the block) to the entropy coding unit 270 in the form of syntax elements 266 to be included in the encoded image data 21, so that (for example) the video decoder 30 can receive and use the prediction parameters for decoding.
[0235] Inter prediction
[0236] The set of (possible) inter prediction modes depends on the available reference picture (i.e. at least part of the decoded picture stored in the DPB 230 as described above) and other inter prediction parameters, e.g. on whether the entire reference picture or only a part of the reference picture (e.g. a search window area around the area of the current block) is used to search for the best matching reference block, and / or e.g. on whether pixel interpolation is performed (e.g. half / half pixel interpolation and / or quarter pixel interpolation).
[0237] In addition to the above-mentioned prediction modes, skip mode and / or direct mode may also be used.
[0238] The inter-frame prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (neither of which is in Figure 2 ). The motion estimation unit may be configured to receive or acquire an image block 203 (a current image block 203 of a current image 17) and a decoded image 231, or at least one or more previously reconstructed blocks (e.g., reconstructed blocks of one or more other / different previously decoded images 231) for motion estimation. For example, a video sequence may include a current image and a previously decoded image 231. In other words, the current image and the previously decoded image 231 may be part of a sequence of images constituting the video sequence or images constituting the sequence.
[0239] For example, the encoder 20 may be configured to select a reference block from a plurality of reference blocks of the same or different images in a plurality of other images, and provide the reference image (or reference image index) and / or the offset (spatial offset) between the position (x coordinate, y coordinate) of the reference block and the position of the current block as an inter-frame prediction parameter to the motion estimation unit. This offset is also referred to as a motion vector (MV).
[0240] The motion compensation unit is used to obtain, such as receiving, inter-frame prediction parameters, and perform inter-frame prediction based on or using the inter-frame prediction parameters to obtain an inter-frame prediction block 265. The motion compensation performed by the motion compensation unit may include extracting or generating a prediction block based on a motion / block vector determined by motion estimation, and may also include interpolating sub-pixel precision. When performing interpolation filtering, other pixel samples can be generated based on known pixel samples, thereby potentially increasing the number of candidate prediction blocks that can be used to decode the image block. The following will be described in more detail: interpolation filtering can be performed using one or more alternative interpolation filters based on motion vector precision. After receiving the motion vector corresponding to the PU of the current image block, the motion compensation unit can locate the prediction block pointed to by the motion vector in one of the reference image lists.
[0241] The motion compensation unit may also generate syntax elements associated with the blocks and video slices for use by video decoder 30 when decoding image blocks of the video slices.
[0242] Entropy Coding
[0243] The entropy coding unit 270 is used to apply or not apply an entropy coding algorithm or scheme (e.g., a variable length coding (VLC) scheme, a context adaptive VLC (CAVLC) scheme, an arithmetic coding scheme, binarization, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding or other entropy coding methods or techniques) to (uncompressed) quantization coefficients 209, inter-frame prediction parameters, intra-frame prediction parameters, loop filter parameters and / or other syntax elements to obtain encoded image data 21 that can be output through an output terminal 272 in the form of an encoded bitstream 21, etc., so that the video decoder 30, etc. can receive and use these parameters for decoding. The encoded bitstream 21 can be transmitted to the video decoder 30, or stored in a memory for subsequent transmission or for subsequent retrieval by the video decoder 30.
[0244] Other structural variations of the video encoder 20 may be used to encode the video stream. For example, a non-transform based encoder 20 may directly quantize the residual signal for certain blocks or frames without the transform processing unit 206. In another implementation, the encoder 20 may include the quantization unit 208 and the inverse quantization unit 210 combined into a single unit.
[0245] Decoder and decoding method
[0246] Figure 3 An exemplary video decoder 30 for implementing the technology of the present application is shown. The video decoder 30 is used to receive the encoded image data 21 (e.g., the encoded code stream 21) as encoded by the encoder 20, and obtain a decoded image 331. The encoded image data or code stream includes information for decoding the encoded image data, such as data representing image blocks in the encoded video slice and related syntax elements.
[0247] exist Figure 3In the example of , decoder 30 includes entropy decoding unit 304, inverse quantization unit 310, inverse transform processing unit 312, reconstruction unit 314 (e.g., summer 314), loop filter 320, decoded picture buffer (DPB) 330, inter prediction unit 344, and intra prediction unit 354. Inter prediction unit 344 may be or may include a motion compensation unit. In some examples, video decoder 30 may perform substantially the same as in combination with Figure 2 The video encoder 100 shown is a decoding process that is the reverse of the encoding process described.
[0248] As described in the encoder 20, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 344, and the intra prediction unit 354 constitute the "built-in decoder" of the video encoder 20. Accordingly, the inverse quantization unit 310 may be functionally identical to the inverse quantization unit 210, the inverse transform processing unit 312 may be functionally identical to the inverse transform processing unit 212, the reconstruction unit 314 may be functionally identical to the reconstruction unit 214, the loop filter 320 may be functionally identical to the loop filter 220, and the decoded picture buffer 330 may be functionally identical to the decoded picture buffer 230. Therefore, the explanation of the corresponding units and functions of the video encoder 20 applies accordingly to the corresponding units and functions of the video decoder 30.
[0249] Entropy decoding
[0250] The entropy decoding unit 304 is used to parse the bit stream 21 (or collectively referred to as the encoded image data 21) and perform entropy decoding on the encoded image data 21 to obtain the quantization coefficient 309 and / or the decoded encoding parameter ( Figure 3 The entropy decoding unit 304 may be used to provide the inter-frame prediction parameters, intra-frame prediction parameters and / or other syntax elements to the mode selection unit 360, and to provide other parameters to other units of the decoder 30. The video decoder 30 may receive syntax elements at the video slice level and / or the video block level.
[0251] Dequantization
[0252] The inverse quantization unit 310 may be configured to receive a quantization parameter (QP) (or generally referred to as inverse quantization related information) and a quantization coefficient from the encoded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304, etc.), and inverse quantize the decoded quantization coefficient 309 according to the quantization parameters to obtain a dequantized coefficient 311. The dequantized coefficient 311 may also be referred to as a transform coefficient 311. The inverse quantization process may include determining a degree of quantization using a quantization parameter determined by the video encoder 20 for each video block in the video slice, and also determining a degree of inverse quantization that needs to be performed.
[0253] Inverse Transform
[0254] The inverse transform processing unit 312 can be used to receive the dequantized coefficients 311 (also called transform coefficients 311) and transform the dequantized coefficients 311 to obtain a reconstructed residual block 313 in the sample domain. The reconstructed residual block 213 can also be called a transform block 313. The transform can be an inverse transform, such as an inverse DCT, an inverse DST, an inverse integer transform, or a conceptually similar inverse transform process. The inverse transform processing unit 312 can also be used to receive transform parameters or corresponding information from the encoded image data 21 (for example, parsed and / or decoded by the entropy decoding unit 304, etc.) to determine the transform to be performed on the dequantized coefficients 311.
[0255] reconstruction
[0256] The reconstruction unit 314 (eg, adder or summer 314 ) may be configured to add the reconstructed residual block 313 to the prediction block 365 by, for example, adding sample values of the reconstructed residual block 313 and sample values of the prediction block 365 to obtain the reconstructed block 315 in the sample domain.
[0257] Filtering
[0258] The loop filter unit 320 (in the decoding loop or after the decoding loop) is used to filter the reconstructed block 315 to obtain a filter block 321, so as to smooth the sudden change of pixels or improve the video quality in other ways. The loop filter unit 320 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as a bilateral filter, an adaptive loop filter (ALF), a sharpening filter, a smoothing filter, a collaborative filter, or any combination thereof. Although the loop filter unit 320 is in Figure 3 3. Although shown as an in-loop filter in FIG. 3, in other configurations, the loop filter unit 320 may be implemented as a post-loop filter.
[0259] Decode image buffer
[0260] The decoded video block 321 of one picture is stored in a decoded picture buffer 330, which stores the decoded picture 331 as a reference picture for subsequent motion compensation of other pictures and / or corresponding output for display.
[0261] The decoder 30 is used to output the decoded image 311 through an output terminal 312 or the like, so as to be presented to a user or for the user to watch.
[0262] predict
[0263] The inter-frame prediction unit 344 may be functionally identical to the inter-frame prediction unit 244 (particularly the motion compensation unit), and the intra-frame prediction unit 354 may be functionally identical to the intra-frame prediction unit 254, and may determine the division or segmentation and perform prediction based on the segmentation and / or prediction parameters or corresponding information received from the coded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304, etc.). The mode selection unit 360 may be configured to perform prediction (intra-frame prediction or inter-frame prediction) by block based on the reconstructed image, the reconstructed block or the corresponding sample (filtered or unfiltered) to obtain a prediction block 365.
[0264] When the video slice is decoded as an intra-coded (I) slice, the intra prediction unit 354 in the mode selection unit 360 is used to generate a prediction block 365 for the image block of the current video slice based on the indicated intra prediction mode and data from a previously decoded block of the current image. When the video image is decoded as an inter-coded (B or P) slice, the inter prediction unit 344 (e.g., a motion compensation unit) in the mode selection unit 360 is used to generate a prediction block 365 for the video block of the current video slice based on the motion vector and other syntax elements received from the entropy decoding unit 304. For inter prediction, these prediction blocks can be generated based on one of the reference images in one of the reference image lists. The video decoder 30 can construct reference frame list 0 and reference frame list 1 using a default construction technique based on the reference images stored in the DPB 330.
[0265] The mode selection unit 360 is used to determine prediction information of the video blocks of the current video slice by parsing the motion vectors and other syntax elements, and use the prediction information to generate a prediction block for the current video block being decoded. For example, the mode selection unit 360 uses some of the received syntax elements to determine the prediction mode (e.g., intra prediction or inter prediction) for decoding the video blocks of the video slice, the inter prediction slice type (e.g., B slice, P slice, or GPB slice), construction information of one or more reference picture lists for the slice, the motion vector for each inter-coded video block of the slice, the inter prediction state for each inter-coded video block of the slice, and other information to decode the video blocks in the current video slice.
[0266] Other variations of the video decoder 30 may be used to decode the encoded image data 21. For example, the decoder 30 may be capable of generating an output video stream without the loop filter unit 320. For example, a non-transform based decoder 30 may directly dequantize the residual signal for certain blocks or frames without the inverse transform processing unit 312. In another implementation, the video decoder 30 may include the dequantization unit 310 and the inverse transform processing unit 312 combined into a single unit.
[0267] It should be understood that in the encoder 20 and the decoder 30, the processing result of the current step may be further processed and then output to the next step. For example, after interpolation filtering, motion vector derivation or loop filtering, the processing result of interpolation filtering, motion vector derivation or loop filtering may be further subjected to operations such as clipping or shifting.
[0268] It should be noted that further operations can be performed on the derived motion vector of the current block (including but not limited to the control point motion vector in affine mode, the sub-block motion vector in affine mode, the sub-block motion vector in plane mode, the sub-block motion vector in ATMVP mode, the time domain motion vector, etc.). For example, the value of the motion vector is limited to a predefined range according to the representation bit of the motion vector. If the representation bit of the motion vector is bitDepth, the range is –2^(bitDepth–1)~2^(bitDepth–1)–1, where the “^” symbol represents the power. For example, if bitDepth is set to 16, the range is –32768~32767; if bitDepth is set to 18, the range is –131072~131071.
[0269] Figure 4Schematic diagram of a video decoding device 400 provided in an embodiment of the present invention. The video decoding device 400 is suitable for implementing the disclosed embodiments described herein. In one embodiment, the video encoding device 400 may be a decoder, such as Figure 1A The video decoder 30 in the embodiment may also be an encoder, for example Figure 1A The video encoder 20 in.
[0270] The video decoding device 400 includes: an input port 410 (or input port 410) and a receiver unit (Rx) 420 for receiving data; a processor, a logic unit or a central processing unit (CPU) 430 for processing data; a transmitter unit (Tx) 440 and an output port 450 (or output port 450) for transmitting data; and a memory 460 for storing data. The video decoding device 400 may also include an optical-to-electrical (OE) component and an electrical-to-optical (EO) component coupled to the input port 410, the receiving unit 420, the transmitting unit 440 and the output port 450, which are used as an outlet or an inlet of an optical signal or an electrical signal.
[0271] The processor 430 is implemented by hardware and software. The processor 430 can be implemented as one or more CPU chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. The processor 430 communicates with the input port 410, the receiving unit 420, the sending unit 440, the output port 450, and the memory 460. The processor 430 includes a decoding module 470. The decoding module 470 implements the disclosed embodiments described above. For example, the decoding module 470 performs, processes, prepares, or provides various decoding operations. Therefore, the inclusion of the decoding module 470 provides substantial improvements to the functionality of the video decoding device 400 and affects the transition of the video decoding device 400 to different states. Alternatively, the decoding module 470 is implemented with instructions stored in the memory 460 and executed by the processor 430.
[0272] The memory 460 may include one or more disks, tape drives, and solid-state hard disks, and may be used as an overflow data storage device to store programs when they are selected for execution and to store instructions and data read during the execution of the programs. For example, the memory 460 may be volatile and / or non-volatile, and may be a read-only memory (ROM), a random access memory (RAM), a ternary content-addressable memory (TCAM), and / or a static random-access memory (SRAM).
[0273] Figure 5 A simplified block diagram of an apparatus 500 is provided for an exemplary embodiment. The apparatus 500 may be used as Figure 1A Either or both of the source device 12 and the destination device 14 in .
[0274] The processor 502 in the device 500 may be a central processing unit. Alternatively, the processor 502 may be any other type of device or devices that are currently available or will be developed in the future and are capable of manipulating or processing information. Although the disclosed implementation may be implemented using a single processor such as the processor 502 shown in the figure, the speed and efficiency may be improved when implemented using more than one processor.
[0275] In one implementation, the memory 504 in the apparatus 500 may be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as the memory 504. The memory 504 may include code and data 506 accessed by the processor 502 via a bus 512. The memory 504 may also include an operating system 508 and an application 510, the application 510 including at least one program that causes the processor 502 to perform the methods described herein. For example, the application 510 may include applications 1 to N, which include a video decoding application that performs the methods described herein.
[0276] The apparatus 500 may also include one or more output devices, such as a display 518. In one example, the display 518 may be a touch-sensitive display that combines a display with a touch-sensitive element, which can be used to sense touch input. The display 518 may be coupled to the processor 502 via the bus 512.
[0277] Although the bus 512 of the device 500 is also shown as a single bus in the figure, there may be multiple buses 512. In addition, the auxiliary memory 514 may be directly coupled to other components in the device 500 or may be accessed through a network, and may include a single integrated unit (e.g., one memory card) or multiple units (e.g., multiple memory cards). Therefore, the device 500 may have a variety of configurations.
[0278] The concepts presented herein are described in more detail below.
[0279] Motion Vector Prediction
[0280] Spatial motion vector prediction is used in the current VVC design. Spatial motion vector prediction refers to predicting the motion vector of the current inter-frame block using the motion information of spatially adjacent blocks during the inter-frame prediction process. Specifically, the motion vectors of the adjacent spatially adjacent blocks of the current block are used in the fusion mode and the skip mode. In the fusion mode and the skip mode, the so-called HMVP candidate can be used. The HMVP candidate includes motion information of historical spatially adjacent blocks. "History-based" means that the motion information of blocks arranged in decoding order before the current block is used. These previous blocks are from the same frame as the current block and are located in some spatial neighborhood around the current block, but are not necessarily adjacent blocks like the general spatial fusion candidate.
[0281] Construction of fusion candidate list
[0282] The fusion candidate list is constructed based on the following candidates:
[0283] • Up to four spatial fusion candidates derived from five spatial neighboring blocks, e.g. Figure 6 As shown;
[0284] • A temporal fusion candidate derived from two temporally collocated blocks;
[0285] • Other fused candidates, including combined bi-directional prediction candidates and zero motion vector candidates. Fig.12 The construction of the fusion candidate list is described in more detail.
[0286] Space Candidates
[0287] The first candidate set in the fusion candidate list includes Figure 6 The spatial neighboring blocks shown. For the fusion of inter-frame prediction blocks, up to four candidates are inserted in the fusion list in this order by checking A1, B1, B0, A0 and B2 in sequence. Before all the motion data of the neighboring blocks are taken as fusion candidates, some additional redundancy checks are performed instead of just checking whether the neighboring blocks are available and contain motion information. These redundancy checks can be divided into two categories according to the following two different purposes:
[0288] • Avoid including candidates with redundant motion data in the HMI list, and
[0289] • Prevents merging two parts (partitions) that could be expressed in other ways, thus producing redundant syntax.
[0290] History-based Motion Vector Prediction
[0291] In order to further improve motion vector prediction, techniques using motion information of non-adjacent CUs (motion information includes reference image index and motion vector) are proposed. One of these techniques is history-based motion vector prediction (HMVP). HMVP uses a look-up table (LUT) consisting of motion information of previously encoded CUs. The HMVP method basically consists of two main parts:
[0292] 1. The construction and update method of HMVP lookup table (HMVP LUT) is as follows Fig.10 and Fig.11 shown.
[0293] 2.HMVP LUT is used to build a fusion candidate list (or AMVP candidate list), such as Fig.12 shown.
[0294] HMVP LUT construction and update method
[0295] The LUT is maintained during the encoding and / or decoding process. When there is a new slice, the LUT is emptied. Whenever the current CU is inter-coded, the relevant motion information is added to the last entry of the table as a new HMVP candidate. The size of the LUT (denoted as N) is a parameter in the HMVP method. If the number of HMVP candidates for the previously encoded CU is greater than the size of this LUT, a lookup table update method is applied so that this LUT always contains no more than the N latest previously encoded motion candidates. Therefore, two lookup table update methods are provided:
[0296] 1. First-In-First-Out (FIFO) LUT update method, such as Fig.10 As shown;
[0297] 2. Constrain the FIFO LUT update method, such as Fig.11 shown.
[0298] FIFO LUT Update Method
[0299] In the FIFO LUT update method, the oldest candidate (at entry 0 in the lookup table) is removed from the table lookup before inserting a new candidate. This process is as follows Fig.10 As shown in the figure, H0 is the earliest HMVP candidate (ie, the 0th HMVP candidate), and X is the new HMVP candidate.
[0300] The complexity of this updating method is relatively small, but when this method is applied, some LUT elements may be the same (i.e. contain the same motion information). Therefore, the data in the LUT may be redundant, and the diversity of motion information in the LUT is worse than that in the method of removing duplicate candidates.
[0301] Constraint FIFO LUT Update Method
[0302] In order to further improve the decoding efficiency, a constrained FIFO LUT update method is provided. In this method, before inserting a new HMVP candidate into the lookup table, a redundancy check is performed. The redundancy check refers to checking whether the motion information of the new candidate X is consistent with the existing candidate HMVP in the LUT. m The motion information contained is the same. If the candidate H is not found m , a simple FIFO method is used; otherwise, the following process is performed:
[0303] 1. Place the m All subsequent LUT entries are shifted one position to the left (ie, toward the beginning of the lookup table), thereby moving candidate H m Delete from the lookup table, freeing up a position at the end of the LUT.
[0304] 2. Add the new candidate X to the first vacant position in the lookup table.
[0305] Fig.11 An example of using the constrained FIFO LUT update method is shown.
[0306] Motion vector decoding using HMVP LUT
[0307] The HMVP candidate may be used in a process of constructing a fusion candidate list and / or a process of constructing an AMVP candidate list.
[0308] Use HMVP LUT to build a fusion candidate list
[0309] In some examples, the HMVP candidates are sorted in order from the last entry to the first entry after the temporal fusion candidates (eg, H N–1 , H N–2 , ..., H0) are inserted into the fusion list. The LUT traversal order is as follows Fig.12As shown. If the HMVP candidate is the same as a candidate already in the fusion list, the HMVP candidate will not be added to the HMVP list. Due to the limited size of the fusion list, some HMVP candidates, especially those at the beginning of the LUT, may not be used in the construction process of the fusion list of the current CU.
[0310] Use HMVP LUT in the construction process of AMVP candidate list
[0311] The HMVP LUT built for fusion mode can also be used to build the AMVP candidate list. The difference is that only a small number of entries in the LUT are used to build the AMVP candidate list. More specifically, only the last M entries in the HMVP LUT are used (for example, M equals 4). In the process of building the AMVP candidate list, the HMVP candidates are sorted from the last entry to the (N–K)th entry after the TMVP candidate (i.e., in the order of Fig.12 H N–1 , H N–2 ,……,H N–K LUT traversal order) into the AMVP candidate list.
[0312] Only HMVP candidates with the same reference image as the AMVP target reference image are used. If an HMVP candidate is the same as a candidate already in the HMVP list, the AMVP candidate list will not be constructed using that HMVP candidate. Due to the limited size of the AMVP candidate list, some HMVP candidates may not be used in the construction process of the AMVP list of the current CU.
[0313] Switchable interpolation filter
[0314] The motion vector difference of the translational inter-frame prediction block can be encoded with 3 different precisions (i.e., 1 / 4 pixel precision, full pixel precision, and 4 pixel precision). The interpolation filter (IF) used for each fractional position is fixed. In the present invention, the switchable interpolation filter (SIF) technology allows the use of one or two alternative brightness interpolation filters for half-pixel positions. CU-level switching can be performed between the available brightness interpolation filters. In order to reduce the indication overhead, the switching depends on the motion vector precision used. To achieve this goal, the Adaptive Motion Vector Resolution (AMVR) scheme is extended to support half-pixel brightness motion vector precision. Only when the half-pixel motion vector precision mode is adopted, the alternative half-pixel interpolation filter can be applied, and the alternative half-pixel interpolation filter is represented by an additional syntax element indicating which interpolation filter to use. In the skip mode or fusion mode containing spatial fusion candidates, the value of this syntax element can be inherited from the adjacent block.
[0315] Half-Pixel AMVR Mode
[0316] An additional AMVR mode for non-affine non-fused inter-coded CUs is introduced, which allows the motion vector difference to be indicated with half-pixel precision. The existing AMVR scheme of the current VVC draft is directly extended in the following way: directly after the syntax element amvr_flag, if amvr_flag==1, there is a new context modeling binary syntax element hpel_amvr_flag, indicating that the new half-pixel AMVR mode is used when hpel_amvr_flag==1. Otherwise, that is, when hpel_amvr_flag==0, the full-pixel AMVR mode or the 4-pixel AMVR mode is selected according to the indication of the syntax element amvr_precision_flag as in the current VVC draft.
[0317] Alternative luma half-pixel interpolation filter
[0318] For non-affine non-fused inter-coded CUs using half-pixel motion vector precision (i.e., half-pixel AMVR mode), it is possible to switch between the HEVC / VVC half-pixel luma interpolation filter and one or more alternative half-pixel interpolation filters based on the value of the new syntax element if_idx (interpolation filter index). The syntax element if_idx is only indicated when half-pixel AMVR mode is adopted. When spatial fused candidates are used in skip mode / fused mode, the value of the interpolation filter index is inherited from the neighboring block.
[0319] It is understood that the fractional position of a motion vector may be represented by, for example, a fractional sample unit (xFrac L ,yFrac L ) is represented by the brightness position in . The motion vector of the selected fusion candidate can be represented by refMvLX[0] and refMvLX[1], where mvLX=mvL0 or mvL1.
[0320] In one example,
[0321] xFrac L =refMvLX[0]&15(8-738),
[0322] yFrac L =refMvLX[1]&15(8-739),
[0323] If xFrac L (or yFrac L ) is equal to zero (meaning that the MV points to an integer position), no interpolation is performed. Otherwise, (xFrac L In the range [1, 15]), use f L [xFrac L The interpolation filter with coefficients specified in ] is shown in Table 1. The luma interpolation filter coefficients f for each fractional sample position p (p is in the range [1, 15]) L [p].
[0324] Table 1 is an example of an interpolation filter set. An interpolation filter is selected according to the fractional position. An interpolation filter (interpolation filter coefficient) may be shown as a row in Table 1. In an example, the interpolation filter set in the present invention may include the same interpolation filter for all positions except the half sample position (fractional position: 1 / 2).
[0325] Table 1 shows the HEVC / VVC interpolation filter coefficients f for each fractional sample position p (p is in the range [1, 15] with an accuracy of 1 / 16 pixel (pixel)) L [p] In this table, when p=8, the interpolation filter coefficient f L [p] is the half-pixel interpolation filter coefficient. As described above, additional interpolation filters can be added as alternative interpolation filter values for the half-pixel interpolation filters to allow switching between these half-pixel interpolation filters. Some examples of alternative half-pixel interpolation filters are described below.
[0326] Table 1 Luminance interpolation filter coefficient specifications
[0327]
[0328] Use an alternative 6-tap half-pixel interpolation filter implementation
[0329] In one example, a 6-tap interpolation filter may be used as an alternative interpolation filter to the normal HEVC / VVC half-pixel interpolation filter as shown in Table 1. Table 2 below shows the mapping between the value of the syntax element if_idx (or the derived IF index) and the selected half-pixel luma interpolation filter:
[0330] Table 2
[0331]
[0332] Implementation using two alternative 8-tap half-pixel interpolation filters
[0333] In another example, two 8-tap interpolation filters may be used as an alternative interpolation filter to the normal HEVC / VVC half-pixel interpolation filter as shown in Table 1. The following Table 3 shows the mapping between the value of the syntax element if_idx and the selected half-pixel luma interpolation filter:
[0334] Table 3
[0335]
[0336] Implementation using two alternative 6-tap half-pixel interpolation filters
[0337] In another example, two 6-tap interpolation filters may be used as an alternative interpolation filter to the normal HEVC / VVC half-pixel interpolation filter as shown in Table 1. Table 4 below shows the mapping between the value of the syntax element if_idx and the selected half-pixel luma interpolation filter:
[0338] Table 4
[0339]
[0340] In the present invention, the interpolation filters for half-pixel positions (such as the row marked as "8" in Table 5) in the interpolation filters shown in Table 5 can be switched. In the present invention, when the corresponding MV points to a half-sample position, an alternative or switchable half-sample interpolation filter can be used to interpolate the half-sample value.
[0341] Table 5 Luma interpolation filter coefficient f for each 1 / 16 fractional sample position L [p] Specifications
[0342]
[0343] Describe more specifically in the following areas:
[0344] 1. Modify the construction / update method of the history-based motion information candidate list (i.e., HMI list). In addition to the motion information of one or more coded / decoded blocks before a block, the interpolation filter (IF) index (e.g., half-pixel interpolation filter index (hpelIfIdx)) of the previous block is stored in the HMI list. In particular, the IF index is stored in the HMI candidate or record of the HMI list. In this way, the IF index can be propagated through the HMI list, thereby achieving decoding consistency and improving decoding efficiency.
[0345] 2. Derivation process of interpolation filter index (half-pixel interpolation filter index) in fusion mode: If a block has a fused candidate index corresponding to a history-based candidate, the IF index (half-pixel interpolation filter index) of the history-based candidate is used for the current block.
[0346] 3. SIF index propagation is across CTU boundaries. Based on the current SIF design, when SIF technology is applied to the mode of inheriting motion information from the upper spatial neighboring blocks, if the current block is at the upper boundary of the CTU, the row memory will be increased. As described in this article, the position of the current block is checked. If the current block is at the upper boundary of the CTU, when inheriting motion information from the upper left neighboring block (B0), the upper neighboring block (B1), and the upper right neighboring block (B2), the IF index is not inherited, but the default value is used, thereby reducing the cost of row memory.
[0347] Figure 7 An example of SIF index propagation across CTU boundaries is shown. In this example, motion information is inherited from an upper adjacent block B1 included in a CTU different from the CTU containing the current block 700. At this time, the SIF index of block B1 must be stored in the row buffer in the prior art. In this case, the present invention prevents SIF index propagation, thereby reducing the capacity requirement of the row buffer. The position of the current block is checked during the construction of the fusion list. Figure 7 As shown in the figure, if the current block is located at the upper boundary of the CTU, when inheriting motion information from the upper left neighboring block (B0), the upper neighboring block (B1), and the upper right neighboring block (B2), the IF index is not inherited, but the default value is used, thereby reducing the cost of row memory. Figure 8 and Fig. 9 Describe in detail.
[0348] Fig.13A 1 is a flow chart of a method 1300 for constructing a history-based motion information candidate list (ie, HMI list). The method comprises the following steps:
[0349] Step 1301: Obtain a history-based motion information candidate list, wherein the HMI list is N history-based motion information candidate HMIs. k An ordered list of (k=0, ..., N-1) of N history-based motion information candidates H k Associated with (or containing) motion information of N previous blocks before a block, N being an integer greater than 0, each history-based motion information candidate includes the elements:
[0350] (i) one or more motion vectors MV of a previous block,
[0351] (ii) one or more reference picture indices corresponding to the MV of the previous block,
[0352] (iii) The interpolation filter index of the previous block.
[0353] Step 1303: Update the HMI list according to the motion information of the block, where the motion information of the block includes the following elements:
[0354] (i) one or more motion vectors MV of the block,
[0355] (ii) one or more reference picture indices corresponding to the MV of the block,
[0356] (iii) The interpolation filter index of the block.
[0357] It should be noted that the one or more MVs of the block refer to the MVs corresponding to the reference picture lists L0 and L1. The same is true for the reference picture index.
[0358] like Fig. 13B As shown, step 1301 may be step 1311 including loading a history-based motion information candidate list (HMI table), and step 1303 may be step 1313 including updating the history-based motion information candidate list (table) using the motion information of the decoded block. The HMI table containing multiple HMVP candidates is maintained during the encoding / decoding process. When there is a new slice, the HMI table is vacated. When the slice has an inter-frame decoding block, the block is decoded according to the motion information candidate list including the history-based motion information candidate (step 1301), and the associated motion information of the block is added as a new HMVP candidate to the last table entry of the HMI table (step 1303).
[0359] Fig.14 The method comprises the following steps: Step 1401: constructing a history-based motion information candidate list (i.e., HMVP list), wherein the HMVP list is N history-based motion information candidate HMVP lists.k An ordered list of (k=0, ..., N-1) of N history-based motion information candidates H k Associated with motion information of multiple blocks before the block, N is an integer greater than 0, and each history-based motion information candidate includes the elements:
[0360] (i) one or more motion vectors MV,
[0361] (ii) one or more reference image indices corresponding to the MV,
[0362] (iii) interpolation filter index;
[0363] Step 1402: Add one or more history-based motion information candidates in the HMI list to the motion information candidate list of the block.
[0364] Step 1403: Derive the motion information of the block according to the motion information candidate list.
[0365] It can be understood that the motion information candidate list refers to the following fusion candidate list.
[0366] It can be understood that the history-based fusion candidate is included in the motion information candidate list in step 1403 .
[0367] Fig.15 The present invention is a flowchart of a method for constructing and updating a history-based motion information candidate list (i.e., HMI list). Step 1501: Construct an HMI list. Step 1502: Compare at least one of the elements (i) and (ii) of each history-based motion information candidate in the HMVP list with the corresponding element of the current block. Step 1502 includes: comparing whether the motion vectors of the history-based motion information candidates in the history-based motion information candidate list are the same as the corresponding motion vectors of the block, and comparing whether the reference image indexes of the history-based motion information candidates are the same as the corresponding reference image indexes of the block. In an alternative design, step 1502 includes: comparing whether at least one of the motion vectors of each history-based motion information candidate is different from the corresponding motion vector of the block, and comparing whether at least one of the reference image indexes of each HMVP candidate is different from the corresponding reference image index of the block. The result of the element-based comparison is called Fig.15 The comparison results in .
[0368] If the comparison result is that at least one of the elements (i) and (ii) of each history-based motion information candidate in the history-based motion information candidate list is different from the corresponding element in the motion information of the block, the motion information of the current block is added to the last position of the HMVP list (step 1503). Otherwise, if the elements (i) and (ii) of a history-based motion information candidate in the history-based motion information candidate list are the same as the corresponding element in the motion information of the block, the history-based motion information candidate is deleted from the history-based motion information candidate list, and the history-based motion information candidate HMVP list containing the motion information of the block is added. k (k=N−1) is added to the last position of the history-based motion information candidate list (step 1504 ).
[0369] In the above comparison process, only the difference between the MV and reference image indexes is checked, and no comparison is made on the IF index.
[0370] Additional embodiments are outlined in the following aspects:
[0371] According to a first aspect of the present invention, a method for deriving an interpolation filter index (or an interpolation filter set index) of a current block is provided, comprising:
[0372] Construct a history-based motion information list (HMIL table or HMVP table), wherein the history-based motion information list is N motion records H k An ordered list of (k=0, ..., N-1), the N motion records H k Respectively associated with N previous blocks of a frame, N is greater than or equal to 1, each motion record includes one or more motion vectors, one or more reference image indexes corresponding to the one or more motion vectors, and interpolation filter indexes (or interpolation filter set indexes) corresponding to the one or more motion vectors (for example, two MVs correspond to the same filter index or the same filter set index);
[0373] A history-based motion information candidate (such as an HMVP candidate) of the current block is determined according to the history-based motion information list (eg, the HMVP candidate of the current block is determined from the HMVP list or the HMVP table).
[0374] According to the method of the first aspect, in a possible implementation manner, determining the history-based motion information candidate of the current block according to the history-based motion information list includes:
[0375] Derive or infer or determine record H kas the interpolation filter index (or interpolation filter set index) of the current block, wherein the determined or selected history-based motion information candidate (eg, HMVP candidate) corresponds to the recorded H k .
[0376] According to the method of any one of the first aspect or the implementation manner of the first aspect, in a possible implementation manner, the motion records in the history-based motion information list are sorted according to the order in which the motion records of the previous block are obtained from the bitstream.
[0377] According to the method of any one of the first aspect or the implementation manner of the first aspect, in a possible implementation manner, the length of the history-based motion information list is N, and N is 5.
[0378] According to the method in any one of the first aspect or the implementation manner of the first aspect, in a possible implementation manner, constructing a history-based motion information list (HMVL) includes:
[0379] Before adding the motion information of the current block to the HMVL, checking whether each element of the HMVL is different from the motion information of the current block;
[0380] The motion information of the current block is added to the HMVL only when every element of the HMVL is different from the motion information of the current block.
[0381] According to the method in any one of the first aspect or the implementation manner of the first aspect, in a possible implementation manner, checking whether each element of the HMVL is different from the motion information of the current block includes:
[0382] comparing corresponding motion vectors;
[0383] Compare the corresponding reference image index.
[0384] According to the method in any one of the first aspect or the implementation manner of the first aspect, in a possible implementation manner, checking whether each element of the HMVL is different from the motion information of the current block includes:
[0385] Compare interpolation filter indices.
[0386] According to the method in any one of the first aspect or the implementation manner of the first aspect, in a possible implementation manner, the method further includes: deriving motion information based on motion information of a first block, wherein the first block has a preset spatial or temporal position relationship with the current block.
[0387] According to the method in any one of the first aspect or the implementation manner of the first aspect, in a possible implementation manner, the method further includes:
[0388] The motion information is derived according to the motion information of a second block, wherein the second block is reconstructed before the current block.
[0389] According to the method described in the first aspect or any implementation manner of the first aspect, in a possible implementation manner, when the current block is in fusion mode, the history-based motion information list (HMIL table or HMVP table) is a subset of the candidate motion information list of the current block; or, when the current block is in AMVP mode, the history-based motion information list (HMIL table or HMVP table) is a subset of the candidate prediction motion information list of the current block.
[0390] According to the method of any one of the first aspect or the implementation manner of the first aspect, in a possible implementation manner, only one interpolation filter set index corresponds to one or more motion vectors of the HMVP candidate (for example, two MVs correspond to the same filter set index); or
[0391] The one or more interpolation filter set indexes respectively correspond to the one or more motion vectors of the HMVP candidate.
[0392] A second aspect of the present invention provides a method for performing inter-frame prediction on a current block, comprising:
[0393] Performing inter-frame prediction on the current block, including deriving an interpolation filter index (or an interpolation filter set index) for the current block;
[0394] Wherein, deriving an interpolation filter index (or an interpolation filter set index) for the current block includes:
[0395] Determine an HMVP candidate of the current block from an HMVP list (e.g., an HMVP table), wherein the HMVP candidate includes at least one motion vector, at least one reference image index corresponding to the at least one motion vector, and at least one interpolation filter index (or interpolation filter set index) corresponding to the at least one motion vector (e.g., the entire candidate corresponds to only one interpolation filter index or one interpolation filter set index);
[0396] Derivation or inference or determination of the interpolation filter index (or interpolation filter set index) of the determined or selected HMVP candidate as the interpolation filter index (or interpolation filter set index) of the current block;
[0397] The one or more candidates (eg, each candidate) of the HMVP list include at least one motion vector and an interpolation filter index (or at least one interpolation filter set index) corresponding to the at least one motion vector.
[0398] According to the method of the second aspect, in a possible implementation manner, an interpolation filter index (or an interpolation filter set index) corresponds to the one or more motion vectors of the HMVP candidate; or
[0399] One or more interpolation filter indices (or one or more interpolation filter set indices) correspond to one or more motion vectors of the HMVP candidate.
[0400] A third aspect of the present invention provides a method for deriving an interpolation filter for a coding unit encoded in a fusion mode according to a position of a current coding unit in a CTU, comprising:
[0401] Parsing or deriving a first fusion index from a bitstream;
[0402] Selecting a fusion candidate from the fusion candidate list according to the first fusion index;
[0403] Determining whether the current coding unit overlaps with an upper boundary or a left boundary of the CTU;
[0404] If the current coding unit overlaps with an upper boundary or a left boundary of the CTU, setting an interpolation filter index (or an interpolation filter set index) of the current coding unit to a predefined value;
[0405] Otherwise, setting the interpolation filter index (or interpolation filter set index) of the current coding unit to be equal to the interpolation filter index (or interpolation filter set index) of the selected fusion candidate;
[0406] Selecting a first interpolation filter set from N interpolation filter sets (e.g., N predefined interpolation filter sets) according to the interpolation filter index (or interpolation filter set index), where N is an integer greater than or equal to 2;
[0407] For each motion vector of the selected fusion candidate, the fractional position (eg, fractional sample unit (xFrac) of the motion vector) is calculated. L ,yFrac L )'s brightness position) selects an interpolation filter from the first interpolation filter set.
[0408] According to the method of the third aspect, in a possible implementation manner, the method further includes:
[0409] The fusion candidate list is constructed, wherein each candidate includes one or more motion vectors, an interpolation filter index (or an interpolation filter set index for representing one interpolation filter set among N interpolation filter sets (eg, N predefined interpolation filter sets)).
[0410] According to the method described in any one of the third aspect or the implementation manner of the third aspect, in one possible implementation manner, determining whether the current coding unit overlaps with the upper boundary or left boundary of the CTB or CTU includes: determining whether the upper left corner of the current block (for example, used to represent the luminance position (xCb, yCb) of the upper left corner sample of the current coding block relative to the upper left luminance sample of the current image) overlaps with the upper boundary of the CTU including the current coding unit.
[0411] According to the method in any one of the third aspect or the implementation manner of the third aspect, in a possible implementation manner, determining whether the upper left corner of the current block overlaps with the upper boundary of the CTU including the current coding unit includes:
[0412] Get the y coordinate position (y coordinate) of the upper left corner of the current block;
[0413] Divide the obtained ordinate position (y coordinate) by the height of the CTU to obtain a remainder;
[0414] If the calculated remainder is equal to zero, it is inferred that the upper left corner of the current block overlaps with the upper boundary of the current CTU; otherwise, it is inferred that the upper left corner of the current coding unit does not overlap with the upper boundary of the current CTU.
[0415] According to the method in any one of the third aspect or the implementation manner of the third aspect, in a possible implementation manner, determining whether the upper left corner of the current coding unit overlaps with the upper boundary of the CTU including the current coding unit includes:
[0416] Calculate a first value as a maximum floor-rounded value, wherein the maximum floor-rounded value is obtained by dividing the upper left ordinate (coordinate y) of the current block by the height of the CTU;
[0417] Calculate a second value as a maximum floor-rounded value, wherein the maximum floor-rounded value is obtained by dividing the upper left vertical coordinate of the inherited adjacent block by the height of the CTU;
[0418] If the second value is equal to the first value, it is inferred that the upper left corner of the current coding unit overlaps with the upper boundary of the current CTU.
[0419] The method according to any one of the third aspect or the implementations of the third aspect, in a possible implementation, determining whether the upper left corner of the current coding unit overlaps with the upper boundary of the CTU including the current coding unit includes:
[0420] Calculate a third value as (yCb >> CtbLog2SizeY) << CtbLog2SizeY, where yCb represents the upper left vertical coordinate (coordinate y) of the current block, ">>" represents logical or arithmetic bit right shift, "<< " represents logical or arithmetic bit left shift, and CtbLog2SizeY represents the binary logarithm scale of the size of the CTU;
[0421] If (yCb – 1) is less than the third value, infer that the upper left corner of the current coding unit overlaps with the upper boundary of the current CTU.
[0422] The method according to any one of the third aspect or the implementations of the third aspect, in a possible implementation, the second predefined region includes or covers the upper left corner of the CTU including the current block.
[0423] The method according to any one of the third aspect or the implementations of the third aspect, in a possible implementation, determining whether the upper left corner of the current block overlaps with the left boundary of the CTU including the current block includes:
[0424] Obtain the abscissa position (x coordinate) of the upper left corner of the current coding unit;
[0425] Divide the obtained abscissa position (x coordinate) by the width of the CTU to calculate the remainder;
[0426] If the calculated remainder is equal to zero, infer that the upper left corner of the current coding unit overlaps with the left boundary of the current CTU;
[0427] Otherwise, infer that the upper left corner of the current coding unit does not overlap with the left boundary of the current CTU.
[0428] The method according to any one of the third aspect or the implementations of the third aspect, in a possible implementation, determining whether the upper left corner of the current coding unit overlaps with the left boundary of the CTU including the current coding unit includes:
[0429] Calculate a fourth value as the maximum floor value, where the maximum floor value is obtained by dividing the upper left abscissa (coordinate x) of the current block by the width of the CTU;
[0430] Calculate a fifth value as the maximum floor value, where the maximum floor value is obtained by dividing the upper left horizontal coordinate of an inherited adjacent block by the width of the CTU;
[0431] If the fifth value is equal to the fourth value, infer that the upper left corner of the current coding unit overlaps with the left boundary of the current CTU.
[0432] According to the method described in any one of the third aspect or the implementations of the third aspect, in a possible implementation, the determining whether the upper left corner of the current coding unit overlaps with the left boundary of the CTU including the current coding unit includes:
[0433] Calculate a sixth value as (xCb >> CtbLog2SizeX) << CtbLog2SizeX, where xCb represents the upper left horizontal coordinate (coordinate x) of the current block, ">>" represents a logical or arithmetic bit right shift, "<<" represents a logical or arithmetic bit left shift, and CtbLog2SizeX represents the binary logarithm scale of the size of the CTU;
[0434] If (xCb – 1) is less than the sixth value, infer that the upper left corner of the current coding unit overlaps with the left boundary of the current CTU.
[0435] According to the method described in any one of the third aspect or the implementations of the third aspect, in a possible implementation, it is associated with a bit mask of bit 0 including N least significant positions and a bit mask of bit 1 including other positions, rather than combining left shift and right shift operations of N bits (for example, (yCb >> CtbLog2SizeY) << CtbLog2SizeY can be calculated as the value of yCb associated with a bit mask of bit 0 including CtbLog2SizeY least significant positions and a bit mask of bit 1 including other positions).
[0436] According to the method described in any one of the third aspect or the implementations of the third aspect, in a possible implementation, the selected interpolation filter is used for reference samples to generate prediction samples at fractional positions between the reference samples; or
[0437] The selected interpolation filter is used to generate prediction samples within the current coding unit (for example, generate prediction samples of sub - blocks of the current coding unit).
[0438] A fourth aspect of the present invention provides a method for performing inter - frame prediction on a current block, including:
[0439] When conditions are met, at least two luminance positions have the same half - sample interpolation filter index and the same bi - directional prediction weight index.
[0440] According to the method described in the fourth aspect, in one possible implementation, when availableA1 is equal to true, the luma position (xNbA1, yNbA1) and (xNbB1, yNbB1), or the luma position (xNbA1, yNbA1) and (xNbB0, yNbB0), or the luma position (xNbA1, yNbA1) and (xNbA0, yNbA0), or the luma position (xNbA1, yNbA1) and (xNbA0, yNbA0), or the luma position (xNbA1, yNbA1) and (xNbA0, yNbA0), or the luma position (xNbA1, yNbA1) and (xNbB2, yNbB2) have the same bidirectional prediction weight index and the same half-sample interpolation filter index.
[0441] According to the method described in any one of the fourth aspect or the implementation manner of the fourth aspect, in one possible implementation manner, when availableB1 is equal to true, the luma position (xNbB1, yNbB1) and (xNbB0, yNbB0), or the luma position (xNbB1, yNbB1) and (xNbA0, yNbA0), or the luma position (xNbB1, yNbB1) and (xNbB2, yNbB2) have the same bidirectional prediction weight index and the same half-sample interpolation filter index.
[0442] According to the method described in any one of the fourth aspect or the implementation manner of the fourth aspect, in one possible implementation manner, when availableB0 is equal to true, the luma position (xNbB0, yNbB0) and (xNbA0, yNbA0), or the luma position (xNbB0, yNbB0) and (xNbB2, yNbB2) have the same bidirectional prediction weight index and the same half-sample interpolation filter index.
[0443] According to the method described in any one of the fourth aspect or the implementation manner of the fourth aspect, in one possible implementation manner, when availableA0 is equal to true, the luminance positions (xNbA0, yNbA0) and (xNbB2, yNbB2) have the same bidirectional prediction weight index and the same half-sample interpolation filter index.
[0444] A fifth aspect of the present invention provides a method for performing inter-frame prediction on a current block, comprising:
[0445] When the conditions are met, the MVP candidate and the fusion candidate have the same motion vector and the same reference index.
[0446] According to the fifth aspect, in a possible implementation manner, the method further includes: obtaining a half-sample interpolation filter index;
[0447] When the conditions are met, the MVP candidate and the fusion candidate have the same half-sample interpolation filter index, the same motion vector, and the same reference index.
[0448] According to the method in any one of the fifth aspect or the implementation manner of the fifth aspect, in a possible implementation manner, the method further includes: obtaining a half-sample interpolation filter index;
[0449] When the condition involving the half-sample interpolation filter index is met, the MVP candidate and the fused candidate have the same motion vector and the same reference index.
[0450] The following is a detailed description of a possible implementation of deriving motion information including an interpolation filter set index of the block according to the fusion candidate list in the proposed method in a format modified from the specification of the VVC working draft (see Fig.14 1403 of method 1400 shown in FIG. 1404 ). The modification is shown in a highlighted manner.
[0451] 8.5.2 Derivation of motion vector components and reference indices
[0452] 8.5.2.1 Overview
[0453] Inputs to this process include:
[0454] – The luma position (xCb, yCb) of the top-left sample of the current luma coding block relative to the top-left luma sample of the current image;
[0455] – The variable cbWidth indicates the width of the current coding block in luma samples;
[0456] – The variable cbHeight represents the height of the current coded block in luma samples.
[0457] The outputs of this process include:
[0458] – Luma motion vectors mvL0[0][0] and mvL1[0][0] with 1 / 16 fractional sample accuracy,
[0459] – reference indexes refIdxL0 and refIdxL1,
[0460] – The prediction list uses flags predFlagL0[0][0] and predFlagL1[0][0],
[0461] – Half-sample interpolation filter index hpelIfIdx,
[0462] – Bidirectional prediction weight index bcwIdx.
[0463] Let variable LX be the RefPicList[X] of the current image, where X is 0 or 1.
[0464] For the derivation of the variables mvL0[0][0], mvL1[0][0], refIdxL0, refIdxL1, and predFlagL0[0][0] and predFlagL1[0][0], the following applies:
[0465] – If general_merge_flag[xCb][yCb] is equal to 1, the derivation process of the luma motion vector in merge mode as specified in clause 8.5.2.2 is called with the luma position (xCb, yCb), the input variables cbWidth and cbHeight, the output luma motion vectors mvL0[0][0] and mvL1[0][0], the reference indices refIdxL0 and refIdxL1, the prediction list usage flags predFlagL0[0][0] and predFlagL1[0][0], the half-sample interpolation filter index hpelIfIdx, the bidirectional prediction weight index bcwIdx and the merge candidate list mergeCandList.
[0466] – Otherwise, the following applies:
[0467] – For variables predFlagLX[0][0], mvLX[0][0] and X in refIdxLX, PRED_LX, and syntax elements ref_idx_lX and MvdLX, where X is replaced by 0 or 1, the following ordered steps apply:
[0468] 1. The variables refIdxLX and predFlagLX[0][0] are derived as follows:
[0469] – If inter_pred_idc[xCb][yCb] is equal to PRED_LX or PRED_BI,
[0470] refIdxLX=ref_idx_lX[xCb][yCb] (8-292),
[0471] predFlagLX[0][0] = 1 (8-293),
[0472] – Otherwise, the variables refIdxLX and predFlagLX[0][0] are represented as:
[0473] refIdxLX = –1 (8-294),
[0474] predFlagLX[0][0] = 0 (8-295),
[0475] 2. The variable mvdLX is derived as follows:
[0476] mvdLX[0]=MvdLX[xCb][yCb][0] (8-296),
[0477] mvdLX[1]=MvdLX[xCb][yCb][1] (8-297),
[0478] 3. When predFlagLX[0][0] is equal to 1, the derivation process of the luma motion vector prediction specified in clause 8.5.2.8 is called simultaneously with the luma coding block position (xCb, yCb), coding block width cbWidth, coding block height cbHeight, variable refIdxLX as input and mvpLX as output.
[0479] 4. When predFlagLX[0][0] is equal to 1, the luminance motion vector mvLX[0][0] is derived as follows:
[0480] uLX[0]=(mvpLX[0]+mvdLX[0]+2 18 )%2 18 (8-298),
[0481] mvLX[0][0][0]=(uLX[0]>=2 17 ) ? (uLX[0]–2 18 ): uLX[0] (8-299),
[0482] uLX[1]=(mvpLX[1]+mvdLX[1]+2 18 )%2 18 (8-300),
[0483] mvLX[0][0][1]=(uLX[1]>=2 17 ) ? (uLX[1]–2 18 ): uLX[1](8-301)
[0484] Note 1: The result values of mvLX[0][0][0] and mvLX[0][0][1] set above are always between -2. 17 to (2 17 –1) (including the first and last digits).
[0485] – The half-sample interpolation filter index hpelIfIdx is derived as follows:
[0486] hpelIfIdx=AmvrShift== 3 ? 1: 0 (8-302),
[0487] – The bidirectional prediction weight index bcwIdx is set equal to bcw_idx[xCb][yCb].
[0488] refIdxL1 is set equal to –1, predFlagL1 is set equal to 0, and bcwIdx is set equal to 0 when all of the following conditions are true:
[0489] –predFlagL0[0][0] is equal to 1;
[0490] –predFlagL1[0][0] is equal to 1;
[0491] – The value of (cbWidth + cbHeight) is equal to 12.
[0492] The update process of the history-based motion vector prediction list specified in clause 8.5.2.16 is consistent with the luminance motion vectors mvL0[0][0] and mvL1[0][0], the reference indices refIdxL0 and refIdxL1, the prediction list usage flags predFlagL0[0][0] and predFlagL1[0][0], the bidirectional prediction weight index bcwIdx, and Half-sample interpolation filter Index hpelIfIdx are called at the same time.
[0493] It is understood that this method is applicable to both unidirectional prediction and bidirectional prediction. It is understood that two reference indexes and two prediction list usage flags are proposed in the VVC working draft specification. However, when the unidirectional prediction predFlagL1 is set equal to 0 (indicating that L1 prediction is not used), refIdxL1 is set equal to –1.
[0494] The following is a detailed description of possible implementations of the proposed method for deriving history-based fusion candidates in the form of modifications to the VVC draft specification (see Fig.14 1402 of method 1400 shown in FIG. 1403 ). The modification is shown in a highlighted manner.
[0495] 8.5.2.6 Derivation of History-Based Fusion Candidates
[0496] Inputs to this process include:
[0497] –Merge candidate list mergeCandList,
[0498] – The number of available merge candidates in the list, numCurrMergeCand.
[0499] The outputs of this process include:
[0500] – The modified fusion candidate list mergeCandList,
[0501] – The number of modified fusion candidates numCurrMergeCand in the list.
[0502] The variables isPrunedA1 and isPrunedB1 are both set equal to false.
[0503] For each candidate with index hMvpIdx=1..NumHmvpCand in HmvpCandList[hMvpIdx], repeat the following ordered steps until numCurrMergeCand equals (MaxNumMergeCand–1).
[0504] 1. The variable sameMotion is derived as follows:
[0505] – For any fusion candidate N (where N is A1 or B1), if all of the following conditions are true, then sameMotion and isPrunedN are both set equal to true:
[0506] –hMvpIdx is less than or equal to 2.
[0507] –HmvpCandList[NumHmvpCandh–MvpIdx] and fusion candidate N have the same motion vector and the same Reference Index .
[0508] –isPrunedN is equal to false.
[0509] – Otherwise, sameMotion is set equal to false.
[0510] 2. When sameMotion is equal to false, the candidate HmvpCandList[NumHmvpCand–hMvpIdx] is added to the fusion candidate list:
[0511] mergeCandList[numCurrMergeCand++]=HmvpCandList[NumHmvpCand–hMvpIdx](8-381)
[0512] The following is a detailed description of the first possible implementation of the proposed method for updating the history-based motion information (HMVP) candidate list in the format modified from the VVC draft specification (see Fig.13Aand Fig. 13B The modification is displayed in a highlighted manner.
[0513] 8.5.2.16 Update process of motion vector prediction candidate list based on history
[0514] Inputs to this process include:
[0515] – Luma motion vectors mvL0 and mvL1 with 1 / 16 fractional sample accuracy,
[0516] – reference indexes refIdxL0 and refIdxL1,
[0517] – The prediction list uses the flags predFlagL0 and predFlagL1,
[0518] – Bidirectional prediction weight index gbiIdx,
[0519] – Half-sample interpolation filter set index hpelIfIdx.
[0520] The MVP candidate hMvpCand is composed of the luma motion vectors mvL0 and mvL1, the reference indexes refIdxL0 and refIdxL1, the prediction list usage flags predFlagL0 and predFlagL1, the bidirectional prediction weight index gbiIdx, and the Half-sample interpolation filter set index hpelIfIdx composition.
[0521] Use the candidate hMvpCand to modify the candidate list HmvpCandList through the following ordered steps:
[0522] 1. The variable identicalCandExist is set equal to false, and the variable removeIdx is set equal to 0.
[0523] 2. When NumHmvpCand is greater than 0, for each index hMvpIdx of hMvpIdx=0..NumHmvpCand–1, perform the following steps until identicalCandExist is equal to true:
[0524] – When hMvpCand is equal to HmvpCandList[hMvpIdx], identicalCandExist is set equal to true and removeIdx is set equal to hMvpIdx.
[0525] 3. The candidate list HmvpCandList is updated as follows:
[0526] – If identicalCandExist equals true, or NumHmvpCand equals (MaxNumMergeCand–1), then the following applies:
[0527] – For each index i where i = (removeIdx+1)..(NumHmvpCand–1), HmvpCandList[i–1] is set equal to HmvpCandList[i].
[0528] –HmvpCandList[NumHmvpCand–1] is set equal to mvCand.
[0529] – Otherwise (identicalCandExist is equal to false and NumHmvpCand is less than (MaxNumMergeCand–1)), the following applies:
[0530] –HmvpCandList[NumHmvpCand++] is set equal to mvCand.
[0531] The second possible implementation of updating the history-based motion information (HMVP) candidate list in the proposed method is described in detail below in the format of modifying the VVC draft specification (see Fig.13A and Fig. 13B Steps 1303 and 1313 shown, and see Fig.15 The modification is shown in steps 1502 to 1504.
[0532] 8.5.2.16 Update process of motion vector prediction candidate list based on history
[0533] Inputs to this process include:
[0534] – Luma motion vectors mvL0 and mvL1 with 1 / 16 fractional sample accuracy,
[0535] – reference indexes refIdxL0 and refIdxL1,
[0536] – The prediction list uses the flags predFlagL0 and predFlagL1,
[0537] – Bidirectional prediction weight index bcwIdx,
[0538] – Half-sample interpolation filter index hpelIfIdx .
[0539] The MVP candidate hMvpCand is composed of the luma motion vectors mvL0 and mvL1, the reference indexes refIdxL0 and refIdxL1, the prediction list usage flags predFlagL0 and predFlagL1, the bidirectional prediction weight index bcwIdx, and the Half-sample interpolation filter index hpelIfIdx composition.
[0540] Use the candidate hMvpCand to modify the candidate list HmvpCandList through the following ordered steps:
[0541] 1. The variable identicalCandExist is set equal to false, and the variable removeIdx is set equal to 0.
[0542] 2. When NumHmvpCand is greater than 0, for each index hMvpIdx of hMvpIdx=0..NumHmvpCand–1, perform the following steps until identicalCandExist is equal to true:
[0543] – When hMvpCand and HmvpCandList[hMvpIdx] have the same motion vector and the same reference index hour , identicalCandExist is set equal to true, and removeIdx is set equal to hMvpIdx.
[0544] 3. The candidate list HmvpCandList is updated as follows:
[0545] – If identicalCandExist equals true, or NumHmvpCand equals 5, then the following applies:
[0546] – For each index i where i = (removeIdx+1)..(NumHmvpCand–1), HmvpCandList[i–1] is set equal to HmvpCandList[i].
[0547] –HmvpCandList[NumHmvpCand–1] is set equal to hMvpCand.
[0548] – Otherwise (identicalCandExist is equal to false and NumHmvpCand is less than 5), the following applies:
[0549] –HmvpCandList[NumHmvpCand++] is set equal to hMvpCand.
[0550] From the above description, it can be seen that the second implementation specifies the comparison elements (i) and (ii) of the HMVP candidate, while the first implementation specifies all the comparison elements of the HMVP candidate (eg, elements (i), (ii), and (iii)).
[0551] The embodiments of the present invention and exemplary embodiments have respective methods and corresponding devices.
[0552] Fig.16 1 is a schematic diagram of an apparatus 1600 for constructing a history-based motion information candidate list. The apparatus comprises an HMI list obtaining unit 1601 and an HMI list updating unit 1603 .
[0553] The history-based motion information (HMI) candidate list obtaining unit 1601 is used to obtain a history-based motion information candidate list, wherein the HMI list is N history-based motion information candidate HMIs. k An ordered list of (k=0, ..., N-1) of N history-based motion information candidates H k Associated with motion information of multiple blocks before a block, N is an integer greater than 0, and each history-based motion information candidate includes the elements:
[0554] (iv) one or more motion vectors MV,
[0555] (v) one or more reference image indices corresponding to the MV,
[0556] (vi) Interpolation filter index.
[0557] The history-based motion information candidate list updating unit 1603 is used to update the HMI list according to the motion information of the block, wherein the motion information of the block includes the following elements:
[0558] (iv) one or more motion vectors MV,
[0559] (v) one or more reference image indices corresponding to the MV,
[0560] (vi) Interpolation filter index.
[0561] It can be understood that the HMI list obtaining unit 1601 and the HMI list updating unit 1603 (corresponding to the inter-frame prediction module) in the encoder 20 or the decoder 30 provided in the embodiment of the present application are functional entities that implement various execution steps included in the above-mentioned corresponding method, that is, functional entities that fully implement the steps in the method of the present application and the extension and variation of these steps. For details, please refer to the description in the above-mentioned corresponding method. For the sake of brevity, it will not be repeated here.
[0562] Fig.17 Schematic diagram of an inter-frame prediction device 1700 provided in an embodiment of the present invention. The device 1700 is used to determine motion information of a current block within a frame. The device 1700 includes:
[0563] The list management unit 1701 is used to construct the HMVP list, wherein the HMVP list is a list of N candidate HMVPs based on history. k An ordered list of (k=0, ..., N-1) of N history-based candidates H k Associated with motion information of N previous blocks in a frame before the current block, N is greater than or equal to 1, each or at least one history-based candidate includes motion information, and the motion information includes elements: (i) one or more motion vectors MV, (ii) one or more reference image indexes corresponding to the MV, (iii) an interpolation filter index (such as a half-pixel interpolation filter index) or an interpolation filter set index; the HMVP list management unit 1701 is also used to add one or more history-based candidates in the HMVP list to the motion information candidate list of the current block; the information derivation unit 1703 is used to derive the motion information according to the motion information candidate list.
[0564] In one implementation, the list management unit 1701 is used to compare at least one element of each history-based candidate element in the HMI list with a corresponding element of the current block. The motion information adding unit is used to: if the result of the comparison is that at least one element of each history-based candidate element in the HMI list is different from a corresponding element in the motion information of the current block, add the motion information of the current block to the HMI list.
[0565] Correspondingly, in one example, the example structure of the apparatus 1700 may correspond to Figure 2 In another example, the example structure of the device 1700 may correspond to Figure 3 The decoder 30 in FIG.
[0566] In another example, the example structure of the apparatus 1700 may correspond to Figure 2 In another example, the example structure of the apparatus 1700 may correspond to Figure 3 The inter-frame prediction unit 344 in.
[0567] It can be understood that the list management unit 1701 and the information derivation unit 1703 (corresponding to the inter-frame prediction module) in the encoder 20 or decoder 30 provided in the embodiment of the present application are functional entities that implement various execution steps included in the above-mentioned corresponding method, that is, functional entities that fully implement the steps in the method of the present application and the extension and variation of these steps. For details, please refer to the description in the above-mentioned corresponding method. For the sake of brevity, it will not be repeated here.
[0568] More specifically, the following describes aspects related to SIF index propagation across CTU boundaries:
[0569] As described above, based on the current SIF design, when the SIF technology is applied to the mode of inheriting motion information from the upper spatial neighboring blocks, if the current block is located at the upper boundary of the CTU / CTB, the row memory is increased. As described in this article, the position of the current block is checked. If the current block is located at the upper boundary of the CTU / CTB, when inheriting motion information from the upper left neighboring block (B0), the upper neighboring block (B1), and the upper right neighboring block (B2), the IF index is not inherited from the neighboring blocks, but the default value is used, thereby reducing the occupancy of the row memory.
[0570] One aspect of the present invention provides a method for performing inter-frame prediction on a current block, comprising:
[0571] Inter-frame prediction is performed on the block, including deriving an interpolation filter index for the current block based on a position of the current block (e.g., coding unit or coding block) within a coding tree block (CTB) or coding tree unit (CTU) and an interpolation filter index inherited from a selected fusion candidate.
[0572] Figure 8 A flowchart of a method for deriving an interpolation filter set index for a current block (e.g., a coding unit or a coding block) within a coding tree block (CTB) or a coding tree unit (CTU). The method comprises:
[0573] Step 803: The method includes: determining whether the current block overlaps with a predefined area of the CTB or CTU (eg, an upper boundary or a left boundary of the CTB or CTU).
[0574] Step 804: The method includes: if the current block does not overlap with a predefined region of the CTU (e.g., the current block does not overlap with the upper boundary or left boundary of the CTB or CTU), setting the interpolation filter set index of the current block to the selected candidate interpolation filter set index. The selected candidate can be, for example, a selected merge candidate or a selected MVP candidate. The selected candidate can also be an adjacent block corresponding to the selected merge candidate.
[0575] Step 805: The method includes: if the current block overlaps with a predefined region of the CTB or CTU (e.g., the upper boundary or left boundary of the CTB or CTU), setting the interpolation filter set index of the current block to a predefined value.
[0576] Further, steps 801 to 802: The method includes constructing a candidate list. For the sake of brevity, it will not be elaborated here.
[0577] As Fig. 9 shown, to determine whether the current block is located at the upper boundary of CTU 900, the ordinate (yCb) of the upper left corner of the current block is checked. Assuming the size of CTU 900 is equal to (1<<CtbLog2SizeY) x (1<<CtbLog2SizeY), if ((yCb >> CtbLog2SizeY) << CtbLog2SizeY) is not equal to yCb, then the current block is not located at the upper boundary of CTU 900 (Scenario 1). Otherwise (if ((yCb >> CtbLog2SizeY) << CtbLog2SizeY) is equal to yCb), then the current block is located at the upper boundary of CTU 900 (Scenario 2).
[0578] In an embodiment of the present invention, the selected candidate (e.g., the selected merge candidate) is a spatial merge candidate.
[0579] In an embodiment of the present invention, the ordinate position associated with the spatial merge candidate is less than the ordinate position of the current block, or the ordinate position of the adjacent block corresponding to the spatial merge candidate is less than the ordinate position of the current block.
[0580] In an embodiment of the present invention, the spatial merge candidate is a top - right candidate (such as B0 shown in Figure 6 ), a top candidate (such as B1 shown in Figure 6 ), or a top - left candidate (such as B2 shown in Figure 6 ).
[0581] In one embodiment of the present invention, the fusion candidate is an affine fusion candidate. The affine fusion candidate is an inherited affine fusion candidate, where "inherited" means: (i) the candidate is derived from adjacent affine blocks; (ii) the affine model of the current block is inherited from the affine model of the adjacent affine block; or (iii) the affine parameters of the current block are derived from the affine parameters of the adjacent affine blocks.
[0582] In one embodiment of the present invention, the inherited affine fusion candidate is derived based on one of the spatially adjacent blocks, wherein the spatially adjacent blocks include the lower left block (e.g. Figure 6 A0 shown), left block (such as Figure 6 A1 shown), the upper right block (such as Figure 6 B0 shown), upper block (such as Figure 6 B1 as shown) or the upper left block (as Figure 6 B2 shown).
[0583] In one embodiment of the present invention, the inherited affine fusion candidate is derived based on a block whose ordinate position is smaller than the ordinate position of the current block.
[0584] In one embodiment of the present invention, the inherited affine fusion candidate is based on the upper right block (such as Figure 6 B0 shown), upper block (such as Figure 6 B1 as shown) or the upper left block (as Figure 6 B2) shown in the figure is derived.
[0585] In one embodiment of the present invention, the selected candidate (eg, the selected fusion candidate) is a sub-block fusion candidate.
[0586] In one embodiment of the present invention, the predefined area of the CTU overlaps with the CTB or CTU.
[0587] In one embodiment of the present invention, whether the current block overlaps with the predefined area is determined according to the coordinate position of the upper left corner of the coding unit (for example, the horizontal coordinate position and the vertical coordinate position of the upper left sample of the current block) (for example, used to represent the brightness position (xCb, yCb) of the upper left sample of the current block relative to the upper left brightness sample of the current image).
[0588] In one embodiment of the present invention, if the upper left corner of the current block overlaps with a second predefined area (eg, the upper left corner of the CTU), it is inferred that the current block overlaps with the predefined area (eg, the upper boundary of the CTU).
[0589] In one embodiment of the present invention, the second predefined area includes or covers the upper boundary of the CTU including the current block (for example, the upper left corner of the CTU includes or covers the upper boundary or left boundary of the CTU including the current block).
[0590] In one embodiment of the present invention, determining whether the current block overlaps with a predefined area of the CTB or CTU includes determining whether the upper left corner of the current block (for example, a luminance position (xCb, yCb) used to represent the upper left sample of the current block relative to the upper left luminance sample of the current image) overlaps with an upper boundary of the CTU containing the current coding unit.
[0591] In one embodiment of the present invention, determining whether the upper left corner of the current block overlaps with the upper boundary of the CTU including the current coding unit includes:
[0592] Get the y coordinate position (y coordinate) of the upper left corner of the current block;
[0593] Divide the obtained ordinate position (y coordinate) by the height of the CTU to obtain a remainder;
[0594] If the calculated remainder is equal to zero, it is inferred that the upper left corner of the current block overlaps with the upper boundary of the current CTU; otherwise, it is inferred that the upper left corner of the current coding unit does not overlap with the upper boundary of the current CTU.
[0595] In one embodiment of the present invention, determining whether the upper left corner of the current coding unit overlaps with the upper boundary of the CTU including the current coding unit includes:
[0596] Calculate a first value as a maximum floor-rounded value, wherein the maximum floor-rounded value is obtained by dividing the upper left ordinate (coordinate y) of the current block by the height of the CTU;
[0597] Calculate a second value as a maximum floor-rounded value, wherein the maximum floor-rounded value is obtained by dividing the upper left vertical coordinate of the inherited adjacent block by the height of the CTU;
[0598] If the second value is equal to the first value, it is inferred that the upper left corner of the current coding unit overlaps with the upper boundary of the current CTU.
[0599] In one embodiment of the present invention, determining whether the upper left corner of the current coding unit overlaps with the upper boundary of the CTU including the current coding unit includes:
[0600] Calculate the third value as (yCb >> CtbLog2SizeY) << CtbLog2SizeY, where yCb represents the upper left vertical coordinate (coordinate y) of the current block, ">>" represents logical or arithmetic bitwise right shift, "<< " represents logical or arithmetic bitwise left shift, and CtbLog2SizeY represents the binary logarithm scale of the size of the CTU or CTB;
[0601] If (yCb – 1) is less than the third value, infer that the upper left corner of the current coding unit overlaps with the upper boundary of the current CTU.
[0602] In one embodiment of the present invention, the second predefined region includes or covers the upper left corner of the CTU containing the current block.
[0603] In one embodiment of the present invention, determining whether the upper left corner of the current block overlaps with the left boundary of the CTU containing the current block includes:
[0604] Obtain the horizontal coordinate position (x coordinate) of the upper left corner of the current coding unit;
[0605] Divide the obtained horizontal coordinate position (x coordinate) by the width of the CTU to calculate the remainder;
[0606] If the calculated remainder is equal to zero, infer that the upper left corner of the current coding unit overlaps with the left boundary of the current CTU;
[0607] Otherwise, infer that the upper left corner of the current coding unit does not overlap with the left boundary of the current CTU.
[0608] In one embodiment of the present invention, determining whether the upper left corner of the current coding unit overlaps with the left boundary of the CTU containing the current coding unit includes:
[0609] Calculate the fourth value as the maximum floor value, where the maximum floor value is obtained by dividing the upper left horizontal coordinate (coordinate x) of the current block by the width of the CTU;
[0610] Calculate the fifth value as the maximum floor value, where the maximum floor value is obtained by dividing the upper left vertical coordinate of the inherited adjacent block by the width of the CTU;
[0611] If the fifth value is equal to the fourth value, infer that the upper left corner of the current coding unit overlaps with the left boundary of the current CTU.
[0612] In one embodiment of the present invention, determining whether the upper left corner of the current coding unit overlaps with the left boundary of the CTU containing the current coding unit includes:
[0613] Calculate the sixth value as (xCb >> CtbLog2SizeX) << CtbLog2SizeX, where xCb represents the abscissa of the upper left corner (coordinate x) of the current block, ">>" represents logical or arithmetic bit right shift, "<< " represents logical or arithmetic bit left shift, and CtbLog2SizeX represents the binary logarithm scale of the size of the CTU;
[0614] If (xCb – 1) is less than the sixth value, infer that the upper left corner of the current coding unit overlaps with the left boundary of the current CTU.
[0615] In an embodiment of the present invention, it is associated with a bit mask of bit 0 including N least significant positions and a bit mask of bit 1 including other positions, rather than combining N-bit left shift and right shift operations (for example, (yCb >> CtbLog2SizeY) << CtbLog2SizeY can be calculated as the value of yCb associated with a bit mask of bit 0 including CtbLog2SizeY least significant positions and a bit mask of bit 1 including other positions). For example, if yCb is in the range of [0, 2 32 –1], and CtbLog2SizeY is equal to 7, the value of yCb & 0xFFFFFF80 can be calculated instead of the value of (yCb >> CtbLog2SizeY) << CtbLog2SizeY. Here, 0xFFFFFF80 is a bit mask, which includes 0 in 7 least significant positions and 1 in other positions.
[0616] In an embodiment of the present invention, logical shift or arithmetic shift is used to calculate the maximum floor value of the division result. (For example, the maximum floor value of a / 2 n can be calculated as a >> n).
[0617] In an embodiment of the present invention, the second predefined region only includes or covers the upper boundary of the CTU containing the current block.
[0618] In an embodiment of the present invention, the second predefined region only includes or covers the left boundary of the CTU containing the current block.
[0619] In an embodiment of the present invention, the second predefined region only includes or covers the upper boundary and the left boundary of the CTU containing the current block.
[0620] In an embodiment of the present invention, setting the interpolation filter set index of the current block to a predefined value includes:
[0621] The interpolation filter set index of the current block is set to a seventh value, wherein the seventh value is determined before constructing a merge list.
[0622] In one embodiment of the present invention, determining the seventh value includes:
[0623] An interpolation filter set index of one of the spatial neighboring blocks of the current block is determined, and the seventh value is set equal to the determined interpolation filter set index.
[0624] In one embodiment of the present invention, "one of the spatially adjacent blocks" refers to the left adjacent block (the block in Figure 6 A1 in the figure).
[0625] The following is a detailed description of possible implementations of the SIF index propagation across CTU boundaries in the proposed method in the form of modifications to the specification of the working draft of the SIF proposal (the process is as follows Figure 8 and 9 The modifications are highlighted.
[0626] 8.5.2.3 Derivation process of spatial fusion candidates
[0627] Inputs to this process include:
[0628] – The luma position (xCb, yCb) of the top-left sample of the current luma coding block relative to the top-left luma sample of the current image;
[0629] – The variable cbWidth indicates the width of the current coding block in luma samples;
[0630] – The variable cbHeight represents the height of the current coded block in luma samples.
[0631] The output of this process includes, where X is either 0 or 1:
[0632] – Available flags availableFlagA0, availableFlagA1, availableFlagB0, availableFlagB1 and availableFlagB2 of adjacent coding units,
[0633] - the reference indices refIdxLXA0, refIdxLXA1, refIdxLXB0, refIdxLXB1 and refIdxLXB2 of the neighboring coding units,
[0634] – the prediction list of the neighboring coding units uses the flags predFlagLXA0, predFlagLXA1, predFlagLXB0, predFlagLXB1 and predFlagLXB2,
[0635] – the motion vectors mvLXA0, mvLXA1, mvLXB0, mvLXB1 and mvLXB2 of the adjacent coding units with 1 / 16 fractional sample accuracy,
[0636] – half-sample interpolation filter indices hpelIfIdxA0, hpelIfIdxA1, hpelIfIdxB0, hpelIfIdxB1, and hpelIfIdxB2,
[0637] – Bidirectional prediction weight indices gbiIdxA0, gbiIdxA1, gbiIdxB0, gbiIdxB1, and gbiIdxB2.
[0638] For the derivation of availableFlagA1, refIdxLXA1, predFlagLXA1, and mvLXA1, the following applies:
[0639] – The luma position (xNbA1, yNbA1) within the adjacent luma coding block is set equal to (xCb–1, yCb+cbHeight–1).
[0640] – The block’s availability derivation procedure specified in clause 6.4 is called with the current luma position (xCurr, yCurr) set equal to (xCb, yCb) and the neighboring luma position (xNbA1, yNbA1) as input, and the output is assigned to the block’s availability flag availableA1.
[0641] – The variables availableFlagA1, refIdxLXA1, predFlagLXA1, and mvLXA1 are derived as follows:
[0642] – If availableA1 is equal to false, availableFlagA1 is set equal to 0, both components of mvLXA1 are set equal to 0, refIdxLXA1 is set equal to –1, predFlagLXA1 is set equal to 0, and gbiIdxA1 is set equal to 0, where X is either 0 or 1.
[0643] – Otherwise, availableFlagA1 is set equal to 1 and the following assignment operation is performed:
[0644] mvLXA1=MvLX[xNbA1][yNbA1] (8-294),
[0645] refIdxLXA1=RefIdxLX[xNbA1][yNbA1] (8-295),
[0646] predFlagLXA1=PredFlagLX[xNbA1][yNbA1] (8-296),
[0647] hpelIfIdxA1=HpelIfIdx[xNbA1][yNbA1] (8-297),
[0648] gbiIdxA1=GbiIdx[xNbA1][yNbA1] (8-298),
[0649] For the derivation of availableFlagB1, refIdxLXB1, predFlagLXB1, and mvLXB1, the following applies:
[0650] – The luma position (xNbB1, xNbB1) within the adjacent luma coding block is set equal to (xCb+cbWidth–1, yCb–1).
[0651] – The block’s availability derivation procedure specified in clause 6.4 is called with the current luma position (xCurr, yCurr) set equal to (xCb, yCb) and the neighboring luma position (xNbB1, yNbB1) as input, and the output is assigned to the block’s availability flag availableB1.
[0652] – The variables availableFlagB1, refIdxLXB1, predFlagLXB1, and mvLXB1 are derived as follows:
[0653] – availableFlagB1 is set equal to 0, both components of mvLXB1 are set equal to 0, refIdxLXB1 is set equal to –1, predFlagLXB1 is set equal to 0, and gbiIdxB1 is set equal to 0 if one or more of the following conditions are true, where X is 0 or 1:
[0654] –availableB1 is equal to false.
[0655] –availableA1 is equal to true, and luma positions (xNbA1, yNbA1) and (xNbB1, yNbB1) have the same motion vector and the same reference index.
[0656] – Otherwise, availableFlagB1 is set equal to 1 and the following assignment operation is performed:
[0657] mvLXB1=MvLX[xNbB1][yNbB1] (8-299),
[0658] refIdxLXB1=RefIdxLX[xNbB1][yNbB1] (8-300),
[0659] predFlagLXB1=PredFlagLX[xNbB1][yNbB1] (8-301),
[0660] If (yCb–1)<((yCb>>CtbLog2SizeY)< <CtbLog2SizeY),
[0661] hpelIfIdxB1=2 (8-302),
[0662] otherwise,
[0663] hpelIfIdxB1=HpelIfIdx[xNbB1][yNbB1] (8-303),
[0664] gbiIdxB1=GbiIdx[xNbB1][yNbB1] (8-304),
[0665] For the derivation of availableFlagB0, refIdxLXB0, predFlagLXB0, and mvLXB0, the following applies:
[0666] – The luma position (xNbB0, yNbB0) within the adjacent luma coding block is set equal to (xCb+cbWidth, yCb–1).
[0667] – The block’s availability derivation procedure specified in clause 6.4 is called with the current luma position (xCurr, yCurr) set equal to (xCb, yCb) and the neighboring luma position (xNbB0, yNbB0) as input, and the output is assigned to the block’s availability flag availableB0.
[0668] – The variables availableFlagB0, refIdxLXB0, predFlagLXB0, and mvLXB0 are derived as follows:
[0669] – availableFlagB0 is set equal to 0, both components of mvLXB0 are set equal to 0, refIdxLXB0 is set equal to –1, predFlagLXB0 is set equal to 0, and gbiIdxB0 is set equal to 0 if one or more of the following conditions are true, where X is 0 or 1:
[0670] –availableB0 is equal to false.
[0671] –availableB1 is equal to true, and luma positions (xNbB1, yNbB1) and (xNbB0, yNbB0) have the same motion vector and the same reference index.
[0672] – availableA1 is equal to true, luma positions (xNbA1, yNbA1) and (xNbB0, yNbB0) have the same motion vector and the same reference index, and merge_triangle_flag[xCb][yCb] is equal to 1.
[0673] – Otherwise, availableFlagB0 is set equal to 1 and the following assignment operation is performed:
[0674] mvLXB0=MvLX[xNbB0][yNbB0] (8-305),
[0675] refIdxLXB0=RefIdxLX[xNbB0][yNbB0] (8-306),
[0676] predFlagLXB0=PredFlagLX[xNbB0][yNbB0] (8-307),
[0677] If (yCb–1)<((yCb>>CtbLog2SizeY)< <CtbLog2SizeY),
[0678] hpelIfIdxB0=2 (8-308),
[0679] otherwise,
[0680] hpelIfIdxB0=HpelIfIdx[xNbB0][yNbB0] (8-309),
[0681] gbiIdxB0=GbiIdx[xNbB0][yNbB0] (8-310),
[0682] For the derivation of availableFlagA0, refIdxLXA0, predFlagLXA0, and mvLXA0, the following applies:
[0683] – The luma position (xNbA0, yNbA0) within the adjacent luma coding block is set equal to (xCb–1, yCb+cbWidth).
[0684] – The block’s availability derivation procedure specified in clause 6.4 is called with the current luma position (xCurr, yCurr) set equal to (xCb, yCb) and the neighboring luma position (xNbA0, yNbA0) as input, and the output is assigned to the block’s availability flag availableA0.
[0685] – The variables availableFlagA0, refIdxLXA0, predFlagLXA0, and mvLXA0 are derived as follows:
[0686] – availableFlagA0 is set equal to 0, both components of mvLXA0 are set equal to 0, refIdxLXA0 is set equal to –1, predFlagLXA0 is set equal to 0, and gbiIdxA0 is set equal to 0 if one or more of the following conditions are true, where X is 0 or 1:
[0687] –availableA0 is equal to false.
[0688] –availableA1 is equal to true, and luma positions (xNbA1, yNbA1) and (xNbA0, yNbA0) have the same motion vector and the same reference index.
[0689] – availableB1 is equal to true, luma positions (xNbB1, yNbB1) and (xNbA0, yNbA0) have the same motion vector and the same reference index, and merge_triangle_flag[xCb][yCb] is equal to 1.
[0690] – availableB0 is equal to true, luma positions (xNbB0, yNbB0) and (xNbA0, yNbA0) have the same motion vector and the same reference index, and merge_triangle_flag[xCb][yCb] is equal to 1.
[0691] – Otherwise, availableFlagA0 is set equal to 1 and the following assignment operation is performed:
[0692] mvLXA0=MvLX[xNbA0][yNbA0] (8-311),
[0693] refIdxLXA0=RefIdxLX[xNbA0][yNbA0] (8-312),
[0694] predFlagLXA0=PredFlagLX[xNbA0][yNbA0] (8-313),
[0695] hpelIfIdxA0=HpelIfIdx[xNbA0][yNbA0] (8-314),
[0696] gbiIdxA0=GbiIdx[xNbA0][yNbA0] (8-315),
[0697] For the derivation of availableFlagB2, refIdxLXB2, predFlagLXB2, and mvLXB2, the following applies:
[0698] – The luma position (xNbB2, yNbB2) within the adjacent luma coding block is set equal to (xCb–1, yCb–1).
[0699] – The block’s availability derivation procedure specified in clause 6.4 is called with the current luma position (xCurr, yCurr) set equal to (xCb, yCb) and the neighboring luma position (xNbB2, yNbB2) as input, and the output is assigned to the block’s availability flag availableB2.
[0700] – The variables availableFlagB2, refIdxLXB2, predFlagLXB2, and mvLXB2 are derived as follows:
[0701] – availableFlagB2 is set equal to 0, both components of mvLXB2 are set equal to 0, refIdxLXB2 is set equal to –1, predFlagLXB2 is set equal to 0, and gbiIdxB2 is set equal to 0 if one or more of the following conditions are true, where X is 0 or 1:
[0702] –availableB2 is equal to false.
[0703] –availableA1 is equal to true, and luma positions (xNbA1, yNbA1) and (xNbB2, yNbB2) have the same motion vector and the same reference index.
[0704] –availableB1 is equal to true, and luma positions (xNbB1, yNbB1) and (xNbB2, yNbB2) have the same motion vector and the same reference index.
[0705] – availableB0 is equal to true, luma positions (xNbB0, yNbB0) and (xNbB2, yNbB2) have the same motion vector and the same reference index, and merge_triangle_flag[xCb][yCb] is equal to 1.
[0706] – availableA0 is equal to true, luma positions (xNbA0, yNbA0) and (xNbB2, yNbB2) have the same motion vector and the same reference index, and merge_triangle_flag[xCb][yCb] is equal to 1.
[0707] –availableFlagA0+availableFlagA1+availableFlagB0+availableFlagB1=4, merge_triangle_flag[xCb][yCb]=0.
[0708] – Otherwise, availableFlagB2 is set equal to 1 and the following assignment operation is performed:
[0709] mvLXB2=MvLX[xNbB2][yNbB2] (8-316),
[0710] refIdxLXB2=RefIdxLX[xNbB2][yNbB2] (8-317),
[0711] predFlagLXB2=PredFlagLX[xNbB2][yNbB2] (8-318),
[0712] If (yCb–1)<((yCb>>CtbLog2SizeY)< <CtbLog2SizeY),
[0713] hpelIfIdxB2=2 (8-319)
[0714] otherwise,
[0715] hpelIfIdxB2=HpelIfIdx[xNbB2][yNbB2] (8-320),
[0716] gbiIdxB2=GbiIdx[xNbB2][yNbB2] (8-321),
[0717] As can be seen from the above, the half-pixel interpolation filter index of the neighboring block of the current block is determined according to whether the boundary of the current block overlaps with the CTU. For example: Equations (8-299) to (8-304) show the steps for determining the half-pixel interpolation filter index of the neighboring block B1 (such as Figure 8 and Fig. 9 Equations (8-305) to (8-310) show the steps for determining the half-pixel interpolation filter index of the neighboring block B0 (as shown in Figure 8 and Fig. 9 Equations (8-316) to (8-321) show the steps for determining the half-pixel interpolation filter index of the neighboring block B2 (as shown in Figure 8 and Fig. 9 as shown).
[0718] Based on the above description, the present invention aims to store the SIF index in the HMVP table (or propagate the SIF index through the HMVP table), and use the SIF index for the HMVP candidate in the process of building the fusion list. The SIF method is used to select a suitable interpolation filter (IF) according to the following: for areas with sharp edges, a conventional DCT-based interpolation filter is used, and for smooth areas (or areas where sharp edges do not need to be retained), an alternative 6-tap interpolation filter (Gaussian filter) is used. For conventional inter-frame prediction, the IF index is explicitly indicated; for the fusion mode, not only the MV and reference image index are borrowed from the corresponding fusion space candidate (HMVP fusion candidate), but also the IF index is borrowed from the corresponding fusion space candidate. This is in contrast to the existing design, in which the IF index is not propagated through the HMVP table. Therefore, in the traditional design, it is impossible to use an alternative interpolation filter for blocks encoded in the fusion mode and fusion candidates obtained from the HMVP table. The HMVP table is used to store motion information of adjacent blocks (but not necessarily adjacent blocks like conventional spatial fusion candidates). The idea of HMVP is to use the motion information of blocks that are close to the current block in space, but not necessarily adjacent to the current block (such as blocks in certain spatial neighborhoods). Therefore, for example, if the current block includes smooth content, and the adjacent blocks mostly include sharp content, the method of borrowing IF indexes from adjacent blocks will fail. However, the smooth content can be located in blocks in certain spatial neighborhoods of the current block, and the motion information of this block can be stored in the HMVP table. By propagating the IF index through the HMVP table presented in this article, it is allowed to use a suitable IF for the current block (for example, a Gaussian filter can be selected for smooth content or when sharp edges do not need to be retained), which is conducive to improving decoding efficiency. If the technology of the present invention is not adopted, the default IF index (corresponding to an 8-tap DCT-based interpolation filter) is always used for the HMVP fusion candidate, and the specific content of the current block (such as whether sharp edges need to be retained) will not be considered.
[0719] Furthermore, the present invention also aims to use only MV and reference picture indexes (but not SIF indexes) during pruning when updating the HMVP table.
[0720] When adding a new element to an HMVP record, it is necessary to determine whether the new element is used for record comparison. The direct method is to use all elements of the HMVP record for record comparison (default C-style structure comparison). However, in the present invention, the IF index is not used for HMVP record comparison. This design is based on two reasons:
[0721] The first reason is to avoid adding extra computational complexity. Each comparison operation will generate extra computational operations in the updating process of the HMVP table and the construction process of the fusion candidate. Therefore, if the comparison operation can be reduced or eliminated, the computational complexity can be reduced, thereby improving the decoding efficiency. From the perspective of implementation, if unnecessary comparison operations can be avoided here, the purpose of the present invention can be better achieved. Therefore, the default C-style structure comparison is not performed on the HMVP record, but the elements of the HMVP record are divided into two subsets: elements used for record comparison and elements not used for record comparison.
[0722] The second reason is to maintain the diversity of HMVP records. For example, if two HMVP records contain the same MV and the same reference index, but only different IF indexes, the two records are invalid because they are not "completely different". On the contrary, in the update process of the HMVP table, it is more effective to consider the two HMVP records to be the same. At this time, if a new record has only a different IF index compared to the existing records in the HMVP table, the new record will not be added to the HMVP table. Therefore, "old" records that are "completely different" from other records (with different MVs or different reference indexes) will be retained. In other words, if a new record is to be added to the HMVP table, the new record should not only be different from the existing records bit by bit, but also must be "substantially different". In terms of decoding efficiency, it is more effective for the HMVP table to contain two records with different MVs or different reference indexes than to contain two records with only different IF indexes.
[0723] Furthermore, the present invention also aims to limit the parameters of the fused switchable interpolation filter (SIF) to save row memory. Compared with previous SIF designs, the present invention introduces a method of applying SIF on top of the motion information inheritance tool without increasing row memory, thereby saving row memory bandwidth. In the case of high resolution, saving row memory will significantly reduce the cost of on-chip memory.
[0724] Since a more appropriate IF index is used for the CU encoded in the fusion mode and a fusion index corresponding to the history-based fusion candidate is provided, the improved IF index derivation method improves decoding efficiency.
[0725] The mathematical operators used in this application are similar to those used in the C programming language, and can be found in the HEVC standard specification for mathematical operators. However, the results of integer division and arithmetic shift operations are accurately defined here, and other operations such as power operations and real-valued division operations are also defined. Numbering and counting specifications usually start from zero, for example, "first" is equivalent to the 0th, "second" is equivalent to the 1st, and so on.
[0726] The following describes applications of the encoding method and decoding method shown in the above embodiments and systems using these methods.
[0727] Fig.18 31 is a block diagram of a content providing system 3100 for implementing a content distribution service. The content providing system 3100 includes a capture device 3102, a terminal device 3106, and optionally a display 3126. The capture device 3102 communicates with the terminal device 3106 via a communication link 3104. The communication link may include the communication channel 13 described above. The communication link 3104 includes, but is not limited to, Wi-Fi, Ethernet, cable, wireless (3G / 4G / 5G), USB, or any combination thereof.
[0728] Capture device 3102 generates data, and can encode the data by the encoding method shown in the above embodiment. Alternatively, capture device 3102 can distribute data to a streaming media server (not shown in the figure), which encodes the data and sends the encoded data to terminal device 3106. Capture device 3102 includes but is not limited to a camera, a smart phone or a tablet computer, a computer or a notebook computer, a video conferencing system, a PDA, a vehicle-mounted device or any combination thereof. For example, capture device 3102 may include source device 12 as described above. When data includes video, the video encoder 20 included in capture device 3102 can actually perform video encoding processing. When data includes audio (i.e., sound), the audio encoder included in capture device 3102 can actually perform audio encoding processing. For some actual scenes, capture device 3102 distributes encoded video and encoded audio data by multiplexing the encoded video data and the encoded audio data together. For other actual scenes, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. The capture device 3102 distributes the encoded audio data and the encoded video data to the terminal device 3106, respectively.
[0729] In the content providing system 3100, the terminal device 3106 receives and reproduces the encoded data. The terminal device 3106 may be a device having data receiving and recovery capabilities, such as a smart phone or tablet computer 3108, a computer or laptop computer 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a television 3114, a set top box (STB) 3116, a video conferencing system 3118, a video surveillance system 3120, a personal digital assistant (PDA) 3122, a vehicle-mounted device 3124, or any combination of the above devices capable of decoding the above-mentioned encoded data. For example, the terminal device 3106 may include the destination device 14 described above. When the encoded data includes video, the video decoder 30 included in the terminal device prioritizes video decoding. When the encoded data includes audio, the audio decoder included in the terminal device prioritizes audio decoding processing.
[0730] For a terminal device with a display, such as a smartphone or tablet 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a television 3114, a personal digital assistant (PDA) 3122, or a vehicle-mounted device 3124, the terminal device can feed the decoded data to the display of the terminal device. For a terminal device without a display, such as a STB 3116, a video conferencing system 3118, or a video surveillance system 3120, the terminal device is connected to an external display 3126 to receive and display the decoded data.
[0731] When each device in the system performs encoding or decoding, the image encoding device or the image decoding device as shown in the above-mentioned embodiments can be used.
[0732] Fig.1931 is a structural diagram of an example of a terminal device 3106. After the terminal device 3106 receives the stream from the capture device 3102, the protocol processing unit 3202 analyzes the transmission protocol of the stream. The protocol includes but is not limited to Real Time Streaming Protocol (RTSP), Hyper Text Transfer Protocol (HTTP), HTTP Live streaming protocol (HLS), MPEG-DASH, Real-time Transport Protocol (RTP), Real-time Messaging Protocol (RTMP), or any combination thereof.
[0733] After the protocol processing unit 3202 processes the stream, a stream file is generated. The file is output to the demultiplexing unit 3204. The demultiplexing unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As described above, for some actual scenarios, such as in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. In this case, the encoded data is sent to the video decoder 3206 and the audio decoder 3208 without passing through the demultiplexing unit 3204.
[0734] Through the demultiplexing process, a video elementary stream (ES) and an audio ES are generated, and subtitles are optionally generated. The video decoder 3206, including the video decoder 30 as described in the above embodiment, decodes the video ES to generate video frames through the decoding method shown in the above embodiment, and feeds this data to the synchronization unit 3212. The audio decoder 3208 decodes the audio ES to generate audio frames, and feeds this data to the synchronization unit 3212. Alternatively, before feeding the video frames to the synchronization unit 3212, the video frames can be stored in a buffer ( Fig.19 Similarly, before feeding the audio frames to the synchronization unit 3212, the audio frames can be stored in the buffer ( Fig.19 not shown).
[0735] The synchronization unit 3212 synchronizes the video frames and the audio frames and provides the video / audio to the video / audio display 3214. For example, the synchronization unit 3212 synchronizes the presentation of the video information and the audio information. The information can be decoded into the syntax using the timestamps associated with the presentation of the decoded audio and visual data and the timestamps associated with the distribution of the data stream.
[0736] If subtitles are included in the stream, the subtitle decoder 3210 decodes the subtitles, synchronizes the subtitles with the video frames and the audio frames, and provides the video / audio / subtitles to the video / audio / subtitle display 3216.
[0737] The present invention is not limited to the above-mentioned system, and the image encoding device or the image decoding device in the above-mentioned embodiments can be used in other systems such as automobile systems.
[0738] Although the embodiments of the present invention are mainly described in terms of video decoding, it should be noted that the embodiments of the decoding system 10, the encoder 20 and the decoder 30 (respectively, the system 10) and other embodiments described herein can also be used for still image processing or decoding, that is, processing or decoding a single image in video decoding that is independent of any previous or consecutive images. In general, if the image processing decoding is limited to a single image 17, only the inter-frame prediction unit 244 (encoder) and the inter-frame prediction unit 344 (decoder) are not available. All other functions (also called tools or techniques) of the video encoder 20 and the video decoder 30 can also be used for still image processing, such as residual calculation 204 / 304, transformation 206, quantization 208, inverse quantization 210 / 310, (inverse) transformation 212 / 312, segmentation 262 / 362, intra-frame prediction 254 / 354 and / or loop filtering 220 / 320, entropy coding 270 and entropy decoding 304.
[0739] Embodiments of encoders 20, decoders 30, etc. and the functions described herein in conjunction with encoders 20, decoders 30, etc. may be implemented by hardware, software, firmware, or any combination thereof. If implemented by software, these functions may be stored as one or more instructions or codes in a computer-readable medium or sent via a communication medium and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which corresponds to tangible media (e.g., data storage media), or includes any communication media that facilitates the transfer of computer programs from one place to another according to a communication protocol, etc. In this manner, computer-readable media may generally correspond to (1) non-transitory tangible computer-readable storage media or (2) communication media such as signals or carrier waves. Data storage media may be any available media that is accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the techniques described in the present invention. A computer program product may include a computer-readable medium.
[0740] As an example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage or other magnetic storage devices, flash memory or any other medium that can be used to store the required program code in the form of instructions or data structures and can be accessed by a computer. In addition, any connection can be appropriately referred to as a computer-readable medium. For example, if a coaxial cable, optical fiber cable, twisted pair, digital subscriber line (DSL) or wireless technologies such as infrared, radio and microwave are used to transmit instructions from a website, server or other remote source, then the coaxial cable, optical fiber cable, twisted pair, DSL or wireless technologies such as infrared, radio and microwave are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carriers, signals or other transient media, but rather involve non-transient tangible storage media. The disks and optical disks used herein include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks and blue-ray discs, wherein disks usually reproduce data magnetically, while optical discs reproduce data optically by lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0741] Instructions may be executed by one or more processors, for example, one or more digital signal processors (DSPs), one or more general-purpose microprocessors, one or more application specific integrated circuits (ASICs), one or more field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the term "processor" used herein may refer to any of the above structures or any other structure suitable for implementing the techniques described herein. In addition, in some aspects, the various functions described herein may be provided in dedicated hardware and / or software modules for encoding and decoding, or incorporated in a combined codec. Moreover, these techniques may be fully implemented in one or more circuits or logic elements.
[0742] The technology of the present invention can be implemented in a variety of devices or apparatuses, including wireless mobile phones, integrated circuits (ICs) or a group of ICs (e.g., chipsets). Various components, modules, or units are described in the present invention to emphasize the functional aspects of the devices used to perform the disclosed technology, but they do not necessarily need to be implemented by different hardware units. Instead, as described above, the various units can be combined in a codec hardware unit in combination with appropriate software and / or firmware, or provided by a collection of interoperable hardware units including one or more processors as described above.
Claims
1. A method for decoding a block in a video signal frame, characterized in that include: Receive and parse the bitstream to obtain a reconstructed residual block; Obtain a history-based motion information candidate list, wherein the history-based motion information candidate list includes N history-based motion information candidate H k An ordered list of the N historical motion information candidates H k Including motion information of N previous blocks before a block, N is an integer greater than 0, k is equal to 0, ..., N-1, each history-based motion information candidate includes motion information, and the motion information includes elements: (i) one or more motion vectors MV, (ii) one or more reference picture indices corresponding to the one or more MVs, (iii) interpolation filter index; The history-based motion information candidate list is updated according to the motion information of the block, wherein the motion information of the block includes elements: (i) one or more motion vectors MV of the block, (ii) one or more reference picture indices corresponding to the MV of the block, (iii) interpolation filter index; Based on the updated history-based motion information candidate list, inter-predict the block to obtain a predicted block; The reconstructed residual block and the prediction block are added to obtain a reconstructed block.
2. The method according to claim 1, characterized in that: The updating of the history-based motion information candidate list comprises: If at least one of the following elements of each history-based motion information candidate in the history-based motion information candidate list is different from the corresponding element in the motion information of the block, the history-based motion information candidate H containing the motion information of the block N Added to the history-based motion information candidate list, wherein the at least one element is: (i) the one or more motion vectors MV, (ii) the one or more reference picture indexes corresponding to the one or more MVs.
3. The method according to claim 1, characterized in that The updating of the history-based motion information candidate list comprises: If the following elements of the history-based motion information candidate in the history-based motion information candidate list are the same as the corresponding elements in the motion information of the block, the history-based motion information candidate is deleted from the history-based motion information candidate list, and the history-based motion information candidate H containing the motion information of the block is replaced N-1 Added to the history-based motion information candidate list, wherein the elements are: (i) one or more motion vectors MV, (ii) one or more reference picture indexes corresponding to the one or more MVs.
4. The method according to any one of claims 1 to 3, characterized in that The updating of the history-based motion information candidate list comprises: If N is equal to a predefined value, the history-based motion information candidate H0 is deleted from the history-based motion information candidate list, and the motion information of the block is used as the history-based motion information candidate H N-1 Added to the history-based motion information candidate list.
5. The method according to claim 2 or 3, characterized in that: Also includes: comparing whether a motion vector of a history-based motion information candidate in the history-based motion information candidate list is the same as a corresponding motion vector of the block; A comparison is made as to whether the reference picture index of the history-based motion information candidate is the same as the corresponding reference picture index of the block.
6. The method according to claim 2 or 3, characterized in that: Also includes: comparing whether at least one of the motion vectors of each history-based motion information candidate is different from a corresponding motion vector of the block; At least one of the reference picture indexes of each history-based motion information candidate is compared to see whether it is different from a corresponding reference picture index of the block.
7. The method according to any one of claims 1 to 3, characterized in that The interpolation filter index included in the history-based motion information candidate represents a half-sample interpolation filter in a half-sample interpolation filter set; the half-sample interpolation filter is applied to interpolate half-sample values only when at least one MV among the one or more MVs of the history-based motion information candidate points to a half-sample position.
8. A method for decoding a block in a video signal frame, characterized in that include: Receive and parse the bitstream to obtain a reconstructed residual block; Construct a history-based motion information candidate list, wherein the history-based motion information candidate list includes N history-based motion information candidate H k An ordered list of the N historical motion information candidates H k Including motion information of N previous blocks before the block, N is an integer greater than 0, k is equal to 0, ..., N-1, and each history-based motion information candidate includes the elements: (i) one or more motion vectors MV, (ii) one or more reference image indices corresponding to the MV, (iii) interpolation filter index; adding one or more history-based motion information candidates in the history-based motion information candidate list to a motion information candidate list for the block; The motion information of the block is derived according to the motion information candidate list, where the motion information of the block includes the following elements: (i) one or more motion vectors MV of the block, (ii) one or more reference picture indices corresponding to the MV of the block, (iii) interpolation filter index; According to the motion information, inter-frame prediction is performed on the block to obtain a predicted block; The reconstructed residual block and the prediction block are added to obtain a reconstructed block.
9. The method according to claim 8, characterized in that Only when at least one MV among the one or more MVs of the derived motion information points to a half-sample position, a substitute half-sample interpolation filter is applied, the substitute half-sample interpolation filter being indicated by an interpolation filter index included in the derived motion information.
10. The method according to claim 8, characterized in that The interpolation filter index included in the history-based motion information candidate represents a half-sample interpolation filter in a half-sample interpolation filter set; the half-sample interpolation filter is applied to interpolate half-sample values only when at least one MV among the one or more MVs of the history-based motion information candidate points to a half-sample position.
11. The method according to any one of claims 8 to 10, characterized in that Also includes: If at least one of the following elements of each history-based motion information candidate in the history-based motion information candidate list is different from the corresponding element in the motion information of the block, the history-based motion information candidate H containing the motion information of the block N Added to the history-based motion information candidate list, wherein the at least one element is: (i) the one or more motion vectors MV, (ii) the one or more reference picture indexes corresponding to the one or more MVs.
12. The method according to any one of claims 8 to 10, characterized in that Also includes: If the following elements of the history-based motion information candidate in the history-based motion information candidate list are the same as the corresponding elements in the motion information of the block, the history-based motion information candidate is deleted from the history-based motion information candidate list, and the history-based motion information candidate H containing the motion information of the block is replaced N-1 Added to the history-based motion information candidate list, wherein the elements are: (i) one or more motion vectors MV, (ii) one or more reference picture indexes corresponding to the MV.
13. The method according to any one of claims 8 to 10, characterized in that Also includes: If N is equal to a predefined value, the history-based motion information candidate H0 is deleted from the history-based motion information candidate list, and the motion information of the block is used as the history-based motion information candidate H N-1 Added to the history-based motion information candidate list.
14. The method according to claim 11, characterized in that Also includes: comparing whether a motion vector of a history-based motion information candidate in the history-based motion information candidate list is the same as a corresponding motion vector of the block; A comparison is made as to whether the reference picture index of the history-based motion information candidate is the same as the corresponding reference picture index of the block.
15. The method according to claim 11, characterized in that Also includes: comparing whether at least one of the motion vectors of each history-based motion information candidate is different from a corresponding motion vector of the block; At least one of the reference picture indexes of each history-based motion information candidate is compared to see whether it is different from a corresponding reference picture index of the block.
16. The method according to any one of claims 8 to 10, characterized in that The length of the history-based motion information candidate list is N, where N is 5 or 6.
17. The method according to any one of claims 8 to 10, characterized in that The motion information candidate list is used for merge mode or skip mode.
18. The method according to any one of claims 8 to 10, characterized in that The deriving the motion information of the block according to the motion information candidate list comprises: The motion information indicated by the candidate index is derived from the motion information candidate list as the motion information of the block.
19. The method according to any one of claims 8 to 10, characterized in that Also includes: When at least one of the one or more motion vectors MV included in the derived motion information points to a half-sample position, a predicted sample value of the block is obtained by applying a half-sample interpolation filter to a sample value of the reference image pointed to by the motion vector, wherein the half-sample interpolation filter is represented by a half-sample interpolation filter index included in the derived motion information, and the reference image is represented by the one or more reference image indexes included in the derived motion information.
20. A decoder (30), characterized in that comprising processing circuitry for performing the method according to any one of claims 1 to 19.
21. A computer program product, characterized in that Comprising program code for executing the method according to any one of claims 1 to 19.
22. A decoder, characterized in that: include: one or more processors; A non-transitory computer-readable storage medium, the non-transitory computer-readable storage medium is coupled to the processor and stores a program executed by the processor, when the processor executes the program, the decoder performs the method according to any one of claims 1 to 19.
23. A non-transitory computer-readable storage medium, characterized in that: The device carries a program code, which, when executed by a computer device, causes the computer device to execute the method according to any one of claims 1 to 19.
Citation Information
Patent Citations
Image encoding / decoding method and recording medium for same
CN109196864A
Method and apparatus for video coding with automatic motion information refinement
CN109417630A