Method and device for video processing and medium
By determining the reference template based on motion information at the sub-block level in video encoding and decoding technology, the problem of insufficient encoding and decoding efficiency in the prior art is solved, and a more efficient encoding and decoding process is achieved.
Patent Information
- Application Number
- CN202380074133.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-20
- Filing Date
- 2023-10-19
- Publication Date
- 2025-05-30
AI Technical Summary
Existing video codec technology has challenges in improving codec efficiency, especially in motion candidate list construction and motion vector prediction.
By determining the reference template based on motion information at the sub-block level, the reference template of the video block is improved, thereby improving the encoding and decoding efficiency and effectiveness.
This method improves the efficiency and effectiveness of video encoding and codec by improving the determination of reference templates, and solves the problem of insufficient encoding and codec efficiency in the prior art.
Smart Images

Figure CN120077650A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to video processing technologies, and more particularly, to motion candidate list construction. Background Art
[0002] Nowadays, digital video capabilities are being applied to all aspects of people's lives. For video encoding / decoding, various types of video compression technologies have been proposed, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T H.265 High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard. However, there is generally a desire to further improve the encoding / decoding efficiency of video encoding / decoding technologies. Summary of the Invention
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method includes: for the conversion between a current video block of a video and a bitstream of the video, determining a reference template for the current video block based on a current template of the current video block and motion information of at least one sub-block of the current video block; and performing the conversion based on the reference template. The method according to the first aspect of the present disclosure determines a reference template based on sub-block level motion information. Therefore, the determined reference template can be improved. In this way, the encoding / decoding efficiency and encoding / decoding effectiveness can be improved.
[0005] In a second aspect, an apparatus for video processing is proposed. The apparatus includes a processor and a non-transitory memory having instructions thereon. These instructions, when executed by the processor, cause the processor to execute the method according to the first aspect of the present disclosure.
[0006] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by an apparatus for video processing. The method includes: determining a reference template for a current video block of the video based on a current template of the current video block and motion information of at least one sub-block of the current video block; and generating a bitstream based on the reference template.
[0008] In a fifth aspect, a method for storing a bitstream of a video is provided. The method includes: determining a reference template for a current video block of the video based on a current template of the current video block and motion information of at least one sub-block of the current video block; generating a bitstream based on the reference template; and storing the bitstream in a non-transitory computer-readable recording medium.
[0009] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0011] Figure 1 A block diagram showing an example video codec system according to some embodiments of the present disclosure;
[0012] Figure 2 A block diagram showing a first example video encoder according to some embodiments of the present disclosure;
[0013] Figure 3 A block diagram showing an example video decoder according to some embodiments of the present disclosure;
[0014] Figure 4 Shows the positions of spatial and temporal neighboring blocks used in advanced motion vector prediction (AMVP) or Merge candidate list construction;
[0015] Figure 5 An example diagram showing the positions of non-adjacent candidates in ECM;
[0016] Figure 6 An example diagram showing template matching performed on a search area around an initial MV;
[0017] Figure 7 An example diagram showing a template and a corresponding reference template;
[0018] Figure 8 An example diagram showing a template and a reference template of a block with sub-block motion using motion information of sub-blocks of a current block;
[0019] Figure 9 An example diagram showing an example of the positions of non-adjacent temporal motion vector prediction (TMVP) candidates;
[0020] Figure 10 An example diagram showing an example of a template is shown;
[0021] Figure 11 An example of a template for a block with sub-block level motion information is shown;
[0022] Figure 12 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown; and
[0023] Figure 13 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.
[0024] Throughout the drawings, the same or similar reference numerals generally refer to the same or similar elements. Detailed Description
[0025] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is for illustrative purposes only and to assist those skilled in the art in understanding and implementing the present disclosure, and does not imply any limitation on the scope of the present disclosure. The disclosure described herein may be implemented in various ways other than those described below.
[0026] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0027] References herein to "one embodiment", "an embodiment", "example embodiment", etc., indicate that the embodiment described may include a particular feature, structure, or characteristic, but not every embodiment must include that particular feature, structure, or characteristic. Moreover, these phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is submitted that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of one of ordinary skill in the art in relation to other embodiments.
[0028] It should be understood that although terms such as "first" and "second" may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0029] The terms used in this document are for the purpose of describing specific embodiments only and are not intended to limit the example embodiments. As used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "comprise", "comprising", "have", "having", "include" and / or "including" when used herein indicate the presence of the stated features, elements and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.
[0030] Example environment
[0031] Figure 1 is a block diagram showing an example video codec system 100 that can utilize the techniques of the present disclosure. As shown, the video codec system 100 can include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 can be configured to generate encoded video data, and the destination device 120 can be configured to decode the encoded video data generated by the source device 110. The source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0032] The video source 112 can include sources such as video capture devices. Examples of video capture devices include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or combinations thereof.
[0033] The video data can include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream can include a sequence of bits that form an encoded representation of the video data. The bitstream can include encoded pictures and associated data. An encoded picture is an encoded representation of a picture. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 can include a modulator / demodulator and / or a transmitter. The encoded video data can be directly transmitted to the destination device 120 via the I / O interface 116 through the network 130A. The encoded video data can also be stored on a storage medium / server 130B for access by the destination device 120.
[0034] Destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from a source device 110 or a storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120 or may be external to the destination device 120, which is configured to interface with an external display device.
[0035] The video encoder 114 and the video decoder 124 may operate according to video compression standards (such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other existing and / or future standards).
[0036] Figure 2 is a block diagram showing an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be Figure 1 an example of the video encoder 114 in the system 100 shown.
[0037] The video encoder 200 may be configured to implement any or all of the techniques of the present disclosure. In Figure 2 the example, the video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video encoder 200. In some examples, a processor may be configured to execute any or all of the techniques described in the present disclosure.
[0038] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206.
[0039] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0040] In addition, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for purposes of explanation, these components are Figure 2is shown separately in the example of
[0041] The segmentation unit 201 may segment the picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0042] The mode selection unit 203 may select, for example, one coding mode from multiple coding modes (intra coding or inter coding) based on an error result, and provide the resulting intra-coded block or inter-coded block to the residual generation unit 207 to generate residual block data, and provide it to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combined intra and inter prediction (CIIP) mode in which the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 may also select a resolution for the motion vector for the block (e.g., sub-pixel accuracy or integer pixel accuracy).
[0043] To perform inter prediction on a current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from the cache 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and the decoded samples of a picture from the cache 213 other than the picture associated with the current video block.
[0044] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, e.g., depending on whether the current video block is in an I-slice, a P-slice, or a B-slice. As used herein, an "I-slice" may refer to a portion of a picture composed of macroblocks, all of which are based on macroblocks within the same picture. Additionally, as used herein, in some aspects, a "P-slice" and a "B-slice" may refer to portions of a picture composed of macroblocks independent of the macroblocks in the same picture.
[0045] In some examples, the motion estimation unit 204 may perform uni-directional prediction on the current video block, and the motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. The motion estimation unit 204 may then generate a reference index and a motion vector, the reference index indicating the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, a prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0046] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 to find a reference video block for the current video block, and may also search the reference pictures in list 1 to find another reference video block for the current video block. The motion estimation unit 204 may then generate a plurality of reference indices and a plurality of motion vectors, the plurality of reference indices indicating the plurality of reference pictures in list 0 and list 1 that contain the plurality of reference video blocks, and the plurality of motion vectors indicating the plurality of spatial displacements between the plurality of reference video blocks and the current video block. The motion estimation unit 204 may output the plurality of reference indices and the plurality of motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the plurality of reference video blocks indicated by the motion information of the current video block.
[0047] In some examples, the motion estimation unit 204 may output a complete set of motion information for use in the decoding process of the decoder. Alternatively, in some embodiments, the motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is similar enough to the motion information of a neighboring video block.
[0048] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block, the value indicating to the video decoder 300 that the current video block has the same motion information as another video block.
[0049] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0050] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0051] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.
[0052] The residual generation unit 207 can generate residual data for a current video block by subtracting (e.g., indicated by a minus sign) the (multiple) predicted video blocks of the current video block from the current video block. The residual data of the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.
[0053] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.
[0054] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0055] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0056] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0057] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation can be performed to reduce the block effect artifacts in the video block.
[0058] The entropy coding unit 214 can receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives the data, the entropy coding unit 214 can perform one or more entropy coding operations to generate the entropy-coded data and output a bitstream including the entropy-coded data.
[0059] Figure 3 is a block diagram showing an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 can be Figure 1 an example of the video decoder 124 in the system 100 shown.
[0060] The video decoder 300 can be configured to perform any or all of the techniques of the present disclosure. In Figure 3In the example of, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0061] In Figure 3 the example of, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a cache 307. In some examples, the video decoder 300 can perform a decoding process that is generally opposite to the encoding process described with respect to the video encoder 200.
[0062] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream can include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, the motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and the Merge mode. AMVP is used, including deriving several most likely candidates based on data from adjacent PBs and reference pictures. The motion information generally includes a horizontal motion vector displacement value and a vertical motion vector displacement value, one or two reference picture indices, and in the case of a prediction region in a B slice, also an indication of which reference picture list is associated with each index. As used herein, in some aspects, the "Merge mode" can refer to deriving motion information from spatially adjacent blocks or temporally adjacent blocks.
[0063] The motion compensation unit 302 can generate motion-compensated blocks, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter used at sub-pixel precision can be included in the syntax element.
[0064] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of the video block to calculate the interpolated values for sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate a prediction block.
[0065] The motion compensation unit 302 may use at least part of the syntax information to determine the size of the blocks for encoding the (multiple) frames and / or (multiple) slices of the encoded video sequence, the partitioning information describing how each macroblock of the pictures of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame encoded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a "slice" may refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding / decoding, signal prediction, and residual signal reconstruction. A slice may be the entire picture or may also be a region of the picture.
[0066] The intra prediction unit 303 may use, for example, the intra prediction mode received in the bitstream to form a prediction block from spatially adjacent blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0067] The reconstruction unit 306 may obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If needed, a deblocking filter may also be applied to filter the decoded block to remove block effect artifacts. The decoded video blocks are then stored in the cache 307, and the cache 307 provides reference blocks for subsequent motion compensation / intra prediction, and the cache 307 also produces the decoded video for presentation on a display device.
[0068] Some exemplary embodiments of the present disclosure will be described in detail below. It should be noted that the use of section headings in this document is for ease of understanding and does not limit the embodiments disclosed in the section to that section. Additionally, although some embodiments are described with reference to the multi-functional video codec or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Furthermore, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. Additionally, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another or at a different compression bit rate.
[0069] 1. Brief Overview
[0070] The present disclosure relates to video codec techniques. Specifically, it is about a method for constructing motion vector prediction (MVP) in video coding and decoding. This concept can be applied alone or in various combinations to any video coding and decoding standard or non-standard video codec.
[0071] 2. Introduction
[0072] The exponential growth of multimedia data has posed severe challenges to video coding and decoding. To meet the growing demand for more efficient compression technologies, ITU-T and ISO / IEC have developed a series of video coding and decoding standards in the past few decades. Specifically, ITU-T has developed H.261 and H.263, ISO / IEC has developed MPEG-1 and MPEG-4 video, and the two organizations have jointly developed H.262 / MPEG-2 video, H.264 / MPEG-4 Advanced Video Coding (AVC), H.265 / HEVC, and the latest VVC standard. Since H.262 / MPEG-2, a hybrid video coding and decoding framework has been adopted, in which intra / inter-frame prediction plus transform coding and decoding are used.
[0073] Figure 4 Fig. 400 shows the positions of spatial and temporal neighboring blocks used in AMVP / Merge candidate list construction.
[0074] 2.1. MVP in Video Coding and Decoding
[0075] Inter-frame prediction aims to remove the temporal redundancy between adjacent frames, which is an essential part of the hybrid video coding and decoding framework. Specifically, inter-frame prediction uses the content specified by the motion vector (MV) as the predicted version of the current block to be coded and decoded, so that only the residual signal and motion information are transmitted in the bitstream. To reduce the cost of MV signaling, Motion Vector Prediction (MVP) emerged as an effective mechanism for transmitting motion information. Early strategies simply used the MV of the specified neighboring block or the median MV of the neighboring blocks as the MVP. In H.265 / HEVC, a competitive mechanism is involved, where Rate-Distortion Optimization (RDO) selects the optimal MVP from multiple candidates. Specifically, the Advanced MVP (AMVP) mode and the Merge mode are designed with different motion information signaling strategies. In the AMVP mode, the MVP candidate index of the reference index, the reference AMVP candidate list, and the Motion Vector Difference (MVD) is signaled. As for the Merge mode, only the Merge index of the reference Merge candidate list is signaled, and all the motion information associated with the Merge candidate is inherited. Both the AMVP mode and the Merge mode need to construct an MVP candidate list, and the details of the construction processes of these two modes are described below.
[0076] AMVP Mode: AMVP utilizes the spatio-temporal correlation of motion vectors with neighboring blocks for explicit transmission of motion parameters. For each reference picture list, a motion vector candidate list is constructed by first checking the availability of left and upper temporal neighboring positions, removing redundant candidates and adding zero vectors to make the candidate list length constant. For spatial motion vector candidate derivation, as Figure 4 shown, two motion vector candidates are finally derived based on the motion vectors of blocks located at five different positions. The five neighboring blocks located at B0, B1, B2 and A0, A1 are divided into two groups, where Group A includes the three upper spatial neighboring blocks and Group B includes the two left spatial neighboring blocks. These two motion vector candidates are derived from the first available candidates in Group A and Group B respectively in a predefined order. For temporal motion vector candidate derivation, as Figure 4 shown, one motion vector candidate is derived based on two different co-located positions (lower-right (C0) and center (C1)) checked in sequence. To avoid MV candidate redundancy, repeated motion vector candidates in the list are discarded. If the number of potential candidates is less than two, additional zero motion vector candidates are added to the list.
[0077] Figure 5 Fig. 500 shows the positions of non-adjacent candidates in ECM.
[0078] Merge Mode: Similar to the AMVP mode, the MVP candidate list for the Merge mode also includes spatial candidates and temporal candidates. For spatial motion vector candidate derivation, after performing availability and redundancy checks, up to four candidates are selected in the order of A1, B1, B0, A0 and B2. For temporal Merge candidate (TMVP) derivation, up to one candidate is selected from two temporal neighboring blocks (C0 and C1). When there are not enough Merge candidates using spatial and temporal candidates, combined bidirectional prediction Merge candidates and zero MV candidates are added to the MVP candidate list. Once the number of available Merge candidates reaches the maximum allowed number transmitted by signal, the Merge candidate list construction process is terminated.
[0079] In VVC, the construction process for the Merge mode is further improved by introducing history-based MVP (HMVP), which combines the motion information of previously encoded / decoded blocks that may be far from the current block. In VVC, the HMVP Merge candidates are appended to the Merge list, after the spatial MVP and TMVP. In this method, the motion information of previously encoded / decoded blocks is stored in a table and used as the MVP for the current CU. During the encoding / decoding process, a first-in-first-out strategy is used to maintain the table with multiple HMVP candidates. Whenever there is a non-sub-block inter-frame encoded / decoded CU, the associated motion information is added to the last entry of the table as a new HMVP candidate.
[0080] During the VVC standardization process, non-adjacent MVP was proposed to facilitate better motion information derivation by adopting non-adjacent regions. In the ECM software, non-adjacent MVP is inserted between TMVP and HMVP, where the distance between the non-adjacent spatial candidates and the current encoded / decoded block is based on the width and height of the current encoded / decoded block, as Figure 5 shown. 2.2. Interpolation Filters in VVC
[0081] In VVC, interpolation filters are used in both intra- and inter-frame encoding / decoding processes. Intra-frame encoding / decoding uses interpolation filters to generate fractional positions in angular prediction modes. In HEVC, a 2-tap linear interpolation filter is used to generate intra-frame prediction blocks in the direction prediction mode (i.e., excluding the planar and DC predictors). While in VVC, a 4-tap intra-frame interpolation filter is used to improve the angular intra-frame prediction accuracy. Specifically, two sets of 4-tap interpolation filters are used in VVC intra-frame encoding / decoding, which are the DCT-based interpolation filter (DCTIF) and the smoothing interpolation filter (SIF). The DCTIF is constructed in the same way as used for chrominance component motion compensation in HEVC and VVC. The SIF is obtained by convolving a 2-tap linear interpolation filter with a [1 2 1] / 4 filter.
[0082] In VVC, the highest precision of the explicitly signaled motion vector is a quarter of a luma sample. In some inter-frame prediction modes (such as the affine mode), the motion vector is derived with 1 / 16 luma sample precision and motion compensation prediction is performed with 1 / 16 sample precision. VVC allows different MVD precisions ranging from 1 / 16 luma sample to 4 luma samples. For half-luma sample precision, a 6-tap interpolation filter is used. While for other fractional precisions, the default 8-tap filter is used. In addition, a bilinear interpolation filter is used to generate fractional samples for the search process of decoder-side motion vector refinement (DMVR) in VVC.
[0083] 2.3. Template Matching Merge / AMVP Mode in ECM
[0084] The template matching (TM) merge / AMVP mode is a decoder-side MV derivation method that refines the motion information of the current CU by finding the closest match between a template in the current picture (i.e., the neighboring blocks above and / or to the left of the current CU) and a block in a reference picture (i.e., of the same size as the template). As Figure 6 shown, a better MV is searched within the [-8, +8] pixel search range around the initial motion of the current CU.
[0085] Figure 6 Fig. 600 shows an example of template matching performed on the search area around the initial MV.
[0086] In the AMVP mode, the MVP candidates are determined based on the template matching error to select the MVP candidate with the minimum difference between the current block and the reference block template, and then TM performs MV refinement only for this specific MVP candidate. TM refines this MVP candidate starting from the full pixel MVD accuracy (or 4-pixel AMVR mode) within the [-8, +8] pixel search range by using iterative diamond search. The AMVP candidate can be further refined by using cross search with full pixel MVD accuracy (or 4-pixel AMVR mode) and then sequentially using half pixel and quarter pixel searches according to the AMVR mode. This search process ensures that the MVP candidate still maintains the same MV accuracy as indicated by the adaptive motion vector resolution (AMVR) mode after TM processing.
[0087] In the Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. TMMerge can be performed all the way to 1 / 8 pixel MVD accuracy or skip the accuracy beyond half pixel MVD accuracy, depending on whether an alternative interpolation filter (used when AMVR is in half pixel mode) is used to merge the motion information. In addition, when the TM mode is enabled, template matching can be an independent process or an additional MV refinement process between the block-based and sub-block-based bilateral matching (BM) methods, depending on whether BM can be enabled according to the enabled conditions check. When both BM and TM are enabled on a CU, the search process of TM will stop at half pixel MVD accuracy, and the resulting MV is further refined by using the same model-based MVD derivation method as in DMVR.
[0088] 2.4. Adaptive Reordering of Merge Candidates (ARMC)
[0089] Inspired by reconstructing the spatial correlation between adjacent pixels and the current coding / decoding block, Adaptive Reordering of Merge Candidates (ARMC) is proposed to refine the candidate order in a given candidate list. The basic assumption is that candidates with lower template matching costs have a higher probability of being selected through the RDO process, and thus should be placed in the front positions in the list to reduce signaling costs.
[0090] This reordering method is applied to the regular Merge mode, the Template Matching (TM) Merge mode, and the Affine Merge mode (excluding SbTMVP candidates). For the TM Merge mode, the Merge candidates are reordered before the refinement process.
[0091] After constructing the Merge candidate list, the Merge candidates are divided into several subgroups. The subgroup size is set to 5. The Merge candidates in each subgroup are reordered in ascending order according to the cost values based on template matching. For simplicity, the Merge candidates in the last subgroup but not in the first subgroup are not reordered.
[0092] The template matching cost is measured by the Sum of Absolute Differences (SAD) between the samples of the template of the current block and its corresponding reference template. The template includes a set of reconstructed samples adjacent to the current block, while the reference template is located by the same motion information of the current block, as Figure 7 shown. When a Merge candidate uses bidirectional prediction, the reference samples of the template of the Merge candidate are also generated by bidirectional prediction.
[0093] Figure 7 Fig. 700 shows the template and the corresponding reference template.
[0094] Figure 8 Fig. 800 shows the template and the reference template of a block with sub-block motion using the motion information of the sub-blocks of the current block.
[0095] For sub-block-based Merge candidates with a sub-block size equal to Wsub*Hsub, the above template includes several sub-templates with a size of Wsub×1, and the left template includes several sub-templates with a size of 1×Hsub. As Figure 8 shown, the motion information of the sub-block in the first row and first column of the current block is used to derive the reference samples of each sub-template.
[0096] Figure 9 Fig. 900 shows an example diagram for the positions of non-adjacent TMVP candidates.
[0097] 2.5. Enhanced MVP Candidate Derivation (EMCD)
[0098] The EMCD based on template matching cost reordering is proposed. A method of optimizing the MVP selection by utilizing the matching cost in the reconstructed template region is studied, so as to include more suitable candidates in the list instead of constructing the MVP list based on a predefined traversal order.
[0099] It should be noted that the proposed MVP list construction strategy can be used in the normal Merge and AMVP list construction processes, and can also be easily extended to other modules that require MVP derivation, such as Merge with Motion Vector Difference (MMVD), affine motion compensation, sub-block based temporal motion vector prediction (SbTMVP), etc.
[0100] Non - adjacent TMVP
[0101] 1. It is proposed to utilize TMVP in non-adjacent regions to further improve the effectiveness of the MVP list.
[0102] a) In one example, the non-adjacent region can be any block (such as a 4×4 block) in the reference picture, and it is neither inside the co-located block in the reference picture of the current block nor adjacent to the co-located block of the current block.
[0103] b) In one example, the positions of the non-adjacent TMVP candidates are as Figure 9 shown, where the black blocks represent potential non-adjacent TMVP positions. It should be noted that this figure only provides an example for non-adjacent TMVP, and the positions are not limited to the indicated blocks. In other cases, the non-adjacent TMVP can be located at any other position in one or more reconstructed frames.
[0104] 2. The maximum number of non-adjacent TMVP allowed in the MVP list can be signaled in the bitstream.
[0105] a) In one example, the maximum allowed number can be signaled in the SPS or PPS.
[0106] 3. The non-adjacent TMVP candidates can be located in the closest reconstructed frame, but they can also be located in other reconstructed frames.
[0107] a) Alternatively, the non-neighboring TMVP candidates can be located in the co-located picture.
[0108] b) Alternatively, it can be signaled in which picture the non-adjacent TMVP candidates are located.
[0109] 4. The non-adjacent TMVP candidates can be located in multiple reference pictures.
[0110] Figure 10 Fig. 1000 shows an example diagram of the template.
[0111] 5. The distance between a non - adjacent region associated with a TMVP candidate and the current coded block can be related to the properties of the current block.
[0112] a) In one example, the distance depends on the width and height of the current coded block.
[0113] b) In other cases, the distance can be signaled as a constant in the bitstream.
[0114] Definition of the template
[0115] 6. The template represents a reconstructed region, which can be used to estimate the priority of MVP candidates. The region can be located in different positions and have a variable shape.
[0116] a) In one example, the template can include reconstructed regions at three positions, namely, the upper pixel, the left pixel, and the upper - left pixel, as Figure 10 shown.
[0117] b) It should be noted that the template is not necessarily rectangular in shape and can be of any shape, such as triangular or polygonal.
[0118] c) In one example, the template regions can be utilized individually or in combination.
[0119] d) The template can include samples from only one component (e.g., luminance) or from multiple components (e.g., luminance and chrominance).
[0120] 7. The template can be located not only in the current frame but also in any other reconstructed frame.
[0121] 8. In one example, the MV can be used to locate a reference template region having the same shape as the template of the current block, as Figure 7 shown.
[0122] 9. In one example, the template can be located not in an adjacent region but in a non - adjacent region far from the current block.
[0123] 10. In one example, the template can include not all the pixels in a specific region but a part of the pixels in the region.
[0124] MVP candidate ranking based on template matching
[0125] 11. In the present disclosure, the template matching cost associated with a certain MVP candidate is used as a metric to evaluate the consistency of the candidate with the true motion information. Based on this metric, a more efficient order is generated by sorting the priorities of each MVP candidate.
[0126] a) In one example, the template matching cost C is evaluated using the mean squared error (MSE) and is calculated as follows:
[0127]
[0128] where T represents the template region, RT represents the corresponding reference template region specified by the MV within the MVP candidate (Figure
[0129] 7), and N is the number of pixels within the template.
[0130] b) In one example, the template matching cost can be evaluated using the sum of squared errors (SSE), sum of absolute differences (SAD), sum of absolute transform differences (SATD), or any other criterion that can measure the difference between two regions.
[0131] 12. Sort all candidate MVPs in ascending order according to their corresponding template matching costs, and traverse the candidate sequence in the sorted order to construct the MVP list until the number of MVPs reaches the maximum allowed number. In this way, candidates with lower matching costs have a higher priority to be included in the final MVP list.
[0132] a) In one example, the sorting process can be performed on all MVP candidates.
[0133] b) Alternatively, this process can also be applied to a part of the candidates, e.g., non - adjacent MVP candidates, HMVP candidates, or any other candidate group.
[0134] c) Alternatively, in addition, the category of MVP candidates that should be reordered (e.g., non - adjacent MVP candidates belong to one category, HMVP candidates belong to another category) and / or the type of candidate group that should be reordered can depend on the decoded information, such as block size / coding - decoding method (e.g., CIIP / MMVD) and / or how many MVP candidates are available before reordering for a given kind / group.
[0135] 1. In one example, the sorting process can be performed on a combined group of MVP candidates that contains only one category.
[0136] 2. In one example, the sorting process can be performed on a combined group of MVP candidates that contains more than one category.
[0137] a) In one example, for the first coding - decoding method (e.g., conventional
[0138] / CIIP / MMVD / GPM / TPM / Sub-block Merge mode), the sorting process can be performed on a combined group of non-adjacent MVPs, non-adjacent TMVP, and HMVP candidates. For the second-pass decoding method (e.g., template matching Merge mode), the sorting process can be performed on a combined group of adjacent MVPs, non-adjacent TMVP, non-adjacent MVPs, and HMVP candidates.
[0139] b) Alternatively, for the first-pass decoding method (e.g., conventional / CIIP / MMVD / GPM / TPM / Sub-block Merge mode), the sorting process can be performed on a combined group of non-adjacent MVPs and HMVP candidates. For the second-pass decoding method (e.g., template matching Merge mode), the sorting process can be performed on a combined group of adjacent MVPs, non-adjacent MVPs, and HMVP candidates.
[0140] 3. In one example, the sorting process can be performed on a combined group that includes some available MVP candidates within a category.
[0141] a) In one example, for the conventional / CIIP / MMVD / TM / GPM / TPM / Sub-block Merge mode, or for the conventional / affine AMVP mode, the sorting process can be performed on a combined group of all or some candidates from one or more categories.
[0142] 4. In the above examples, the categories can be:
[0143] i. Adjacent neighboring MVPs;
[0144] ii. Adjacent neighboring MVPs at (a) specific location(s);
[0145] iii. TMVP MVPs;
[0146] iv. HMVP MVPs;
[0147] v. Non-adjacent MVPs;
[0148] vi. Constructed MVPs (e.g., paired MVPs);
[0149] vii. Inherited affine MV candidates;
[0150] viii. Constructed affine MV candidates;
[0151] ix. SbTMVP candidates.
[0152] d) In one example, the process can be performed multiple times on different candidate sets.
[0153] 1. For example, a candidate set (such as non - adjacent MVP candidates) can be sorted, and the N non - adjacent MVP candidates with the lowest cost can be placed in a candidate list. After the entire candidate list is constructed, the cost of the candidates in the list can be calculated, and the candidates can be re - sorted based on the cost.
[0154] 13. It is proposed that the MVP list construction process can involve re - sorting of a single group / class and re - sorting of a combined group that includes candidates from more than one category.
[0155] a) In one example, the combined group can include candidates from a first category and a second category.
[0156] 1. Alternatively, in addition, the first category and the second category can be defined as non - adjacent MVP categories and HMVP categories.
[0157] 2. Alternatively, in addition, the first category and the second category can be defined as non - adjacent MVP categories and HMVP categories, and the combined group can include candidates from a third category (e.g., TMVP category).
[0158] b) In one example, the single group can include candidates from a fourth category.
[0159] 1. Alternatively, in addition, the fourth category can be defined as an adjacent MVP category.
[0160] 14. Multiple groups or categories can be re - sorted separately to construct the MVP list.
[0161] a) In one example, during the MVP list construction process, only one single group is constructed and re - sorted (all candidates belong to one category, e.g., adjacent MVP, non - adjacent MVP, HMVP, etc.).
[0162] b) In one example, during the MVP list construction process, only one combined group is constructed and re - sorted (including some or all candidates from multiple categories).
[0163] c) In one example, during the MVP list construction process, more than one group (whether single group or combined group) is constructed and re - sorted separately.
[0164] 1. In one example, during the MVP list construction process, two or more single groups are constructed and re - sorted separately.
[0165] 2. In one example, during the MVP list construction process, two or more combined groups are constructed and re - sorted separately.
[0166] 3. In one example, during the MVP list construction process, one or more single groups and one or more combined groups are re - sorted separately.
[0167] a) In one example, a single group and a combined group are respectively constructed and reordered to construct
[0168] an MVP list.
[0169] b) In one example, a single group and multiple combined groups are respectively constructed and reordered to construct
[0170] an MVP list.
[0171] c) In one example, multiple single groups and a combined group are respectively constructed and reordered to construct
[0172] an MVP list.
[0173] d) In one example, multiple single groups and multiple combined groups are respectively constructed and reordered to construct
[0174] an MVP list.
[0175] d) In one example, candidates belonging to the same category can be divided into different groups and reordered respectively in the corresponding groups.
[0176] e) In one example, only some of the candidates in a specific category are put into a single group or a combined group, and the remaining candidates in this category are not reordered.
[0177] f) In the above examples, the categories can be:
[0178] 1. Adjacent neighboring MVPs;
[0179] 2. (Multiple) Adjacent neighboring MVPs at specific positions;
[0180] 3. TMVP MVPs;
[0181] 4. HMVP MVPs;
[0182] 5. Non - adjacent MVPs;
[0183] 6. Constructed MVPs (such as paired MVPs);
[0184] 7. Inherited affine MV candidates;
[0185] 8. Constructed affine MV candidates;
[0186] 9. SbTMVP candidates.
[0187] 15. The proposed sorting method can also be applied to the AMVP mode.
[0188] a) In one example, the MVP in the AMVP mode can be extended by using non - adjacent MVPs, non - adjacent TMVPs, and HMVPs.
[0189] b) In one example, the MVP list for the AMVP mode includes K candidates selected from M categories such as adjacent MVPs, non - adjacent MVPs, non - adjacent TMVPs, and HMVPs, where K and M are integers.
[0190] 1. In one example, K can be less than M, or equal to M, or greater than M.
[0191] 2. In one example, one candidate is selected from each category.
[0192] 3. Alternatively, for a given category, no candidate is selected.
[0193] 4. Alternatively, for a given category, more than 1 candidate is selected.
[0194] 5. In one example, the MVP list for the AMVP mode includes 4 candidates selected from adjacent MVPs, non - adjacent MVPs, non - adjacent TMVPs, and HMVPs.
[0195] 6. In one example, the MVP candidates for each category are sorted separately using the template matching cost, and the MVP candidate with the minimum cost in the corresponding category is selected and included in the MVP list.
[0196] 7. Alternatively, the combined group of non - adjacent MVP, non - adjacent TMVP, and HMVP candidates and adjacent MVP candidates are sorted using the template matching cost separately. One adjacent candidate with the minimum template matching cost is selected from the adjacent MVP candidates, and the other three candidates are derived by traversing the candidates in the combined group in ascending order of the template matching cost.
[0197] 8. In one example, the MVP list for the AMVP mode includes 2 candidates, one from adjacent MVPs and the other from non - adjacent MVPs, non - adjacent TMVPs, or HMVPs. Specifically, the combined group of non - adjacent MVPs, non - adjacent TMVPs, and HMVPs and adjacent MVP candidates are sorted together using the template matching cost, and the MVP candidate with the minimum cost in the corresponding category (or group) is included in the MVP list.
[0198] 16. The proposed sorting method can be applied to other coding and decoding methods, for example, for constructing a block vector list for blocks coded and decoded by IBC.
[0199] a) In one example, it can be used for blocks coded and decoded by affine.
[0200] b) Alternatively, moreover, how to define the template cost can depend on the coding and decoding method.
[0201] 17. The use of this method can utilize different coding and decoding level syntax controls, including but not limited to one or more of the PU, CU, CTU, stripe, picture, and sequence levels.
[0202] 18. Regarding how to insert the sorted candidates into the MVP list.
[0203] a) In one example, which candidates within a combined group or a separate group are included in the MVP list depends on the sorting result of the template matching cost.
[0204] b) In one example, whether to put the candidates within a separate group or a combined group into the MVP list depends on the sorting result of the template matching cost.
[0205] c) In one example, how many candidates within a separate group or a combined group are included in the MVP list depends on the sorting result of the template matching cost.
[0206] 1. In one example, only one candidate with the minimum template matching cost is included in the MVP list.
[0207] 2. In a group, the top N candidates in ascending order of the template matching cost are included in the MVP list, where N is the maximum allowable number of candidates that can be inserted into the MVP list corresponding to a single group or a combined group.
[0208] a) In one example, for each single group or combined group, N can be a predefined constant.
[0209] b) Alternatively, N can be adaptively derived based on the template matching cost within a single group or a combined group.
[0210] c) Alternatively, N can be signaled in the bitstream.
[0211] d) In one example, different candidate groups share the same N value.
[0212] e) Alternatively, different single groups or combined groups can have different N values.
[0213] Duplicate removal for MVP candidates
[0214] 19. The deduplication for MVP candidates aims to increase the diversity within the MVP list, which can be achieved by using an appropriate threshold TH.
[0215] a) In one example, if two candidates point to the same reference frame, then they can both be included in the MVP list only if the absolute difference between the corresponding X and Y components is greater than (or not less than) TH.
[0216] 20. The deduplication threshold can be signaled in the bitstream.
[0217] b) In one example, the deduplication threshold can be signaled at the PU, CU, CTU, or slice level.
[0218] 21. The deduplication threshold can depend on the characteristics of the current block.
[0219] c) In one example, the threshold can be derived by analyzing the diversity between candidates.
[0220] d) In one example, the optimal threshold can be derived by RDO.
[0221] 22. Deduplication for MVP candidates can be first performed within a single group or a combined group before classification.
[0222] a) Alternatively, in addition, for two candidates belonging to two different groups or one belonging to a combined group and the other not belonging to the combined group, deduplication between these two MVP candidates is not performed before sorting.
[0223] b) Alternatively, in addition, deduplication between multiple groups can be applied after sorting.
[0224] 23. Deduplication for MVP candidates can be first performed between multiple groups, and sorting can be further applied to one or more single / combined groups.
[0225] a) Alternatively, the MVP list can be first constructed using deduplication between the available MVP candidates involved.
[0226] Then, sorting can be further applied to reorder one or more single / combined groups.
[0227] b) Alternatively, in addition, for two MVP candidates belonging to two different groups or one belonging to a combined group and the other not belonging to the combined group, deduplication between these two MVP candidates is performed before sorting.
[0228] Interaction with other codec tools
[0229] 24. After applying the sorting method to the MVP list, the Adaptive Reordering and Merging Candidates (ARMC) process can also be applied.
[0230] a) In one example, the template cost used during the sorting process during MVP list construction can be further utilized in the ARMC.
[0231] b) In another example, different template costs can be used during the classification process and the ARMC process.
[0232] 1. In one example, the template can be different for the sorting and ARMC processes.
[0233] 25. Whether and / or how to enable the sorting process can depend on the codec tool.
[0234] a) In one example, when a certain tool (e.g., MMVD or affine mode) is enabled for a block, the sorting is disabled.
[0235] b) In one example, for two different tools, the sorting rules can be different (e.g., applied to different groups or different template settings).
[0236] 2.6. Simplification of the video codec method based on template matching
[0237] The video codec method based on template matching is optimized in two aspects. First, the reference template derivation process is modified, and the interpolation process in the prediction block generation process is replaced in a different way. Second, several fast strategies are designed to accelerate the tools related to template matching.
[0238] It should be noted that the proposed method can be used for ARMC, EMCD, and template matching MV refinement, and can also be easily extended to other potential uses that require a template matching process, such as template matching-based candidate reordering for Merge with motion vector difference (MMVD), affine motion compensation, sub-block-based temporal motion vector prediction (SbTMVP), etc. In another example, the proposed method can be applied to other codec tools that require a motion information refinement process, such as codec tools based on bilateral matching.
[0239] The following detailed implementation examples should be regarded as examples for explaining general concepts. These embodiments should not be interpreted narrowly. In addition, these embodiments can be combined in any way. Combinations between this patent application and other documents are also applicable.
[0240] 1. It is proposed to replace the interpolation filtering process involved in the motion compensation process of generating the inter-frame prediction signal with other methods during the reference template generation process.
[0241] a) It is proposed to exclude the interpolation filtering process to generate the reference template, even if the motion vector points to a fractional position.
[0242] i. In one example, it is proposed to use integer precision to generate a reference template.
[0243] ii. In one example, if the motion vector points to a fractional position, it is first rounded to an integer MV.
[0244] 1. In one example, the fractional position is rounded towards zero (i.e., a negative motion vector prediction value is rounded towards positive infinity, and a positive motion vector prediction value is rounded towards negative infinity).
[0245] 2. In one example, the rounding step size can be greater than 1.
[0246] b) It is proposed to use different interpolation filters to generate a reference template for a motion vector pointing to a fractional position.
[0247] i. In one example, a simplified interpolation filter can be applied.
[0248] 1. In one example, the simplified interpolation filter can be a 2-tap bilinear one. Alternatively, it can also be a 4-tap, 6-tap, or 8-tap filter belonging to DCT, DST, Lanczos, or any other interpolation type.
[0249] ii. In one example, a more complex interpolation filter (e.g., with longer filter taps) can be applied.
[0250] c) The above methods can be used to reorder Merge candidates for the template matching Merge mode.
[0251] i. In one example, integer precision can be used in ARMC, EMCD, LIC, and any other potential scenarios.
[0252] ii. The above methods can be used to reorder candidates for the regular Merge mode.
[0253] 1. In one example, integer precision can be used to reorder candidates for the regular Merge mode.
[0254] d) In one example, whether to use the above methods (e.g., integer precision, different interpolation filters) and / or how to use the above methods can be determined in the bitstream by signaling or at runtime according to the decoded information.
[0255] i. In one example, the method to be applied can depend on the codec tool.
[0256] ii. In one example, the method to be applied can depend on the block size.
[0257] iii. In one example, integer precision can be used for a given color component (e.g., luminance only).
[0258] iv. Alternatively, integer precision can be used for all three components.
[0259] 2. Whether and / or how to perform EMCD can be based on the maximum allowable number of candidates within the candidate list and / or the number of available candidates before being added to the candidate list.
[0260] a) In one example, assuming the number of available candidates (valid candidates that can be used to construct the candidate list) is NAVAL, and the maximum allowable number of candidates is NMAX (i.e., at most NMAX candidates can be included in the final Merge list), then EMCD is enabled only when NAVAL - NMAX is greater than a constant or an adaptively derived threshold T.
[0261] 3. It is proposed to organize the available Merge candidates into subgroups.
[0262] a) In one example, the available candidates can be classified into subgroups, each subgroup containing a fixed or adaptively derived number of candidates, and each subgroup selects a fixed number of candidates into the list. On the decoder side, only the candidates within the selected subgroups need to be reordered.
[0263] b) In one example, the candidates can be classified into subgroups according to the category of the candidates, such as non - adjacent MVP, temporal MVP (TMVP), or HMVP, etc.
[0264] 4. A piece of information calculated by a first codec tool using at least one template cost can be reused by a second codec tool using at least one template cost.
[0265] a) It is proposed to construct a unified storage shared by ARMC, EMCD, and any other potential tools to store information for each Merge candidate.
[0266] b) In one example, this storage can be a map, table, or other data structure.
[0267] c) In one example, the information stored can be the template matching cost.
[0268] d) In one example, EMCD first traverses all the MVs associated with the available candidates and stores the corresponding information (including but not limited to the template matching cost) in this storage. Then, ARMC and / or other potential tools can simply access the required information from this shared storage without performing repeated calculations.
[0269] 2.7. Extension of Motion Vector Prediction List Construction Based on Template Matching Cost Sorting
[0270] The present disclosure proposes an optimized MVP list derivation method based on template matching cost sorting. An optimized MVP selection method is studied by utilizing the matching cost in the reconstructed template region, and the MVP list is constructed without relying on a predefined traversal order, such that more suitable candidates are included in the list.
[0271] It should be noted that the proposed strategy for MVP list construction can be used in the conventional Merge and AMVP list construction processes, and can also be easily extended to other modules that require MVP derivation, such as Merge with Motion Vector Difference (MMVD), affine motion compensation, Sub - block - based Temporal Motion Vector Prediction (SbTMVP), and so on.
[0272] In the following discussion, category represents the attribution of MVP candidates. For example, non - adjacent MVP candidates belong to one category, and HMVP candidates belong to another category. Group represents a set of MVP candidates that contains one or more MVP candidates. In one example, a single group represents a set of MVP candidates where all candidates belong to one category, such as adjacent MVPs, non - adjacent MVPs, HMVP, etc. In another example, a combined group represents a set of MVP candidates that contains candidates from multiple categories.
[0273] The following detailed embodiments should be regarded as examples for explaining general concepts. These embodiments should not be interpreted narrowly. In addition, these embodiments can be combined in any way. Combinations with other documents for this patent application are also applicable.
[0274] 5. Multiple thresholds can be utilized in the candidate deduplication process to determine whether a candidate can be added to the candidate list.
[0275] a) The threshold can be used to determine whether a potential candidate can be put into the candidate list.
[0276] i. For example, if the absolute difference between at least one component of the MV of the potential candidate and at least one component of the MV of a candidate existing in the candidate list is less than the threshold, the potential candidate is not put into the list.
[0277] ii. For example, if the absolute difference between all components of the MV of the potential candidate and all components of the MV of a candidate existing in the candidate list is less than the threshold, the potential candidate is not put into the list.
[0278] b) In one example, the candidate is an MVP candidate, the candidate deduplication process is an MVP candidate deduplication process, and the candidate list is a motion candidate list.
[0279] i. In one example, the motion candidate list is a Merge candidate list.
[0280] ii. In one example, the motion candidate list is the AMVP candidate list.
[0281] iii. In one example, the motion candidate list is an extended Merge or AMVP list, such as a sub-block Merge candidate list, an affine Merge candidate list, an MMVD list, a GPM list, a template matching Merge list, a bilateral matching Merge list, etc.
[0282] c) In one example, for two groups, the deduplication threshold can be different, where the group can be a single group (containing candidates of only one category) or a combined group (containing candidates of at least two categories).
[0283] d) Alternatively, for all potential MVP candidates, only one threshold is used, regardless of category and / or group.
[0284] e) In one example, N (e.g., N = 2) thresholds are used during the deduplication process.
[0285] i. Assume A is the MVP set containing all available MVP candidates, regardless of category. In one example, the first threshold is used for a first subset of the candidates in set A, and the second threshold is used for a second subset of the candidates in set A (e.g., the remaining candidates excluding those in the first subset).
[0286] ii. In one example, the first threshold is used for a single group represented by A, and the second threshold is used for another group (single group or combined group) / multiple other groups / remaining candidates that do not have the same category as those in A.
[0287] 1) In one example, the first threshold is used for adjacent candidates of a single group, and the second threshold is used for the remaining candidates, including but not limited to non-adjacent MVPs, HMVPs, paired MVPs, and zero MVPs.
[0288] iii. The first threshold can be greater than or less than the second threshold.
[0289] f) Alternatively, in addition, the threshold for an MVP category or group can depend on the decoded information, such as block size / encoding method (e.g., CIIP / MMVD) and / or the variance of the motion information within the category or group.
[0290] 6. Multiple passes of reordering can be performed to construct the MVP list.
[0291] a) In one example, the multiple passes can involve different reordering criteria.
[0292] b) In one example, multiple single groups / union groups can be reordered in multiple passes, where at least two single / union groups can have (or not have) overlapping MVP candidates.
[0293] c) In one example, K-pass (e.g., K = 2) reordering is used to build the MVP list.
[0294] i. In one example, in the first pass, first, based on the first cost (e.g., template matching cost)
[0295] reorder single group / union group A, and identify the candidate with the maximum cost (CL) in A, then transfer it to another single group / union group B (e.g., B can include the remaining candidates that do not have the same category as the candidates in A). Subsequently, group B performs 2 to K-pass reordering based on the first cost (or other cost metric) sorting. Finally, according to the sorted order, include the candidates in group A (except CL) and B (including CL) in the MVP list.
[0296] ii. In one example, group A in the above case is a single group of adjacent candidates, and group B is a union group of non-adjacent candidates and HMVP.
[0297] iii. Alternatively, groups A and B can be any other single or union candidate groups.
[0298] iv. In one example, in the first pass, first, based on the first cost (e.g., template matching cost)
[0299] reorder one or more single / union groups. Then, build a preliminary MVP list by inserting some candidates from each group into a list with a sorted order. Subsequently, the preliminary MVP list performs a second pass reordering to select some candidates into the final MVP list.
[0300] 1) In one example, different single / union groups can have (or not have) overlapping candidates.
[0301] 2) In one example, select all candidates in the preliminary MVP list from the sorted single / union groups.
[0302] 3) Alternatively, select some candidates in the preliminary MVP list from the sorted groups, and include the remaining candidates in a list with other rules.
[0303] 4) In one example, in the second pass, sort all candidates in the preliminary list based on the cost (e.g., template matching cost), regardless of the corresponding category, and include a limited number of candidates in the final MVP list only based on the sorted order.
[0304] a) Alternatively, in addition, all candidates in the preliminary MVP list are included in the final MVP list according to the sorted order.
[0305] 5) The costs calculated in previous passes (e.g., template matching costs) can be reused in later passes.
[0306] a) In one example, when calculating the cost for a certain candidate in a previous pass, it is saved in a variable or any other data structure in case the same cost is needed in a later pass.
[0307] b) In one example, in a later pass, if the cost for a certain candidate is needed,
[0308] then it will first be checked whether the cost has been calculated before. If the cost has been calculated and / or saved before the current pass, and / or is accessible in the current pass, it will be retrieved in the current pass instead of being calculated again.
[0309] 7. At least one virtual candidate (e.g., paired MVP and zero MVP) can be involved in at least one group.
[0310] a) In one example, all virtual candidates are processed using a combined group.
[0311] i. Alternatively, virtual candidates of each category are considered as a single group.
[0312] ii. In one example, paired MVP and / or zero MVP are included in a single / combined group.
[0313] iii. Alternatively, in addition, the groups containing virtual candidates are reordered and then put into the candidate list.
[0314] b) Alternatively, virtual candidates (e.g., paired MVP and / or zero MVP) are not included in any single / combined group.
[0315] i. Alternatively, in addition, no reordering process is applied to virtual candidates.
[0316] 6) Alternatively, in addition, they can be further appended to the candidate list.
[0317] ii. In one example, one or more single / combined groups are constructed, and some or all of the groups are reordered. In this case, at least one position in the MVP list is reserved for virtual candidates (e.g., paired MVP and / or zero MVP), which are appended to the MVP list as the last entry or any other entry.
[0318] iii. In one example, in addition, a single group of adjacent candidates is first included in the MVP list, and then a combined group of non - adjacent candidates and HMVP is reordered and subsequently appended to the MVP list. In this case, at least one position is reserved for virtual candidates (e.g., paired MVP and / or zero MVP), and the virtual candidates are appended to the MVP list as the last entry or any other entry.
[0319] iv. In addition, in one example, a combined group of adjacent candidates, non - adjacent and HMVP is reordered and subsequently appended to the MVP list, and virtual candidates (e.g., paired MVP and / or zero MVP)
[0320] are appended to the MVP list as the last entry or any other entry.
[0321] c) Alternatively, virtual candidates of one category (e.g., paired MVP) are included in the single / combined group, and virtual candidates of another category are not included.
[0322] d) In one example, when a reordering operation is performed for the construction of the MVP list, virtual candidates (e.g., paired MVP and / or zero MVP) do not appear in the final MVP list.
[0323] 8. The number of candidates in the single / combined group may not be allowed to exceed the maximum number of candidates.
[0324] a) In one example, the single / combined group is constructed to have a finite number of candidates constrained by a maximum number N i where i ∈ [0, 1, …, K] is the index of the corresponding group. For different i, N i may be the same, or may not be the same.
[0325] b) In one example, some of the candidates in the single / combined group are restricted by a maximum number N i of.
[0326] i. In one example, candidates of one or more categories in the group are constructed to have a finite number N i while other categories in the same group may include any number.
[0327] 7) In one example, the categories include but are not limited to adjacent candidates, non - adjacent candidates, HMVP,
[0328] paired candidates, etc.
[0329] c) Alternatively, the first single / combined group may be constructed to have at most N i MVP candidates, while the second single
[0330] The combined group may not have this constraint.
[0331] d) In one example, N i is a fixed value shared by both the encoder and the decoder.
[0332] i. Alternatively, N i is determined by the encoder and signaled in the bitstream. And the decoder decodes the N i value and then constructs the corresponding i-th single / combined group with at most N i candidates.
[0333] ii. Alternatively, N is derived in both the encoder and the decoder with the same operation i such that signaling of the N i value is not required.
[0334] 1) In one example, the encoder and the decoder may derive the N i value based on the variance of all available motion information of the i-th group.
[0335] 2) Alternatively, the encoder and the decoder may derive the N i value based on the number of all available candidates of the i-th group.
[0336] 3) In one example, the encoder and the decoder may derive the N i value based on the number of available neighboring candidates.
[0337] a) In one example, N i is set to N - N ADJ where N is a constant and N ADJ is the number of available neighboring candidates.
[0338] 4) Alternatively, in addition, the encoder and the decoder may derive the N i value based on any information that the encoder / decoder can access when constructing the MVP list.
[0339] e) In one example, all or part of the single / combined groups may share the same maximum number of candidates N.
[0340] 9. The construction of the single / combined group may depend on the maximum number constraint N i .
[0341] a) In one example, all available MVP candidates for the i-th group are included in the group in a specific order. Once the number of candidates in the current group reaches N i , the construction of group i is terminated.
[0342] b) In one example, in the above case, the order for group construction can be derived based on the distance between the CU to be coded / decoded and the MVP candidates, where the closer MVP candidate is assigned a higher priority.
[0343] c) Alternatively, the order can be derived based on the cost (e.g., template matching) cost, where the MVP with a lower cost has a higher priority.
[0344] d) In one example, the construction of a single / union group is performed using at least one deduplication operation within or between at least one group.
[0345] e) In one example, the constructed single / union group is further reordered based on at least one cost method (e.g., template matching cost), and then some or all of the candidates in the group can be included in the MVP list.
[0346] i. Alternatively, the candidates in the constructed single / union group will not be further reordered, and some or all of the candidates in the group are included in the MVP list in the same order as they are included in the group.
[0347] 10. Regarding how to deduplicate MVP candidates.
[0348] a) In one example, K passes (e.g., K = 2) of deduplication are performed to construct the MVP list.
[0349] 1) In one example, the first pass of deduplication can be performed within at least one single / union group, and the second pass of deduplication can be performed between at least two candidates belonging to different groups.
[0350] a) In one example, in the first pass of deduplication, the deduplication thresholds for two single / union groups can be the same or different.
[0351] b) In one example, in addition, in the first pass of deduplication, some of the single / union groups can share the same threshold, while other single / union groups can use different thresholds.
[0352] 2) In one example, in addition, the threshold for a specific pass or group is determined by the decoding information, including but not limited to the block size, coding / decoding tools used (e.g., TM, DMVR, adaptive DMVR, CIIP, AFFINE, AMVP-Merge).
[0353] a) Alternatively, the threshold can be determined by at least one syntax element signaled to the decoder.
[0354] 3. Problem
[0355] 1) The goal of existing MVP list construction methods is to build a subset with a constant number of MVPs from a given candidate set, which is typically achieved by selecting available candidates in a predefined order. However, this strategy does not utilize the prior information generated during the encoding / decoding process, which may lead to a mismatch between the true motion information and the motion information of the candidates in the constructed MVP list.
[0356] Existing deduplication processes for MVP candidates only consider identical MVs as redundant. Therefore, the constructed MVP list may contain very similar MVs, resulting in limited diversity within the list.
[0357] 4. Detailed solutions
[0358] In the present disclosure, an enhanced MVP list derivation method based on template matching cost sorting is proposed. Instead of constructing the MVP list based on a predetermined traversal order, the method studies and optimizes the MVP selection method by utilizing the matching cost in the reconstructed template region, such that more suitable candidates are included in the list.
[0359] It should be noted that the proposed strategy for MVP list construction can be used in the normal Merge and AMVP list construction processes, and can also be easily extended to other modules that require MVP derivation, such as Merge with Motion Vector Difference (MMVD), affine motion compensation, Sub - block based Temporal Motion Vector Prediction (SbTMVP), etc.
[0360] In the following discussion, a category represents the belongings of an MVP candidate. For example, non - adjacent MVP candidates belong to one category, and HMVP candidates belong to another category. A group represents a set of MVP candidates that contains one or more MVP candidates. In one example, a single group represents an MVP candidate set in which all candidates belong to one category, such as adjacent MVPs, non - adjacent MVPs, HMVP, etc. In another example, a combined group represents an MVP candidate set that contains candidates from multiple categories.
[0361] In the following discussion, functions such as SAD / SATD / SSD / MR - SAD (Mean Removed SAD) can be utilized to derive the "cost" of a candidate based on template matching or bilateral matching.
[0362] The following detailed embodiments should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow manner. In addition, these embodiments can be combined in any way. Combinations between this patent application and other patent applications are also applicable.
[0363] 1. During the candidate deduplication process, multiple thresholds can be utilized to determine whether a candidate can be added to the candidate list.
[0364] a) A threshold can be used to determine whether a potential candidate can be placed in the candidate list.
[0365] i. For example, if the absolute difference between at least one component of the MV of a potential candidate and at least one component of a candidate present in the candidate list is less than the threshold, the potential candidate is not placed in the list.
[0366] ii. For example, if the absolute difference between all components of the MV of a potential candidate and all components of a candidate present in the candidate list is less than the threshold, the potential candidate is not placed in the list.
[0367] b) In one example, the candidate is an MVP candidate, the candidate deduplication process is an MVP candidate deduplication process, and the candidate list is a motion candidate list.
[0368] i. In one example, the motion candidate list is a Merge candidate list.
[0369] ii. In one example, the motion candidate list is an AMVP candidate list.
[0370] iii. In one example, the motion candidate list is an extended Merge or AMVP list, such as a sub-block Merge candidate list, an affine Merge candidate list, an MMVD list, a GPM list, a template matching Merge list, a bilateral matching Merge list, etc.
[0371] iv. In one example, the motion candidate list is an IBC Merge candidate list.
[0372] v. In one example, the motion candidate list is an IBC AMVP candidate list.
[0373] vi. In one example, the motion candidate list is an extended IBC Merge or IBC AMVP list, such as an IBC-MMVD list.
[0374] c) In one example, for two groups, the deduplication threshold can be different, where the group can be a single group (containing candidates of only one category) or a combined group (containing candidates of at least two categories).
[0375] d) In one example, N (e.g., N = 2) thresholds are used in the deduplication process.
[0376] i. Suppose A is an MVP set containing all available MVP candidates (regardless of category). In one example, the first threshold is used for a first subset of the candidates in set A, and the second threshold is used for a second subset of the candidates in set A (e.g., the remaining candidates excluding those in the first subset).
[0377] ii. In one example, the first threshold is used for a single group represented by A, and the second threshold is used for another group (single group or combined group) / multiple other groups / the remaining candidates that do not have the same class as those candidates in A.
[0378] 1) In one example, the first threshold is used for a single group of adjacent candidates, and the second threshold is used for the remaining candidates, including but not limited to non - adjacent MVPs, HMVPs, paired MVPs, and zero MVPs.
[0379] iii. The first threshold can be greater than or less than the second threshold.
[0380] e) In one example, K passes (e.g., K = 4) of deduplication are performed to build the MVP list.
[0381] i. In one example, the first pass of deduplication (referred to as P1) is performed within a single group or combined group to avoid duplicate candidates.
[0382] 1) In one example, some or all of the groups can be sorted (i.e., ARMC) after P1.
[0383] ii. In one example, when multiple groups are merged into one or more hybrid groups, the second pass of deduplication (referred to as P2) is performed.
[0384] 1) In one example, after P2, the (multiple) hybrid groups may or may not perform sorting.
[0385] iii. In one example, some new candidates can be inserted into the hybrid group, and the third pass of deduplication is triggered to ensure that there are no duplicates after adding the new candidates.
[0386] iv. In one example, the fourth pass of deduplication (referred to as P4) is performed to further increase the
[0387] diversity within the (multiple) hybrid group.
[0388] v. In one example, the above multi - pass deduplication can be utilized in a separate or combined manner.
[0389] 1) In one example, only some of the passes are used to build the MVP list, i.e., P1 -> P2 -> p4,
[0390] P1 -> P2 -> p3, P1 -> P2, P1 -> P3, P1 -> P3 -> p4, P1 -> p4, etc.
[0391] 2) In one example, the order of each pass can be changed during the construction process, i.e., later pass deduplication can be performed before earlier pass deduplication.
[0392] 3) In one example, a certain deduplication process can be performed multiple times during construction.
[0393] a) In one example, deduplication can be performed in the order of P1 -> P4 -> P2 -> P3 -> p4. vi. In one example, the thresholds used in different passes can be the same or different.
[0394] 1) In one example, the threshold in a certain pass deduplication can be a constant.
[0395] 2) In one example, the threshold in a certain pass deduplication can be derived from the bitstream.
[0396] a) In one example, all available threshold values can be stored in a lookup table or any other data structure, and the index of the selected threshold is signaled in the bitstream.
[0397] The decoder can first parse the threshold index and then obtain the threshold from the corresponding lookup table or other data structure.
[0398] b) In one example, the threshold can be derived based on the information of the current block (i.e., the QP or Lagrange multiplier (Lamda) used in the RDO process).
[0399] 2. Regarding how to construct the MVP list.
[0400] a) One or more groups can be constructed first, where each group includes candidates belonging to one or more categories.
[0401] i. In one example, the categories can include but are not limited to adjacent MVPs, non - adjacent MVPs,
[0402] HMVP, paired MVPs, constructed MVPs, etc.
[0403] ii. In one example, the number of candidates in each group can be not allowed to exceed a certain value.
[0404] 1) In one example, the maximum allowed number for each group can be a constant or determined instantaneously.
[0405] 2) In one example, the maximum allowed number for each group can be different.
[0406] iii. In one example, if only one group is constructed, the candidates belonging to different categories are inserted into the group based on a predefined order.
[0407] 1) In one example, specifically, in the constructed group, the number of candidates of a specific one or more categories cannot exceed a constant or a value determined during operation.
[0408] iv. The deduplication operation can be performed during the construction of each group.
[0409] 1) In one example, specifically, deduplication is performed within the group, that is, there are no duplicates for any two candidates from any one group.
[0410] 2) Alternatively, deduplication is performed between groups, that is, there are no duplicates for any two candidates from any one or two groups.
[0411] v. The deduplication threshold for any two groups can be the same or different.
[0412] b) If multiple groups have been constructed, some or all of the groups can be merged into a mixed group.
[0413] i. In one example, if only one group is constructed in a), no merging process is performed, and
[0414] this group will be regarded as a special case of the mixed group.
[0415] ii. In one example, specifically, if deduplication has not been performed on the group(s) before merging or in-group deduplication has been performed, a second pass of deduplication is performed during the merging process.
[0416] iii. Alternatively, specifically, if inter-group deduplication has been performed on the group(s) before merging, no deduplication is performed during the merging process.
[0417] c) Subsequently, the mixed group can be sorted based on ARMC or any other metric.
[0418] i. In one example, specifically, before or after sorting the mixed group, all or part of the candidates within the group can be refined by template or bilateral matching.
[0419] ii. In one example, specifically, zero MVPs are excluded during the sorting process, and zero MVPs can be forced to the end of the sorted list.
[0420] d) The constructed candidates (i.e., paired candidates) can be generated and / or inserted into the mixed group, and / or another round of sorting can be invoked to re-sort the extended group.
[0421] i. In one example, the constructed candidates can be generated based on the sorted groups.
[0422] 1) In one example, specifically, the constructed candidates can be paired candidates.
[0423] 2) In one example, specifically, the constructed candidates are inserted into the mixed group (along with the deduplication operation).
[0424] e) Finally, perform a final round of deduplication to further increase the diversity within the (multiple) larger groups.
[0425] i. In one example, calculate the template matching cost for all candidates in the sorted list, and determine the minimum cost difference between a candidate and its previous candidate among all candidates in the list. If the minimum cost difference is less than TH, the candidate will be discarded and moved to another position in the list. This other position is the first position where the cost difference relative to its previous candidate is greater than TH. The algorithm stops after a finite number of iterations, or the remaining number of candidates reaches the target value for the MVP list.
[0426] a) In one example, TH can be derived based on the information of the current block (i.e., the QP or Lagrange multiplier (Lam da) used in the RDO process).
[0427] 3. The disclosed method can be applied to potential candidates before they are put into the candidate list, or can be applied to candidates after they are put into the candidate list.
[0428] 4. General claims
[0429] 1) Whether and / or how to apply the method disclosed above can be signaled at the sequence level / picture group level / picture level / strip level / slice group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / slice group header.
[0430] 2) Whether and / or how to apply the method disclosed above can be signaled in PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / strip / slice / sub-picture / other types of regions containing more than one sample or pixel.
[0431] 3) Whether and / or how to apply the method disclosed above can depend on the decoded information, such as block size, color format, single / double-tree segmentation, color component, strip / picture type.
[0432] 5. Regarding how to generate a reference template for a block with sub - block level motion information.
[0433] a) The template includes the reconstructed region in the upper and / or left position adjacent to the current block.
[0434] b) In one example, to obtain the reference template for the current template T, T can first be divided into one or more template segments.
[0435] i. In one example, the width of each segment (for the upper template segment) and / or the height (for the left template segment) is equal to the size of the sub-blocks in the current block, as Figure 11 shown, Figure 11 FIG. 1100 showing an example of a template for a block with sub-block level motion information.
[0436] c) In one example, for each template segment, the motion information of at least one adjacent / non-adjacent sub-block is used to obtain the reference template segment. For example, the motion information of sub-block A can be used to generate the reference template segments for T U_A and T L_A , and / or the motion information of sub-block C can be used to generate the reference template segment for T U_C , as Figure 11 shown.
[0437] i. In one example, the reference template segment can be unidirectionally predicted or bidirectionally predicted.
[0438] 1) In one example, all the reference template segments of the current block are unidirectionally predicted (or bidirectionally predicted).
[0439] 2) In one example, for the template segments belonging to one coding block, some of them can be unidirectionally predicted while others can be bidirectionally predicted.
[0440] ii. In one example, for any template segment T _SEG , if the adjacent sub-block is bidirectionally predicted (i.e., has two sets of motion information, each set including the corresponding MV, reference, etc.), then either both sets of information or any one of them can be used to generate the reference template (segment).
[0441] 1) In one example, the motion information in one set corresponds to a specific reference list.
[0442] 2) In one example, specifically, both sets of motion information can be used to generate the prediction of the template. In particular, two reference segment predictions of the current template segment specified by the corresponding sets of motion information are generated respectively, and then the weighted average of the two predictions is used as the reference template segment.
[0443] 3) In one example, only one of the two sets of motion information is used to generate the reference template segment.
[0444] a) In one example, which set of motion information to use can depend on the motion information of the central sub-block of the current block.
[0445] 4) In one example, before generating the prediction or the reference template, (multiple)
[0446] The MV is scaled to a specific reference frame.
[0447] iii. In one example, for any template segment T _SEG , if the adjacent sub-block is unidirectionally predicted (i.e., has only one set of motion information), the reference region of the current template segment specified by the same set of motion information is used as the reference template region.
[0448] 1) In one example, alternatively, for any template segment T _SEG , even if the corresponding sub-block is unidirectionally predicted, the reference template segment of T can be generated using bidirectional prediction. _SEG of T.
[0449] a) In one example, in addition to the unidirectional motion information of the adjacent sub-block, one or more of the following methods can be used to construct an additional set of motion information:
[0450] i. Use a zero MV.
[0451] ii. Scale an existing MV to another reference frame.
[0452] iii. Mirror an existing MV.
[0453] iv. Obtain from non-adjacent sub-blocks.
[0454] 2) In one example, alternatively, for any template segment T _SEG , if the adjacent sub-block is unidirectionally predicted, a zero MV can be used instead of the motion information of the adjacent sub-block to generate the reference template.
[0455] d) In one example, for each template segment of the current template, whether the corresponding reference template is unidirectionally predicted or bidirectionally predicted, and / or which reference list is used to provide the motion information for the reference template segment can depend on the motion information of a specific sub-block (such as the central sub-block of the current coded block or a sub-block located at any other position).
[0456] i. In one example, if a specific sub-block of the current coded block is bidirectionally predicted, all reference template segments can be bidirectionally predicted.
[0457] 1) In one example, specifically, for each template segment, the motion information of the adjacent sub-block is obtained to generate the reference template segment. If the adjacent sub-block has only unidirectional motion information (i.e., the motion information comes from one reference list), additional motion information (which can be associated with other reference lists) is constructed based on one or more of the following methods:
[0458] i. Use a zero MV.
[0459] ii. Scale the existing MV to another reference frame.
[0460] iii. Mirror the existing MV.
[0461] iv. Obtain from non - adjacent sub - blocks.
[0462] 2) In one example, in the above case, the reference index of the constructed motion information can be the same as the reference index of the motion information of a specific sub - block associated with the same reference list.
[0463] ii. In one example, alternatively, if a specific sub - block of the current coded block is unidirectionally predicted, all reference template segments are unidirectionally predicted.
[0464] 1) In one example, specifically, for each template segment, the motion information associated with the same reference list as that of the specific sub - block obtained from an adjacent sub - block is used to generate the reference template segment. If the corresponding motion information does not exist in the adjacent sub - block, the corresponding motion information can be constructed based on one or more of the following methods:
[0465] i. Use a zero MV.
[0466] ii. Scale the existing MV to another reference frame.
[0467] iii. Mirror the existing MV.
[0468] iv. Obtain from non - adjacent sub - blocks.
[0469] 2) In one example, in the above case, the reference index of the constructed motion information can be the same as the reference index of the motion information of a specific sub - block associated with the same reference list.
[0470] e) The above method can be applied to any coding tool with sub - block level motion information, including but not limited to AFFINE, SbTMVP, etc.
[0471] 5. Embodiments
[0472] In one example, when the encoder / decoder starts building the MVP candidate list, multiple small groups are first built, where each small group includes candidates from one or more categories. Specifically, the number of candidates in each group should not exceed the maximum allowed number, where the maximum number can vary from one group to another. In addition, in-group deduplication operations with a constant threshold are performed together with the construction of each group. After each group is built, all or part of them will be further merged into a hybrid group, where a second pass deduplication is triggered to exclude redundant candidates in the larger group. Then, all or part of the candidates in the hybrid group are sorted based on the ARMC method, and it should be noted that before ARMC, all or part of the candidates can be refined first by template matching or bilateral matching. Based on the sorted hybrid group, some constructed candidates (i.e., paired candidates) can be generated and then inserted into the hybrid group (along with a third pass deduplication operation). And the extended hybrid group performs ARMC again, and all candidates are sorted based on the TM cost. Finally, if the number of candidates in the hybrid group is greater than the maximum allowed value for the MVP list, a final pass deduplication operation is performed. Specifically, the template matching cost for all candidates in the sorted group is calculated, and the minimum cost difference between a candidate and its previous candidate among all candidates is determined. If this minimum cost difference is less than a constant TH, the candidate will be discarded and moved to a farther position in the list. This farther position is the first position where the cost difference relative to its previous candidate is greater than TH. The algorithm stops after a finite number of iterations or after the remaining number of candidates reaches the target value for the MVP list.
[0473] Figure 11 FIG. 1100 is a flowchart of a method 1100 for video processing according to an embodiment of the present disclosure. Method 1100 can be implemented for conversion between a current video block of a video and a bitstream of the video.
[0474] In block 1210, a reference template for the current video block is determined based on a current template of the current video block and motion information of at least one sub-block of the current video block. In block 1220, the conversion is performed based on the reference template.
[0475] Method 1200 enables determination of a reference template for a block based on sub-block level motion information. The determined reference template can be more accurate. In this way, codec effectiveness and codec efficiency can be improved.
[0476] In some embodiments, the current template of the current video block includes at least one of the following: a first set of reconstructed regions above and adjacent to the current video block, or a second set of reconstructed regions to the left and adjacent to the current video block. In other words, the current template of the current video block (also referred to as "template") may include reconstructed regions adjacent to the current video block at the upper or / and left position, as Figure 11 shown.
[0477] In some embodiments, determining a reference template includes: determining a plurality of template segments of the current template; determining a plurality of reference template segments based on the plurality of template segments and the motion information of at least one sub-block; and determining a reference template based on the plurality of reference template segments. For example, the template segments of the current template may include segment T U_A , T U_B , T U_C , T U_D , T L_A , T L_E , T L_F , T L_G , T L_A , T L_A , T L_A , T L_A , as Figure 11 shown. A plurality of reference template segments can be determined based on these template segments and the motion information of the sub-blocks associated with these template segments.
[0478] In some embodiments, determining a plurality of reference template segments includes: for a first template segment among the plurality of template segments, determining a first sub-block associated with the first template segment from at least one sub-block; and determining a first reference template segment among the plurality of reference template segments based on the motion information of the first sub-block. For example, the first sub-block can be a sub-block adjacent to the first template segment or a sub-block not adjacent to the first template segment. That is, the motion information of the adjacent or non-adjacent sub-blocks of the template segment can be used to determine the corresponding reference template segment. For example, the motion information of sub-block A in Figure 11 can be used to generate reference template segments for T U_A , T L_A or for T U_B . Another example is that Figure 11 the motion information of sub-block C in U_C can be used to generate a reference template segment for T
[0479] In some embodiments, the first sub-block adjacent to the first template segment is bi-predicted, and at least one of the two motion information sets of the first sub-block is used to determine the first reference template segment.
[0480] In some embodiments, the motion information set of the first sub-block is associated with a reference list, and the motion information set includes at least one of a motion vector or a reference frame in the reference list.
[0481] In some embodiments, determining the first reference template segment includes: determining two predictions of the first reference segment based on two motion information sets; and determining the first reference template segment based on a weighted average of the two predictions. For example, two reference segment predictions can be generated by obtaining the reference regions of the current template segment specified by the corresponding motion information sets. Then, the weighted average of the two predictions can be used as the reference template segment.
[0482] In some embodiments, determining two predictions of the first reference segment includes: updating the motion information sets of the two motion information sets by scaling the motion vectors in the motion information set to the reference frame; and determining the corresponding predictions of the first reference segment based on the updated motion information sets. For example, the (multiple) MVs can be scaled to a specific reference frame first before generating the predictions or the reference template.
[0483] In some embodiments, determining the first reference template segment includes: selecting a motion information set from the two motion information sets based on the motion information of the second sub-block of the current video block; and determining the first reference template segment based on the selected motion information set.
[0484] In some embodiments, the second sub-block includes the central sub-block of the current video block. Alternatively, the second sub-block can be a sub-block at a certain position of the current video block.
[0485] In some embodiments, the first sub-block adjacent to the first template segment is unidirectionally predicted. In some embodiments, the region of the first reference template segment is determined based on a single motion information set of the first sub-block. For example, for any template segment T _SEG , if the adjacent sub-block is unidirectionally predicted, that is, has only one motion information set, the reference region of the current template segment specified by the same motion information set is used as the reference template region.
[0486] In some embodiments, a first set of motion information is associated with a first sub-block, and determining a first reference template segment includes: determining a second set of motion information based on at least one of: a zero motion vector, a scaled motion vector of a first motion vector in the first set of motion information, the scaled motion vector being associated with a second reference frame different from the first reference frame of the first motion vector, a mirrored motion vector of a second motion vector in the first set of motion information, a third set of motion information of a non-adjacent sub-block of the current video block; and determining the first reference template segment based on the first set of motion information and the second set of motion information. In other words, for any template segment, even if the corresponding sub-block is unidirectionally predicted, the reference template segment of the template segment can be generated using bidirectional prediction. For example, in addition to the unidirectional motion information of adjacent sub-blocks, an additional set of motion information can be constructed based on one or more of the methods described above.
[0487] In some embodiments, a first reference index of the second set of motion information is the same as a second reference index of the motion information of a second sub-block of the current video block, and the first set of motion information and the motion information of the second sub-block are from the same reference list. That is, the reference index of the constructed motion information can be the same as the reference index of the motion information of a specific sub-block associated with the same reference list. The specific sub-block can be a central sub-block or a sub-block at a specific position.
[0488] In some embodiments, the motion information associated with the reference list of the first sub-block is unavailable, and determining the first reference template segment includes: determining the motion information of the first sub-block based on at least one of: a zero motion vector, a scaled motion vector of an existing motion vector associated with the current video block, the scaled motion vector being associated with a second reference frame different from the existing motion vector, a mirrored motion vector of an existing motion vector associated with the current video block, additional motion information of a non-adjacent sub-block of the current video block; and determining the first reference template segment based on the motion information of the first sub-block.
[0489] In some embodiments, if a specific sub-block of the current video block is unidirectionally predicted, then all reference template segments are unidirectionally predicted. In one example, for each template segment, the motion information associated with the same reference list as the specific sub-block obtained from adjacent sub-blocks is used to generate the reference template segment. If the corresponding motion information does not exist in the adjacent sub-blocks, the corresponding motion information can be constructed based on one or more of the methods described above.
[0490] In some embodiments, the reference list of the first sub-block is the same as the reference list of the second sub-block, and a first reference index of the motion information of the first sub-block is the same as a second reference index of the motion information of the second sub-block. For example, the reference index of the constructed motion information can be the same as the reference index of the motion information of a specific sub-block associated with the same reference list.
[0491] In some embodiments, the second sub-block includes one of the following: the central sub-block of the current video block, or a sub-block at a predetermined position of the current video block.
[0492] In some embodiments, the first reference template segment is determined based on a zero motion vector. That is, for any template segment, if the adjacent sub-block is unidirectionally predicted, a zero MV can be used instead of the motion information of the adjacent sub-block to generate the reference template segment of the reference template.
[0493] In some embodiments, the template segment among the multiple template segments is above the current video block, and the width of the template segment is equal to the width of the sub-block in the current video block.
[0494] In some embodiments, the template segment among the multiple template segments is to the left of the current video block, and the height of the template segment is equal to the height of the sub-block in the current video block.
[0495] In some embodiments, the motion information of a single sub-block of the current video block is used to determine one or more reference template segments.
[0496] In some embodiments, the reference template segments among the multiple reference template segments are unidirectionally predicted or bidirectionally predicted.
[0497] In some embodiments, the multiple reference template segments are unidirectionally predicted or bidirectionally predicted.
[0498] In some embodiments, for a set of template segments among the multiple template segments, the set of template segments belongs to a coded block, the first subset of the set of template segments is unidirectionally predicted, and the second subset of the set of template segments is bidirectionally predicted.
[0499] In some embodiments, for the first template segment among the multiple template segments, whether the corresponding reference template segment is unidirectionally predicted or bidirectionally predicted is based on the motion information of a predetermined sub-block of the current video block.
[0500] In some embodiments, for the first template segment among the multiple template segments, the reference list associated with the motion information of the corresponding reference template segment is determined based on the motion information of a predetermined sub-block of the current video block.
[0501] In some embodiments, the predetermined sub-block includes one of the following: the central sub-block of the current video block, or a sub-block at a predetermined position of the current video block.
[0502] In some embodiments, the predetermined sub-block is bidirectionally predicted, and the multiple reference template segments are bidirectionally predicted.
[0503] In some embodiments, the predetermined sub-block is unidirectionally predicted, and the multiple reference template segments are bidirectionally predicted.
[0504] In some embodiments, the current video block is encoded and decoded using an encoding / decoding tool with sub-block level motion information. For example, the encoding / decoding tool may include at least one of the following: an affine encoding / decoding tool, or a sub-block based temporal motion vector prediction (SbTMVP) encoding / decoding tool.
[0505] In some embodiments, information regarding the application of the method is included in the bitstream.
[0506] In some embodiments, the information is included in at least one of the following: sequence level, group of pictures level, picture level, slice level, slice group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header, or slice group header.
[0507] In some embodiments, the information is included in a region containing more than one sample or pixel. For example, the region includes one of the following: prediction block (PB), transform block (TB), coding block (CB), prediction unit (PU), transform unit (TU), coding unit (CU), virtual pipeline data unit (VPDU), coding tree unit (CTU), CTU row, slice, picture, sub-picture.
[0508] In some embodiments, the information is based on the encoded / decoded information of the current video block. In some embodiments, the encoded / decoded information includes at least one of the following: encoding / decoding mode, block size, color format, single-tree or dual-tree segmentation, color component, slice type, or picture type.
[0509] In some embodiments, the transformation includes encoding the current video block into the bitstream. Alternatively or additionally, in some embodiments, the transformation includes decoding the current video block from the bitstream.
[0510] According to further embodiments of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by a device for video processing. In the method, a reference template for the current video block is determined based on a current template of the current video block of the video and motion information of at least one sub-block of the current video block. The bitstream is generated based on the reference template.
[0511] According to still other embodiments of the present disclosure, a method for storing a bitstream of a video is provided. In this method, a reference template for a current video block is determined based on a current template of the current video block of the video and motion information of at least one sub-block of the current video block. A bitstream is generated based on the reference template. The bitstream is stored in a non-transitory computer-readable recording medium.
[0512] The embodiments of the present disclosure may be described according to the following items, and the features may be combined in any reasonable manner.
[0513] Item 1. A method for video processing, including: for the conversion between a current video block of a video and the bitstream of the video, determining a reference template for the current video block based on a current template of the current video block and motion information of at least one sub-block of the current video block; and performing the conversion based on the reference template.
[0514] Item 2. The method according to Item 1, wherein the current template of the current video block includes at least one of the following: a first set of reconstructed regions, the first set of reconstructed regions being above and adjacent to the current video block, or a second set of reconstructed regions, the second set of reconstructed regions being to the left and adjacent to the current video block.
[0515] Item 3. The method according to Item 1 or Item 2, wherein determining the reference template includes: determining a plurality of template segments of the current template; determining a plurality of reference template segments based on the plurality of template segments and the motion information of the at least one sub-block; and determining the reference template based on the plurality of reference template segments.
[0516] Item 4. The method according to Item 3, wherein determining a plurality of reference template segments includes: for a first template segment among the plurality of template segments, determining a first sub-block associated with the first template segment from the at least one sub-block; and determining a first reference template segment among the plurality of reference template segments based on the motion information of the first sub-block.
[0517] Item 5. The method according to Item 4, wherein the first sub-block includes at least one of the following: a sub-block adjacent to the first template segment, or a sub-block not adjacent to the first template segment.
[0518] Item 6. The method according to Item 4 or Item 5, wherein the first sub-block adjacent to the first template segment is bi-directionally predicted, and at least one of two motion information sets of the first sub-block is used to determine the first reference template segment.
[0519] Item 7. The method according to item 6, wherein the set of motion information of the first sub-block is associated with a reference list, and the set of motion information includes at least one of the following: a motion vector, or a reference frame in the reference list.
[0520] Item 8. The method according to item 6 or item 7, wherein determining the first reference template segment includes: determining two predictions of the first reference segment based on the two sets of motion information; and determining the first reference template segment based on a weighted average of the two predictions.
[0521] Item 9. The method according to item 8, wherein determining the two predictions of the first reference segment includes: updating the set of motion information in the two sets of motion information by scaling the motion vector in the set of motion information to a reference frame; and determining the corresponding prediction of the first reference segment based on the updated set of motion information.
[0522] Item 10. The method according to item 6 or item 7, wherein determining the first reference template segment includes: selecting one set of motion information from the two sets of motion information based on the motion information of the second sub-block of the current video block; and determining the first reference template segment based on the selected set of motion information.
[0523] Item 11. The method according to item 10, wherein the second sub-block includes the central sub-block of the current video block.
[0524] Item 12. The method according to item 4 or item 5, wherein the first sub-block adjacent to the first template segment is unidirectionally predicted.
[0525] Item 13. The method according to item 12, wherein the region of the first reference template segment is determined based on a single set of motion information of the first sub-block.
[0526] Item 14. The method according to item 12, wherein a first set of motion information is associated with the first sub-block, and determining the first reference template segment includes: determining a second set of motion information based on at least one of the following: a zero motion vector, a scaled motion vector of a first motion vector in the first set of motion information, the scaled motion vector being associated with a second reference frame different from the first reference frame of the first motion vector, a mirrored motion vector of a second motion vector in the first set of motion information, a third set of motion information of a non-adjacent sub-block of the current video block; and determining the first reference template segment based on the first set of motion information and the second set of motion information.
[0527] Item 15. The method according to Item 14, wherein a first reference index of the second motion information set is the same as a second reference index of the motion information of a second sub-block of the current video block, and the first motion information set and the motion information of the second sub-block are from the same reference list.
[0528] Item 16. The method according to Item 12, wherein motion information associated with a reference list of the first sub-block is unavailable, and determining the first reference template segment includes: determining the motion information of the first sub-block based on at least one of: a zero motion vector, a scaled motion vector of an existing motion vector associated with the current video block, the scaled motion vector being associated with a second reference frame different from a first reference frame of the existing motion vector, a mirror motion vector of an existing motion vector associated with the current video block, additional motion information of a non-adjacent sub-block of the current video block; and determining the first reference template segment based on the motion information of the first sub-block.
[0529] Item 17. The method according to Item 16, wherein the reference list of the first sub-block is the same as the reference list of a second sub-block, and a first reference index of the motion information of the first sub-block is the same as a second reference index of the motion information of the second sub-block.
[0530] Item 18. The method according to Item 15 or Item 17, wherein the second sub-block includes one of: a central sub-block of the current video block, or a sub-block at a predetermined position of the current video block.
[0531] Item 19. The method according to Item 12, wherein the first reference template segment is determined based on a zero motion vector.
[0532] Item 20. The method according to any one of Items 3 to 19, wherein a template segment among the plurality of template segments is above the current video block, and a width of the template segment is equal to a width of a sub-block in the current video block.
[0533] Item 21. The method according to any one of Items 3 to 19, wherein a template segment among the plurality of template segments is to the left of the current video block, and a height of the template segment is equal to a height of a sub-block in the current video block.
[0534] Item 22. The method according to any one of Items 3 to 21, wherein motion information of a single sub-block of the current video block is used to determine one or more reference template segments.
[0535] Item 23. The method according to any one of Items 3 to 22, wherein the reference template segments among the plurality of reference template segments are unidirectionally predicted or bidirectionally predicted.
[0536] Item 24. The method according to any one of Items 3 to 22, wherein the plurality of reference template segments are unidirectionally predicted or bidirectionally predicted.
[0537] Item 25. The method according to any one of Items 3 to 24, wherein for a set of template segments among the plurality of template segments, the set of template segments belongs to an encoding / decoding block, a first subset of the set of template segments is unidirectionally predicted, and a second subset of the set of template segments is bidirectionally predicted.
[0538] Item 26. The method according to any one of Items 3 to 25, wherein for a first template segment among the plurality of template segments, whether the corresponding reference template segment is unidirectionally predicted or bidirectionally predicted is based on the motion information of a predetermined sub-block of the current video block.
[0539] Item 27. The method according to any one of Items 3 to 25, wherein for a first template segment among the plurality of template segments, the reference list associated with the motion information of the corresponding reference template segment is determined based on the motion information of a predetermined sub-block of the current video block.
[0540] Item 28. The method according to Item 26 or Item 27, wherein the predetermined sub-block includes one of the following: the central sub-block of the current video block, or a sub-block at a predetermined position of the current video block.
[0541] Item 29. The method according to any one of Items 26 to 28, wherein the predetermined sub-block is bidirectionally predicted, and the plurality of reference template segments are bidirectionally predicted.
[0542] Item 30. The method according to any one of Items 26 to 28, wherein the predetermined sub-block is unidirectionally predicted, and the plurality of reference template segments are bidirectionally predicted.
[0543] Item 31. The method according to any one of Items 1 to 30, wherein the current video block is encoded / decoded using an encoding / decoding tool with sub-block level motion information.
[0544] Item 32. The method according to Item 31, wherein the encoding / decoding tool includes at least one of the following: an affine encoding / decoding tool, or a sub-block based temporal motion vector prediction (SbTMVP) encoding / decoding tool.
[0545] Item 33. The method according to any one of Items 1 to 32, wherein information about applying the method is included in the bitstream.
[0546] Item 34. The method according to item 33, wherein the information is included in at least one of the following: sequence level, picture group level, picture level, slice level, slice group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header, or slice group header.
[0547] Item 35. The method according to item 33, wherein the information is included in a region containing more than one sample or pixel.
[0548] Item 36. The method according to item 35, wherein the region includes one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, slice, picture, sub-picture.
[0549] Item 37. The method according to any one of items 33 to 36, wherein the information is based on the decoded and encoded information of the current video block.
[0550] Item 38. The method according to item 37, wherein the decoded and encoded information includes at least one of the following: codec mode, block size, color format, single-tree or double-tree segmentation, color component, slice type, or picture type.
[0551] Item 39. The method according to any one of items 1 to 38, wherein the conversion includes encoding the current video block into the bitstream.
[0552] Item 40. The method according to any one of items 1 to 38, wherein the conversion includes decoding the current video block from the bitstream.
[0553] Item 41. An apparatus for video processing, including a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to execute the method according to any one of items 1 to 40.
[0554] Item 42. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of items 1 to 40.
[0555] Item 43. A non-transitory computer-readable recording medium stores a bitstream of video, the bitstream being generated by a method executed by a device for video processing, wherein the method includes: determining a reference template for a current video block based on a current template of the current video block of the video and motion information of at least one sub-block of the current video block; and generating the bitstream based on the reference template.
[0556] Item 44. A method for storing a bitstream of video includes: determining a reference template for a current video block based on a current template of the current video block of the video and motion information of at least one sub-block of the current video block; generating the bitstream based on the reference template; and storing the bitstream in a non-transitory computer-readable recording medium.
[0557] Example device
[0558] Figure 13 A block diagram of a computing device 1300 in which various embodiments of the present disclosure may be implemented is shown. The computing device 1300 may be implemented as the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300), or may be included in the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300).
[0559] It should be understood that Figure 13 the computing device 1300 shown is for illustrative purposes only and does not imply any limitation to the functionality and scope of the embodiments of the present disclosure in any way.
[0560] As Figure 13 shown, the computing device 1300 includes a general-purpose computing device 1300. The computing device 1300 may include at least one or more processors or processing units 1310, a memory 1320, a storage unit 1330, one or more communication units 1340, one or more input devices 1350, and one or more output devices 1360.
[0561] In some embodiments, computing device 1300 may be implemented as any user terminal or server terminal having computing capabilities. The server terminal may be a server provided by a service provider, a large computing device, etc. The user terminal may be, for example, any type of mobile terminal, fixed terminal or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / video cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is contemplated that computing device 1300 may support any type of interface to the user (such as "wearable" circuitry, etc.).
[0562] Processing unit 1310 may be a physical processor or a virtual processor, and may implement various processes based on programs stored in memory 1320. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 1300. Processing unit 1310 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0563] Computing device 1300 generally includes various computer storage media. Such media may be any media accessible by computing device 1300, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. Memory 1320 may be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. Storage unit 1330 may be any removable or non-removable media, and may include machine-readable media, such as memory, flash drives, magnetic disks, or other media that can be used to store information and / or data and can be accessed in computing device 1300.
[0564] Computing device 1300 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although not shown in Figure 13 a disk drive for reading from and / or writing to a removable non-volatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable non-volatile optical disk may be provided. In such a case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0565] The communication unit 1340 communicates with another computing device via a communication medium. Additionally, the functionality of the components in the computing device 1300 can be implemented by a single computing cluster or multiple computer machines, which can communicate via a communication connection. Thus, the computing device 1300 can operate in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs), or other general network nodes.
[0566] The input device 1350 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, and so on. The output device 1360 can be one or more of various output devices, such as a display, speaker, printer, and so on. With the aid of the communication unit 1340, the computing device 1300 can also communicate with one or more external devices (not shown), such as storage devices and display devices, the computing device 1300 can also communicate with one or more devices that enable a user to interact with the computing device 1300, or if needed, the computing device 1300 can also communicate with any device (such as a network card, modem, etc.) that enables the computing device 1300 to communicate with one or more other computing devices. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0567] In some embodiments, some or all of the components of the computing device 1300 can also be arranged in a cloud computing architecture rather than being integrated in a single device. In a cloud computing architecture, the components can be provided remotely and work together to implement the functions described in the present disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services, which do not require the end user to be aware of the physical location or configuration of the system or hardware providing these services. In various embodiments, cloud computing uses suitable protocols to provide services via a wide area network (such as the Internet). For example, a cloud computing provider provides an application via a wide area network, and the application can be accessed via a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on a server at a remote location. The computing resources in a cloud computing environment can be consolidated or distributed at the locations of remote data centers. The cloud computing infrastructure can provide services through a shared data center, although to the user, they appear as a single access point. Thus, the cloud computing architecture can be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein can be provided by a conventional server or installed directly or otherwise on a client device.
[0568] In an embodiment of the present disclosure, the computing device 1300 may be used to implement video encoding / decoding. The memory 1320 may include one or more video codec modules 1325 having one or more program instructions. These modules are accessible and executable by the processing unit 1310 to perform the functions of the various embodiments described herein.
[0569] In an example embodiment of performing video encoding, the input device 1350 may receive video data as the input 1370 to be encoded. The video data may be processed, for example, by the video codec module 1325 to generate an encoded bitstream. The encoded bitstream may be provided as the output 1380 via the output device 1360.
[0570] In an example embodiment of performing video decoding, the input device 1350 may receive the encoded bitstream as the input 1370. The encoded bitstream may be processed, for example, by the video codec module 1325 to generate decoded video data. The decoded video data may be provided as the output 1380 via the output device 1360.
[0571] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, those skilled in the art will understand that various changes may be made in form and detail without departing from the spirit and scope of the present application as defined by the appended claims. These variations are intended to be covered by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for video processing, comprising: Based on the current template of the current video block and the motion information of at least one sub-block of the current video block, determining a reference template for the current video block for the conversion between the current video block of the video and the bitstream of the video; and Performing the conversion based on the reference template.
2. The method according to claim 1, wherein the current template of the current video block comprises at least one of the following: A first set of reconstructed regions, which are above the current video block and adjacent to the current video block, or A second set of reconstructed regions, which are to the left of the current video block and adjacent to the current video block.
3. The method according to claim 1 or claim 2, wherein determining the reference template comprises: Determining a plurality of template segments of the current template; Based on the plurality of template segments and the motion information of the at least one sub-block, determining a plurality of reference template segments; and Based on the plurality of reference template segments, determining the reference template.
4. The method according to claim 3, wherein determining a plurality of reference template segments comprises: For a first template segment among the plurality of template segments, determining a first sub-block associated with the first template segment from the at least one sub-block; and Based on the motion information of the first sub-block, determining a first reference template segment among the plurality of reference template segments.
5. The method according to claim 4, wherein the first sub-block comprises at least one of the following: A sub-block adjacent to the first template segment, or A sub-block not adjacent to the first template segment.
6. The method according to claim 4 or claim 5, wherein the first sub-block adjacent to the first template segment is bi-directionally predicted, and at least one of the two motion information sets of the first sub-block is used to determine the first reference template segment.
7. The method according to claim 6, wherein the motion information set of the first sub-block is associated with a reference list, and the motion information set comprises at least one of the following: A motion vector, or A reference frame in the reference list.
8. The method according to claim 6 or claim 7, wherein determining the first reference template segment comprises: Based on the two motion information sets, determining two predictions of the first reference segment; and Based on the weighted average of the two predictions, determining the first reference template segment.
9. The method according to claim 8, wherein determining the two predictions of the first reference segment comprises: Updating the motion information set in the two motion information sets by scaling the motion vector in the motion information set to the reference frame; and Based on the updated motion information set, determining the corresponding prediction of the first reference segment.
10. The method according to claim 6 or claim 7, wherein determining the first reference template segment comprises: Based on the motion information of a second sub-block of the current video block, selecting one motion information set from the two motion information sets; and Based on the selected motion information set, determining the first reference template segment.
11. The method according to claim 10, wherein the second sub-block comprises a central sub-block of the current video block.
12. The method according to claim 4 or claim 5, wherein the first sub-block adjacent to the first template segment is unidirectionally predicted.
13. The method according to claim 12, wherein the region of the first reference template segment is determined based on a single set of motion information of the first sub-block.
14. The method according to claim 12, wherein a first set of motion information is associated with the first sub-block, and determining the first reference template segment comprises: determining a second set of motion information based on at least one of: a zero motion vector, a scaled motion vector of a first motion vector in the first set of motion information, the scaled motion vector being associated with a second reference frame different from a first reference frame of the first motion vector, a mirrored motion vector of a second motion vector in the first set of motion information, a third set of motion information of a non-adjacent sub-block of the current video block; and determining the first reference template segment based on the first set of motion information and the second set of motion information.
15. The method according to claim 14, wherein a first reference index of the second set of motion information is the same as a second reference index of motion information of a second sub-block of the current video block, and the first set of motion information and the motion information of the second sub-block are from the same reference list.
16. The method according to claim 12, wherein motion information associated with a reference list of the first sub-block is unavailable, and determining the first reference template segment comprises: determining the motion information of the first sub-block based on at least one of: a zero motion vector, a scaled motion vector of an existing motion vector associated with the current video block, the scaled motion vector being associated with a second reference frame different from a first reference frame of the existing motion vector, a mirrored motion vector of an existing motion vector associated with the current video block, additional motion information of a non-adjacent sub-block of the current video block; and determining the first reference template segment based on the motion information of the first sub-block.
17. The method according to claim 16, wherein the reference list of the first sub-block is the same as the reference list of a second sub-block, and a first reference index of the motion information of the first sub-block is the same as a second reference index of the motion information of the second sub-block.
18. The method according to claim 15 or claim 17, wherein the second sub-block comprises one of the following: a central sub-block of the current video block, or a sub-block at a predetermined position of the current video block.
19. The method according to claim 12, wherein the first reference template segment is determined based on a zero motion vector.
20. The method according to any one of claims 3 to 19, wherein the template segment among the plurality of template segments is above the current video block, and the width of the template segment is equal to the width of the sub-blocks in the current video block.
21. The method according to any one of claims 3 to 19, wherein a template segment among the plurality of template segments is on the left side of the current video block, and the height of the template segment is equal to the height of a sub-block in the current video block.
22. The method according to any one of claims 3 to 21, wherein the motion information of a single sub-block of the current video block is used to determine one or more reference template segments.
23. The method according to any one of claims 3 to 22, wherein a reference template segment among the plurality of reference template segments is unidirectionally predicted or bidirectionally predicted.
24. The method according to any one of claims 3 to 22, wherein the plurality of reference template segments are unidirectionally predicted or bidirectionally predicted.
25. The method according to any one of claims 3 to 24, wherein for a set of template segments among the plurality of template segments, the set of template segments belongs to a coded block, a first subset of the set of template segments is unidirectionally predicted, and a second subset of the set of template segments is bidirectionally predicted.
26. The method according to any one of claims 3 to 25, wherein for a first template segment among the plurality of template segments, whether the corresponding reference template segment is unidirectionally predicted or bidirectionally predicted is based on the motion information of a predetermined sub-block of the current video block.
27. The method according to any one of claims 3 to 25, wherein for a first template segment among the plurality of template segments, the reference list associated with the motion information of the corresponding reference template segment is determined based on the motion information of a predetermined sub-block of the current video block.
28. The method according to claim 26 or claim 27, wherein the predetermined sub-block includes one of the following: The central sub-block of the current video block, or A sub-block at a predetermined position of the current video block.
29. The method according to any one of claims 26 to 28, wherein the predetermined sub-block is bidirectionally predicted, and the plurality of reference template segments are bidirectionally predicted.
30. The method according to any one of claims 26 to 28, wherein the predetermined sub-block is unidirectionally predicted, and the plurality of reference template segments are bidirectionally predicted.
31. The method according to any one of claims 1 to 30, wherein the current video block is coded and decoded using coding and decoding tools with sub-block level motion information.
32. The method according to claim 31, wherein the coding and decoding tools include at least one of the following: An affine coding and decoding tool, or A sub-block based temporal motion vector prediction (SbTMVP) coding and decoding tool.
33. The method according to any one of claims 1 to 32, wherein information about applying the method is included in the bitstream.
34. The method according to claim 33, wherein the information is included in at least one of the following: Sequence level, Group of pictures level, Picture level, Slice level, Slice group level, Sequence header. Picture header, Sequence parameter set (SPS), Video parameter set (VPS), Decoding parameter set (DPS), Decoding capability information (DCI), Picture parameter set (PPS), Adaptive Parameter Set (APS), strip head, or slice group head.
35. The method according to claim 33, wherein the information is included in a region containing more than one sample or pixel.
36. The method according to claim 35, wherein the region includes one of the following: Prediction Block (PB), Transform Block (TB), Coding Block (CB), Prediction Unit (PU), Transform Unit (TU), Coding Unit (CU), Virtual Pipeline Data Unit (VPDU), Coding Tree Unit (CTU), CTU row, strip, slice, sub-picture.
37. The method according to any one of claims 33 to 36, wherein the information is based on the coded and decoded information of the current video block.
38. The method according to claim 37, wherein the coded and decoded information includes at least one of the following: coding mode, block size, color format, single-tree or double-tree segmentation, color component, strip type, or picture type.
39. The method according to any one of claims 1 to 38, wherein the conversion includes encoding the current video block into the bitstream.
40. The method according to any one of claims 1 to 38, wherein the conversion includes decoding the current video block from the bitstream.
41. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 40.
42. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 40.
43. A non-transitory computer-readable recording medium storing a bitstream of a video, the bitstream being generated by a method executed by an apparatus for video processing, wherein the method comprises: determining a reference template for the current video block based on a current template of the current video block of the video and motion information of at least one sub-block of the current video block; and generating the bitstream based on the reference template.
44. A method for storing a bitstream of a video, comprising: determining a reference template for the current video block based on a current template of the current video block of the video and motion information of at least one sub-block of the current video block; generating the bitstream based on the reference template; and storing the bitstream in a non-transitory computer-readable recording medium.