Method and device for video processing and medium
By combining the motion vector difference (MVD) tool with other prediction tools, and utilizing the motion information of spatial and temporal neighboring blocks of video blocks, a variety of motion vector prediction candidate lists are constructed. This solves the problem of high motion vector signaling cost in inter-frame prediction in existing video encoding and decoding technologies, and improves encoding and decoding efficiency and prediction accuracy.
Patent Information
- Application Number
- CN202480064421.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-04
- Filing Date
- 2024-10-01
- Publication Date
- 2026-05-08
AI Technical Summary
Existing video coding and decoding technologies have room for improvement in coding and decoding efficiency, especially in hybrid video coding and decoding frameworks. The motion vector signaling cost of inter-frame prediction is high, and existing prediction methods are difficult to effectively utilize the spatiotemporal redundancy between video blocks.
By combining motion vector difference (MVD) based encoding and decoding tools with other prediction tools, and through a hybrid prediction method, multiple motion vector prediction candidate lists are constructed using the motion information of spatial and temporal neighboring blocks of video blocks, including Advanced Motion Vector Prediction (AMVP), Merge mode and affine motion compensation prediction, thereby improving prediction accuracy and efficiency.
By combining MVD tools with other prediction tools, the encoding and decoding efficiency of video codecs is improved, the cost of motion vector signaling is reduced, and the accuracy and compression efficiency of inter-frame prediction are enhanced.
Smart Images

Figure CN122003864A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure generally relate to video processing techniques, and more specifically, to intra- and inter-frame joint prediction based on motion vector difference (MVD). Background Technology
[0002] Today, digital video capabilities are being applied to all aspects of people's lives. Various video compression technologies have been proposed for video encoding / decoding, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-TH.265 High Efficiency Video Codec (HEVC) standard, and Multi-Functional Video Codec (VVC) standard. However, the encoding and decoding efficiency of video encoding and decoding technologies is generally expected to be further improved. Summary of the Invention
[0003] Embodiments of this disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method includes: for a conversion between a current video block and a video bitstream, determining a first prediction of the current video block based on a first encoding / decoding tool, the first encoding / decoding tool including a motion vector difference (MVD) related encoding / decoding tool; determining a third prediction of the current video block based on the first prediction and a second prediction of the current video block, the second prediction being determined based on a second encoding / decoding tool different from the first encoding / decoding tool; and performing a conversion based on the third prediction. The method according to the first aspect of this disclosure is capable of mixing predictions from MVD related encoding / decoding tools with another prediction.
[0005] In a second aspect, an apparatus for video processing is provided. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform the method according to the first aspect of this disclosure.
[0006] In a third aspect, a non-transitory computer-readable storage medium is proposed. This non-transitory computer-readable storage medium stores instructions that cause a processor to perform the method according to the first aspect of this disclosure.
[0007] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining a first prediction of a current video block based on a first encoding / decoding tool, the first encoding / decoding tool including a motion vector difference (MVD) related encoding / decoding tool; determining a third prediction of the current video block based on the first prediction and a second prediction of the current video block, the second prediction being determined based on a second encoding / decoding tool different from the first encoding / decoding tool; and generating a bitstream based on the third prediction.
[0008] In a fifth aspect, a method for storing a bitstream of video is proposed. The method includes: determining a first prediction of a current video block based on a first encoding / decoding tool, the first encoding / decoding tool including a motion vector difference (MVD) related encoding / decoding tool; determining a third prediction of the current video block based on the first prediction and a second prediction of the current video block, the second prediction being determined based on a second encoding / decoding tool different from the first encoding / decoding tool; generating a bitstream based on the third prediction; and storing the bitstream in a non-transitory computer-readable recording medium.
[0009] This synopsis aims to present, in a simplified form, the selected concepts further described below in the detailed embodiments. This synopsis is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0010] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0011] Figure 1 A block diagram of an example video codec system according to some embodiments of the present disclosure is shown; Figure 2 A block diagram of a first example video encoder according to some embodiments of the present disclosure is shown; Figure 3 A block diagram of an example video decoder according to some embodiments of the present disclosure is shown; Figure 4 This shows the locations of spatial and temporal neighbor blocks used in the construction of the AMVP / Merge candidate list; Figure 5 This shows the positions of non-adjacent candidates in the ECM; Figure 6 An affine motion model based on control points is shown; Figure 7 An example affine MVF for each sub-block is shown; Figure 8 The location of the inherited affine motion prediction values is shown; Figure 9 This demonstrates the inheritance of control point motion vectors; Figure 10 The locations of candidate positions for constructing the affine Merge pattern are shown; Figure 11 The spatial nearest neighbor is shown for deriving the affine Merge candidate; Figure 12 The affine Merge candidates from non-nearest neighbors to the constructed ones are shown; Figure 13 An example of generating HAPC is shown; Figure 14 A schematic diagram of regression-based affine Merge candidate derivation is shown; Figure 15 This demonstrates template matching execution over the search area surrounding the initial MV; Figure 16 The template and the corresponding reference template are shown; Figure 17 The template and reference template of a block with sub-block motion information using the motion information of the current block are shown; Figure 18 The derivation of the sub-CU motion field obtained by applying motion displacement based on neighbor motion information is shown; Figure 19 The top and left neighbor blocks used in the CIIP weight derivation are shown; Figure 20 A flowchart of CIIP_PDPC using the extended CIIP mode with PDPC is shown; Figure 21 The method for dividing angle patterns is shown; Figure 22 This demonstrates the generation of sub-block templates for SbTMVP; Figure 23A and Figure 23B The MMVD search points are shown respectively; Figure 24 The diamond-shaped area in the search region is shown; Figure 25 A flowchart of a method for video processing according to embodiments of the present disclosure is shown; and Figure 26 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.
[0012] In all accompanying drawings, the same or similar reference numerals usually refer to the same or similar elements. Detailed Implementation
[0013] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.
[0014] In the following description and claims, unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0015] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Additionally, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, whether explicitly described or not, it is believed that such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.
[0016] It should be understood that although the terms “first” and “second”, etc., can be used to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0017] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0018] Example Environment Figure 1This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0019] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.
[0020] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded / decoded representation of the video data. The bitstream may include encoded / decoded images and associated data. The encoded / decoded images are the encoded / decoded representations of the images. The associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. The encoded video data may be transmitted directly to destination device 120 via network 130A through I / O interface 116. The encoded video data may also be stored on storage medium / server 130B for access by destination device 120.
[0021] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.
[0022] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or further standards.
[0023] Figure 2This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... Figure 1 An example of a video encoder 114 in system 100 is shown.
[0024] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. Figure 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0025] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.
[0026] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0027] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... Figure 2 The examples are shown separately.
[0028] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0029] The mode selection unit 203 can select one of several codec modes (intra-frame codec or inter-frame codec) based, for example, on the error result, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-frame / inter-frame joint prediction (CIIP) mode, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).
[0030] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.
[0031] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks that do not depend on macroblocks within the same image.
[0032] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0033] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for reference images in list 0 to find a reference video block for the current video block, and can also search for reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate reference indices indicating the reference images containing the reference video blocks in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0034] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0035] In one example, the motion estimation unit 204 may indicate a value to the video decoder 300 in the syntax structure associated with the current video block, which indicates that the current video block has the same motion information as another video block.
[0036] In another example, motion estimation unit 204 may identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0037] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0038] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0039] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0040] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform subtraction operations.
[0041] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0042] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0043] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0044] After the video block is reconstructed in reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0045] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0046] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.
[0047] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0048] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.
[0049] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded video data blocks). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine this information, for example, by performing AMVP and Merge pattern. Using AMVP, several most likely candidates are derived based on data from adjacent PBs and reference pictures. Motion information typically includes horizontal and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge pattern" may refer to deriving motion information from spatially or temporally adjacent blocks.
[0050] The motion compensation unit 302 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. The identifier of the interpolation filter to be used, with sub-pixel accuracy, can be included in the syntax element.
[0051] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of the video block to calculate the interpolation for sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate the prediction block.
[0052] Motion compensation unit 302 may use at least some of the syntax information to determine the size of the blocks used to encode the encoded video sequence (multiple frames) and / or (multiple stripes), segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be an entire image or a region of an image.
[0053] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies the inverse transform.
[0054] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding predicted block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be used to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0055] Some exemplary embodiments of this disclosure will be described in detail below. It should be understood that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section only. Furthermore, although specific embodiments are described with reference to multi-functional video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Furthermore, although some embodiments describe video codec steps in detail, it should be understood that the corresponding decoding steps of the inverse codec will be implemented by the decoder. Additionally, the term "video processing" includes video codec or compression, video decoding or decompression, and video transcoding, wherein video pixels are represented from one compression format to another or at different compression bitrates.
[0056] 1. Brief Overview This disclosure relates to video codec techniques. Specifically, it relates to joint prediction methods in video codecs. These ideas can be applied individually or in various combinations to any video codec standard or non-standard video codec.
[0057] 2. Introduction The exponential growth of multimedia data has posed significant challenges to video encoding and decoding. To meet the ever-increasing demand for more efficient compression technologies, the ITU-T and ISO / IEC have developed a series of video encoding and decoding standards over the past few decades. Specifically, the ITU-T developed the H.261 and H.263 standards, and ISO / IEC developed the MPEG-1 and MPEG-4 visual standards. The two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, the H.265 / HEVC standard, and the latest VVC standard. Starting with H.262 / MPEG-2, a hybrid video encoding and decoding framework has been adopted, utilizing intra / inter-frame prediction plus transform encoding and decoding.
[0058] Figure 4 The location of spatial and temporal neighbor blocks used in the construction of the AMVP / Merge candidate list is shown.
[0059] 2.1. MVP in Video Encoding and Decoding Inter-frame prediction aims to eliminate temporal redundancy between adjacent frames, an indispensable component in hybrid video codec frameworks. Specifically, inter-frame prediction utilizes the content specified by motion vectors (MVs) as the predicted version of the current block to be encoded / decoded, thus transmitting only residual signals and motion information in the bitstream. To reduce the cost of MV signaling, motion vector prediction (MVP) emerged as an efficient mechanism for conveying motion information. Early strategies simply used the MV of a specified neighboring block or the median MV of neighboring blocks as the MVP. In H.265 / HEVC, a contention mechanism is involved, where rate-distortion optimization (RDO) selects the best MVP from multiple candidates. Specifically, Advanced MVP (AMVP) mode and Merge mode with different motion information signaling strategies were designed. Using AMVP mode, a reference index, an MVP candidate index referencing the AMVP candidate list, and motion vector difference (MVD) are transmitted via signaling. Regarding Merge mode, only the Merge index referencing the Merge candidate list is transmitted via signaling, and all motion information associated with the Merge candidate is inherited. Both the AMVP and Merge modes require building an MVP candidate list. The details of the construction process for these two modes are described below.
[0060] AMVP mode: AMVP utilizes the spatial-temporal correlation of motion vectors with neighboring blocks for explicit transfer of motion parameters. For each list of reference images, a motion vector candidate list is constructed by first checking the availability of temporally adjacent locations to the left and top, removing redundant candidates, and adding zero vectors to make the candidate list a constant length. For the derivation of spatial motion vector candidates, the final result is based on locations such as... Figure 4 The motion vector derivation for the five blocks at different locations shown derives two motion vector candidates. The five neighboring blocks located at B0, B1, B2, and A0, A1 are classified into two groups: group A includes the three spatially adjacent blocks above, and group B includes the two spatially adjacent blocks to the left. The two motion vector candidates are derived separately using the first available candidates from groups A and B in a predefined order. For the temporal motion vector candidate derivation, based on... Figure 4 The diagram shows the sequential examination of two distinct co-locations (bottom right (C0) and center (C1)) to derive a motion vector candidate. To avoid redundant MV candidates, duplicate motion vector candidates in the list are discarded. If the number of potential candidates is less than 2, additional zero motion vector candidates are added to the list.
[0061] Figure 5 The positions of non-adjacent candidates in the ECM are shown.
[0062] Merge modeSimilar to the AMVP mode, the MVP candidate list for the Merge mode also includes spatial and temporal candidates. For spatial motion vector candidate derivation, after performing availability and redundancy checks, up to four candidates are selected in the order A1, B1, B0, A0, and B2. For temporal Merge candidate (TMVP) derivation, a candidate is selected from at most two temporally neighboring blocks (C0 and C1). When there are not enough Merge candidates using both spatial and temporal candidates, joint bidirectional prediction Merge candidates and zero MV candidates are added to the MVP candidate list. The Merge candidate list construction process terminates once the number of available Merge candidates reaches the maximum allowed number for signal transmission.
[0063] In VVC, the construction process for the Merge mode is further improved by introducing a history-based MVP (HMVP), where the HMVP incorporates motion information from previously encoded / decoded blocks that can be far removed from the current block. In VVC, HMVP Merge candidates are appended to the Merge list, following the Spatial MVP and TMVP. In this method, motion information from previously encoded / decoded blocks is stored in a table and used as the MVP for the current CU. During the encoding / decoding process, the table with multiple HMVP candidates is maintained using a first-in, first-out (FIFO) strategy. Whenever a non-sub-block inter-frame encoded / decoded CU is present, the associated motion information is added to the last entry of the table as a new HMVP candidate.
[0064] During the standardization of VVC, a non-adjacent MVP was proposed to facilitate better motion information derivation by utilizing non-adjacent regions. In ECM software, the non-adjacent MVP is inserted between the TMVP and HMVP, where the distance between the non-adjacent spatial candidate and the current codec block is based on the width and height of the current codec block, such as... Figure 5 As shown.
[0065] 2.2. Affine Motion Compensation Prediction In HEVC, only a translational motion model is applied for motion compensation prediction (MCP). In the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transformation motion compensation prediction is applied. Figure 6 As shown, the affine motion field of a block is described by motion information from two control points (4 parameters) or three control point motion vectors (6 parameters).
[0066] Figure 6 Affine motion models based on control points are shown, including (a) a 4-parameter affine model and (b) a 6-parameter affine model.
[0067] For the 4-parameter affine motion model, the motion vector at the sample point position (x, y) in the block is derived as: (1) For the 6-parameter affine motion model, the motion vector at the sample point position (x, y) in the block is derived as: (2) in( mv0x, mv0y ) is the motion vector of the upper left control point, ( mv1x, mv1y ) is the motion vector of the upper right control point, and ( mv2x, mv2y ) is the motion vector of the lower left control point.
[0068] To simplify motion compensation prediction, a block-based affine transformation prediction is applied. To derive the motion vector for each 4×4 lumen sub-block, the motion vector of the center sample point of each sub-block is calculated according to the above equation (e.g., ...). Figure 7 (as shown), and rounded to 1 / 16 fractional precision. Then, a motion-compensated interpolation filter is applied to generate a prediction for each sub-block with a derived motion vector. The sub-block size for the chroma component is also set to 4×4. The MV of the 4×4 chroma sub-block is calculated as the average of the MV of the upper-left luminance sub-block and the lower-right luminance sub-block in the corresponding 8×8 luminance region.
[0069] Figure 7 The affine MVF for each sub-block is shown.
[0070] Similar to translational motion inter-frame prediction, there are two affine motion inter-frame prediction modes: affine Merge mode and affine AMVP mode.
[0071] 2.2.1. Affine Merge Prediction The Affine Merge pattern can be applied to CUs with a width and height greater than or equal to 8. In this pattern, the CPVM of the current CU is generated based on the motion information of spatially neighboring CUs. There can be up to five CPVM candidates, and the one to be used for the current CU is indicated by a signal transmission index. In VVC, the following three types of CPVM candidates are used to form the Affine Merge candidate list: – Inherited affine Merge candidates inferred from the CPMV of neighboring CUs; – An affine Merge candidate CPMVP constructed using translational MV derivation of neighboring CUs; and – Zero MV.
[0072] In VVC, there are at most two inherited affine candidates, which are derived from the affine motion model of neighboring blocks: one from the left neighboring CU and one from the upper neighboring CU. Candidate blocks are as follows: Figure 8As shown. For the predicted values on the left, the scan order is A0->A1, and for the predicted values above, the scan order is B0->B1->B2. Only candidates from the first inheritance on each side are selected. No deduplication check is performed between candidates from two inheritances. When a neighboring affine CU is identified, its control point motion vector is used to derive the CPMVP candidate in the affine Merge list of the current CU. Figure 9 As shown, if the adjacent lower-left block A is encoded and decoded in affine mode, the motion vectors of the upper-left, upper-right, and lower-left corners of the CU containing block A are obtained. , and When block A is encoded and decoded using a 4-parameter affine model, according to and Calculate the two CPMVs of the current CU. When block A is encoded and decoded using a 6-parameter affine model, according to... , and Calculate the three CPMVs of the current CU.
[0073] Figure 8 The location of the inherited affine motion prediction value is shown.
[0074] Figure 9 The inheritance of control point motion vectors is shown.
[0075] The constructed affine candidate refers to the candidate built by combining the translational motion information of the neighbors of each control point. The motion information for the control points is derived from... Figure 10 The derivation is shown in the specified spatial and temporal nearest neighbors. CPMVk (k=1, 2, 3, 4) represents the k-th control point. For CPMV1, check the B2->B3->A2 block and use the MV of the first available block. For CPMV2, check the B1->B0 block, and for CPMV3, check the A1->A0 block. If available, the TMVP is used as CPMV4.
[0076] After obtaining the motion signatures (MVs) of the four control points, affine merge candidates are constructed based on this motion information. The following combinations of control point MVs are used for sequential construction: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3} Combining three CPMVs constructs a 6-parameter affine merge candidate, and combining two CPMVs constructs a 4-parameter affine merge candidate. To avoid motion scaling, combinations of control point MVs are discarded if the reference indices of the control points are different.
[0077] Figure 10 The locations of candidate positions for constructing the affine Merge pattern are shown.
[0078] After the inherited affine Merge candidate and the constructed affine Merge candidate are checked, if the list is still not full, a zero MV is inserted at the end of the list.
[0079] 2.2.2. Affine AMVP Prediction The affine AMVP mode can be applied to CUs with a width and height both greater than or equal to 16. An affine flag at the CU level is signaled in the bitstream to indicate whether the affine AMVP mode is used, and another flag is signaled to indicate whether it is a 4-parameter affine or a 6-parameter affine. In this mode, the difference between the current CU's CPVM and its predicted CPMVP is signaled in the bitstream. The affine AMVP candidate list is of size 2 and is generated sequentially using the following four types of CPVM candidates: – Inherited affine AMVP candidates inferred from the CPMV of neighboring CUs; – A constructive affine AMVP candidate CPMVP derived using translational MV of neighboring CUs; – Translation MV from the neighboring CU; and – Zero MV.
[0080] The checking order for inherited affine AMVP candidates is the same as that for inherited affine Merge candidates. The only difference is that, for AMVP candidates, only affine CUs with the same reference picture as those in the current block are considered. No deduplication is applied when inserting inherited affine motion predictions into the candidate list.
[0081] The constructed AMVP candidate is from Figure 10 The derivation is based on the specified spatial nearest neighbor. The same inspection order as in the affine Merge candidate construction is used. Additionally, the reference picture index of neighboring blocks is checked. The block that is the first to be inter-coded in the inspection order and has the same reference picture as the current CU is used. This applies when the current CU is encoded / decoded in a 4-parameter affine mode, and... mv0 and mv1When all three CPMVs are available, they are added as candidates in the affine AMVP list. If the current CU is encoded / decoded in a 6-parameter affine mode and all three CPMVs are available, they are added as candidates in the affine AMVP list. Otherwise, the constructed AMVP candidates are set to unavailable.
[0082] Figure 11 Spatial nearest neighbors used to derive affine Merge candidates are shown: (a) for deriving inherited affine Merge candidates, and (b) for deriving constructed affine Merge candidates.
[0083] If the affine AMVP list still has fewer than 2 candidates after inserting valid inherited affine AMVP candidates and constructed AMVP candidates, then when available, mv0 , mv1 and mv2 They will be added sequentially as translation MVs to predict all control point MVs for the current CU. Finally, if the affine AMVP list is still not full, zero MVs are used to populate the affine AMVP list.
[0084] 2.2.3. New Affine Candidate Derivation Method in ECM-8.0 ECM-6.0 integrates three additional affine Merge and AMVP candidate derivation methods: non-adjacent spatial domain candidates, historical parameter-based candidates, regression-based affine candidates, and pixel-based affine motion compensation.
[0085] 2.2.3.1. Non-adjacent airspace candidates In ECM-6.0, the study of non-adjacent spatial neighbors provides candidates for both affine Merge and affine AMVP. The format for obtaining non-adjacent spatial candidate pairs is as follows: Figure 11 As shown. Similar to non-adjacent regular merge candidates, the distance between non-adjacent spatial candidates and the current codec block is also defined based on the width and height of the current CU.
[0086] Figure 11 Motion information of non-adjacent spatial neighbors is used to generate additional inheritance and construct affine merge candidates. Specifically, to generate inheritance candidates, non-adjacent spatial neighbors are checked based on their distance from the current block (i.e., from nearest to farthest). At a specific distance, only the first available neighbors encoded in affine mode from each side (e.g., left and top) of the current block are included. Figure 11 As shown in (a), the checks of the left and top nearest neighbors are performed from bottom to top and from right to left, respectively. For the constructed candidates, as... Figure 11As shown in (b), the positions of non-adjacent spatial neighbors on the left and top are first determined independently; then, the position of the top-left neighbor can be determined accordingly to form a rectangular virtual block together with the non-adjacent neighbors on the left and top. The motion information of the three non-adjacent neighbors is used to form a CPMV at the top-left (A), top-right (B), and bottom-left (C) of the virtual block, which is projected onto the current CU to generate corresponding construction candidates, such as... Figure 12 As shown. Figure 12 The affine Merge candidates from non-nearest neighbors to the constructed ones are shown.
[0087] 2.2.3.2. Affine Candidates Based on Historical Parameters History-based Affine Model Inheritance (HAMI) allows affine models to inherit from blocks that may not be adjacent to the current block from previous affine codecs. A History-based Table (HPT) is created. Each entry in the HPT stores a set of affine parameters: a, b, c, and d, each represented by a 16-bit signed integer. Entry in the HPT is categorized by reference lists and reference indices. Each reference list in the HPT supports five reference indices. The HPT category (denoted as HPTCat) is calculated in a formulaic manner. HPTCat (RefList, RefIdx) = 5×RefList + min (RefIdx, 4)(3) Here, RefList and RefIdx represent the list of reference images (0 or 1) and the reference index, respectively. A maximum of 7 entries can be stored for each category, resulting in a total of 70 entries in the HPT. At the beginning of each CTU line, the number of entries for each category is initialized to zero. After decoding the affine-encoded CU with reference lists RefListcur and RefIdxcur, the affine parameters are used to update the entries in the category HPTCat(RefListcur, RefIdxcur) in a manner similar to HMVP table updates.
[0088] Candidates based on historical affine parameters (HAPC) are derived from... Figure 13 The MV is derived from a set of affine parameters in the corresponding entries stored in the HPT, represented as the neighboring 4×4 blocks of A0, A1, B0, B1, or B2. The MV of the neighboring 4×4 blocks is used as the base MV. The MV of the current block at position (x, y) is calculated in a formulaic manner as follows: (4) Where (mvhbase, mvvbase) represents the MV of the nearest 4×4 blocks, and (xbase, ybase) represents the center position of the nearest 4×4 blocks. (x, y) can be the top left, top right, and bottom left corners of the current block to obtain the corner position MV (CPMV) for the current block, or it can be the center of the current block to obtain the regular MV for the current block.
[0089] Figure 13 An example of how to derive the HAPC from block A0 is shown. The affine parameters {a0, b0, c0, d0} are obtained directly from an entry in the class HPTIdx(RefListA0, refIdx0A0) in the HPT. The affine parameters from the HPT (with the center position of A0 as the base position and the MV of block A0 as the base MV) are used together to derive the CPMV for either the affine MergeHAPC or the affine AMVP HAPC. They can also be used to derive the MV located at the center of the current block as a regular Merge candidate. The HAPC can be placed into the sub-block-based Merge candidate list, the affine AMVP candidate list, or the regular Merge candidate list. In response to the introduction of the new HAPC, the size of the sub-block-based Merge candidate list is increased from 5 to 10 and 12 for random access and low-latency B configurations, respectively. Furthermore, for the random access configuration, the size of the regular Merge candidate list is increased from 10 to 11 to accommodate the newly added regular Merge candidates.
[0090] Figure 13 An example of generating HAPC is shown.
[0091] 2.2.3.3. Regression-based Affine Candidates In ECM-6.0, regression-based affine merge candidates are derived and added to the affine merge list. The sub-block motion fields from previously encoded and decoded affine CUs and the motion information of neighboring sub-blocks from the current CU are used as inputs to the regression process to derive the proposed affine candidates.
[0092] Previously encoded and decoded affine CUs can be identified by scanning non-adjacent positions and the affine HMVP table. For example... Figure 14 As shown, the neighboring sub-block information of the current CU is obtained from the 4x4 sub-blocks represented by the gray area. For each sub-block, given a reference list, the corresponding motion vector and center coordinates of the sub-block can be used.
[0093] For each affine CU, at most two affine candidates can be derived: one with neighboring subblock information and one without. All candidates generated by linear regression are deduplicated and collected into a candidate subgroup. When ARMC is enabled, an ARMC process based on TM cost is applied. Subsequently, when N affine CUs are found, at most N candidates generated by linear regression are added to the affine merge list.
[0094] Figure 14 A schematic diagram of the regression-based affine Merge candidate derivation is shown.
[0095] 2.2.3.4. Pixel-based Affine Motion Compensation Using pixel-based affine motion compensation, when OBMC is not applied, the minimum affine sub-block size for the luma component is set to 1x1, and the minimum sub-block size for the chroma component is always set to 1x1.
[0096] 2.3. Template Matching Merge / AMVP Pattern in ECM Template Matching (TM) Merge / AMVP mode is a decoder-side MV derivation method used to refine the motion information of the current CU by finding the closest match between a template in the current image (i.e., the block above and / or to the left of the current CU) and a block in the reference image (i.e., the block of the same size as the template). Figure 15 As shown, within the search range of [-8, +8] pixels, a better MV is searched around the initial motion of the current CU.
[0097] Figure 15 This demonstrates template matching execution over the search area surrounding the initial MV.
[0098] In AMVP mode, an MVP candidate is determined based on the template matching error, selecting the one that minimizes the difference between the current block and the reference block template. Then, the TM process performs MV refinement only on that specific MVP candidate. The TM refines the MVP candidate using an iterative diamond search, starting with full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode) within a search range of [-8, +8] pixels. The AMVP candidate can be further refined using a cross search with full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode), followed by half-pixels and quarter-pixels sequentially depending on the AMVR mode. This search process ensures that the MVP candidate maintains the same MV precision as indicated by the Adaptive Motion Vector Resolution (AMVR) mode after the TM process.
[0099] In Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. TMMerge can proceed up to 1 / 8 pixel MVD accuracy, or skip those accuracies beyond half-pixel MVD accuracy, depending on whether an alternative interpolation filter is used based on the merged motion information (i.e., used when AMVR is in half-pixel mode). Furthermore, when TM mode is enabled, template matching can operate as a standalone process, or as an additional MV refinement process between block-based and sub-block-based bilateral matching (BM) methods, depending on whether BM can be enabled according to its enable condition check. When both BM and TM are enabled for the CU, the TM search process stops at half-pixel MVD accuracy, and the resulting MV is further refined using the same model-based MVD derivation method as in DMVR.
[0100] 2.4. Adaptive Reordering of Merge Candidates (ARMC) Inspired by the spatial correlation between reconstructed neighboring pixels and the current codec block, an Adaptive Reordering of Merge Candidates (ARMC) is proposed to refine the order of candidates in a given candidate list. The basic assumption is that candidates with lower template matching costs have a higher probability of being selected through the RDO process and should therefore be placed earlier in the list to reduce signaling costs.
[0101] The reordering method is applied to the regular Merge pattern, the Template Matching (TM) Merge pattern, and the Affine Merge pattern (excluding SbTMVP candidates). For the TM Merge pattern, the Merge candidates are reordered before the refinement process.
[0102] After constructing the Merge candidate list, the Merge candidates are divided into several subgroups. The subgroup size is set to 5. The Merge candidates in each subgroup are reordered in ascending order based on the cost value of template matching. For simplicity, the Merge candidates in the last subgroup (not the first subgroup) are not reordered.
[0103] Template matching cost is measured by the sum of absolute differences (SAD) between the samples of the current block's template and the samples of its corresponding reference template. For example... Figure 16 As shown, the template includes a set of reconstructed samples adjacent to the current block, while the reference template is located using the same motion information of the current block. When the Merge candidate utilizes bidirectional prediction, the reference samples of the Merge candidate's template are also generated through bidirectional prediction.
[0104] For a sub-block-based merge candidate with a sub-block size equal to Wsub * Hsub, the upper template includes several sub-templates of size Wsub × K, and the left template includes several sub-templates of size K × Hsub. For example... Figure 17 As shown, the motion information of the sub-blocks in the first row and first column of the current block is used to derive the reference sample points of each sub-template.
[0105] 2.5. Sub-block-based temporal motion vector prediction (SbTMVP) VVC supports a sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to TMVP, SbTMVP leverages motion fields in co-located images to facilitate more accurate MVP derivation. The same co-located image used by TMVP is used in SbTMVP. SbTMVP differs from TMVP primarily in two ways. First, SbTMVP enables motion prediction at the sub-CU level, while TMVP predicts motion at the CU level. Second, compared to TMVP, which obtains temporal MVs from co-located blocks in a co-located image (where a co-located block is the lower right or center block relative to the current CU), SbTMVP applies motion displacements before obtaining temporal motion information from the co-located image. These motion displacements are obtained by reusing the MV of one of the spatially neighboring blocks from the current CU.
[0106] Figure 16 The template and the corresponding reference template are shown.
[0107] Figure 18 The derivation process of the sub-block level motion field for SbTMVP is shown. Specifically, the motion information of the lower left sub-block A1 is first obtained. If any MV in reference list 0 and list 1 points to the same frame, the corresponding MV will be identified as the motion displacement. Otherwise, zero MV will be used as the motion displacement.
[0108] Once the motion displacement is determined, a designated region within the same frame is used to derive the sub-block level motion field. Assuming... Figure 18 As shown, the motion of A1 is used as the motion displacement. Then, for each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the center sample point) in the co-location image is obtained to provide motion information, wherein the MV scaling operation is first performed to align the reference frame of the temporal motion vector with the reference frame of the current CU.
[0109] Figure 17 The template and reference template of a block with sub-block motion information are shown.
[0110] Figure 18 The derivation of the sub-CU motion field obtained by applying motion displacement based on neighbor motion information is shown.
[0111] In VVC and ECM, in addition to the CU-level MVP candidate list, a sub-CU-level MVP candidate list is constructed to provide more accurate motion predictions for the current CU. This list includes the motion field generated by both the SbTMVP and AFFINE methods. Specifically, only one SbTMVP candidate is included, and this SbTMVP candidate is always placed as the first entry in the constructed sub-CU-level MVP candidate list. Multiple AFFINE candidates are included in the list after performing template matching-based reordering, with those AFFINE candidates with lower costs placed earlier.
[0112] 2.6. Intra-Frame and Inter-Frame Joint Prediction (CIIP) In VVC, when a CU is encoded and decoded in Merge mode, if the CU contains at least 64 luma samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64), and if both the CU width and height are less than 128 luma samples, an additional flag is transmitted via signaling to indicate whether Intra-Inter Joint Prediction (CIIP) mode is applied to the current CU. As the name suggests, CIIP prediction combines inter-frame prediction signals with intra-frame prediction signals. The inter-frame prediction signal in CIIP mode... The inter-frame prediction process is derived using the same procedure as the regular Merge mode; and the intra-frame prediction signal... The conventional intra-frame prediction process with a planar pattern is derived. Then, a weighted average is used to combine the intra-frame and inter-frame prediction signals, where the predictions are based on the upper and left neighboring blocks (e.g., ...). Figure 19 The weight values for the encoding / decoding modes (as shown) are calculated as follows: - If the upper nearest neighbor is available and is intra-coded, set isIntraTop to 1; otherwise, set isIntraTop to 0. – If the left nearest neighbor is available and is intra-coded, set isIntraLeft to 1; otherwise, set isIntraLeft to 0. – If (isIntraLeft + isIntraTop) equals 2, then wt is set to 3; Otherwise, if (isIntraLeft + isIntraTop) equals 1, then wt is set to 2; Otherwise, set wt to 1.
[0113] The CIIP predictions are formed as follows:
[0114] Figure 20 The top and left neighbor blocks used in the CIIP weight derivation are shown.
[0115] 2.7. CIIP with PDPC hybrid In ECM, the CIIP model is extended as described. In this extended model (CIIP_PDPC), the predictions of the regular Merge model are refined using top (Rx, -1) and left (R-1, y) reconstructed samples. This refinement inherits the Location-Related Prediction Combination (PDPC) scheme. The flowchart of the CIIP_PDPC model prediction can be seen as follows... Figure 20 As shown, WT and WL are weighted values that depend on the sample location in the block, as defined in PDPC.
[0116] The CIIP_PDPC mode is transmitted via signaling along with the CIIP mode. When the CIIP flag is true, another flag (i.e., the CIIP_PDPC flag) is further transmitted via signaling to indicate whether CIIP_PDPC is used.
[0117] Figure 20 A flowchart of the CIIP_PDPC process using the extended CIIP mode with PDPC is shown.
[0118] 2.8. Combination of CIIP with TIMD and TM Merge In ECM CIIP mode, prediction samples can be generated by weighting the inter-frame prediction signal using CIIP-TM Merge candidate prediction and the intra-frame prediction signal using the intra-frame prediction mode derived using TIMD. This method is only applied to codec blocks with an area less than or equal to 1024.
[0119] The TIMD derivation method is used to derive intra-prediction modes in CIIP. Specifically, the intra-prediction mode with the smallest SATD value in the TIMD mode list is selected and mapped to one of 67 regular intra-prediction modes.
[0120] Furthermore, it is proposed that if the derived intra-prediction mode is an angular mode, the weights of the two tests (wIntra, wInter) should be modified. For near-horizontal mode (2 <= angular mode index < 34), the current block is vertically divided, as shown below. Figure 21 As shown in (a); for near-vertical mode (34 <= angle mode index <= 66), the current block is divided horizontally, as... Figure 21 As shown in (b).
[0121] The different sub-blocks (wIntra, wInter) are shown in Table 1.
[0122] Figure 21 The method for dividing angle patterns is shown.
[0123] Table 1. Weights used for modifications to the angle mode.
[0124]
[0125] Using CIIP-TM, a CIIP-TM Merge candidate list is constructed for the CIIP-TM pattern. Merge candidates are refined through template matching. CIIP-TM Merge candidates are also reordered as regular Merge candidates using the ARMC method. The maximum number of CIIP-TM Merge candidates is 2.
[0126] 2.9. Combination of intra-block copying and intra-prediction Intra-Block Copy and Intra-Prediction Combination (IBC-CIIP) is an encoding / decoding tool for CUs that uses IBC and intra-prediction to obtain two prediction signals, which are then weighted and summed to generate the final prediction, as shown below:
[0127] in and This represents the IBC prediction signal and the intra-frame prediction signal. For IBC Merge mode and IBCAMVP mode, It is set to equal to (13, 4) and (1, 1).
[0128] An intra-prediction mode (IPM) candidate list is used to generate the intra-prediction signal, and the IPM candidate list size is predefined to 2. The IPM index is transmitted via signaling to indicate which IPM to use.
[0129] Figure 22 This demonstrates the generation of sub-block templates for SbTMVP.
[0130] 2.10. Derivation of Time-Domain Motion in ECM In VVC, temporal motion vector prediction (TMVP) for AMVP and Merge modes is derived by acquiring motion information from the center or lower right of the co-occurrence block in the co-occurrence image transmitted via signal transmission. Similarly, for sub-block-based temporal motion vector prediction (SbTMVP) mode, motion information from the left neighboring location is used as motion displacement, which is then used to obtain the TMVP at the sub-CU level.
[0131] In ECM, two aspects were modified to further improve the encoding and decoding efficiency of TMVP. First, two co-located images are used, which are two reference frames with the minimum POC distance relative to the frame to be encoded / decoded. Second, the motion displacement for locating the TMVP is adaptively determined from multiple positions based on template cost. More specifically, two motion displacement candidate lists are constructed for each of the two co-located frames. The motion displacement with the minimum template matching cost is used to derive either SbTMVP or TMVP candidates. The merge list based on sub-blocks includes a maximum of four SbTMVP candidates. The SbTMVP candidate with the minimum template matching cost derived from the first co-located frame is placed in the first entry without reordering, while other SbTMVP candidates are ordered together with affine candidates. Additionally, the prediction direction of the template for each sub-block is determined based on the central sub-block. Figure 22 As shown, if the central sub-block is unidirectionally predicted, then all sub-block templates are unidirectionally predicted, and vice versa. If the motion vector of the corresponding adjacent sub-block at a defined reference list is unavailable for a sub-block template, then zero MV is used for that sub-block template.
[0132] 2.11. Merge Schema with MVD (MMVD) In addition to the Merge mode (where implicitly derived motion information is directly used for generating prediction samples for the current CU), a Merge mode with motion vector difference (MMVD) is also introduced in VVC. The MMVD flag is transmitted via signaling immediately after the skip flag and the Merge flag are sent to indicate whether the MMVD mode is used for the CU.
[0133] In MMVD, after selecting a Merge candidate, the candidate is further refined using MVD information transmitted via signals. This further information includes a Merge candidate flag, an index specifying the motion amplitude, and an index indicating the motion direction. In MMVD mode, one of the top two candidates in the Merge list is selected as the MV basis. The Merge candidate flag is transmitted via signals to specify which one to use.
[0134] Figure 23A and Figure 23B The MMVD search points are shown separately.
[0135] The distance index specifies motion amplitude information and indicates a predefined offset from the starting point. For example... Figure 23A and Figure 23B As shown, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 2.
[0136] Table 2. Relationship between distance index and predefined offset
[0137] The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent the four directions shown in Table 3. It is important to note that the meaning of the MVD symbol can vary depending on the information of the starting MV. When the starting MV is a unidirectional or bidirectional prediction MV (where both lists point to the same side of the current image, i.e., both reference POCs are greater than or less than the current image's POC), the symbols in Table 3 specify the sign of the MV offset added to the starting MV. When the starting MV is a bidirectional prediction MV (where the two MVs point to different sides of the current image, i.e., one reference POC is greater than the current image's POC, and the other reference POC is less than the current image's POC), the symbols in Table 3 specify the sign of the MV offset added to the list 0 MV component of the starting MV, and the signs of the list 1 MV have the opposite values.
[0138] Table 3. Sign of MV offset specified by direction index
[0139] 2.12. Multi-pass decoder-side motion vector refinement (DMVR) Multi-pass decoder-side motion vector refinement is integrated into the ECM. In the first pass, bilateral matching (BM) is applied to the codec block. In the second pass, BM is applied to each 16x16 sub-block within the codec block. In the third pass, the motion vectors (MVs) in each 8x8 sub-block are refined by applying bidirectional optical flow (BDOF). The refined MVs are stored for both spatial and temporal motion vector prediction.
[0140] 2.12.1 First pass – Block-based bilateral matching MV refinement In the first pass, the refined MV is derived by applying BM to the codec block. Similar to decoder-side motion vector refinement (DMVR), in the bidirectional prediction operation, a refined MV is searched around the two initial MVs (MV0 and MV1) in the reference picture lists L0 and L1. Based on the minimum bilateral matching cost between the two reference blocks in L0 and L1, the refined MV (MV0_pass1 and MV1_pass1) is derived around the initial MVs.
[0141] BM performs a local search to derive the integer sample precision intDeltaMV. The local search applies a 3×3 square search pattern to iterate through the search range [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block dimension, and the maximum value of sHor and sVer is 8.
[0142] The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW * cbH is greater than 64, the mean-removed SAD (MRSAD) cost function is applied to remove the DC effect of distortion between reference blocks. The local search intDeltaMV is terminated when bilCost at the center point of the 3×3 search pattern has the minimum cost. Otherwise, the current minimum cost search point becomes the new center point of the 3×3 search pattern, and the search for the minimum cost continues until the end of the search range is reached.
[0143] We further refine the existing fractional sample points to derive the final deltaMV. Then, the refined MV after the first pass is derived as follows: · MV0_pass1 = MV0 + deltaMV · MV1_pass1 = MV1 – deltaMV 2.12.2 Second pass – Sub-block-based bilateral matching MV refinement In the second pass, the refined MV is derived by applying BM to 16×16 grid sub-blocks. For each sub-block, a refined MV is searched around the two MVs (MV0_pass1 and MV1_pass1) obtained in the first pass in the reference image lists L0 and L1. The refined MV (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) is derived based on the minimum bilateral matching cost between the two reference sub-blocks in L0 and L1.
[0144] For each sub-block, BM performs a full search to derive the integer sample precision intDeltaMV. The full search has a search range of [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block dimension, and the maximum value of sHor and sVer is 8.
[0145] The bilateral matching cost is calculated by applying a cost factor to the SATD cost between the two reference subblocks, as follows: bilCost = satdCost * costFactor. The search region (2*sHor + 1) * (2*sVer + 1) is divided into Figure 24A maximum of five diamond-shaped search regions are shown. Each search region is assigned a costFactor, determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond region is processed sequentially starting from the center of the search region. Within each region, search points are processed in raster scan order, starting from the top left corner and ending at the bottom right corner. A full pixel search terminates when the minimum bilCost within the current search region is less than a threshold (equal to sbW * sbH); otherwise, the full pixel search continues to the next search region until all search points have been checked. Additionally, the search process terminates if the difference between the previous minimum cost and the current minimum cost in an iteration is less than or equal to a threshold representing the area of the block.
[0146] Figure 24 The diamond-shaped area in the search area is shown.
[0147] Further refinement using existing VVC DMVR fractional samples is applied to derive the final deltaMV(sbIdx2). Then, the refined MV for the second pass is derived as follows: · MV0_pass2(sbIdx2) = MV0_pass1 + deltaMV(sbIdx2) · MV1_pass2(sbIdx2) = MV1_pass1 – deltaMV(sbIdx2) 2.12.3 Third pass – Sub-block based bidirectional optical flow MV refinement In the third pass, the refined bioMv is derived by applying BDOF to 8×8 grid sub-blocks. For each 8×8 sub-block, BDOF refinement is applied to derive scaled Vx and Vy without clipping, starting from the refined bioMv of the parent block in the second pass. The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample precision and clipped between -32 and 32.
[0148] The refined MVs (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) of the third pass are derived as follows: · MV0_pass3(sbIdx3) = MV0_pass2(sbIdx2) + bioMv · MV1_pass3(sbIdx3) = MV0_pass2(sbIdx2) – bioMv In all of the aforementioned sub-items, when surround motion compensation is enabled, the motion vector must be limited to account for surround offset.
[0149] 2.13. Sub-block-based CIIP in video encoding and decoding In this disclosure, a further improvement to CIIP is proposed by allowing sub-block-based predictions as inter-frame components. Specifically, intra-frame predictions and sub-block-based inter-frame predictions can be mixed to form CIIP predictions.
[0150] The detailed disclosures below should be considered as examples for explaining general concepts. These disclosures should not be interpreted in a narrow sense. Furthermore, these disclosures can be combined in any way.
[0151] The terms "video unit," "code-decoder unit," or "block" can refer to code-decoder tree block (CTB), code-decoder tree unit (CTU), code-decoder block (CB), CU, PU, TU, PB, and TB. The term "sub-block-based codec tool" can refer to affine codecs, SbTMVP, and their corresponding variants.
[0152] In this disclosure, with respect to "blocks encoded and decoded in mode N", "mode N" can be a predictive mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a encoding / decoding technique (e.g., DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-inter-inter, GPM intra-intra-intra, GPM inter-intra-intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, ALF, deblocking, SAO, bilateral filter, LMCS and corresponding variants, etc.).
[0153] It should be noted that the terms mentioned below are not limited to the specific terms defined in existing standards. Any changes to encoding / decoding tools also apply.
[0154] 1. The prediction generated by the first codec tool can be mixed with the prediction of the second codec to form a third prediction.
[0155] a) In one example, the first codec tool could be a sub-block-based approach, such as affine, SbTMVP, etc.
[0156] b) In one example, the second codec tool can be a regular intra-frame (plane / DC / angle) / TIMD / DIMD / ISP / PDPC / MIP / IBC / regular inter-frame, etc.
[0157] 2. A method for generating CIIP inter-frame components using a sub-block-based encoding / decoding tool is proposed, namely, sub-block-based CIIP.
[0158] a) In one example, a first prediction can be generated using affine motion or SbTMVP, and a second prediction can be generated using intra-frame mode. The two predictions are then mixed with a weighted average to produce a final prediction for sub-block-based CIIP.
[0159] b) In one example, the intra-frame mode can be plane / DC / angle / MIP / ISP / IBC / intra-frame TMP, etc.
[0160] c) In one example, the intra-frame mode can be derived based on TIMD / DIMD / intra-frame TMP, etc.
[0161] d) In one example, intra-frame components of CIIP based on sub-blocks can be processed via PDPC.
[0162] 3. A first sub-block-based motion list can be constructed to provide motion information for sub-block-based CIIP.
[0163] a) In one example, the first sub-block-based motion list may include at least one affine and / or at least one (or more) SbTMVP candidates.
[0164] i. In one example, specifically, the first sub-block-based motion list may include at least one adjacent / non-adjacent / history-based / regression-based / zero affine candidate.
[0165] ii. In one example, specifically, the first sub-block-based motion list may include at least one SbTMVP candidate.
[0166] 1) In one example, multiple SbTMVP candidates can appear in the list, which can be collected from at least one co-position frame.
[0167] b) In one example, alternatively, the first sub-block-based motion list may include only affine or SbTMVP candidates.
[0168] c) In one example, the number of candidates in the list may not exceed a constant or an adaptively determined number.
[0169] d) In one example, the first sub-block-based motion list can be the same as the second sub-block-based motion list used in the sub-block-based Merge / AMVP mode.
[0170] i. Alternatively, the first sub-block-based motion list and the second sub-block-based motion list may be different.
[0171] 1) In one example, the first sub-block-based motion list and the second sub-block-based motion list can be constructed using different maximum allowed number of candidates.
[0172] 2) In one example, the first sub-block-based motion list and the second sub-block-based motion list can be constructed using different candidate types.
[0173] 3) In one example, the first sub-block-based motion list can be constructed by selecting at least one candidate from the second sub-block-based motion list.
[0174] 4. The list of motions based on sub-blocks can be reordered based on specific metrics after construction.
[0175] a) In one example, template matching or bilateral matching costs can be used to reorder a list.
[0176] i. In one example, specifically, if a reconstructed template region for the current block exists, the template matching cost can be used to reorder the list.
[0177] 1) In one example, alternatively, if the reconstruction template region for the current block does not exist, the list may not be reordered.
[0178] 5. Indicates that the index of a specific candidate in the motion list based on a sub-block can be transmitted via signal in the bitstream.
[0179] a) In one example, the candidate specified by the index can be used to provide sub-block-based motion information and / or motion prediction.
[0180] i. In one example, if the candidate specified by the index is an affine candidate, it can first be refined by TM or DMVR and then used to generate inter-frame predictions.
[0181] 1) In one example, if the specified candidate is an affine candidate for one-way prediction, then TM can be used to refine the candidate, such as TM-based CPMV refinement, and the refined candidate is used to provide predictions.
[0182] 2) In one example, alternatively, if the specified candidate is an affine candidate for bidirectional prediction, DMVR can be used to refine the candidate, and the refined candidate is used to provide predictions.
[0183] ii. In one example, alternatively, no TM or DMVR processing is used to refine the specified affine candidates.
[0184] b) In one example, no candidate index is transmitted via signal in the bitstream.
[0185] i. In one example, the encoder and decoder can generate the same sub-block-based motion candidates based on predefined rules, which are used to provide inter-frame prediction.
[0186] 1) In one example, specifically, after the list of motions based on sub-blocks is constructed, it can be reordered based on a specific metric, and candidates in fixed positions (e.g., first / last / middle or any other position) can be used by default.
[0187] 6. At least one syntax (or flag) indicating the use of sub-block-based CIIP can be transmitted via signaling in the bitstream.
[0188] a) In one example, at least one block-level CIIP flag based on sub-blocks can be transmitted via signaling in the bitstream.
[0189] b) In one example, at least one sub-block-based CIIP flag at the sequence level / picture group level / picture level / strip level / piece group level (such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header) can be transmitted via signaling in the bitstream.
[0190] c) In one example, whether the sub-block-based CIIP flag is signaled or whether the sub-block-based CIIP is applied may depend on the value of at least one other syntax.
[0191] i. In one example, whether the sub-block-based CIIP flag is transmitted via signaling, or whether the sub-block-based CIIP is applied, may depend on the value of at least one other sequence-level / picture-group-level / picture-level / strip-level / piece-group-level syntax.
[0192] 1) In one example, whether the CIIP flag based on the sub-block is transmitted via signaling may depend on the value of the syntax indicating whether CIIP and / or affine and / or SbTMVP is enabled.
[0193] ii. In one example, whether the sub-block-based CIIP flag is signaled or whether the sub-block-based CIIP is applied may depend on the value of at least one other block-level syntax.
[0194] 1) In one example, specifically, the CIIP flag based on the sub-block can only be transmitted via signaling if the CIIP / CIIP-PDPC / CIIP-TM / Affine / SbTMVP flag is true or false.
[0195] 2) In one example, a first flag indicating the use of CIIP (i.e., the CIIP flag) can be transmitted via signaling. If this flag is true, a second flag indicating the use of sub-block-based CIIP (i.e., the sub-block-based CIIP flag) can be transmitted via signaling. If the sub-block-based CIIP flag is true, the index of the sub-block-based motion candidate can be indicated via signaling. Otherwise, if the sub-block-based CIIP flag is true, a third flag (i.e., CIIP-TM) can be transmitted via signaling.
[0196] a) In one example, alternatively, a first flag indicating the use of CIIP (i.e., the CIIP flag) can be signaled, and if this flag is true, a second flag indicating the use of CIIP-TM (i.e., the CIIP-TM flag) can be signaled. If the CIIP-TM flag is false, a third flag indicating the use of sub-block-based CIIP (i.e., the sub-block-based CIIP flag) can be signaled. If the sub-block-based CIIP flag is true, the index of the sub-block-based motion candidate can be signaled.
[0197] iii. In one example, whether a sub-block-based CIIP flag is transmitted via signaling, or whether a sub-block-based CIIP is applied, may depend on whether the value of at least one of the following syntaxes is true or false: sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0198] d) In one example, whether another flag is transmitted via signaling may depend on the value of the CIIP flag based on the sub-block.
[0199] i. In one example, whether a flag indicating the use of CIIP-TM / CIIP-PDPC and / or any other codec tool is transmitted via signaling may depend on the value of the sub-block-based CIIP flag.
[0200] 1) In one example, when the CIIP flag based on the sub-block is true, the CIIP-TM flag is not transmitted via signaling.
[0201] a) In one example, alternatively, if the CIIP flag based on the sub-block is false, the CIIP-TM flag is transmitted via signaling.
[0202] b) In one example, alternatively, the CIIP-TM flag is transmitted via signaling regardless of whether the CIIP flag based on the sub-block is true or false.
[0203] 2) In one example, when the CIIP flag based on the sub-block is true, the CIIP-PDPC flag is not transmitted via signaling, and / or the regular intra-frame or PDPC intra-frame is used by default.
[0204] a) In one example, alternatively, if the CIIP flag based on the sub-block is false, the CIIP-PDPC flag is transmitted via signaling.
[0205] b) In one example, alternatively, the CIIP-PDPC flag is transmitted via signaling regardless of whether the CIIP flag based on the sub-block is true or false.
[0206] e) In one example, whether a sub-block-based CIIP flag is transmitted via signaling may depend on the usage conditions of at least one other codec tool.
[0207] i. In one example, a sub-block-based CIIP flag can be signaled or applied only if CIIP and / or affine and / or SbTMVP are appropriate for the current block.
[0208] ii. In one example, a sub-block-based CIIP flag is transmitted or applied only when the block dimension (e.g., width / height / width-to-height ratio / block area) meets a specific condition.
[0209] 7. On the hybridization of sub-block-based CIIP modes.
[0210] a) In one example, sub-block-based CIIP can use the same hybrid approach used by regular CIIP / CIIP-TM / CIIP-PDPC.
[0211] i. In one example, alternatively, sub-block-based CIIP can use different hybrid methods.
[0212] b) In one example, the mixed weights can be position-dependent within a block.
[0213] i. In one example, specifically, different locations within a block can have different blending weights.
[0214] ii. In one example, a constant blending weight can be applied to all locations within a block.
[0215] c) In one example, the blending weight matrix can be intra-mode dependent.
[0216] i. In one example, specifically, different blending weight matrices can be used depending on whether the intra-frame angle mode is near vertical or near horizontal.
[0217] 8. For example, how to apply sub-block-based predictions can be conditionalized based on whether they are used in sub-block-based CIIP.
[0218] a) For example, the size of a sub-block can be conditionalized.
[0219] b) For example, whether and / or how an operation is applied can be conditionalized.
[0220] i. The operation can be PROF.
[0221] ii. The operation can be an interlaced affine.
[0222] iii. The operation can be OBMC.
[0223] General items 9. In the above examples, a video unit can refer to a color component / sub-picture / strip / piece / code-decode tree unit (CTU) / CTU line / CTU group / code-decode unit (CU) / prediction unit (PU) / transform unit (TU) / code-decode tree block (CTB) / code-decode block (CB) / prediction block (PB) / transform block (TB) / block / sub-block of a block / sub-region within a block / any other region containing more than one sample or pixel.
[0224] 10. Whether and / or how the methods disclosed above can be applied can be transmitted via signaling at the sequence level / picture group level / picture level / strip level / piece group level (such as in sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header).
[0225] 11. Whether and / or how to apply the above methods may depend on the following information: a) Messages transmitted via signals in DPS / SPS / VPS / PPS / APS / Picture Header / Strip Header / Group Header / Maximum Codec Unit (LCU) / Codec Unit (CU) / LCU Line / LCU Group / TU / PU Block / Video Codec Unit b) Location of CU / PU / TU / block / video codec unit c) Block dimensions of the current block and / or its neighboring blocks d) The block shape of the current block and / or its neighboring blocks e) The encoding / decoding mode of the block, such as IBC or non-IBC inter-frame mode or non-IBC sub-block mode. f) Indication of color format (such as 4:2:0, 4:4:4) g) Encoder / decoder tree structure h) Strip / panel type and / or image type i) Color components (e.g., can be applied only to the chromaticity component or the luminance component) j) Temporal layer ID k) Standard grade / level / layer 3. Problem Existing CIIP methods have the following problems: 1) In VVC and ECM, CIIP only allows regular Merge candidates as inter-frame components, lacking the utilization of MVD, which may lead to suboptimal inter-frame prediction.
[0226] 2) How to mix MVD-assisted inter-frame prediction with intra-frame prediction, and how the proposed method interacts with other codec tools, can be further specified.
[0227] 4. Detailed Solution In this disclosure, a further improvement to CIIP is proposed by allowing the generation of inter-frame components using MVD (MV difference). Specifically, intra-frame prediction and inter-frame prediction generated using MVD information can be mixed to form CIIP prediction.
[0228] The detailed embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.
[0229] The terms "video unit," "code-decoder unit," or "block" can refer to code-decoder tree block (CTB), code-decoder tree unit (CTU), code-decoder block (CB), CU, PU, TU, PB, and TB. The term "sub-block-based codec tool" can refer to affine codecs, SbTMVP, and their corresponding variants.
[0230] In this disclosure, with respect to "blocks encoded and decoded in mode N", "mode N" can be a predictive mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a encoding / decoding technique (e.g., DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-inter-inter, GPM intra-intra-intra, GPM inter-intra-intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, ALF, deblocking, SAO, bilateral filter, LMCS and corresponding variants, etc.).
[0231] It should be noted that the terms mentioned below are not limited to the specific terms defined in existing standards. Any variants of encoding / decoding tools also apply.
[0232] 1. The predictions of the first codec tool can be mixed with the predictions of the second codec tool to form a third prediction.
[0233] a) In one example, the first encoding / decoding tool could be an MVD-related method, such as regular AMVP, MMVD, etc.
[0234] b) In one example, the second codec tool can be a regular intra-frame (plane / DC / angle) / TIMD / DIMD / ISP / PDPC / MIP / IBC / regular inter-frame, etc.
[0235] c) Two predictions can be mixed in a weighted sum manner.
[0236] 2. A method for generating CIIP inter-frame prediction using MVD-related encoding and decoding tools was proposed, called CIIP-MVD.
[0237] a) In one example, a first prediction can be generated using regular AMVD or MMVD, and a second prediction can be generated using intra-frame mode. The two predictions are then mixed with a weighted average to produce the final prediction for CIIP-MVD.
[0238] b) In one example, the intra-frame mode can be plane / DC / angle / MIP / ISP / IBC / intra-frame TMP, etc.
[0239] c) In one example, the intra-frame mode can be derived based on TIMD / DIMD / intra-frame TMP, etc.
[0240] d) In one example, intra-frame prediction of CIIP-MVD can be processed by PDPC.
[0241] Regarding CIIP-MMVD 3. When using CIIP-MMVD, predictions generated by intra-frame mode and predictions generated by MMVD can be mixed to produce a third prediction for further processing.
[0242] 4. When using CIIP-MMVD, you can construct a list of MMVD candidates that is the same as (or different from) the regular MMVD pattern.
[0243] a) In one example, when CIIP-MMVD is used for blocks, the same (or different) underlying MV candidate derivation method as the regular MMVD pattern is used to construct the MMVD candidate list.
[0244] b) In one example, when CIIP-MMVD is used for blocks, the same (or different) number of base MV candidates as the regular MMVD pattern are used to construct the MMVD candidate list.
[0245] c) In one example, when CIIP-MMVD is used for a block, the same (or different) number of MV offsets as the regular MMVD mode are used to construct the MMVD candidate list.
[0246] d) In one example, when CIIP-MMVD is used for a block, the same (or different) MMVD candidate signaling method as the regular MMVD mode can be used to specify a particular candidate.
[0247] e) In one example, in CIIP-MMVD mode, only unidirectional prediction is applied to generate inter-frame predictions for blending.
[0248] 5. The MMVD candidate list of the CIIP-MMVD pattern can be reordered based on specific metrics after construction.
[0249] a) In one example, template matching or bilateral matching costs can be used to reorder a list.
[0250] i. In one example, specifically, if a reconstructed template region for the current block exists, the template matching cost can be used to reorder the list.
[0251] 1) In one example, alternatively, if the reconstruction template region for the current block does not exist, the list may not be reordered.
[0252] b) In one example, alternatively, no reordering is performed on the CIIP-MMVD motion candidate list.
[0253] c) In one example, during the reordering process, intra-frame prediction of CIIP can be considered to compute the cost of the candidates.
[0254] i. For example, a prediction that is a mixture of inter-frame prediction and intra-frame prediction can be computed as a candidate reference template, and the distortion (such as SAD) between the reference template and the template of the current block can be derived as a candidate cost.
[0255] 6. Whether a BCW (Bilibili Combining Weighted Bidirectional Prediction) index is reordered by a specific metric (i.e., TM cost) can depend on the CIIP pattern.
[0256] a) In one example, if CIIP is used, the BCW index may (or may not) be reordered.
[0257] b) In one example, the BCW index can be reordered for some CIIP sub-patterns, but not for others.
[0258] i. In one example, the BCW index can be reordered if one of the patterns in CIIP subgroup A is used; otherwise, the BCW index is not reordered.
[0259] 1) In one example, CIIP subgroup A includes at least one mode among regular CIIP / CIIP-TM / CIIP-PDPC / CIIP-MMVD / CIIP-affine, etc.
[0260] 7. When using CIIP-MMVD, the index of a specific candidate in the MMVD candidate list is indicated to be transmitted via signal in the bitstream.
[0261] a) In one example, the candidate specified by the index can be used to provide inter-frame prediction.
[0262] i. In one example, TM or DMVR processing can be used to refine specified MMVD candidates.
[0263] b) In one example, no candidate index is transmitted via signal in the bitstream.
[0264] i. In one example, the encoder and decoder can generate the same MMVD candidate based on predefined rules, which is used to provide inter-frame prediction.
[0265] 1) In one example, specifically, after the MMVD candidate list is constructed, it can be reordered based on a specific metric, and candidates in fixed positions (e.g., first / last / middle or any other position) can be used by default.
[0266] 8. At least one syntax (or flag) indicating the use of CIIP-MMVD can be transmitted via signaling in the bitstream.
[0267] a) In one example, at least one block-level CIIP-MMVD flag can be transmitted via signaling in the bitstream.
[0268] b) In one example, at least one CIIP-MMVD flag at the sequence level / picture group level / picture level / strip level / piece group level (such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header) can be transmitted via signaling in the bitstream.
[0269] i. In one example, whether the CIIP-MMVD flag is transmitted via signaling, or whether CIIP-MMVD is applied, may depend on whether the value of at least one of the following syntaxes is true or false: sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0270] c) In one example, whether the CIIP-MMVD flag is transmitted via signaling, or whether CIIP-MMVD is applied, may depend on the value of at least one other syntax.
[0271] i. In one example, whether the CIIP-MMVD flag is transmitted via signaling, or whether CIIP-MMVD is applied, may depend on the value of at least one other sequence-level / picture-group-level / picture-level / strip-level / piece-group-level syntax.
[0272] 1) In one example, whether the CIIP-MMVD flag is transmitted via signaling may depend on the value of the syntax indicating whether CIIP and / or MMVD are enabled.
[0273] ii. In one example, whether the CIIP-MMVD flag is transmitted via signaling, or whether CIIP-MMVD is applied, may depend on the value of at least one other block-level syntax.
[0274] 1) In one example, specifically, the CIIP-MMVD flag can only be transmitted via signaling if the CIIP / CIIP-PDPC / CIIP-TM / MMVD / CIIP-affine flag is true or false.
[0275] 2) In one example, a first flag indicating the use of CIIP (i.e., the CIIP flag) can be signaled. If this flag is true, a second flag indicating the use of CIIP-MMVD (i.e., the CIIP-MMVD flag) can be signaled. If the CIIP-MMVD flag is true, the index of a specific MMVD candidate can be signaled. Otherwise, if the CIIP-MMVD flag is false, a third flag (i.e., CIIP-TM) can be signaled.
[0276] a) In one example, alternatively, a first flag indicating the use of CIIP (i.e., the CIIP flag) can be signaled, and if this flag is true, a second flag indicating the use of CIIP-TM (i.e., the CIIP-TM flag) can be signaled. If the CIIP-TM flag is false, a third flag indicating the use of CIIP-MMVD (i.e., the CIIP-MMVD flag) can be signaled. If the CIIP-MMVD flag is true, the index of the MMVD candidate can be signaled.
[0277] d) In one example, whether another flag is transmitted via signaling may depend on the value of the CIIP-MMVD flag.
[0278] i. In one example, whether a flag indicating the use of CIIP-TM / CIIP-PDPC / CIIP-affine and / or any other codec tool is transmitted via signaling may depend on the value of the CIIP-MMVD flag.
[0279] 1) In one example, when the CIIP-MMVD flag is true, the CIIP-TM flag is not transmitted via signaling.
[0280] a) In one example, alternatively, if the CIIP-MMVD flag is false, the CIIP-TM flag can be transmitted via signaling.
[0281] b) In one example, alternatively, the CIIP-TM flag can be transmitted via signaling regardless of whether the CIIP-MMVD flag is true or false.
[0282] 2) In one example, when the CIIP-MMVD or CIIP-affine flag is true, the CIIP-PDPC flag is not transmitted via signaling, and / or the regular intra-frame or PDPC intra-frame default is used.
[0283] a) In one example, alternatively, if the CIIP-MMVD flag is false, the CIIP-PDPC flag is transmitted via signaling.
[0284] b) In one example, alternatively, the CIIP-PDPC flag is transmitted via signaling regardless of whether the CIIP-MMVD flag is true or false.
[0285] e) In one example, whether the CIIP-MMVD flag is transmitted via signaling may depend on the usage conditions of at least one other codec tool.
[0286] i. In one example, the CIIP-MMVD flag can only be signaled or CIIP-MMVD can only be applied if CIIP and / or MMVD are suitable for the current block.
[0287] ii. In one example, the CIIP-MMVD flag is transmitted or CIIP-MMVD is applied only when the block dimension (e.g., width / height / width-to-height ratio / block area) meets a specific condition.
[0288] 9. Regarding CIIP derivation mode signaling.
[0289] a) In one example, a first flag (i.e., the CIIP flag) can be signaled to indicate whether CIIP or CIIP derivative mode is used.
[0290] b) In one example, if the first flag indicates true, the second flag indicates whether CIIP-MMVD or another (or more) CIIP derivative modes (i.e., CIIP-affine) are used.
[0291] i. In one example, if another (or more) CIIP derivative modes are not suitable for the current block, the second flag may not need to be transmitted via signaling.
[0292] c) In one example, if the second flag indicates true, or if the second flag is not signaled, then the third flag is signaled to indicate whether the CIIP-MMVD mode is used for a particular block.
[0293] 10. On the mixing of CIIP-MMVD modes.
[0294] a) In one example, CIIP-MMVD can use the same hybrid approach used by regular CIIP / CIIP-TM / CIIP-PDPC.
[0295] i. In one example, alternatively, CIIP-MMVD can use different hybrid methods.
[0296] b) In one example, the mixed weights can be position-dependent within a block.
[0297] i. In one example, specifically, different locations within a block can have different blending weights.
[0298] ii. In one example, a constant blending weight can be used at all locations within a block.
[0299] c) In one example, the blending weight matrix can be intra-mode dependent.
[0300] i. In one example, specifically, different blending weight matrices can be used depending on whether the intra-frame angle mode is near vertical or near horizontal.
[0301] General Information 11. In the above examples, a video unit can refer to a color component / sub-picture / strip / piece / code-decode tree unit (CTU) / CTU line / CTU group / code-decode unit (CU) / prediction unit (PU) / transform unit (TU) / code-decode tree block (CTB) / code-decode block (CB) / prediction block (PB) / transform block (TB) / block / sub-block of a block / sub-region within a block / any other region containing more than one sample or pixel.
[0302] 12. Whether and / or how the methods disclosed above can be signaled at the sequence level / picture group level / picture level / strip level / film group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / film group header.
[0303] 13. Whether and / or how to apply the above methods may depend on the following information: a) Messages transmitted via signals in DPS / SPS / VPS / PPS / APS / Picture Header / Strip Header / Group Header / Maximum Codec Unit (LCU) / Codec Unit (CU) / LCU Line / LCU Group / TU / PU Block / Video Codec Unit b) Location of CU / PU / TU / block / video codec unit c) Block dimensions of the current block and / or its neighboring blocks d) The block shape of the current block and / or its neighboring blocks e) The encoding / decoding mode of the block, such as IBC or non-IBC inter-frame mode or non-IBC sub-block mode. f) Indication of color format (such as 4:2:0, 4:4:4) g) Encoder / decoder tree structure h) Strip / panel type and / or image type i) Color components (e.g., can be applied only to the chromaticity component or the luminance component) j) Temporal layer ID k) Standard grade / level / layer Figure 25 A flowchart of a method 2500 for video processing according to an embodiment of the present disclosure is shown. Method 2500 is implemented during the conversion between video units of a video and a bitstream of a video.
[0304] At box 2510, for the conversion between the current video block and the video bitstream, the first prediction of the current video block is determined based on a first codec tool. The first codec tool includes a motion vector difference (MVD) related codec tool.
[0305] At box 2520, the third prediction for the current video block is determined based on the first prediction and the second prediction for the current video block. The second prediction is determined based on a second codec tool that is different from the first codec tool.
[0306] At box 2530, the transformation is performed based on a third prediction. In some embodiments, the transformation includes encoding the current video block into a bitstream. Alternatively or additionally, in some embodiments, the transformation includes decoding the current video block from the bitstream.
[0307] Method 2500 is able to mix predictions from MVD-related codec tools with another prediction.
[0308] In some embodiments, the first codec tool includes at least one of a conventional Advanced Motion Vector Prediction (AMVP) codec tool or a Merge Mode with Motion Vector Difference (MMVD) codec tool, and the second codec tool includes at least one of the following: a conventional intra-frame codec tool, a conventional intra-frame plane codec tool, a conventional intra-frame DC codec tool, a conventional intra-frame angle codec tool, a template-based intra-frame mode derivation (TIMD) codec tool, a decoder-side intra-frame mode derivation (DIMD) codec tool, an intra-frame sub-segmentation codec (ISP) codec tool, a position-dependent (intra-frame) prediction combination (PDPC) codec tool, a matrix-based intra-frame prediction (MIP) inter-frame codec tool, an intra-frame block copy (IBC) codec tool, or a conventional inter-frame codec tool.
[0309] In some embodiments, the third prediction is determined based on a weighted sum of the first and second predictions.
[0310] In some embodiments, the inter-frame components of intra-inter joint prediction (CIIP) are determined by a first codec tool.
[0311] In some embodiments, the first prediction is generated using conventional Advanced Motion Vector Difference (AMVD) or MMVD, the second prediction is generated using intra-frame mode, and the third prediction is determined by using a weighted average to mix the first and second predictions.
[0312] In some embodiments, the intra-frame mode includes at least one of the following: planar mode, DC mode, angle mode, matrix-based intra-frame prediction (MIP) mode, intra-frame sub-segmentation coding and decoding (ISP) mode, intra-frame block copying (IBC) mode, or intra-frame template matching prediction (intra-frame TMP) mode.
[0313] In some embodiments, the intra-frame mode is derived based on at least one of the following: template-based intra-frame mode derivation (TIMD), decoder-side intra-frame mode derivation (DIMD), or intra-frame template matching prediction (intra-frame TMP).
[0314] In some embodiments, the intra-frame components of CIIP are processed by location-dependent (intra-frame) prediction combination (PDPC).
[0315] In some embodiments, if CIIP-Merge mode with motion vector difference (MMVD) is used, the first prediction is generated using MMVD, the second prediction is generated using intra-frame mode, and the third prediction is determined by mixing the first and second predictions.
[0316] In some embodiments, the MMVD candidate list may be different from or the same as the MMVD candidate list of a regular MMVD.
[0317] In some embodiments, a candidate derivation scheme for the basic motion vector (MV) is used to construct a candidate list for MMVD, wherein the basic MV is different from or the same as the basic MV of a regular MMVD.
[0318] In some embodiments, a certain number of basic MV candidates are used to construct the MMVD candidate list, and the number of basic MV candidates may be different from or the same as the number of basic MV candidates in a regular MMVD.
[0319] In some embodiments, a certain number of MV offsets are used to construct the MMVD candidate list, and the number of MV offsets may be different from or the same as the number of MV offsets in a regular MMVD.
[0320] In some embodiments, MMVD candidate indicators are used to construct an MMVD candidate list, and the MMVD candidate indicators may be different from or the same as those of regular MMVD.
[0321] In some embodiments, in the CIIP-MMVD mode, only one-way prediction is applied to generate the first prediction.
[0322] In some embodiments, the MMVD candidate list for the CIIP-MMVD mode is reordered based on at least one metric after it is constructed.
[0323] In some embodiments, at least one metric includes at least one of template matching cost or bilateral matching cost.
[0324] In some embodiments, if a reconstructed template region exists for the current video block, the template matching cost is used to reorder the MMVD candidate list.
[0325] In some embodiments, if the reconstructed template region for the current video block does not exist, the template matching cost is not used to reorder the MMVD candidate list.
[0326] In some embodiments, the MMVD candidate list for the CIIP-MMVD mode is not reordered.
[0327] In some embodiments, the cost of MMVD candidates in the MMVD candidate list is determined based on a second prediction for CIIP.
[0328] In some embodiments, the third prediction is determined as a reference template for the MMVD candidate, and the distortion between the reference template and the template of the current video block is determined as the cost of the MMVD candidate. For example, the distortion may include the sum of absolute differences (SAD).
[0329] In some embodiments, whether an index utilizing weighted bidirectional prediction (BCW) is reordered by a metric is determined based on the CIIP pattern.
[0330] In some embodiments, if CIIP mode is used, the BCW index is reordered. Alternatively or additionally, if CIIP mode is used, the BCW index may not be reordered.
[0331] In some embodiments, the BCW index is reordered for a portion of the CIIP subpattern, while the BCW index is not reordered for another portion of the CIIP subpattern.
[0332] In some embodiments, if at least one pattern in the CIIP subgroup is used, the BCW index is reordered; and if none of the patterns in the CIIP subgroup are used, the BCW index is not reordered.
[0333] In some embodiments, the CIIP subgroup includes at least one of the following: regular CIIP mode, CIIP-template matching (TM) mode, CIIP-PDPC mode, CIIP-MMVD mode, or CIIP-affine mode.
[0334] In some embodiments, the index of an MMVD candidate in the MMVD candidate list is indicated in the bitstream.
[0335] In some embodiments, the MMVD candidates specified by the index are used to provide inter-frame prediction.
[0336] In some embodiments, TM or decoder-side motion vector refinement (DMVR) is used to refine MMVD candidates specified by the index.
[0337] In some embodiments, the index of the MMVD candidate in the MMVD candidate list is not included in the bitstream.
[0338] In some embodiments, the encoder and decoder generate the same MMVD candidates based on predefined rules, and the MMVD candidates are used to provide inter-frame prediction.
[0339] In some embodiments, the MMVD candidate list is reordered based on a metric after it is constructed, and MMVD candidates in predefined positions of the MMVD candidate list are used by default. The predefined positions include at least one of the following: start position, end position, or middle position.
[0340] In some embodiments, at least one syntax element indicating the use of CIIP-MMVD is indicated in the bitstream. In one example, at least one flag indicating the use of CIIP-MMVD is indicated in the bitstream.
[0341] In some embodiments, at least one syntax element includes at least one block-level CIIP-MMVD flag.
[0342] In some embodiments, at least one syntax element includes at least one CIIP-MMVD flag from at least the following: sequence level, picture group level, picture level, strip level, slice group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header.
[0343] In some embodiments, whether at least one syntax element for CIIP-MMVD is included in the bitstream or whether CIIP-MMVD is applied depends on at least one value of at least one syntax element, which is at least one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.
[0344] In some embodiments, whether at least one syntax element for CIIP-MMVD is included in the bitstream or whether CIIP-MMVD is applied depends on at least one value of at least one other syntax element.
[0345] In some embodiments, at least one other syntax element includes at least one of the following: sequence-level syntax element, picture group-level syntax element, picture-level syntax element, strip-level syntax element, or slice group-level syntax element.
[0346] In some embodiments, whether the CIIP-MMVD flag is included in the bitstream depends on at least one value of at least one other syntax element. In one example, whether the CIIP-MMVD flag is included in the bitstream may depend on the value of at least one value of at least one other syntax element that indicates whether CIIP and / or MMVD are enabled.
[0347] In some embodiments, at least one other syntax element includes at least one other block-level syntax element.
[0348] In some embodiments, the CIIP-MMVD flag is indicated in the bitstream if at least one of the CIIP flag, CIIP-Location-Related (Intra-Frame) Prediction Combination (PDPC) flag, CIIP-Template Matching (TM) flag, MMVD flag, or CIIP-Affine flag is true or false.
[0349] In some embodiments, a first flag indicating the use of CIIP is indicated in the bitstream, and if the first flag is true, a second flag indicating the use of CIIP-MMVD is indicated in the bitstream, wherein if the second flag is true, an index indicating an MMVD candidate is indicated in the bitstream, and if the second flag is false, a third flag indicating CIIP-template matching (TM) is indicated in the bitstream. In one example, the first flag may include the CIIP flag. Furthermore, the second flag may include the CIIP-MMVD flag. The third flag may include the CIIP-TM flag.
[0350] In some embodiments, a first flag indicating the use of CIIP is indicated in the bitstream, and if the first flag is true, a second flag indicating the use of CIIP-MMVD is indicated in the bitstream, and if the second flag is false, a third flag indicating the use of CIIP-MMVD is indicated in the bitstream, and if the third flag is true, an index indicating an MMVD candidate is indicated in the bitstream. In one example, the first flag may include the CIIP flag. Furthermore, the second flag may include the CIIP-MMVD flag. The third flag may include the CIIP-TM flag.
[0351] In some embodiments, another flag is indicated in the bitstream based on the value of the CIIP-MMVD flag.
[0352] In some embodiments, whether another flag indicating the use of at least one of CIIP-template matching (TM), CIIP-position-related (intra-frame) prediction combination (PDPC), CIIP-affine, or another codec tool is indicated in the bitstream is based on the value of the CIIP-MMVD flag.
[0353] In some embodiments, if the CIIP-MMVD flag is true, then no CIIP-TM flag is indicated in the bitstream, and if the CIIP-MMVD flag is false, then the CIIP-TM flag is indicated in the bitstream.
[0354] In some embodiments, the CIIP-template matching (TM) flag is indicated in the bitstream regardless of whether the CIIP-MMVD flag is true or false.
[0355] In some embodiments, if the CIIP-MMVD flag is true, no CIIP-PDCP flag is indicated in the bitstream and / or the regular intra-frame or PDCP intra-frame default is used.
[0356] In some embodiments, if the CIIP-MMVD flag is false, the CIIP-PDPC flag is indicated in the bitstream.
[0357] In some embodiments, the CIIP-PDPC flag is indicated in the bitstream regardless of whether the CIIP-MMVD flag is true or false.
[0358] In some embodiments, whether the CIIP-MMVD flag is indicated in the bitstream is based on the use of at least one other codec tool.
[0359] In some embodiments, if at least one of CIIP, affine, or sub-block-based temporal motion vector prediction (SbTMVP) is applicable to the current video block, the CIIP-MMVD flag is indicated in the bitstream or CIIP-MMVD is applied.
[0360] In some embodiments, if the block dimension meets the conditions, the CIIP-MMVD flag is indicated in the bitstream or CIIP-MMVD is applied, and the block dimension is associated with at least one of the following: width, height, the ratio of width to height, or the block area of the current video block.
[0361] In some embodiments, a first flag indicating whether a CIIP mode or a CIIP-related mode is used is included in the bitstream, and if the first flag is true, a second flag indicating whether a CIIP-MMVD mode or at least one other CIIP-related mode is used, including CIIP-affine. In one example, the first flag may include a CIIP flag.
[0362] In some embodiments, if at least one other CIIP-related mode is not suitable for the current video block, the second flag is not indicated.
[0363] In some embodiments, if the second flag is true or not indicated, a third flag indicating whether the CIIP-MMVD mode is used for the block is indicated.
[0364] In some embodiments, a hybrid scheme using at least one of conventional intra-inter joint prediction (CIIP), CIIP-template matching (CIIP-TM), or CIIP-location-related (intra) prediction combination (PDPC) is used by CIIP-MMVD.
[0365] In some embodiments, the hybrid scheme used by at least one of conventional intra-inter joint prediction (CIIP), CIIP-template matching (CIIP-TM), or CIIP-location-related (intra) prediction combination (PDPC) is different from the hybrid scheme used by CIIP-MMVD.
[0366] In some embodiments, the hybrid weights for CIIP-MMVD are position-dependent within a block.
[0367] In some embodiments, different locations within a block have different blending weights.
[0368] In some embodiments, predefined hybrid weights are used for position within a block.
[0369] In some embodiments, the hybrid weight matrix for CIIP-MMVD is intra-mode dependent.
[0370] In some embodiments, different hybrid weight matrices are used based on whether the intra-frame angle pattern is near vertical or near horizontal.
[0371] In some embodiments, the current video block or video unit includes at least one of the following: color component, sub-picture, strip, slice, codec tree unit (CTU), CTU row, CTU group, codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), prediction block (PB), transform block (TB), block, sub-block of block, sub-region within block, or region containing more than one sample point or pixel.
[0372] In some embodiments, the indication of whether and / or how to apply method 2500 is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.
[0373] In some embodiments, the indication of whether and / or how to apply method 2500 is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.
[0374] In some embodiments, whether and / or how method 2500 is applied is based on at least one of the following: messages including one of the following: Dependency Parameter Set (DPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Picture Header, Strip Header, Piece Group Header, Maximum Codec Unit (LCU), Codec Unit (CU), LCU Row, LCU Group, Transform Unit (TU), Prediction Unit (PU), Block or Video Codec Unit, Position of CU, PU, TU, Block or Video Codec Unit, Block Dimension of the current video block and / or neighboring blocks of the current video block, Block Shape of the current video block and / or neighboring blocks of the current video block, Codec Mode of the block, Indicator of Color Format, Codec Tree Structure, Strip, Piece Group Type and / or Picture Type, Color Components, Temporal Layer Identifier (ID) or Standard Level, Level or Layer.
[0375] In some embodiments, the encoding / decoding mode includes one of the following: intra-block copy (IBC), non-IBC inter-frame mode, or non-IBC sub-block mode, or the color format includes one of the following: 4:2:0 or 4:4:4.
[0376] In some embodiments, the conversion includes encoding the current video block into a bitstream.
[0377] In some embodiments, the conversion includes decoding the current video block from the bitstream.
[0378] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores video data generated as a bitstream by a method performed by an apparatus for video processing. The method includes: determining a first prediction of a current video block based on a first encoding / decoding tool, the first encoding / decoding tool including a motion vector difference (MVD) related encoding / decoding tool; determining a third prediction of the current video block based on the first prediction and a second prediction of the current video block, the second prediction being determined based on a second encoding / decoding tool different from the first encoding / decoding tool; and generating a bitstream based on the third prediction.
[0379] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: determining a first prediction of a current video block based on a first encoding / decoding tool, the first encoding / decoding tool including a motion vector difference (MVD) related encoding / decoding tool; determining a third prediction of the current video block based on the first prediction and a second prediction of the current video block, the second prediction being determined based on a second encoding / decoding tool different from the first encoding / decoding tool; generating a bitstream based on the third prediction; and storing the bitstream in a non-transitory computer-readable recording medium.
[0380] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.
[0381] Item 1. A method for video processing, comprising: for a conversion between a current video block of a video and a bitstream of the video; determining a first prediction of the current video block based on a first encoding / decoding tool, the first encoding / decoding tool including a motion vector difference (MVD) related encoding / decoding tool; determining a third prediction of the current video block based on the first prediction and a second prediction of the current video block, the second prediction being determined based on a second encoding / decoding tool different from the first encoding / decoding tool; and performing the conversion based on the third prediction.
[0382] Item 2. The method according to Item 1, wherein the first codec tool includes at least one of a conventional Advanced Motion Vector Prediction (AMVP) codec tool or a Merge Mode with Motion Vector Difference (MMVD) codec tool, and wherein the second codec tool includes at least one of the following: a conventional intra-frame codec tool, a conventional intra-frame plane codec tool, a conventional intra-frame DC codec tool, a conventional intra-frame angle codec tool, a template-based intra-frame mode derivation (TIMD) codec tool, a decoder-side intra-frame mode derivation (DIMD) codec tool, an intra-frame sub-segmentation codec (ISP) codec tool, a position-dependent (intra-frame) prediction combination (PDPC) codec tool, a matrix-based intra-frame prediction (MIP) inter-frame codec tool, an intra-frame block copy (IBC) codec tool, or a conventional inter-frame codec tool.
[0383] Item 3. The method according to Item 1, wherein the third prediction is determined based on a weighted sum of the first prediction and the second prediction.
[0384] Item 4. The method according to any one of items 1 to 3, wherein the inter-frame components of the intra-inter-frame joint prediction (CIIP) are determined by the first codec tool.
[0385] Item 5. The method according to Item 4, wherein the first prediction is generated using conventional advanced motion vector difference (AMVD) or MMVD, the second prediction is generated using intra-frame mode, and the third prediction is determined by using a weighted average to mix the first prediction and the second prediction.
[0386] Item 6. The method according to Item 5, wherein the intra-frame mode includes at least one of the following: planar mode, DC mode, angle mode, matrix-based intra-frame prediction (MIP) mode, intra-frame sub-segmentation coding and decoding (ISP) mode, intra-frame block copying (IBC) mode, or intra-frame template matching prediction (intra-frame TMP) mode.
[0387] Item 7. The method according to Item 5, wherein the intra-frame mode is derived based on at least one of the following: template-based intra-frame mode derivation (TIMD), decoder-side intra-frame mode derivation (DIMD), or intra-frame template matching prediction (intra-frame TMP).
[0388] Item 8. The method according to Item 5, wherein the intra-frame components of the CIIP are processed by position-related (intra-frame) prediction combination (PDPC).
[0389] Item 9. The method according to any one of Items 1 to 8, wherein if CIIP-Merge mode with motion vector difference (MMVD) is used, the first prediction is generated using MMVD, the second prediction is generated using intra-frame mode, and the third prediction is determined by mixing the first prediction and the second prediction.
[0390] Item 10. The method according to Item 9 further includes: determining an MMVD candidate list, wherein the MMVD candidate list is different from or the same as the MMVD candidate list of a conventional MMVD.
[0391] Item 11. The method according to Item 10, wherein a basic motion vector (MV) candidate derivation scheme is used to construct the MMVD candidate list, the basic MV being different from or the same as the basic MV of a conventional MMVD.
[0392] Item 12. The method according to Item 10, wherein a certain number of basic MV candidates are used to construct the MMVD candidate list, the number of basic MV candidates being different from or the same as the number of basic MV candidates in a regular MMVD.
[0393] Item 13. The method according to Item 10, wherein a certain number of MV offsets are used to construct the MMVD candidate list, the number of MV offsets being different from or the same as the number of MV offsets in a regular MMVD.
[0394] Item 14. The method according to Item 10, wherein an MMVD candidate indicator is used to construct the MMVD candidate list, the MMVD candidate indicator being different from or the same as the MMVD candidate indicator of a regular MMVD.
[0395] Item 15. The method according to Item 9 or 10, wherein in CIIP-MMVD mode, only unidirectional prediction is applied to generate the first prediction.
[0396] Item 16. The method according to any one of items 10 to 15, wherein the MMVD candidate list for the CIIP-MMVD mode is reordered based on at least one metric after it is constructed.
[0397] Item 17. The method according to Item 16, wherein the at least one metric includes at least one of template matching cost or bilateral matching cost.
[0398] Item 18. The method according to Item 17, wherein if a reconstructed template region exists for the current video block, the template matching cost is used to reorder the MMVD candidate list.
[0399] Item 19. The method according to Item 17, wherein if the reconstructed template region of the current video block does not exist, the template matching cost is not used to reorder the MMVD candidate list.
[0400] Item 20. The method according to any one of items 10 to 15, wherein the MMVD candidate list for the CIIP-MMVD mode is not reordered.
[0401] Item 21. The method according to Item 16, wherein the cost of the MMVD candidates in the MMVD candidate list is determined based on the second prediction for the CIIP.
[0402] Item 22. The method according to Item 21, wherein the third prediction is determined as a reference template for the MMVD candidate, and the distortion between the reference template and the template of the current video block is determined as the cost of the MMVD candidate.
[0403] Item 23. The method according to any one of items 4 to 22, wherein whether an index utilizing weighted bidirectional prediction (BCW) is reordered by a metric is determined based on the CIIP pattern.
[0404] Item 24. The method according to Item 23, wherein if the CIIP mode is used, the index of BCW is reordered.
[0405] Item 25. The method according to Item 23, wherein the index of BCW is reordered for a portion of the CIIP sub-pattern, and the index of BCW is not reordered for another portion of the CIIP sub-pattern.
[0406] Item 26. The method according to Item 25, wherein if at least one pattern in the CIIP subgroup is used, the index of BCW is reordered; and if all patterns in the CIIP subgroup are not used, the index of BCW is not reordered.
[0407] Item 27. The method according to Item 26, wherein the CIIP subgroup comprises at least one of the following: regular CIIP mode, CIIP-template matching (TM) mode, CIIP-PDPC mode, CIIP-MMVD mode, or CIIP-affine mode.
[0408] Item 28. The method according to any one of items 10 to 27, wherein the index indicating the MMVD candidate in the MMVD candidate list is indicated in the bitstream.
[0409] Item 29. The method according to Item 28, wherein the MMVD candidate specified by the index is used to provide inter-frame prediction.
[0410] Item 30. The method according to Item 29 or 29, wherein TM or decoder-side motion vector refinement (DMVR) is used to refine the MMVD candidate specified by the index.
[0411] Item 31. The method according to any one of items 10 to 27, wherein the index of the MMVD candidate in the MMVD candidate list is not included in the bitstream.
[0412] Item 32. The method according to Item 31, wherein the encoder and decoder generate the same MMVD candidate based on predefined rules, the MMVD candidate being used to provide inter-frame prediction.
[0413] Item 33. The method according to Item 31 or 32, wherein the MMVD candidate list is reordered based on a metric after being constructed, and MMVD candidates in predefined positions of the MMVD candidate list are used by default, the predefined positions including at least one of the following: a start position, an end position, or a middle position.
[0414] Item 34. The method according to any one of items 1 to 33, wherein at least one syntax element indicating the use of CIIP-MMVD is indicated in the bitstream.
[0415] Item 35. The method according to Item 34, wherein the at least one syntax element includes at least one block-level CIIP-MMVD flag.
[0416] Item 36. The method according to Item 34 or 35, wherein the at least one syntax element includes at least one CIIP-MMVD flag of at least one of the following: sequence level, picture group level, picture level, strip level, slice group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header or slice group header.
[0417] Item 37. The method according to any one of items 34 to 36, wherein whether the at least one syntax element for the CIIP-MMVD is included in the bitstream or whether the CIIP-MMVD is applied depends on at least one value of the at least one syntax element, said at least one syntax element being at least one of: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptive parameter set (APS), a strip header, or a slice header.
[0418] Item 38. The method according to any one of items 34 to 36, wherein whether the at least one syntax element for the CIIP-MMVD is included in the bitstream or whether the CIIP-MMVD is applied depends on at least one value of at least one other syntax element.
[0419] Item 39. The method according to Item 38, wherein the at least one other syntax element includes at least one of the following: sequence-level syntax element, picture group-level syntax element, picture-level syntax element, strip-level syntax element, or slice group-level syntax element.
[0420] Item 40. The method according to Item 38, wherein whether the CIIP-MMVD flag is included in the bitstream depends on the at least one value of at least one other syntax element.
[0421] Item 41. The method according to Item 38, wherein the at least one other syntax element includes at least one other block-level syntax element.
[0422] Item 42. The method according to Item 39, wherein the CIIP-MMVD flag is indicated in the bitstream if at least one of the CIIP flag, CIIP-Location-Related (Intra-Frame) Prediction Combination (PDPC) flag, CIIP-Template Matching (TM) flag, MMVD flag, or CIIP-Affine flag is true or false.
[0423] Item 43. The method according to Item 41, wherein a first flag indicating the use of CIIP is indicated in the bitstream, and if the first flag is true, a second flag indicating the use of CIIP-MMVD is indicated in the bitstream, and wherein if the second flag is true, an index indicating an MMVD candidate is indicated in the bitstream, and if the second flag is false, a third flag indicating CIIP-template matching (TM) is indicated in the bitstream.
[0424] Item 44. The method according to Item 41, wherein a first flag indicating the use of CIIP is indicated in the bitstream, and if the first flag is true, a second flag indicating the use of CIIP-MMVD is indicated in the bitstream, and wherein if the second flag is false, a third flag indicating the use of CIIP-MMVD is indicated in the bitstream, and if the third flag is true, an index indicating an MMVD candidate is indicated in the bitstream.
[0425] Item 45. The method according to any one of items 34 to 44, wherein whether another flag is indicated in the bitstream is based on the value of the CIIP-MMVD flag.
[0426] Item 46. The method according to Item 45, wherein whether the other flag indicating the use of at least one of CIIP-template matching (TM), CIIP-position-related (intra-frame) prediction combination (PDPC), CIIP-affine or another codec tool is indicated in the bitstream is based on the value of the CIIP-MMVD flag.
[0427] Item 47. The method according to Item 46, wherein if the CIIP-MMVD flag is true, then no CIIP-TM flag is indicated in the bitstream, and wherein if the CIIP-MMVD flag is false, then the CIIP-TM flag is indicated in the bitstream.
[0428] Item 48. The method according to Item 34, wherein the CIIP-template matching (TM) flag is indicated in the bitstream regardless of whether the CIIP-MMVD flag is true or false.
[0429] Item 49. The method according to Item 46, wherein if the CIIP-MMVD flag is true, no CIIP-PDCP flag is indicated in the bitstream and / or the default is used in regular intra-frame or PDCP intra-frame.
[0430] Item 50. The method according to Item 46, wherein if the CIIP-MMVD flag is false, the CIIP-PDPC flag is indicated in the bitstream.
[0431] Item 51. The method according to Item 46, wherein the CIIP-PDPC flag is indicated in the bitstream regardless of whether the CIIP-MMVD flag is true or false.
[0432] Item 52. The method according to Item 34, wherein whether the CIIP-MMVD flag is indicated in the bitstream is based on the use of at least one other codec tool.
[0433] Item 53. The method according to Item 52, wherein if at least one of CIIP, affine, or sub-block-based temporal motion vector prediction (SbTMVP) is applicable to the current video block, then the CIIP-MMVD flag is indicated in the bitstream or the CIIP-MMVD is applied.
[0434] Item 54. The method according to Item 52, wherein the CIIP-MMVD flag is indicated in the bitstream or the CIIP-MMVD is applied if the block dimension satisfies the following conditions: width, height, the ratio of width to height, or the block area of the current video block.
[0435] Item 55. The method according to any one of items 1 to 33, wherein a first flag indicating whether a CIIP mode or a CIIP-related mode is used is indicated in the bitstream, and if the first flag is true, a second flag indicating whether a CIIP-MMVD mode or at least one other CIIP-related mode is used, the at least one other CIIP-related mode including CIIP-affine.
[0436] Item 56. The method according to Item 55, wherein the second flag is not indicated if the at least one other CIIP-related mode is not suitable for the current video block.
[0437] Item 57. The method according to Item 55, wherein if the second flag is true or not indicated, a third flag indicating whether the CIIP-MMVD mode is used for the block is indicated.
[0438] Item 58. The method according to any one of Items 1 to 57, wherein the hybrid scheme used by at least one of conventional intra-inter joint prediction (CIIP), CIIP-template matching (CIIP-TM), or CIIP-location-related (intra) prediction combination (PDPC) is used by CIIP-MMVD.
[0439] Item 59. The method according to any one of Items 1-57, wherein the hybrid scheme used by at least one of conventional intra-inter joint prediction (CIIP), CIIP-template matching (CIIP-TM), or CIIP-location-related (intra) prediction combination (PDPC) is different from the hybrid scheme used by CIIP-MMVD.
[0440] Item 60. The method according to any one of items 1 to 59, wherein the hybrid weights for CIIP-MMVD are position-dependent within the block.
[0441] Item 61. The method according to Item 60, wherein different positions within the block have different mixing weights.
[0442] Item 62. The method according to Item 60, wherein predefined blending weights are used for the position within the block.
[0443] Item 63. The method according to any one of items 1 to 59, wherein the hybrid weight matrix for CIIP-MMVD is intra-mode dependent.
[0444] Item 64. The method according to Item 63, wherein different hybrid weight matrices are used based on whether the intra-frame angle mode is near vertical or near horizontal.
[0445] Item 65. The method according to any one of items 1 to 64, wherein the current video block or video unit includes at least one of the following: color component, sub-picture, strip, slice, codec tree unit (CTU), CTU row, CTU group, codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), prediction block (PB), transform block (TB), block, sub-block of block, sub-region within block, or region containing more than one sample point or pixel.
[0446] Item 66. The method according to any one of items 1 to 65, wherein an indication of whether and / or how the method is applied is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.
[0447] Item 67. The method according to any one of items 1 to 65, wherein an indication of whether and / or how to apply the method is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.
[0448] Item 68. The method according to any one of Items 1-65, wherein whether and / or how the method is applied is based on at least one of the following: a message including one of the following: Dependency Parameter Set (DPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Picture Header, Strip Header, Piece Group Header, Maximum Codec Unit (LCU), Codec Unit (CU), LCU Row, LCU Group, Transform Unit (TU), Prediction Unit (PU), Block or Video Codec Unit, Position of CU, PU, TU, Block or Video Codec Unit, Block Dimension of the current video block and / or the neighboring blocks of the current video block, Block Shape of the current video block and / or the neighboring blocks of the current video block, Codec Mode of the Block, Indicator of Color Format, Codec Tree Structure, Strip, Piece Group Type and / or Picture Type, Color Components, Temporal Layer Identifier (ID) or Standard Level, Level or Layer.
[0449] Item 69. The method according to Item 68, wherein the encoding / decoding mode includes one of the following: intra-block copy (IBC), non-IBC inter-frame mode, or non-IBC sub-block mode, or wherein the color format includes one of the following: 4:2:0 or 4:4:4.
[0450] Item 70. The method according to any one of items 1 to 69, wherein the conversion includes encoding the current video block into the bitstream.
[0451] Item 71. The method according to any one of items 1 to 69, wherein the conversion includes decoding the current video block from the bitstream.
[0452] Item 72. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of items 1 to 71.
[0453] Item 73. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of items 1 to 71.
[0454] Item 74. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: determining a first prediction of a current video block of the video based on a first encoding / decoding tool, the first encoding / decoding tool including a motion vector difference (MVD) related encoding / decoding tool; determining a third prediction of the current video block based on the first prediction and a second prediction of the current video block, the second prediction being determined based on a second encoding / decoding tool different from the first encoding / decoding tool; and generating the bitstream based on the third prediction.
[0455] Item 75. A method for storing a bitstream of video, comprising: determining a first prediction of a current video block of the video based on a first encoding / decoding tool, the first encoding / decoding tool including a motion vector difference (MVD) related encoding / decoding tool; determining a third prediction of the current video block based on the first prediction and a second prediction of the current video block, the second prediction being determined based on a second encoding / decoding tool different from the first encoding / decoding tool; generating the bitstream based on the third prediction; and storing the bitstream in a non-transitory computer-readable recording medium.
[0456] Example device Figure 26 A block diagram of a computing device 2600 in which various embodiments of the present disclosure may be implemented is shown. The computing device 2600 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0457] It should be understood that, Figure 26 The computing device 2600 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.
[0458] like Figure 26 As shown, computing device 2600 includes general-purpose computing device 2600. Computing device 2600 may include at least one or more processors or processing units 2610, memory 2620, storage unit 2630, one or more communication units 2640, one or more input devices 2650, and one or more output devices 2660.
[0459] In some embodiments, the computing device 2600 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 2600 can support any type of interface to the user (such as "wearable" circuit systems, etc.).
[0460] Processing unit 2610 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 2620. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 2600. Processing unit 2610 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0461] Computing device 2600 typically includes various computer storage media. Such media can be any media accessible by computing device 2600, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 2620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 2630 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 2600.
[0462] The computing device 2600 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 26 Not shown, but may provide disk drives for reading from and / or writing to removable non-volatile disks, and optical disc drives for reading from and / or writing to removable non-volatile optical discs. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0463] Communication unit 2640 communicates with another computing device via a communication medium. Furthermore, the functionality of components in computing device 2600 can be implemented by a single computing cluster or by multiple computing machines communicating via communication connections. Therefore, computing device 2600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0464] Input device 2650 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 2660 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 2640, computing device 2600 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 2600 can also communicate with one or more devices that enable a user to interact with computing device 2600, or any device that enables computing device 2600 to communicate with one or more other computing devices (e.g., network card, modem, etc.), if needed. Such communication can be performed via an input / output (I / O) interface (not shown).
[0465] In some embodiments, some or all components of computing device 2600 are not integrated into a single device, but may be arranged in a cloud computing architecture. In a cloud computing architecture, components may be provided remotely and work together to achieve the functions described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing is provided via a wide area network (WAN) such as the Internet using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at a remote location. Computing resources in a cloud computing environment may be consolidated or distributed at locations in remote data centers. Cloud computing infrastructure may provide services through shared data centers, although they may appear as a single access point for users. Therefore, cloud computing architectures can be used to provide the components and functions described herein from service providers at remote locations. Alternatively, they may be provided from conventional servers, or directly installed or otherwise installed on client devices.
[0466] In embodiments of this disclosure, computing device 2600 can be used to implement video encoding / decoding. Memory 2620 may include one or more video codec modules 2625 having one or more program instructions. These modules can be accessed and executed by processing unit 2610 to perform the functions of the various embodiments described herein.
[0467] In an example embodiment of performing video encoding, input device 2650 may receive video data as input 2670 to be encoded. The video data may be processed, for example, by video codec module 2625 to generate an encoded bitstream. The encoded bitstream may be provided as output 2680 via output device 2660.
[0468] In an example embodiment of performing video decoding, input device 2650 may receive an encoded bitstream as input 2670. The encoded bitstream may be processed, for example, by a video codec module 2625 to generate decoded video data. The decoded video data may be provided as output 2680 via output device 2660.
[0469] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These changes are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.
Claims
1. A method for video processing, comprising: For the conversion between the current video block and the bitstream of the video, a first prediction of the current video block is determined based on a first encoding and decoding tool, the first encoding and decoding tool including a motion vector difference (MVD) related encoding and decoding tool; A third prediction for the current video block is determined based on the first prediction and a second prediction for the current video block, wherein the second prediction is determined based on a second codec tool different from the first codec tool; and The transformation is performed based on the third prediction.
2. The method of claim 1, wherein the first encoding / decoding tool comprises at least one of a conventional Advanced Motion Vector Prediction (AMVP) encoding / decoding tool or a Merge Mode with Motion Vector Difference (MMVD) encoding / decoding tool, and The second encoding / decoding tool includes at least one of the following: a conventional intra-frame encoding / decoding tool, a conventional intra-frame plane encoding / decoding tool, a conventional intra-frame DC encoding / decoding tool, a conventional intra-frame angle encoding / decoding tool, a template-based intra-frame mode derivation (TIMD) encoding / decoding tool, a decoder-side intra-frame mode derivation (DIMD) encoding / decoding tool, an intra-frame sub-segmentation encoding / decoding (ISP) encoding / decoding tool, a position-dependent (intra-frame) prediction combination (PDPC) encoding / decoding tool, a matrix-based intra-frame prediction (MIP) inter-frame encoding / decoding tool, an intra-frame block copy (IBC) encoding / decoding tool, or a conventional inter-frame encoding / decoding tool.
3. The method of claim 1, wherein the third prediction is determined based on a weighted sum of the first prediction and the second prediction.
4. The method according to any one of claims 1 to 3, wherein the inter-frame components of the intra-inter-frame joint prediction (CIIP) are determined by the first encoding / decoding tool.
5. The method of claim 4, wherein the first prediction is generated using conventional Advanced Motion Vector Difference (AMVD) or MMVD, the second prediction is generated using intra-frame mode, and the third prediction is determined by using a weighted average to mix the first prediction and the second prediction.
6. The method according to claim 5, wherein the intra-frame mode includes at least one of the following: planar mode, DC mode, angle mode, matrix-based intra-frame prediction (MIP) mode, intra-frame sub-segmentation codec (ISP) mode, intra-frame block copy (IBC) mode, or intra-frame template matching prediction (intra-frame TMP) mode.
7. The method of claim 5, wherein the intra-frame mode is derived based on at least one of the following: template-based intra-frame mode derivation (TIMD), decoder-side intra-frame mode derivation (DIMD), or intra-frame template matching prediction (intra-frame TMP).
8. The method of claim 5, wherein the intra-frame components of the CIIP are processed by position-related (intra-frame) prediction combination (PDPC).
9. The method according to any one of claims 1 to 8, wherein if CIIP-Merge mode with motion vector difference (MMVD) is used, the first prediction is generated using MMVD, the second prediction is generated using intra-frame mode, and the third prediction is determined by mixing the first prediction and the second prediction.
10. The method of claim 9, further comprising: Determine the MMVD candidate list. The MMVD candidate list mentioned therein may be different from or the same as the MMVD candidate list of a regular MMVD.
11. The method of claim 10, wherein a basic motion vector (MV) candidate derivation scheme is used to construct the MMVD candidate list, the basic MV being different from or the same as the basic MV of a conventional MMVD.
12. The method of claim 10, wherein a certain number of basic MV candidates are used to construct the MMVD candidate list, the number of basic MV candidates being different from or the same as the number of basic MV candidates in a conventional MMVD.
13. The method of claim 10, wherein a certain number of MV offsets are used to construct the MMVD candidate list, the number of MV offsets being different from or the same as the number of MV offsets in a conventional MMVD.
14. The method of claim 10, wherein the MMVD candidate indicator is used to construct the MMVD candidate list, the MMVD candidate indicator being different from or the same as the MMVD candidate indicator of a regular MMVD.
15. The method of claim 9 or 10, wherein in CIIP-MMVD mode, only one-way prediction is applied to generate the first prediction.
16. The method according to any one of claims 10 to 15, wherein the MMVD candidate list for the CIIP-MMVD mode is reordered based on at least one metric after being constructed.
17. The method of claim 16, wherein the at least one metric includes at least one of template matching cost or bilateral matching cost.
18. The method of claim 17, wherein if a reconstructed template region exists for the current video block, the template matching cost is used to reorder the MMVD candidate list.
19. The method of claim 17, wherein if the reconstructed template region of the current video block does not exist, the template matching cost is not used to reorder the MMVD candidate list.
20. The method according to any one of claims 10 to 15, wherein the MMVD candidate list for the CIIP-MMVD mode is not reordered.
21. The method of claim 16, wherein the cost of the MMVD candidates in the MMVD candidate list is determined based on the second prediction for the CIIP.
22. The method of claim 21, wherein the third prediction is determined as a reference template for the MMVD candidate, and the distortion between the reference template and the template of the current video block is determined as the cost of the MMVD candidate.
23. The method of any one of claims 4 to 22, wherein whether the index using weighted bidirectional prediction (BCW) is reordered by metric is determined based on the CIIP pattern.
24. The method of claim 23, wherein if the CIIP mode is used, the index of BCW is reordered.
25. The method of claim 23, wherein the index of BCW is reordered for a portion of the CIIP subpattern, and the index of BCW is not reordered for another portion of the CIIP subpattern.
26. The method of claim 25, wherein if at least one pattern in the CIIP subgroup is used, the index of BCW is reordered; and If none of the patterns in the CIIP subgroup are used, the index of BCW will not be reordered.
27. The method of claim 26, wherein the CIIP subgroup comprises at least one of the following: Standard CIIP mode, CIIP-Template Matching (TM) Mode CIIP-PDPC mode CIIP-MMVD mode, or CIIP - Affine Mode.
28. The method according to any one of claims 10 to 27, wherein the index indicating the MMVD candidate in the MMVD candidate list is indicated in the bitstream.
29. The method of claim 28, wherein the MMVD candidate specified by the index is used to provide inter-frame prediction.
30. The method of claim 29 or 29, wherein TM or decoder-side motion vector refinement (DMVR) is used to refine the MMVD candidate specified by the index.
31. The method according to any one of claims 10 to 27, wherein the index of the MMVD candidate in the MMVD candidate list is not included in the bitstream.
32. The method of claim 31, wherein the encoder and decoder generate the same MMVD candidate based on predefined rules, the MMVD candidate being used to provide inter-frame prediction.
33. The method of claim 31 or 32, wherein the MMVD candidate list is reordered based on a metric after being constructed, and MMVD candidates in predefined positions of the MMVD candidate list are used by default, the predefined positions including at least one of the following: a start position, an end position, or a middle position.
34. The method according to any one of claims 1 to 33, wherein at least one syntax element indicating the use of CIIP-MMVD is indicated in the bitstream.
35. The method of claim 34, wherein the at least one syntax element comprises at least one block-level CIIP-MMVD flag.
36. The method according to claim 34 or 35, wherein the at least one syntax element comprises at least one CIIP-MMVD flag of at least one of the following: sequence level, picture group level, picture level, strip level, slice group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header or slice group header.
37. The method according to any one of claims 34 to 36, wherein whether the at least one syntax element for the CIIP-MMVD is included in the bitstream or whether the CIIP-MMVD is applied depends on at least one value of the at least one syntax element, said at least one syntax element being at least one of: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptive parameter set (APS), a strip header, or a slice header.
38. The method according to any one of claims 34 to 36, wherein whether the at least one syntax element for the CIIP-MMVD is included in the bitstream or whether the CIIP-MMVD is applied depends on at least one value of at least one other syntax element.
39. The method of claim 38, wherein the at least one other syntax element comprises at least one of the following: a sequence-level syntax element, a picture group-level syntax element, a picture-level syntax element, a strip-level syntax element, or a slice group-level syntax element.
40. The method of claim 38, wherein whether the CIIP-MMVD flag is included in the bitstream depends on the at least one value of at least one other syntax element.
41. The method of claim 38, wherein the at least one other syntax element comprises at least one other block-level syntax element.
42. The method of claim 39, wherein the CIIP-MMVD flag is indicated in the bitstream if at least one of the CIIP flag, CIIP-Location-Related (Intra-Frame) Prediction Combination (PDPC) flag, CIIP-Template Matching (TM) flag, MMVD flag, or CIIP-Affine flag is true or false.
43. The method of claim 41, wherein a first flag indicating the use of CIIP is indicated in the bitstream, and if the first flag is true, a second flag indicating the use of CIIP-MMVD is indicated in the bitstream, and If the second flag is true, the index of the MMVD candidate is indicated in the bitstream, and if the second flag is false, the third flag of CIIP-template matching (TM) is indicated in the bitstream.
44. The method of claim 41, wherein a first flag indicating the use of CIIP is indicated in the bitstream, and if the first flag is true, a second flag indicating the use of CIIP-MMVD is indicated in the bitstream, and If the second flag is false, a third flag indicating the use of CIIP-MMVD is indicated in the bitstream, and if the third flag is true, an index indicating an MMVD candidate is indicated in the bitstream.
45. The method according to any one of claims 34 to 44, wherein whether another flag is indicated in the bitstream is based on the value of the CIIP-MMVD flag.
46. The method of claim 45, wherein whether the other flag indicating the use of at least one of CIIP-template matching (TM), CIIP-position-related (intra-frame) prediction combination (PDPC), CIIP-affine, or another codec tool is indicated in the bitstream is based on the value of the CIIP-MMVD flag.
47. The method of claim 46, wherein if the CIIP-MMVD flag is true, then no CIIP-TM flag is indicated in the bitstream, and If the CIIP-MMVD flag is false, then the CIIP-TM flag is indicated in the bitstream.
48. The method of claim 34, wherein the CIIP-template matching (TM) flag is indicated in the bitstream regardless of whether the CIIP-MMVD flag is true or false.
49. The method of claim 46, wherein if the CIIP-MMVD flag is true, then no CIIP-PDCP flag is indicated in the bitstream and / or the default is used within a regular intra-frame or PDCP intra-frame.
50. The method of claim 46, wherein if the CIIP-MMVD flag is false, the CIIP-PDPC flag is indicated in the bitstream.
51. The method of claim 46, wherein the CIIP-PDPC flag is indicated in the bitstream regardless of whether the CIIP-MMVD flag is true or false.
52. The method of claim 34, wherein whether the CIIP-MMVD flag is indicated in the bitstream is based on the use of at least one other codec tool.
53. The method of claim 52, wherein if at least one of CIIP, affine, or sub-block-based temporal motion vector prediction (SbTMVP) is applicable to the current video block, then the CIIP-MMVD flag is indicated in the bitstream or the CIIP-MMVD is applied.
54. The method of claim 52, wherein the CIIP-MMVD flag is indicated in the bitstream or the CIIP-MMVD is applied if the block dimension satisfies a condition, the block dimension being associated with at least one of the following: width, height, the ratio of width to height, or the block area of the current video block.
55. The method according to any one of claims 1 to 33, wherein a first flag indicating whether a CIIP mode or a CIIP-related mode is used is indicated in the bitstream, and if the first flag is true, a second flag indicating whether a CIIP-MMVD mode or at least one other CIIP-related mode is used, the at least one other CIIP-related mode including CIIP-affine.
56. The method of claim 55, wherein the second flag is not indicated if the at least one other CIIP-related mode is not suitable for the current video block.
57. The method of claim 55, wherein if the second flag is true or not indicated, a third flag indicating whether the CIIP-MMVD mode is used for the block is indicated.
58. The method according to any one of claims 1 to 57, wherein the hybrid scheme used by at least one of conventional intra-inter joint prediction (CIIP), CIIP-template matching (CIIP-TM), or CIIP-location-related (intra) prediction combination (PDPC) is used by CIIP-MMVD.
59. The method according to any one of claims 1 to 57, wherein the hybrid scheme used by at least one of conventional intra-inter joint prediction (CIIP), CIIP-template matching (CIIP-TM), or CIIP-location-related (intra) prediction combination (PDPC) is different from the hybrid scheme used by CIIP-MMVD.
60. The method according to any one of claims 1 to 59, wherein the mixing weights for CIIP-MMVD are position-dependent within the block.
61. The method of claim 60, wherein different locations within the block have different mixing weights.
62. The method of claim 60, wherein a predefined blending weight is used for the position within the block.
63. The method according to any one of claims 1 to 59, wherein the hybrid weight matrix for CIIP-MMVD is intra-mode dependent.
64. The method of claim 63, wherein different hybrid weight matrices are used based on whether the intra-frame angle mode is near vertical or near horizontal.
65. The method according to any one of claims 1 to 64, wherein the current video block or video unit comprises at least one of the following: Color components, sub-images, stripes, slices, codec tree units (CTUs), CTU rows, CTU groups, codec units (CUs), prediction units (PUs), transform units (TUs), codec tree blocks (CTBs), codec blocks (CBs), prediction blocks (PBs), transform blocks (TBs), blocks, sub-blocks of blocks, sub-regions within blocks, or regions containing more than one sample point or pixel.
66. The method according to any one of claims 1 to 65, wherein an indication of whether and / or how the method is applied is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.
67. The method according to any one of claims 1 to 65, wherein an indication of whether and / or how to apply the method is indicated in one of the following: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.
68. The method according to any one of claims 1 to 65, wherein whether and / or how the method is applied is based on at least one of the following: Messages that include one of the following: Dependency Parameter Set (DPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Picture Header, Strip Header, Slice Header, Maximum Codec Unit (LCU), Codec Unit (CU), LCU Row, LCU Group, Transform Unit (TU), Prediction Unit (PU) Block, or Video Codec Unit. The location of CU, PU, TU, block, or the video codec unit. The block dimension of the current video block and / or the neighboring blocks of the current video block. The block shape of the current video block and / or the neighboring blocks of the current video block. Block encoding / decoding modes, Indicators of color format, Encoder tree structure, Strip, slice group type and / or image type, Color components, Temporal layer identifier (ID), or Standard grade, level, or tier.
69. The method of claim 68, wherein the encoding / decoding mode includes one of the following: intra-block copy (IBC), non-IBC inter-frame mode, or non-IBC sub-block mode, or The color format mentioned includes one of the following: 4:2:0 or 4:4:
4.
70. The method of any one of claims 1 to 69, wherein the conversion comprises encoding the current video block into the bitstream.
71. The method according to any one of claims 1 to 69, wherein the conversion comprises decoding the current video block from the bitstream.
72. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 71.
73. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 71.
74. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: A first prediction of the current video block of the video is determined based on a first encoding and decoding tool, the first encoding and decoding tool including a motion vector difference (MVD) related encoding and decoding tool; A third prediction for the current video block is determined based on the first prediction and a second prediction for the current video block, wherein the second prediction is determined based on a second codec tool different from the first codec tool; and The bit stream is generated based on the third prediction.
75. A method for storing a bitstream of video, comprising: A first prediction of the current video block of the video is determined based on a first encoding and decoding tool, the first encoding and decoding tool including a motion vector difference (MVD) related encoding and decoding tool; A third prediction for the current video block is determined based on the first prediction and the second prediction for the current video block, wherein the second prediction is determined based on a second codec tool that is different from the first codec tool; The bitstream is generated based on the third prediction; as well as The bitstream is stored in a non-transitory computer-readable recording medium.