Method and device for video processing and medium
By using affine information based on video blocks and multi-context pair intra- and inter-frame joint prediction techniques, the problem of low efficiency in existing video encoding and decoding is solved, achieving more efficient video data compression and decoding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2024-09-30
- Publication Date
- 2026-05-01
AI Technical Summary
Existing video encoding and decoding technologies have room for improvement in encoding and decoding efficiency, especially in intra-frame and inter-frame joint prediction, where existing methods struggle to effectively utilize various encoding and decoding tools and syntax elements for efficient encoding and decoding.
The intra-frame and inter-frame joint prediction (CIIP) technique, based on affine information and multiple context pairs of video blocks, is used to perform encoding and decoding transformation of video blocks using different encoding and decoding tools and syntax elements, generating predictions and bitstreams.
It improves the efficiency of video encoding and decoding by utilizing affine codec tools and CIIP technology, achieving more efficient video data compression and decoding.
Smart Images

Figure CN121970332A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure generally relate to video processing techniques, and more specifically, to sub-block-based intra- and inter-frame joint prediction. Background Technology
[0002] Today, digital video capabilities are being applied to all aspects of people's lives. Various video compression technologies have been proposed for video encoding / decoding, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-T H.265 High Efficiency Video Codec (HEVC) standard, and Multi-Functional Video Codec (VVC) standard. However, the encoding and decoding efficiency of video encoding and decoding technologies is generally expected to be further improved. Summary of the Invention
[0003] Embodiments of this disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method includes: a conversion between a current video block and a bitstream of the video; determining affine information of the current video block based on first affine information of a first block of the video, the first block being encoded and decoded using a first codec tool different from the affine codec tool; determining a prediction of the current video block based on the affine information; and performing the conversion based on the prediction. The method according to the first aspect of this disclosure enables the generation of predictions having affine information of blocks encoded and decoded using a codec tool different from the affine codec tool.
[0005] In a second aspect, another method for video processing is proposed. This method includes: conversion between a current video block and a video bitstream; encoding / decoding syntax elements related to intra-frame joint prediction (CIIP) based on more than one context; and performing the conversion based on the syntax elements. The method according to the second aspect of this disclosure enables encoding / decoding of CIIP-related syntax elements using at least one context.
[0006] In a third aspect, an apparatus for video processing is proposed. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform a method according to either the first or second aspect of this disclosure.
[0007] In a fourth aspect, a non-transitory computer-readable storage medium is provided. This non-transitory computer-readable storage medium stores instructions that cause a processor to perform a method according to the first or second aspect of this disclosure.
[0008] In a fifth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining affine information of a current video block based on first affine information of a first block of video, the first block being encoded and decoded using a first encoding / decoding tool different from the affine encoding / decoding tool; determining a prediction of the current video block based on the affine information; and generating a bitstream based on the prediction.
[0009] In a sixth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores video data generated as a bitstream by a method performed by an apparatus for video processing. The method includes: encoding and decoding syntax elements related to intra-frame / inter-frame joint prediction (CIIP) based on more than one context; and generating a bitstream based on the syntax elements.
[0010] In the seventh aspect, a method for storing video bitstreams is proposed. The method includes: determining affine information of a current video block based on first affine information of a first block of video, the first block being encoded and decoded using a first encoding / decoding tool different from the affine encoding / decoding tool; determining a prediction of the current video block based on the affine information; generating a bitstream based on the prediction; and storing the bitstream in a non-transitory computer-readable recording medium.
[0011] In the eighth aspect, a method for storing video bitstreams is proposed. This method includes: encoding and decoding syntax elements related to intra-frame / inter-frame joint prediction (CIIP) based on more than one context; generating a bitstream based on the syntax elements; and storing the bitstream in a non-transitory computer-readable recording medium.
[0012] This synopsis aims to present, in a simplified form, the selected concepts further described below in the detailed embodiments. This synopsis is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0013] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0014] Figure 1 A block diagram of an example video codec system according to some embodiments of the present disclosure is shown; Figure 2 A block diagram of a first example video encoder according to some embodiments of the present disclosure is shown; Figure 3A block diagram of an example video decoder according to some embodiments of the present disclosure is shown; Figure 4 This shows the locations of spatial and temporal neighbor blocks used in the construction of the AMVP / Merge candidate list; Figure 5 This shows the positions of non-adjacent candidates in the ECM; Figure 6 An affine motion model based on control points is shown; Figure 7 An example affine MVF for each sub-block is shown; Figure 8 The location of the inherited affine motion prediction value is shown; Figure 9 This demonstrates the inheritance of control point motion vectors; Figure 10 The locations of candidate positions for the constructed affine Merge pattern are shown; Figure 11 The spatial nearest neighbor used to derive the affine Merge candidate is shown; Figure 12 The affine Merge candidates from non-nearest neighbors to the constructed ones are shown; Figure 13 An example of generating HAPC is shown; Figure 14 A diagram illustrating the regression-based affine Merge candidate derivation is shown. Figure 15 This demonstrates template matching execution over the search area surrounding the initial MV; Figure 16 The template and the corresponding reference template are shown; Figure 17 A template and a reference template are shown for a block with sub-block motion that uses motion information of the current block's sub-blocks; Figure 18 The derivation of the sub-CU motion field obtained by applying motion displacement based on neighbor motion information is shown; Figure 19 The top and left neighbor blocks used in the CIIP weight derivation are shown; Figure 20 A flowchart of CIIP_PDPC using the extended CIIP mode with PDPC is shown; Figure 21 The method for dividing angle patterns is shown; Figure 22 This demonstrates the generation of sub-block templates for SbTMVP; Figure 23 The diamond-shaped area in the search region is shown; Figure 24A flowchart of a method for video processing according to an embodiment of the present disclosure is shown; Figure 25 A flowchart of another method for video processing according to embodiments of the present disclosure is shown; and Figure 26 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.
[0015] In all accompanying drawings, the same or similar reference numerals usually refer to the same or similar elements. Detailed Implementation
[0016] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.
[0017] In the following description and claims, unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0018] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Additionally, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, whether explicitly described or not, it is believed that such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.
[0019] It should be understood that although the terms “first” and “second”, etc., can be used to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0020] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0021] Example Environment Figure 1 This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0022] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.
[0023] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded / decoded representation of the video data. The bitstream may include encoded / decoded images and associated data. The encoded / decoded images are the encoded / decoded representations of the images. The associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. The encoded video data may be transmitted directly to destination device 120 via network 130A through I / O interface 116. The encoded video data may also be stored on storage medium / server 130B for access by destination device 120.
[0024] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.
[0025] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or further standards.
[0026] Figure 2 This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... Figure 1 An example of a video encoder 114 in system 100 is shown.
[0027] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. Figure 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0028] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.
[0029] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0030] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... Figure 2 The examples are shown separately.
[0031] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0032] The mode selection unit 203 can select one of several codec modes (intra-frame codec or inter-frame codec) based, for example, on the error result, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-frame / inter-frame joint prediction (CIIP) mode, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).
[0033] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.
[0034] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks that do not depend on macroblocks within the same image.
[0035] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0036] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for reference images in list 0 to find a reference video block for the current video block, and can also search for reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate reference indices indicating the reference images containing the reference video blocks in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0037] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0038] In one example, the motion estimation unit 204 may indicate a value to the video decoder 300 in the syntax structure associated with the current video block, which indicates that the current video block has the same motion information as another video block.
[0039] In another example, motion estimation unit 204 may identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0040] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0041] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0042] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0043] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform subtraction operations.
[0044] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0045] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0046] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block, which is stored in the buffer 213.
[0047] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0048] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0049] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.
[0050] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0051] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.
[0052] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded video data blocks). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information from the entropy-decoded video data, which includes motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine this information, for example, by performing AMVP and Merge pattern. AMVP is used, which involves deriving several most likely candidates based on data from neighboring PBs and reference pictures. Motion information typically includes horizontal and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge pattern" may refer to deriving motion information from spatially or temporally neighboring blocks.
[0053] The motion compensation unit 302 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. The identifier of the interpolation filter to be used, with sub-pixel accuracy, can be included in the syntax element.
[0054] The motion compensation unit 302 can use interpolation filters, such as those used by the video encoder 200 during the encoding of a video block, to calculate interpolations for sub-integer pixels of a reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate a prediction block.
[0055] Motion compensation unit 302 may use at least some of the syntax information to determine the block size of the frames(multiple) and / or stripes(multiple) used to encode the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be the entire image or a region of the image.
[0056] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies the inverse transform.
[0057] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding predicted block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be used to filter the decoded block to eliminate block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0058] Some exemplary embodiments of this disclosure will be described in detail below. It should be understood that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section only. Furthermore, although some embodiments are described with reference to multi-functional video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Furthermore, although some embodiments describe video codec steps in detail, it should be understood that the corresponding decoding steps of the inverse codec will be implemented by the decoder. Additionally, the term video processing includes video codec or compression, video decoding or decompression, and video transcoding, wherein video pixels are represented from one compression format to another compression format or at different compression bit rates.
[0059] 1. Brief Overview This disclosure relates to video codec technology. Specifically, it pertains to joint prediction methods in video codecs. This idea can be applied alone or in various combinations to any standard or non-standard video codec.
[0060] 2. Introduction The exponential growth of multimedia data poses a significant challenge to video encoding and decoding. To meet the ever-increasing demand for more efficient compression technologies, the ITU-T and ISO / IEC have developed a series of video encoding and decoding standards over the past few decades. Specifically, the ITU-T developed the H.261 and H.263 standards, and ISO / IEC developed the MPEG-1 and MPEG-4 visual standards. The two organizations have jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, the H.265 / HEVC standard, and the latest VVC standard. Since H.262 / MPEG-2, a hybrid video encoding and decoding framework has been adopted, utilizing intra / inter-frame prediction plus transform encoding and decoding.
[0061] Figure 4 This shows the locations of spatial and temporal neighbor blocks used in the construction of the AMVP / Merge candidate list.
[0062] 2.1. MVP in Video Encoding and Decoding Inter-frame prediction aims to eliminate temporal redundancy between adjacent frames and is an indispensable component in hybrid video codec frameworks. Specifically, inter-frame prediction utilizes the content specified by motion vectors (MVs) as the predicted version of the current block to be encoded / decoded, thus transmitting only residual signals and motion information in the bitstream. To reduce the cost of MV signaling, motion vector prediction (MVP) emerged as an efficient mechanism for conveying motion information. Early strategies simply used the MV of a specified neighboring block or the median MV of neighboring blocks as the MVP. In H.265 / HEVC, a contention mechanism is involved, where rate-distortion optimization (RDO) selects the best MVP from multiple candidates. Specifically, Advanced MVP (AMVP) mode and Merge mode with different motion information signaling strategies were designed. Using AMVP mode, the reference index, the MVP candidate index referencing the AMVP candidate list, and the motion vector difference (MVD) are transmitted via signaling. Regarding Merge mode, only the Merge index referencing the Merge candidate list is transmitted via signaling, and all motion information associated with the Merge candidate is inherited. Both the AMVP and Merge modes require building an MVP candidate list. The following describes the details of the building process for these two modes.
[0063] AMVP mode: AMVP utilizes the spatial-temporal correlation of motion vectors with neighboring blocks for explicit transfer of motion parameters. For each list of reference images, a motion vector candidate list is constructed by first checking the availability of temporal neighbors to the left and top, removing redundant candidates, and adding zero vectors to bring the candidate list to a constant length. For the derivation of spatial motion vector candidates, the final result is based on locations such as... Figure 4The motion vectors of five blocks at different locations are shown to derive two motion vector candidates. The five neighboring blocks located at B0, B1, B2, and A0, A1 are classified into two groups: group A includes the three spatially adjacent blocks above, and group B includes the two spatially adjacent blocks to the left. The two motion vector candidates are derived using the first available candidates from group A and group B in a predefined order, respectively. For the temporal motion vector candidate derivation, as follows... Figure 4 As shown, a motion vector candidate is derived based on two distinct co-positions examined sequentially (bottom right (C0) and center (C1)). To avoid redundant MV candidates, duplicate motion vector candidates in the list are discarded. If the number of potential candidates is less than 2, additional zero motion vector candidates are added to the list.
[0064] Figure 5 The positions of non-adjacent candidates in the ECM are shown.
[0065] Merge mode Similar to the AMVP model, the MVP candidate list for the Merge model also includes spatial and temporal candidates. For spatial motion vector candidate derivation, after performing availability and redundancy checks, a maximum of four candidates are selected, in the order A1, B1, B0, A0, and B2. For temporal Merge candidate (TMVP) derivation, a candidate is selected from at most two temporally neighboring blocks (C0 and C1). When there are not enough Merge candidates using both spatial and temporal candidates, joint bidirectional prediction Merge candidates and zero MV candidates are added to the MVP candidate list. The Merge candidate list construction process terminates once the number of available Merge candidates reaches the maximum allowed number for signal transmission.
[0066] In VVC, the Merge pattern construction process is further improved by introducing a history-based MVP (HMVP), which incorporates motion information from previously encoded / decoded blocks that can be far removed from the current block. In VVC, HMVP merge candidates are appended to the Merge list, following the spatial MVP and TMVP. In this method, motion information from previously encoded / decoded blocks is stored in a table and used as the MVP for the current CU. During the encoding / decoding process, the table with multiple HMVP candidates is maintained using a first-in, first-out (FIFO) strategy. Whenever a non-sub-block inter-frame encoded / decoded CU is present, the associated motion information is added to the last entry of the table as a new HMVP candidate.
[0067] During the standardization of VVC, a non-adjacent MVP was proposed to facilitate better motion information derivation by utilizing non-adjacent regions. In ECM software, the non-adjacent MVP is inserted between the TMVP and HMVP, where the distance between the non-adjacent spatial candidate and the current codec block is based on the width and height of the current codec block, such as... Figure 5 As shown.
[0068] 2.2. Affine Motion Compensation Prediction In HEVC, only a translational motion model is applied for motion compensation prediction (MCP). In the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transformation motion compensation prediction is applied. Figure 6 As shown, the affine motion field of a block is described by motion information from two control points (4 parameters) or three control point motion vectors (6 parameters).
[0069] Figure 6 Affine motion models based on control points are shown, including (a) a 4-parameter affine model and (b) a 6-parameter affine model.
[0070] For the 4-parameter affine motion model, the motion vector at the sample point position (x, y) in the block is derived as: (1).
[0071] For the 6-parameter affine motion model, the motion vector at the sample point position (x, y) in the block is derived as: (2).
[0072] in( mv0x, mv0y ) is the motion vector of the upper left control point, ( mv1x, mv1y ) is the motion vector of the upper right control point, and ( mv2x, mv2y ) is the motion vector of the lower left control point.
[0073] To simplify motion compensation prediction, a block-based affine transformation prediction is applied. To derive the motion vector for each 4×4 lumen sub-block, the motion vector of the center sample point of each sub-block is calculated according to the above equation (e.g., ...). Figure 7 (as shown), and rounded to 1 / 16 fractional precision. Then, a motion-compensated interpolation filter is applied to generate a prediction for each sub-block with a derived motion vector. The sub-block size for the chroma component is also set to 4×4. The MV of the 4×4 chroma sub-block is calculated as the average of the MV of the upper-left luminance sub-block and the lower-right luminance sub-block in the corresponding 8×8 luminance region.
[0074] Figure 7 The affine MVF for each sub-block is shown.
[0075] Similar to translational motion inter-frame prediction, there are two affine motion inter-frame prediction modes: affine Merge mode and affine AMVP mode.
[0076] 2.2.1. Affine Merge Prediction The Affine Merge pattern can be applied to CUs with a width and height greater than or equal to 8. In this pattern, the CPVM of the current CU is generated based on the motion information of spatially neighboring CUs. There can be up to five CPVM candidates, and the one to be used for the current CU is indicated by a signal transmission index. In VVC, the following three types of CPVM candidates are used to form the Affine Merge candidate list: – Affine Merge candidates inferred from the CPMV of neighboring CUs.
[0077] – A constructive affine Merge candidate CPMVP derived using translational MV of neighboring CUs.
[0078] – Zero MV.
[0079] In VVC, there are at most two inherited affine candidates, which are derived from the affine motion model of neighboring blocks: one from the left neighboring CU and one from the upper neighboring CU. Candidate blocks are as follows: Figure 8 As shown. For the predicted value on the left, the scan order is A0->A1, and for the predicted value above, the scan order is B0->B1->B2. Only the first inherited candidate from each side is selected. No deduplication check is performed between candidates from two inheritances. When a neighboring affine CU is identified, its control point motion vector is used to derive the CPMVP candidate in the affine Merge list of the current CU. Figure 9 As shown, if the adjacent lower-left block A is encoded and decoded in affine mode, the motion vectors of the upper-left, upper-right, and lower-left corners of the CU containing block A are obtained. , and When block A is encoded and decoded using a 4-parameter affine model, according to and Calculate the two CPMVs of the current CU. When block A is encoded and decoded using a 6-parameter affine model, according to... , and Calculate the three CPMVs of the current CU.
[0080] Figure 8 The location of the inherited affine motion prediction value is shown.
[0081] Figure 9 The inheritance of control point motion vectors is shown.
[0082] The constructed affine candidate refers to the candidate built by combining the translational motion information of the neighbors of each control point. The motion information of the control points is derived from... Figure 10The spatial and temporal nearest neighbors shown are derived. CPMVk (k=1, 2, 3, 4) represents the k-th control point. For CPMV1, check the B2->B3->A2 block and use the MV of the first available block. For CPMV2, check the B1->B0 block, and for CPMV3, check the A1->A0 block. If available, the TMVP is used as CPMV4.
[0083] After obtaining the motion signatures (MVs) of the four control points, the affine merge candidate is constructed based on this motion information. The following combinations of control point MVs are used for sequential construction: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}.
[0084] Combining three CPMVs constructs a 6-parameter affine merge candidate, and combining two CPMVs constructs a 4-parameter affine merge candidate. To avoid motion scaling, combinations of control point MVs are discarded if the reference indices of the control points are different.
[0085] Figure 10 The locations of candidate positions for the constructed affine Merge pattern are shown.
[0086] After checking the inherited affine Merge candidates and the constructed affine Merge candidates, if the list is still not full, a zero MV is inserted at the end of the list.
[0087] 2.2.2. Affine AMVP Prediction The affine AMVP mode can be applied to CUs with a width and height greater than or equal to 16. An affine flag at the CU level is signaled in the bitstream to indicate whether the affine AMVP mode is used, and another flag is signaled to indicate whether it is a 4-parameter affine or a 6-parameter affine. In this mode, the difference between the current CU's CPVM and its predicted CPMVP is signaled in the bitstream. The affine AMVP candidate list is of size 2 and is generated by sequentially using the following four types of CPVM candidates: – Inherited affine AMVP candidates inferred from the CPMV of neighboring CUs.
[0088] – A constructive affine AMVP candidate CPMVP derived using the translational MV of neighboring CUs.
[0089] – Translation MV from the neighboring CU.
[0090] – Zero MV.
[0091] The checking order for inherited affine AMVP candidates is the same as that for inherited affine Merge candidates. The only difference is that, for AMVP candidates, only affine CUs with the same reference picture as those in the current block are considered. No deduplication process is applied when inserting inherited affine motion predictions into the candidate list.
[0092] The constructed AMVP candidate is derived from the specified spatial nearest neighbor as shown in the figure. The same checking order as in the affine Merge candidate construction is used. Additionally, the reference picture index of neighboring blocks is checked. The block that is the first in the checking order to be inter-coded and has the same reference picture as the current CU is used. Only one is used. This occurs when the current CU is encoded and decoded using a 4-parameter affine mode, and... mv0 and mv1 If all three CPMVs are available, they are added as candidates to the affine AMVP list. If the current CU is encoded / decoded in a 6-parameter affine mode and all three CPMVs are available, they are added as candidates to the affine AMVP list. Otherwise, the constructed AMVP candidates are set to unavailable.
[0093] Figure 11 Spatial nearest neighbors used to derive affine Merge candidates are shown: (a) for deriving inherited affine Merge candidates, and (b) for deriving constructed affine Merge candidates.
[0094] If the affine AMVP list still has fewer than 2 candidates after inserting valid inherited affine AMVP candidates and constructed AMVP candidates, then when available, mv0 , mv1 and mv2 They will be added sequentially as translation MVs to predict all control point MVs for the current CU. Finally, if the affine AMVP list is still not full, zero MVs are used to populate the affine AMVP list.
[0095] 2.2.3. A Novel Affine Candidate Derivation Method ECM-6.0 integrates three additional affine Merge and AMVP candidate derivation methods: non-adjacent spatial domain candidates, historical parameter-based candidates, regression-based affine candidates, and pixel-based affine motion compensation.
[0096] 2.2.3.1. Non-adjacent airspace candidates In ECM-6.0, the study of non-adjacent spatial neighbors provides candidates for both affine Merge and affine AMVP. The pattern for obtaining non-adjacent spatial candidate pairs is as follows: Figure 11As shown. Similar to non-adjacent regular merge candidates, the distance between non-adjacent spatial candidates and the current codec block is also defined based on the width and height of the current CU.
[0097] Figure 11 Motion information of non-adjacent spatial neighbors is used to generate additional inheritance and construct affine merge candidates. Specifically, to generate inheritance candidates, non-adjacent spatial neighbors are checked based on their distance from the current block (i.e., from nearest to farthest). At a specific distance, only the first available neighbors encoded in affine mode from each side (e.g., left and top) of the current block are included. Figure 11 As shown in (a), the checks of the left and top nearest neighbors are performed from bottom to top and from right to left, respectively. For the constructed candidates, as... Figure 11 As shown in (b), the positions of non-adjacent spatial neighbors on the left and top are first determined independently; then, the position of the top-left neighbor can be determined accordingly to form a rectangular virtual block together with the non-adjacent neighbors on the left and top. The motion information of the three non-adjacent neighbors is used to form a CPMV at the top-left (A), top-right (B), and bottom-left (C) of the virtual block, which is projected onto the current CU to generate corresponding construction candidates, such as... Figure 12 As shown. Figure 12 The affine Merge candidates from non-nearest neighbors to the constructed ones are shown.
[0098] 2.2.3.2. Affine Candidates Based on Historical Parameters History-based Affine Model Inheritance (HAMI) allows affine models to inherit from previously affine-encoded blocks that are not adjacent to the current block. A History Parameter Table (HPT) is established. Each entry in the HPT stores a set of affine parameters: a, b, c, and d, each represented by a 16-bit signed integer. Entries in the HPT are categorized by reference lists and reference indices. Each reference list in the HPT supports five reference indices. The HPT category (denoted as HPTCat) is calculated in a formulaic manner. HPTCat (RefList, RefIdx) = 5×RefList + min (RefIdx, 4)(3) Here, RefList and RefIdx represent the list of reference images (0 or 1) and the reference index, respectively. A maximum of seven entries can be stored for each category, resulting in a total of 70 entries in the HPT. At the beginning of each CTU line, the number of entries for each category is initialized to zero. After decoding the affine-encoded CU with reference lists RefListcur and RefIdxcur, the affine parameters are used to update the entries in the category HPTCat(RefListcur, RefIdxcur) in a manner similar to HMVP table updates.
[0099] Candidates based on historical affine parameters (HAPC) from... Figure 13 The derivation of a set of affine parameters, represented as A0, A1, B0, B1, or B2, and their corresponding entries stored in the HPT, is as follows. The MV of the neighboring 4×4 blocks is used as the base MV. The MV of the current block at position (x, y) is calculated in a formulaic manner as follows: (4) Where (mvhbase, mvvbase) represents the MV of the nearest 4×4 blocks, and (xbase, ybase) represents the center position of the nearest 4×4 blocks. (x, y) can be the top left, top right, and bottom left corners of the current block to obtain the corner position MV (CPMV) for the current block, or it can be the center of the current block to obtain the regular MV for the current block.
[0100] Figure 13 An example of how to derive the HAPC from block A0 is shown. The affine parameters {a0, b0, c0, d0} are directly obtained from an entry in the category HPTIdx(RefListA0, refIdx0A0) in the HPT. The affine parameters from the HPT (with the center position of A0 as the base position and the MV of block A0 as the base MV) are used together to derive the CPMV for either the affine MergeHAPC or the affine AMVP HAPC. They can also be used to derive the MV located at the center of the current block as a regular Merge candidate. The HAPC can be placed into the sub-block-based Merge candidate list, the affine AMVP candidate list, or the regular Merge candidate list. In response to the introduction of new HAPCs, the size of the sub-block-based Merge candidate list is increased from 5 to 10 and 12 for random access and low-latency B configurations, respectively. Furthermore, for the random access configuration, the size of the regular Merge candidate list is increased from 10 to 11 to accommodate the newly added regular Merge candidates.
[0101] Figure 13 An example of generating HAPC is shown.
[0102] 2.2.3.3. Regression-based Affine Candidates In ECM-6.0, regression-based affine merge candidates are derived and added to the affine merge list. The sub-block motion fields from previously encoded and decoded affine CUs and the motion information of neighboring sub-blocks from the current CU are used as inputs to the regression process to derive the proposed affine candidates.
[0103] Previously encoded and decoded affine CUs can be identified by scanning non-adjacent positions and the affine HMVP table. For example... Figure 14As shown, the information about the adjacent sub-blocks of the current CU is obtained from the 4x4 sub-blocks represented by the gray area. For each sub-block, given a reference list, the corresponding motion vector and center coordinates of the sub-block can be used.
[0104] For each affine CU, at most two affine candidates can be derived: one with neighboring subblock information and one without. All candidates generated by linear regression are deduplicated and collected into a candidate subgroup. When ARMC is enabled, an ARMC process based on TM cost is applied. Subsequently, when N affine CUs are found, at most N candidates generated by linear regression are added to the affine merge list.
[0105] Figure 14 A diagram illustrating the regression-based affine Merge candidate derivation is shown.
[0106] 2.2.3.4. Pixel-based Affine Motion Compensation Using pixel-based affine motion compensation, when OBMC is not applied, the minimum affine sub-block size is set to 1x1 for the luma component and is always set to 1x1 for the chroma component.
[0107] 2.3. Template Matching Merge / AMVP Pattern in ECM Template Matching (TM) Merge / AMVP mode is a decoder-side MV derivation method used to refine the motion information of the current CU by finding the closest match between a template in the current image (i.e., the block above and / or to the left of the current CU) and a block in the reference image (i.e., the block of the same size as the template). Figure 15 As shown, within the search range of [-8, +8] pixels, a better MV is searched around the initial motion of the current CU.
[0108] Figure 15 This demonstrates template matching execution over the search area surrounding the initial MV.
[0109] In AMVP mode, an MVP candidate is determined based on the template matching error, selecting the one that minimizes the difference between the current block and the reference block template. Then, the TM process performs MV refinement only on that specific MVP candidate. The TM refines the MVP candidate using an iterative diamond search, starting with full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode) within a search range of [-8, +8] pixels. The AMVP candidate can be further refined using a cross search with full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode), followed by half-pixels and quarter-pixels sequentially depending on the AMVR mode. This search process ensures that the MVP candidate maintains the same MV precision as indicated by the Adaptive Motion Vector Resolution (AMVR) mode after the TM process.
[0110] In Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. TMMerge can proceed up to 1 / 8 pixel MVD accuracy, or skip those accuracies beyond half-pixel MVD accuracy, depending on whether an alternative interpolation filter is used based on the merged motion information (i.e., used when AMVR is in half-pixel mode). Furthermore, when TM mode is enabled, template matching can work as a standalone process or as an additional MV refinement process between block-based and sub-block-based bilateral matching (BM) methods, depending on whether BM is enabled according to its enable condition check. When both BM and TM are enabled for the CU, the TM search process stops at half-pixel MVD accuracy, and the resulting MV is further refined using the same model-based MVD derivation method as in DMVR.
[0111] 2.4. Adaptive Reordering of Merge Candidates (ARMC) Inspired by the spatial correlation between reconstructed neighboring pixels and the current codec block, an Adaptive Reordering of Merge Candidates (ARMC) is proposed to refine the order of candidates in a given candidate list. The basic assumption is that candidates with lower template matching costs have a higher probability of being selected through the RDO process and should therefore be placed earlier in the list to reduce signaling costs.
[0112] The reordering method is applied to the regular Merge pattern, the Template Matching (TM) Merge pattern, and the Affine Merge pattern (excluding SbTMVP candidates). For the TM Merge pattern, the Merge candidates are reordered before the refinement process.
[0113] After constructing the Merge candidate list, the Merge candidates are divided into several subgroups. The subgroup size is set to 5. The Merge candidates in each subgroup are reordered in ascending order based on the cost value of template matching. For simplicity, the Merge candidates in the last subgroup (not the first subgroup) are not reordered.
[0114] Template matching cost is measured by the sum of absolute differences (SAD) between the samples of the current block's template and the samples of its corresponding reference template. For example... Figure 16 As shown, the template includes a set of reconstructed samples adjacent to the current block, while the reference template is located using the same motion information of the current block. When the Merge candidate utilizes bidirectional prediction, the reference samples of the Merge candidate's template are also generated through bidirectional prediction. Figure 16 The template and the corresponding reference template are shown.
[0115] For a sub-block-based merge candidate with a sub-block size equal to Wsub*Hsub, the upper template includes several sub-templates of size Wsub×K, and the left template includes several sub-templates of size K×Hsub. For example... Figure 17 As shown, the motion information of the sub-blocks in the first row and first column of the current block is used to derive the reference sample points of each sub-template.
[0116] 2.5. Sub-block-based temporal motion vector prediction (SbTMVP) VVC supports a sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to TMVP, SbTMVP leverages motion fields in co-located images to facilitate more accurate MVP derivation. The same co-located image used by TMVP is used in SbTMVP. SbTMVP differs from TMVP primarily in two ways. First, SbTMVP enables motion prediction at the sub-CU level, while TMVP predicts motion at the CU level. Second, compared to TMVP, which obtains temporal MVs from co-located blocks in a co-located image (where a co-located block is the lower right or center block relative to the current CU), SbTMVP applies motion displacements before obtaining temporal motion information from the co-located image. These motion displacements are obtained by reusing the MV of one of the spatially neighboring blocks from the current CU.
[0117] Figure 18 The derivation process of the sub-block level motion field for SbTMVP is shown. Specifically, the motion information of the lower left sub-block A1 is first obtained. If any MV in reference list 0 and list 1 points to the same frame, the corresponding MV will be identified as the motion displacement. Otherwise, zero MV will be used as the motion displacement.
[0118] Once the motion displacement is determined, a designated region within the same frame is used to derive the sub-block-level motion field. Assuming... Figure 18As shown, the motion of A1 is used as the motion displacement. Then, for each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the center sample point) in the co-location image is obtained to provide motion information, wherein the MV scaling operation is first performed to align the reference frame of the temporal motion vector with the reference frame of the current CU.
[0119] Figure 17 Templates and reference templates are shown for blocks with sub-block motion that use motion information of the current block's sub-blocks.
[0120] Figure 18 The derivation of the sub-CU motion field obtained by applying motion displacement based on neighbor motion information is shown.
[0121] In VVC and ECM, in addition to the CU-level MVP candidate list, a sub-CU-level MVP candidate list is constructed to provide more accurate motion predictions for the current CU. This list includes the motion field generated by both the SbTMVP and AFFINE methods. Specifically, only one SbTMVP candidate is included, and this SbTMVP candidate is always placed as the first entry in the constructed sub-CU-level MVP candidate list. Multiple AFFINE candidates are included in the list after performing a template matching-based reordering, with those AFFINE candidates having lower costs placed earlier.
[0122] 2.6. Intra-Frame and Inter-Frame Joint Prediction (CIIP) In VVC, when a CU is encoded and decoded in Merge mode, if the CU contains at least 64 luma samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64), and if both the CU width and height are less than 128 luma samples, an additional flag is transmitted via signaling to indicate whether Intra-Inter Joint Prediction (CIIP) mode is applied to the current CU. As the name suggests, CIIP prediction combines inter-frame prediction signals with intra-frame prediction signals. The inter-frame prediction signal in CIIP mode... The inter-frame prediction process is derived using the same procedure as the regular Merge mode; and the intra-frame prediction signal... The conventional intra-frame prediction process with a planar pattern is derived. Then, a weighted average is used to combine the intra-frame and inter-frame prediction signals, where the predictions are based on the upper and left neighboring blocks (in...). Figure 19 The weight values for the encoding / decoding modes (described in the image) are calculated as follows: - If the upper nearest neighbor is available and is intra-coded, set isIntraTop to 1; otherwise, set isIntraTop to 0. – If the left nearest neighbor is available and is intra-coded, set isIntraLeft to 1; otherwise, set isIntraLeft to 0. – If (isIntraLeft + isIntraTop) equals 2, then wt is set to 3; Otherwise, if (isIntraLeft + isIntraTop) equals 1, then wt is set to 2; Otherwise, set wt to 1.
[0123] The CIIP predictions are formed as follows:
[0124] Figure 19 The top and left neighbor blocks used in the CIIP weight derivation are shown.
[0125] 2.7. CIIP with PDPC hybrid In ECM, the CIIP model is extended. In this extended model (CIIP_PDPC), the predictions of the regular Merge model are refined using samples reconstructed from the top (Rx, -1) and left (R-1, y). This refinement inherits the Location-Related Prediction Combination (PDPC) scheme. The flowchart of the CIIP_PDPC model prediction can be seen as follows... Figure 20 As shown, WT and WL are weighted values that depend on the location of the sample points in the block, as defined by PDPC.
[0126] The CIIP_PDPC mode is transmitted via signaling along with the CIIP mode. When the CIIP flag is true, another flag (i.e., the CIIP_PDPC flag) is further transmitted via signaling to indicate whether CIIP_PDPC is used.
[0127] Figure 20 A flowchart of the CIIP_PDPC process using the extended CIIP mode with PDPC is shown.
[0128] 2.8. Combination of CIIP with TIMD and TM Merge In ECM CIIP mode, prediction samples can be generated by weighting the inter-frame prediction signal using CIIP-TM Merge candidate prediction and the intra-frame prediction signal using the intra-frame prediction mode derived using TIMD. This method is only applied to codec blocks with an area less than or equal to 1024.
[0129] The TIMD derivation method is used to derive intra-prediction modes in CIIP. Specifically, the intra-prediction mode with the smallest SATD value in the TIMD mode list is selected and mapped to one of 67 regular intra-prediction modes.
[0130] Furthermore, it is proposed that if the derived intra-prediction mode is an angle mode, the weights (wIntra, wInter) for the two tests should be modified. For near-horizontal modes (2 <= angle mode index < 34), such as... Figure 21 As shown in (a), the current block is divided vertically; for near-vertical mode (34 <= angle mode index <= 66), as Figure 21 As shown in (b), the current block is divided horizontally.
[0131] The different sub-blocks (wIntra, wInter) are shown in Table 1.
[0132] Figure 21 The method for dividing angle patterns is shown.
[0133] Table 1. Weights used for modifying angle modes
[0134] Using CIIP-TM, a CIIP-TM Merge candidate list is constructed for the CIIP-TM pattern. Merge candidates are refined through template matching. CIIP-TM Merge candidates are also reordered using the ARMC method as regular Merge candidates. The maximum number of CIIP-TM Merge candidates is 2.
[0135] 2.9. Combination of intra-block copying and intra-prediction Intra-Block Copy and Intra-Prediction Combination (IBC-CIIP) is an encoding / decoding tool for CUs that uses IBC and intra-prediction to obtain two prediction signals, which are then weighted and summed to generate the following final prediction:
[0136] in and This represents the IBC prediction signal and the intra-frame prediction signal. For IBC Merge mode and IBCAMVP mode, It is set to equal to (13, 4) and (1, 1).
[0137] An intra-prediction mode (IPM) candidate list is used to generate the intra-prediction signal, and the IPM candidate list size is predefined to 2. An IPM index is transmitted via signaling to indicate which IPM to use.
[0138] Figure 22 This demonstrates the generation of sub-block templates for SbTMVP.
[0139] 2.10. Derivation of Time-Domain Motion in ECM In VVC, temporal motion vector prediction (TMVP) for AMVP and Merge modes is derived by acquiring motion information from the center or lower right of the co-occurrence block in the co-occurrence image transmitted via signal transmission. Similarly, for sub-block-based temporal motion vector prediction (SbTMVP) mode, motion information from the left neighboring location is used as motion displacement, which is then used to obtain the TMVP at the sub-CU level.
[0140] In ECM, two aspects were modified to further improve the encoding and decoding efficiency of TMVP. First, two co-located images are used, which are two reference frames with the minimum POC distance relative to the frame to be encoded / decoded. Second, the motion displacement for locating the TMVP is adaptively determined from multiple positions based on template cost. More specifically, two motion displacement candidate lists are constructed for each of the two co-located frames. The motion displacement with the minimum template matching cost is used to derive either SbTMVP or TMVP candidates. The sub-block-based merge list includes a maximum of four SbTMVP candidates. The SbTMVP candidate with the minimum template matching cost derived from the first co-located frame is placed as the first entry without reordering, while other SbTMVP candidates are ordered together with affine candidates. Additionally, the prediction direction of the template for each sub-block is determined based on the central sub-block. Figure 22 As shown, if the central sub-block is unidirectionally predicted, then all sub-block templates are unidirectionally predicted, and vice versa. If the motion vector of the corresponding adjacent sub-block at a defined reference list is unavailable for a sub-block template, then zero MV is used for that sub-block template.
[0141] 2.11. Multi-pass decoder-side motion vector refinement (DMVR) Multi-pass decoder-side motion vector refinement is integrated into the ECM. In the first pass, bilateral matching (BM) is applied to the codec block. In the second pass, BM is applied to each 16x16 sub-block within the codec block. In the third pass, the motion vector (MV) in each 8x8 sub-block is refined by applying bidirectional optical flow (BDOF). The refined MV is stored for both spatial and temporal motion vector prediction.
[0142] 2.11.1 First pass – Block-based bilateral matching MV refinement In the first pass, the refined MV is derived by applying BM to the codec block. Similar to decoder-side motion vector refinement (DMVR), in the bidirectional prediction operation, a refined MV is searched around the two initial MVs (MV0 and MV1) in the reference picture lists L0 and L1. Based on the minimum bilateral matching cost between the two reference blocks in L0 and L1, the refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MVs.
[0143] BM performs a local search to derive the integer sample precision intDeltaMV. The local search applies a 3x3 square search pattern to iterate through the search range [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block dimension, and the maximum value of sHor and sVer is 8.
[0144] The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW * cbH is greater than 64, the mean-removed SAD (MRSAD) cost function is applied to remove the DC effect of distortion between reference blocks. The local search intDeltaMV terminates when bilCost at the center point of the 3x3 search pattern has the minimum cost. Otherwise, the current minimum cost search point becomes the new center point of the 3x3 search pattern, and the search for the minimum cost continues until the end of the search range is reached.
[0145] The existing fractional sample refinement is further applied to derive the final deltaMV. The refined MV after the first pass is then derived as: · MV0_pass1 = MV0 + deltaMV, · MV1_pass1 = MV1 – deltaMV.
[0146] 2.11.2 Second pass – Sub-block-based bilateral matching MV refinement In the second pass, the refined MV is derived by applying BM to 16x16 grid sub-blocks. For each sub-block, a refined MV is searched around the two MVs (MV0_pass1 and MV1_pass1) obtained in the first pass in the reference image lists L0 and L1. The refined MV (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) is derived based on the minimum bilateral matching cost between the two reference sub-blocks in L0 and L1.
[0147] For each subblock, BM performs a full search to derive the integer sample precision intDeltaMV. The full search has a search range of [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block dimension, and the maximum value of sHor and sVer is 8.
[0148] The bilateral matching cost is calculated by applying a cost factor to the SATD cost between the two reference subblocks, as follows: bilCost = satdCost * costFactor. The search region (2*sHor + 1) * (2*sVer + 1) is divided into Figure 23 A maximum of five diamond-shaped search regions are shown. Each search region is assigned a costFactor, determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond region is processed sequentially starting from the center of the search region. Within each region, search points are processed in raster scan order, starting from the top left corner and ending at the bottom right corner. A full pixel search terminates when the minimum bilCost within the current search region is less than a threshold (equal to sbW * sbH); otherwise, the full pixel search continues to the next search region until all search points have been checked. Additionally, the search process terminates if the difference between the previous minimum cost and the current minimum cost in an iteration is less than or equal to a threshold representing the area of the block.
[0149] Figure 23 The diamond-shaped area in the search area is shown.
[0150] The existing VVC DMVR fractional sample refinement is further applied to derive the final deltaMV(sbIdx2). The refined MV at the second pass is then derived as: · MV0_pass2(sbIdx2) = MV0_pass1 + deltaMV(sbIdx2), · MV1_pass2(sbIdx2) = MV1_pass1 – deltaMV(sbIdx2).
[0151] 2.11.3 Third pass – Sub-block based bidirectional optical flow MV refinement In the third pass, the refined MV is derived by applying BDOF to 8x8 grid sub-blocks. For each 8x8 sub-block, BDOF refinement is applied to derive scaled Vx and Vy without clipping, starting from the refined MV of the parent block from the second pass. The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample precision and clipped between -32 and 32.
[0152] The refined MVs (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) at the third pass are derived as follows: · MV0_pass3(sbIdx3) = MV0_pass2(sbIdx2) + bioMv, · MV1_pass3(sbIdx3) = MV0_pass2(sbIdx2) – bioMv.
[0153] Of all the aforementioned sub-items, when surround motion compensation is enabled, surround offset must be considered to limit the motion vector.
[0154] 3. Problem Existing CIIP methods have the following problems: 1) In VVC and ECM, CIIP only allows regular Merge candidates as inter-frame components, while sub-block-based motion candidates (i.e., AFFINE and SbTMVP) can bring additional benefits to blocks with complex motion.
[0155] 2) How to mix sub-block-based motion prediction with intra-frame prediction, and how sub-block-based CIIP interacts with other codec tools, can be further specified.
[0156] 4. Detailed Solution In this disclosure, a further improvement to CIIP is proposed by allowing sub-block-based predictions as inter-frame components. Specifically, intra-frame predictions and sub-block-based inter-frame predictions can be mixed to form CIIP predictions.
[0157] The detailed embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.
[0158] The terms "video unit," "code-decoder unit," or "block" can refer to code-decoder tree block (CTB), code-decoder tree unit (CTU), code-decoder block (CB), CU, PU, TU, PB, and TB. The term "sub-block-based codec tool" can refer to affine codecs, SbTMVP, and their corresponding variants.
[0159] In this disclosure, regarding "blocks encoded and decoded in mode N", "mode N" can be a prediction mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or an encoding / decoding technique (e.g., DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-inter, GPM intra-intra, GPM inter-intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, ALF, deblocking, SAO, bilateral filter, LMCS and its corresponding variants, etc.). In the following discussion, SE can be binarized into a fixed-length code, EG(x) code, unary code, rounded unary code, rounded binary code, etc. It can be signed or unsigned.
[0160] It should be noted that the following terms are not limited to the specific terms defined in existing standards. Any variants of encoding / decoding tools also apply.
[0161] 1. The prediction generated by the first codec tool can be mixed with the prediction of the second codec tool to form a third prediction.
[0162] a) In one example, the first codec tool could be a sub-block-based approach, such as affine, SbTMVP, etc.
[0163] b) In one example, the second codec tool can be a regular intra-frame (plane / DC / angle) / TIMD / DIMD / ISP / PDPC / MIP / IBC / regular inter-frame, etc.
[0164] 2. A method for generating CIIP inter-frame components using a sub-block-based encoding / decoding tool is proposed, namely, sub-block-based CIIP.
[0165] a) In one example, a first prediction can be generated using affine motion or SbTMVP, and a second prediction can be generated using intra-frame mode. The two predictions are then mixed with a weighted average to produce a final prediction for sub-block-based CIIP.
[0166] b) In one example, the intra-frame mode can be plane / DC / angle / MIP / ISP / IBC / intra-frame TMP, etc.
[0167] c) In one example, the intra-frame mode can be derived based on TIMD / DIMD / intra-frame TMP, etc.
[0168] d) In one example, intra-frame components of sub-block-based CIIP can be processed by PDPC.
[0169] 3. A first sub-block-based motion list can be constructed to provide motion information for sub-block-based CIIP.
[0170] a) In one example, the first sub-block-based motion list may include at least one affine and / or at least one (or more) SbTMVP candidates.
[0171] i. In one example, specifically, the first sub-block-based motion list may include at least one adjacent / non-adjacent / history-based / regression-based / zero affine candidate.
[0172] ii. In one example, specifically, the first sub-block-based motion list may include at least one SbTMVP candidate.
[0173] 1) In one example, multiple SbTMVP candidates can appear in the list, which can be collected from at least one co-position frame.
[0174] b) In one example, alternatively, the first sub-block-based motion list may include only affine or SbTMVP candidates.
[0175] c) In one example, the number of candidates in the list may not exceed a constant or an adaptively determined number.
[0176] d) In one example, the first sub-block-based motion list can be the same as the second sub-block-based motion list used in the sub-block-based Merge / AMVP mode.
[0177] i. Alternatively, the first sub-block-based motion list and the second sub-block-based motion list may be different.
[0178] 1) In one example, the first sub-block-based motion list and the second sub-block-based motion list can be constructed using different maximum allowed number of candidates.
[0179] 2) In one example, the first sub-block-based motion list and the second sub-block-based motion list can be constructed using different candidate types.
[0180] 3) In one example, the first sub-block-based motion list can be constructed by selecting at least one candidate from the second sub-block-based motion list.
[0181] 4. The list of motions based on sub-blocks can be reordered based on specific metrics after construction.
[0182] a) In one example, template matching or bilateral matching costs can be used to reorder a list.
[0183] i. In one example, specifically, if a reconstructed template region for the current block exists, the template matching cost can be used to reorder the list.
[0184] 1) In one example, alternatively, if the reconstruction template region for the current block does not exist, the list may not be reordered.
[0185] 5. Indicates that the index of a specific candidate in the motion list based on a sub-block can be transmitted via signal in the bitstream.
[0186] a) In one example, the candidate specified by the index can be used to provide sub-block-based motion information and / or motion prediction.
[0187] i. In one example, if the candidate specified by the index is an affine candidate, it can first be refined by TM or DMVR and then used to generate inter-frame predictions.
[0188] 1) In one example, if the specified candidate is an affine candidate for one-way prediction, then TM can be used to refine the candidate, such as TM-based CPMV refinement, and the refined candidate is used to provide predictions.
[0189] 2) In one example, alternatively, if the specified candidate is an affine candidate for bidirectional prediction, DMVR can be used to refine the candidate, and the refined candidate is used to provide predictions.
[0190] ii. In one example, alternatively, no TM or DMVR processing is used to refine the specified affine candidates.
[0191] b) In one example, no candidate index is transmitted via signal in the bitstream.
[0192] i. In one example, the encoder and decoder can generate the same sub-block-based motion candidates based on predefined rules, which are used to provide inter-frame prediction.
[0193] 1) In one example, specifically, after the list of motions based on sub-blocks is constructed, it can be reordered based on a specific metric, and candidates in fixed positions (e.g., first / last / middle or any other position) can be used by default.
[0194] 6. At least one syntax (or flag) indicating the use of sub-block-based CIIP can be transmitted via signaling in the bitstream.
[0195] a) In one example, at least one block-level CIIP flag based on sub-blocks can be transmitted via signaling in the bitstream.
[0196] b) In one example, at least one sub-block-based CIIP flag at the sequence level / picture group level / picture level / strip level / piece group level, such as a sub-block-based CIIP flag in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header, can be transmitted via signaling in the bitstream.
[0197] c) In one example, whether the sub-block-based CIIP flag is signaled or whether the sub-block-based CIIP is applied may depend on the value of at least one other syntax.
[0198] i. In one example, whether the sub-block-based CIIP flag is transmitted via signaling, or whether the sub-block-based CIIP is applied, may depend on the value of at least one other sequence-level / picture-group-level / picture-level / strip-level / piece-group-level syntax.
[0199] 1) In one example, whether the CIIP flag based on the sub-block is transmitted via signaling may depend on the value of the syntax indicating whether CIIP and / or affine and / or SbTMVP is enabled.
[0200] ii. In one example, whether the sub-block-based CIIP flag is signaled or whether the sub-block-based CIIP is applied may depend on the value of at least one other block-level syntax.
[0201] 1) In one example, specifically, the sub-block-based CIIP flag can only be transmitted via signaling if the CIIP / CIIP-PDPC / CIIP-TM / Affine / SbTMVP flag is either true or false.
[0202] 2) In one example, a first flag indicating the use of CIIP (i.e., the CIIP flag) can be transmitted via signaling. If this flag is true, a second flag indicating the use of sub-block-based CIIP (i.e., the sub-block-based CIIP flag) can be transmitted via signaling. If the sub-block-based CIIP flag is true, the index of the sub-block-based motion candidate can be indicated via signaling. Otherwise, if the sub-block-based CIIP flag is true, a third flag (i.e., CIIP-TM) can be transmitted via signaling.
[0203] a) In one example, alternatively, a first flag indicating the use of CIIP (i.e., the CIIP flag) can be signaled, and if this flag is true, a second flag indicating the use of CIIP-TM (i.e., the CIIP-TM flag) can be signaled. If the CIIP-TM flag is false, a third flag indicating the use of sub-block-based CIIP (i.e., the sub-block-based CIIP flag) can be signaled. If the sub-block-based CIIP flag is true, the index of the sub-block-based motion candidate can be signaled.
[0204] iii. In one example, whether the sub-block-based CIIP flag is transmitted via signaling, or whether the sub-block-based CIIP is applied, may depend on whether the value of at least one of the following syntaxes is true or false: sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0205] d) In one example, whether another flag is transmitted via signaling may depend on the value of the CIIP flag based on the sub-block.
[0206] i. In one example, whether a flag indicating the use of CIIP-TM / CIIP-PDPC and / or any other codec tool is transmitted via signaling may depend on the value of the sub-block-based CIIP flag.
[0207] 1) In one example, when the CIIP flag based on the sub-block is true, the CIIP-TM flag is not transmitted via signaling.
[0208] a) In one example, alternatively, if the CIIP flag based on the sub-block is false, the CIIP-TM flag is signaled.
[0209] b) In one example, alternatively, the CIIP-TM flag is signaled regardless of whether the CIIP flag based on the sub-block is true or false.
[0210] 2) In one example, when the CIIP flag based on the sub-block is true, the CIIP-PDPC flag is not transmitted via signaling, and / or the regular intra-frame or PDPC intra-frame is used by default.
[0211] a) In one example, alternatively, if the CIIP flag based on the sub-block is false, the CIIP-PDPC flag is signaled.
[0212] b) In one example, alternatively, the CIIP-PDPC flag is signaled regardless of whether the CIIP flag based on the sub-block is true or false.
[0213] e) In one example, whether a sub-block-based CIIP flag is transmitted via signaling may depend on the usage conditions of at least one other codec tool.
[0214] i. In one example, a sub-block-based CIIP flag can be signaled or applied only if CIIP and / or affine and / or SbTMVP are appropriate for the current block.
[0215] ii. In one example, a sub-block-based CIIP flag is transmitted or applied only when the block dimension (e.g., width / height / width-to-height ratio / block area) meets a specific condition.
[0216] 7. On the hybridization of sub-block-based CIIP modes.
[0217] a) In one example, the same hybrid approach used by regular CIIP / CIIP-TM / CIIP-PDPC can be used by sub-block-based CIIP.
[0218] i. In one example, alternatively, different mixing methods can be used by sub-block-based CIIP.
[0219] b) In one example, the mixed weights can be position-dependent within a block.
[0220] i. In one example, specifically, different locations within a block can have different blending weights.
[0221] ii. In one example, constant blending weights can be used at all locations within a block.
[0222] c) In one example, the blending weight matrix can be intra-mode dependent.
[0223] i. In one example, specifically, different hybrid weight matrices can be used depending on whether the intra-frame angle pattern is near vertical or near horizontal.
[0224] 8. For example, how to apply sub-block-based predictions can be conditional depending on whether they are used in sub-block-based CIIP.
[0225] a) For example, the size of a sub-block can be conditionalized.
[0226] b) For example, whether and / or how an operation is applied can be conditionalized.
[0227] i. The operation can be PROF.
[0228] ii. The operation can be an interlaced affine.
[0229] iii. The operation can be OBMC.
[0230] 9. It is proposed that at least one affine information may be stored in a first block that is not a block encoded or decoded by affine Merge or affine AMVP.
[0231] a) In one example, the first block can be encoded and decoded using GPM mode.
[0232] i. In one example, affine motion compensation can be applied to generate a GPM prediction for the first block.
[0233] b) In one example, the first block can be encoded or decoded using CIIP or a sub-block-based CIIP mode.
[0234] i. In one example, affine motion compensation can be applied to generate CIIP predictions for the first block.
[0235] c) In one example, the first block can be encoded and decoded using MHP mode.
[0236] i. In one example, affine motion compensation can be applied to generate an MHP prediction for the first block.
[0237] d) In one example, the stored affine model can be utilized by at least one block that is encoded / decoded after the first block.
[0238] e) In one example, affine information may include: i. Inter-frame direction (one-way from list 0, one-way or two-way from list 1); ii. CPMV (such as CPMV in the top left, top right, and bottom left corners); iii. Affine parameters (such as those in equations (1) and (2)) a, b, c, d, e, f ).
[0239] iv. Reference index for list 0; v. Reference index for List 1; vi. Local Lighting Compensation (LIC) sign; vii. Overlapping Block Motion Compensation (OBMC) flag; viii. Utilize codec unit-level weighted bidirectional prediction (BCW) indexing; ix. Affine type (4-parameter or 6-parameter affine); x. Merge type (affine or sbTMVP); xi. Image index.
[0240] 10. Affine information used in the first block encoded / decoded using sub-block-based CIIP can be used by blocks encoded / decoded after the first block.
[0241] a) In one example, the affine information of a block encoded and decoded by sub-block-based CIIP can be placed into the History Parameter Table (HPT).
[0242] i. In one example, the HPT can be updated after encoding / decoding a block that has been encoded / decoded using sub-block-based CIIP.
[0243] b) In one example, the affine information of a block encoded / decoded using sub-block-based CIIP can be obtained and placed into the affine Merge / AMVP candidate list of blocks encoded / decoded after the first block.
[0244] c) In one example, the affine information of a block encoded and decoded by sub-block-based CIIP can be obtained and placed into the affine candidate list of a second block encoded and decoded in sub-block-based CIIP mode or affine-GPM mode.
[0245] 11. In one example, when attempting to add a candidate based on a sub-block CIIP inheritance to a new candidate list, it can be compared with at least one candidate already in the new candidate list.
[0246] a) For example, if a candidate based on CIIP inheritance of a sub-block is the same as at least one candidate already in the new candidate list, it may not be placed in the new candidate list.
[0247] b) For example, if a candidate based on CIIP inheritance of a sub-block is similar to at least one candidate already in the new candidate list, it may not be placed in the new candidate list.
[0248] 12. When constructing an affine list that includes at least one affine candidate, such as a sub-block merge list, an affine merge list, an affine AMVP list, an affine list for GPM, or an affine list for sub-block-based CIIP, specific sub-blocks or blocks covering specific locations adjacent to or not adjacent to the current block (such as in...) can be checked in a specific order and in more than one mode. Figure 8 (A0, A1, B0, B1, and B2 in the original text).
[0249] a) For example, first check whether a specific sub-block or block is affine-coded (such as affine Merge or affine AMVP).
[0250] i. For example, if a particular sub-block or block is affine-coded, the affine information stored in or associated with that sub-block or block can be used to generate affine candidates for the affine list.
[0251] b) For example, it can be checked whether a specific sub-block or block is encoded or decoded using GPM-affine or sub-block-based CIIP.
[0252] i. For example, if a particular sub-block or block is encoded using GPM-affine or sub-block-based CIIP, the affine information stored in or associated with that particular sub-block or block can be used to generate affine candidates for the affine list.
[0253] 13. In one example, when constructing an affine list that includes at least one affine candidate, such as a sub-block merge list, an affine merge list, an affine AMVP list, an affine list for GPM, or an affine list for sub-block-based CIIP, a first sub-block or block covering a specific location adjacent or non-adjacent to the current block can be checked at a specific step in the list construction (such as in...). Figure 8 (A0, A1, B0, B1, and B2 in the code) to generate candidates from blocks encoded and decoded using sub-block-based CIIP.
[0254] a) For example, after checking affine candidates from a specific inheritance of neighboring blocks, the first sub-block or block can be checked immediately to generate candidates from blocks encoded and decoded using sub-block-based CIIP.
[0255] b) For example, after checking affine candidates from a specific inheritance of a non-adjacent block, the first sub-block or block can be checked immediately to generate candidates from a block encoded and decoded using sub-block-based CIIP.
[0256] c) For example, after checking a specific RMVF affine candidate, the first sub-block or block can be checked immediately to generate candidates from the block encoded and decoded by sub-block-based CIIP.
[0257] d) For example, after examining the affine candidates of a particular build, the first sub-block or block can be examined immediately to generate candidates from the blocks encoded and decoded by sub-block-based CIIP.
[0258] e) For example, after checking the affine candidates of temporal inheritance, the first sub-block or block can be checked immediately to generate candidates from the blocks encoded and decoded by sub-block-based CIIP.
[0259] f) For example, after checking history-based affine candidates, the first sub-block or block can be checked immediately to generate candidates from the block encoded and decoded by sub-block-based CIIP.
[0260] 14. A first SE related to CIIP was proposed that can be encoded and decoded using more than one context.
[0261] a) In one example, the first SE can be a CIIP Merge / MMVD index.
[0262] b) In one example, the first SE may indicate the use of CIIP-PDPC / CIIP-TM / subblock-based CIIP / CIIP-MMVD.
[0263] c) In one example, SE can be binarized into N binary bits, denoted as B0, B1, ..., B N-1 In one example, if n != m, then B n and B m It can be encoded and decoded using different contexts.
[0264] i. In one example, B0, B1, ..., B W They can be represented as C0, C1, ..., C W Different contexts are encoded and decoded, while B W B W+1 B N-1 Shared can be represented as C W+1 The same context. W is an integer, such as 1, 2, 3, 4, 5, 6, or 7.
[0265] ii. In one example, B0, B1, ..., B W They can be represented as C0, C1, ..., C W Different contexts are encoded and decoded, while B W B W+1 B N-1 It can be bypassed for encoding and decoding. W is an integer, such as 1, 2, 3, 4, 5, 6, or 7.
[0266] 15. A method is proposed to select at least one context for encoding and decoding a first SE associated with CIIP based on encoding and decoding information.
[0267] a) In one example, the first SE can be a CIIP Merge / MMVD index.
[0268] b) In one example, the first SE may indicate the use of CIIP-PDPC / CIIP-TM / subblock-based CIIP / CIIP-MMVD.
[0269] c) In one example, context selection may depend on whether sub-block-based CIIP is used.
[0270] i. For example, when sub-block-based CIIP is used, the (multiple) contexts of the CIIP Merge / MMVD index used to encode and decode CIIP blocks can be different.
[0271] ii. For example, when sub-block-based CIIP is used, the context(s) used to indicate the use of CIIP-PDPC / CIIP-TM / sub-block-based CIIP / CIIP-MMVD for CIIP blocks can be different.
[0272] d) In one example, the context selection may depend on whether CIIP-PDPC / CIIP-TM / sub-block based CIIP / CIIP-MMVD is used.
[0273] e) In one example, the context selection may depend on QP.
[0274] f) In one example, the context selection can depend on the stripe type.
[0275] 16. It is proposed that at least one context for encoding and decoding a first SE associated with CIIP can be selected based on encoding and decoding information of at least one neighboring block.
[0276] a) In one example, the first SE could be the CIIP Merge index.
[0277] b) In one example, the first SE can be a CIIP-MMVD index.
[0278] c) In one example, the first SE may indicate the use of CIIP-PDPC / CIIP-TM / subblock-based CIIP / CIIP-MMVD.
[0279] d) In one example, the context used to encode and decode the SE can be selected from a candidate set that includes more than one candidate context.
[0280] i. In one example, the candidate set may include three members represented as {C[0], C[1], C[2]}, and the selected context is C[S].
[0281] 1) S may depend on the encoding / decoding mode of at least one neighboring block.
[0282] 2) For example, S is initialized to 0. If the left neighboring block is available and a certain condition is met, S is incremented by 1; if the upper neighboring block is available and a certain condition is met, S is incremented by 1.
[0283] 3) In one example, SE can be a CIIP flag based on a sub-block.
[0284] 4) In one example, a specific condition could be whether neighboring blocks are affine encoded or decoded.
[0285] a) For example, if a block is encoded or decoded using an affine-AMVP mode, it can be considered affine encoded or decoded.
[0286] b) For example, if a block is encoded or decoded using affine-Merge mode, it can be considered affine encoded or decoded.
[0287] c) In one example, if the CIIP of a block based on its sub-blocks is true, then it can be decoded as an affine codec.
[0288] 17. In one example, the context of the affine flag used to encode / decode a block may depend on whether neighboring blocks are being encoded / decoded using a sub-block-based CIIP mode.
[0289] a) In one example, the context used to encode and decode affine flags can be selected from a candidate set that includes more than one candidate context.
[0290] b) In one example, the candidate set may include three members represented as {C[0], C[1], C[2]}, and the selected context is C[S].
[0291] i. For example, S is initialized to 0. If the left neighboring block is available and a certain condition is met, S is incremented by 1; if the upper neighboring block is available and a certain condition is met, S is incremented by 1.
[0292] ii. In one example, a specific condition could be whether neighboring blocks are affine encoded or decoded.
[0293] 1) For example, if a block is encoded or decoded using the affine-AMVP mode, it can be considered affine encoded or decoded.
[0294] 2) For example, if a block is encoded or decoded using affine-Merge mode, it can be considered affine encoded or decoded.
[0295] 3) In one example, if the CIIP flag of a block based on its sub-blocks is true, it can be decoded as affine codec.
[0296] c) For example, the context can be selected based on the number of available affine-coded neighboring blocks (denoted as N).
[0297] i. For example, if N <= T, the first context can be used; if N > T, the second context can be used. T is an integer, such as 0, 1, 2, 3, 4, 5, or 6.
[0298] ii. May include, for example Figure 4 Five neighboring blocks or such Figure 10 The seven neighboring sample points.
[0299] iii. For example, if a block is encoded or decoded using an affine-AMVP mode, it can be considered affine encoded or decoded.
[0300] iv. For example, if a block is encoded or decoded using an affine-Merge mode, it can be considered affine encoded or decoded.
[0301] v. In one example, if the CIIP flag of a block based on its sub-blocks is true, then it can be decoded as an affine codec.
[0302] 18. The maximum allowed value for the first SE may depend on whether CIIP based on sub-blocks is used.
[0303] a) In one example, the first SE could be the CIIP Merge index.
[0304] b) In one example, the first SE can be a CIIP-MMVD index.
[0305] c) In one example, the first SE may indicate the use of CIIP-PDPC / CIIP-TM / subblock-based CIIP / CIIP-MMVD.
[0306] d) If the maximum allowed value is encoded or decoded using a rounding unary code such as a rounding unary code or a rounding binary code, then it can determine the last valid codeword of the SE.
[0307] e) For example, the allowed value of the first SE can be transmitted via signaling separately for affine codec and non-affine codec cases, such as in VPS / SPS / PPS / image header / strip header.
[0308] 19. In one example, the maximum allowed value of the first SE can be determined based on the number of available affine-coded neighboring blocks (denoted as N).
[0309] a) In one example, the first SE could be the CIIP Merge index.
[0310] b) In one example, the first SE can be a CIIP-MMVD index.
[0311] c) In one example, the first SE may indicate the use of CIIP-PDPC / CIIP-TM / subblock-based CIIP / CIIP-MMVD.
[0312] d) For example, if N <= T, the first maximum allowed value can be used; if N > T, the second maximum allowed value can be used. T is an integer, such as 0, 1, 2, 3, 4, 5, or 6.
[0313] e) may include, for example Figure 4 Five neighboring blocks or such Figure 10 The seven neighboring sample points.
[0314] f) For example, if a block is encoded or decoded using an affine-AMVP mode, it can be considered affine encoded or decoded.
[0315] g) For example, if a block is encoded or decoded using affine-Merge mode, it can be considered affine encoded or decoded.
[0316] h) In one example, if the CIIP flag of a block based on its sub-blocks is true, then it can be decoded as affine.
[0317] General Information 20. In the above examples, a video unit can refer to a color component / sub-picture / strip / piece / code-decode tree unit (CTU) / CTU line / CTU group / code-decode unit (CU) / prediction unit (PU) / transform unit (TU) / code-decode tree block (CTB) / code-decode block (CB) / prediction block (PB) / transform block (TB) / block / sub-block of a block / sub-region within a block / any other region containing more than one sample or pixel.
[0318] 21. Whether and / or how the methods disclosed above can be applied to be transmitted via signaling at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0319] 22. Whether and / or how to apply the above methods may depend on the following information: a) Messages transmitted via signals in DPS / SPS / VPS / PPS / APS / Picture Header / Strip Header / Piece Group Header / Maximum Codec Unit (LCU) / Codec Unit (CU) / LCU Line / LCU Group / TU / PU Block / Video Codec Unit.
[0320] b) Location of CU / PU / TU / block / video codec unit.
[0321] c) The block dimension of the current block and / or the blocks adjacent to the current block.
[0322] d) The block shape of the current block and / or the blocks adjacent to the current block.
[0323] e) The encoding / decoding mode of the block, such as IBC or non-IBC inter-frame mode or non-IBC sub-block mode.
[0324] f) Indication of color format (such as 4:2:0, 4:4:4).
[0325] g) Encoding / decoding tree structure.
[0326] h) Strip / panel type and / or image type.
[0327] i) Color components (e.g., can be applied only to the chromaticity component or the luminance component).
[0328] j) Temporal layer ID.
[0329] k) Standard grade / level / layer.
[0330] Figure 24 A flowchart of a method 2400 for video processing according to an embodiment of the present disclosure is shown. Method 2400 is implemented during the conversion between video units of a video and a bitstream of a video.
[0331] At box 2410, for the conversion between the current video block and the video bitstream, the affine information of the current video block is determined based on the first affine information of the first block of the video. The first block is encoded and decoded using a first encoding / decoding tool different from the affine encoding / decoding tool.
[0332] At box 2420, the prediction of the current video block is determined based on affine information.
[0333] At box 2430, the transformation is performed based on prediction. In some embodiments, the transformation includes encoding the current video block into a bitstream. Alternatively or additionally, in some embodiments, the transformation includes decoding the current video block from the bitstream.
[0334] Method 2400 enables the generation of predictions containing affine information of blocks encoded using a different encoding / decoding tool than the affine encoding / decoding tool. In this way, encoding / decoding efficiency and effectiveness can be improved.
[0335] In some embodiments, the first block is encoded and decoded in at least one of intra-frame joint prediction (CIIP) mode or sub-block-based CIIP mode.
[0336] In some embodiments, affine motion compensation is applied to generate the CIIP prediction for the first block.
[0337] In some embodiments, the affine information includes at least one of the following: inter-frame orientation, control point motion vector (CPMV), affine parameters, at least one reference index of at least one list of reference pictures, local illumination compensation (LIC) flag, overlap block motion compensation (OBMC) flag, bidirectional prediction (BCW) index weighted by codec unit level, affine type (e.g., 4-parameter affine or 6-parameter affine), merge type (e.g., affine or sbTMVP), or iso-picture index. For example, the inter-frame orientation may include unidirectional or bidirectional prediction from list 0 or list 1. In one example, the CPMV may include the CPMVs of the top left, top right, and bottom left corners. For example, the affine parameters may be those in equations (1) and (2). a, b, c, d, e f .
[0338] In some embodiments, the first block is encoded and decoded in Geometric Partition Mode (GPM).
[0339] In some embodiments, affine motion compensation is applied to generate the GPM prediction for the first block.
[0340] In some embodiments, the first block is encoded and decoded in a multiple hypothesis prediction (MHP) mode.
[0341] In some embodiments, affine motion compensation is applied to generate the MHP prediction for the first block.
[0342] In some embodiments, the affine model included in the affine information is utilized by at least one block that is encoded and decoded after the first block.
[0343] In some embodiments, affine information is used by at least one block that is encoded and decoded after a first block, the first block being encoded and decoded using sub-block-based CIIP.
[0344] In some embodiments, the affine information of the first block is included in the History Parameter Table (HPT).
[0345] In some embodiments, the HPT is updated after encoding or decoding at least one of the blocks encoded or decoded using sub-block-based CIIP encoding / decoding.
[0346] In some embodiments, the affine information of the first block is included in at least one of the affine merge list or the advanced motion vector prediction (AMVP) candidate list of blocks encoded and decoded after the first block.
[0347] In some embodiments, the affine information of the first block is included in the affine candidate list of the second block, which is encoded and decoded in at least one of the sub-block-based CIIP mode or the affine-GPM mode.
[0348] In some embodiments, method 2400 further includes: determining to include the affine information in the affine candidate list based on a comparison between the affine information and at least one affine candidate in the affine candidate list.
[0349] In some embodiments, affine information is excluded from the affine candidate list in response to the affine information being identical to at least one affine candidate in the affine candidate list.
[0350] In some embodiments, the affine information is excluded from the affine candidate list in response to a difference between the affine information and at least one affine candidate in the affine candidate list being less than or equal to a threshold.
[0351] In some embodiments, method 2400 further includes constructing an affine list based on examining at least one of the sub-blocks or blocks that cover a location related to the current video block. For example, the location related to the current video block can be implemented as follows: Figure 5 A0, A1, B0, B1, and B2 in the example.
[0352] In some embodiments, the affine candidate list includes at least one of the following: a sub-block merge list, an affine merge list, an affine AMVP list, an affine list for GPM, or an affine list for sub-block-based CIIP.
[0353] In some embodiments, method 2400 further includes: determining whether the encoding / decoding tool used for at least one of the sub-blocks or blocks is an affine-based encoding / decoding tool. For example, the encoding / decoding tool may include affine Merge or affine AMVP.
[0354] In some embodiments, method 2400 further includes: if it is determined that the encoding / decoding tool is an affine-based encoding / decoding tool, then using affine information associated with at least one of the sub-blocks or blocks to generate an affine candidate list.
[0355] In some embodiments, method 2400 further includes: determining whether the encoding / decoding tool used for at least one of the sub-blocks or blocks is at least one of a GPM-based affine encoding / decoding tool or a sub-block-based CIIP encoding / decoding tool.
[0356] In some embodiments, method 2400 further includes: if it is determined that the encoding / decoding tool is at least one of a GPM-based affine encoding / decoding tool or a sub-block-based CIIP encoding / decoding tool, then using affine information related to the sub-block or at least one of the blocks to generate an affine candidate list.
[0357] In some embodiments, at least one of the sub-blocks or blocks is examined to generate affine candidates from blocks encoded and decoded using sub-block-based CIIP.
[0358] In some embodiments, after examining affine candidates inherited from blocks adjacent to the current block, at least one of the sub-blocks or blocks is examined to immediately generate affine candidates from blocks encoded and decoded using sub-block-based CIIP.
[0359] In some embodiments, after examining affine candidates inherited from blocks that are not adjacent to the current block, at least one of the sub-blocks or blocks is examined to immediately generate affine candidates from blocks encoded and decoded using CIIP based on sub-blocks.
[0360] In some embodiments, after checking regression-based motion vector field (RMVF) affine candidates, at least one of the sub-blocks or blocks is checked to immediately generate affine candidates from the blocks encoded and decoded by the sub-blocks based CIIP.
[0361] In some embodiments, after examining the constructed affine candidates, at least one of the sub-blocks or blocks is examined to immediately generate affine candidates from the blocks encoded and decoded using sub-block-based CIIP.
[0362] In some embodiments, after examining the affine candidates of temporal inheritance, at least one of the sub-blocks or blocks is examined to immediately generate affine candidates from the blocks encoded and decoded based on the sub-blocks using CIIP.
[0363] In some embodiments, after examining history-based affine candidates, at least one of the sub-blocks or blocks is examined to immediately generate affine candidates from the blocks encoded and decoded using sub-block-based CIIP.
[0364] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by a video processing apparatus. The method includes: determining affine information of a current video block based on first affine information of a first block of video, the first block being encoded and decoded using a first encoding / decoding tool different from the affine encoding / decoding tool; determining a prediction of the current video block based on the affine information; and generating a bitstream based on the prediction.
[0365] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: determining affine information of a current video block based on first affine information of a first block of video, the first block being encoded and decoded using a first encoding / decoding tool different from the affine encoding / decoding tool; determining a prediction of the current video block based on the affine information; generating a bitstream based on the prediction; and storing the bitstream in a non-transitory computer-readable recording medium.
[0366] Figure 25 A flowchart of a method 2500 for video processing according to an embodiment of the present disclosure is shown. Method 2500 is implemented during the conversion between video units of a video and a bitstream of a video.
[0367] At box 2510, for the conversion between the current video block and the video bitstream, the syntax elements related to intra-frame and inter-frame joint prediction (CIIP) are encoded and decoded based on more than one context.
[0368] At box 2520, the conversion is performed based on syntax elements. Alternatively or additionally, in some embodiments, the conversion includes decoding the current video block from the bitstream.
[0369] Method 2500 enables encoding and decoding of CIIP-related syntax elements using at least one context. In this way, encoding / decoding efficiency and effectiveness can be improved.
[0370] In some embodiments, the syntax element is an index of at least one of CIIP Merge or Merge Pattern with Motion Vector Difference (MMVD).
[0371] In some embodiments, the syntax element indicates the use of at least one of the following: CIIP-Location-Related Intra-Prediction Combination (PDPC), CIIP-Template Matching (TM), Sub-Block-Based CIIP, or CIIP-MMVD.
[0372] In some embodiments, syntax elements are binarized into N binary bits, where N is an integer. For example, syntax elements can be binarized to represent B0, B1, ..., B N-1 N binary bits. In one example, if n != m, then B n and B m It can be encoded and decoded using different contexts.
[0373] In some embodiments, a first number of binary bits out of N binary bits are encoded and decoded using a different context, and the remaining binary bits out of N binary bits are encoded and decoded using the same context, wherein the first number is one of the following: 1, 2, 3, 4, 5, 6, or 7. For example, B0, B1, ..., B W They can be represented as C0, C1, ..., C W Different contexts are encoded and decoded. Furthermore, B W B W+1 B N-1 Shared can be represented as C W+1 The same context. Specifically, W is an integer, such as 1, 2, 3, 4, 5, 6, or 7.
[0374] In some embodiments, a first number of binary bits out of N binary bits are encoded and decoded using a different context, and the remaining binary bits out of N binary bits are bypassed and decoded, wherein the first number is one of the following: 1, 2, 3, 4, 5, 6, or 7. For example, B0, B1, ..., B W They can be represented as C0, C1, ..., C W Different contexts are encoded and decoded. Furthermore, B W B W+1 B N-1 It can be bypassed for encoding and decoding. Specifically, W is an integer, such as 1, 2, 3, 4, 5, 6, or 7.
[0375] In some embodiments, at least one context for encoding / decoding syntax elements is determined based on encoding / decoding information.
[0376] In some embodiments, at least one context is determined based on whether a sub-block-based CIIP is used.
[0377] In some embodiments, at least one context of the index of at least one of the CIIP Merge or MMVD used for encoding and decoding CIIP blocks is determined based on whether CIIP based on sub-blocks is used.
[0378] In some embodiments, at least one context for indicating the use of at least one of the following is determined based on whether sub-block-based CIIP is used: CIIP-PDPC, CIIP-TM, sub-block-based CIIP, or CIIP-MMVD.
[0379] In some embodiments, at least one context is determined based on whether at least one of the following is used: CIIP-PDPC, CIIP-TM, sub-block based CIIP, or CIIP-MMVD.
[0380] In some embodiments, at least one context is determined based on a quantization parameter (QP).
[0381] In some embodiments, at least one context is determined based on the stripe type.
[0382] In some embodiments, the encoding / decoding information is associated with at least one neighboring block.
[0383] In some embodiments, at least one context is selected from a candidate set that includes at least one candidate context. For example, the candidate set may include three members denoted as {C[0], C[1], C[2]}, and the selected context is C[S].
[0384] In some embodiments, the index of the selected candidate context is determined based on at least one neighboring block.
[0385] In some embodiments, the index is determined by: initializing the index to 0; and incrementing the index by 1 in response to at least one neighboring block being available and a condition being met, wherein the at least one neighboring block includes at least one of a left neighboring block or an upper neighboring block. For example, S is initialized to 0. In one example, S is incremented by 1 if a left neighboring block is available and a specific condition is met. Alternatively or additionally, S is incremented by 1 if an upper neighboring block is available and a specific condition is met.
[0386] In some embodiments, the condition includes that at least one neighboring block is affine-coded.
[0387] In some embodiments, at least one neighboring block encoded in affine-advanced motion vector prediction (AMVP) mode is considered to be affine encoded.
[0388] In some embodiments, at least one neighboring block encoded in affine-Merge mode is considered to be affine encoded.
[0389] In some embodiments, at least one neighboring block for which CIIP based on a sub-block is true is considered to be affine-coded.
[0390] In some embodiments, the syntax element is a CIIP flag based on sub-blocks.
[0391] In some embodiments, at least one context for the affine tag of a block is determined based on whether at least one neighboring block of the block is encoded or decoded in a sub-block-based CIIP mode.
[0392] In some embodiments, at least one context is selected from a candidate set that includes at least one candidate context. For example, the candidate set may include three members denoted as {C[0], C[1], C[2]}, and the selected context is C[S].
[0393] In some embodiments, at least one context is selected by: incrementing the index of the selected candidate context by 1 in response to at least one neighboring block being available and a condition being met; and selecting at least one context based on the index, wherein the index is initialized to 0, and at least one neighboring block includes at least one of a left neighboring block or an upper neighboring block. For example, S is initialized to 0. In one example, S is incremented by 1 if a left neighboring block is available and a specific condition is met. Alternatively or additionally, S is incremented by 1 if an upper neighboring block is available and a specific condition is met.
[0394] In some embodiments, the condition includes that at least one neighboring block is affine-coded.
[0395] In some embodiments, at least one neighboring block encoded in affine-advanced motion vector prediction (AMVP) mode is considered to be affine encoded.
[0396] In some embodiments, at least one neighboring block encoded in affine-Merge mode is considered to be affine encoded.
[0397] In some embodiments, at least one neighboring block for which CIIP based on a sub-block is true is considered to be affine-coded.
[0398] In some embodiments, at least one context is selected based on the number of available affine-coded neighboring blocks.
[0399] In some embodiments, a first context is used if the number of available affine-coded neighboring blocks is less than or equal to a threshold, and a second context is used if the number is greater than the threshold, wherein the threshold is an integer.
[0400] In some embodiments, the threshold is 0, 1, 2, 3, 4, 5, or 6.
[0401] In some embodiments, the available affine-coded neighbor blocks include five neighbor blocks or seven neighbor samples. For example, an example of a five-neighbor block is... Figure 4 It is shown in the image. Additionally, examples of seven neighboring samples are shown in... Figure 10 It is shown in the middle.
[0402] In some embodiments, at least one neighboring block encoded in affine-advanced motion vector prediction (AMVP) mode is considered to be affine encoded.
[0403] In some embodiments, at least one neighboring block encoded in affine-Merge mode is considered to be affine encoded.
[0404] In some embodiments, at least one neighboring block for which CIIP based on a sub-block is true is considered to be affine-coded.
[0405] In some embodiments, the maximum value of a syntax element is determined based on whether sub-block-based CIIP is used.
[0406] In some embodiments, if a syntax element is encoded or decoded using at least one of rounding unary codes or rounding binary codes, the last valid codeword of the syntax element is determined based on the maximum value.
[0407] In some embodiments, the maximum value of a syntax element is indicated individually for affine-coded blocks and non-affine-coded blocks in at least one of the following: Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture Header, or Strip Header.
[0408] In some embodiments, the maximum value is determined based on the number of available affine-coded neighboring blocks.
[0409] In some embodiments, a first context is used if the number of available affine-coded neighboring blocks is less than or equal to a threshold, and a second context is used if the number is greater than the threshold, wherein the threshold is an integer.
[0410] In some embodiments, the threshold is 0, 1, 2, 3, 4, 5, or 6.
[0411] In some embodiments, the available affine-coded neighbor blocks include five neighbor blocks or seven neighbor samples. For example, an example of a five-neighbor block is... Figure 4 It is shown in the image. Additionally, examples of seven neighboring samples are shown in... Figure 10 It is shown in the middle.
[0412] In some embodiments, at least one neighboring block encoded in affine-advanced motion vector prediction (AMVP) mode is considered to be affine encoded.
[0413] In some embodiments, at least one neighboring block encoded in affine-Merge mode is considered to be affine encoded.
[0414] In some embodiments, at least one neighboring block for which CIIP based on a sub-block is true is considered to be affine-coded.
[0415] In some embodiments, the current video block or video unit includes at least one of the following: color component, sub-picture, strip, slice, codec tree unit (CTU), CTU row, CTU group, codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), prediction block (PB), transform block (TB), block, sub-block of block, sub-region within block, or region containing more than one sample point or pixel.
[0416] In some embodiments, the indication of whether and / or how to apply method 2400 and / or method 2500 is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.
[0417] In some embodiments, the indication of whether and / or how to apply method 2400 and / or method 2500 is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header or slice header.
[0418] In some embodiments, whether and / or how method 2400 and / or method 2500 are applied is based on at least one of the following: messages including one of the following: Dependency Parameter Set (DPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Picture Header, Strip Header, Piece Group Header, Maximum Codec Unit (LCU), Codec Unit (CU), LCU Row, LCU Group, Transform Unit (TU), Prediction Unit (PU), Block or Video Codec Unit, Position of CU, PU, TU, Block or Video Codec Unit, Block Dimension of the current video block and / or Block Dimension of the current video block's neighboring blocks, Block Shape of the current video block and / or Block Shape of the current video block's neighboring blocks, Codec Mode of the block, Indicator of color format, Codec Tree Structure, Strip, Piece Group Type and / or Picture Type, Color Components, Temporal Layer Identifier (ID), or Standard Level, Grade or Layer.
[0419] In some embodiments, the encoding / decoding mode includes one of the following: intra-block copy (IBC), non-IBC inter-frame mode, or non-IBC sub-block mode, or the color format includes one of the following: 4:2:0 or 4:4:4.
[0420] In some embodiments, the conversion includes encoding the current video block into a bitstream.
[0421] In some embodiments, the conversion includes decoding the current video block from the bitstream.
[0422] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: encoding and decoding syntax elements related to intra-frame / inter-frame joint prediction (CIIP) based on more than one context; and generating a bitstream based on the syntax elements.
[0423] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: encoding and decoding syntax elements related to intra-frame / inter-frame joint prediction (CIIP) based on more than one context; generating a bitstream based on the syntax elements; and storing the bitstream in a non-transitory computer-readable recording medium.
[0424] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.
[0425] Item 1. A method for video processing, comprising: a conversion between a current video block and a bitstream of the video; determining affine information of the current video block based on first affine information of a first block of the video, the first block being encoded and decoded using a first codec tool different from the affine codec tool; determining a prediction of the current video block based on the affine information; and performing the conversion based on the prediction.
[0426] Item 2. The method according to Item 1, wherein the first block is encoded or decoded in at least one of intra-frame joint prediction (CIIP) mode or sub-block based CIIP mode.
[0427] Item 3. The method according to Item 2, wherein affine motion compensation is applied to generate the CIIP prediction of the first block.
[0428] Item 4. The method according to Item 1, wherein the affine information includes at least one of the following: inter-frame direction, control point motion vector (CPMV), affine parameters, at least one reference index of at least one list of reference pictures, local illumination compensation (LIC) flag, overlap block motion compensation (OBMC) flag, bidirectional prediction (BCW) index weighted by the codec unit level, affine type, merge type, or isotope index.
[0429] Item 5. The method according to Item 1, wherein the first block is encoded and decoded in Geometric Partition Mode (GPM).
[0430] Item 6. The method according to Item 5, wherein affine motion compensation is applied to generate the GPM prediction of the first block.
[0431] Item 7. The method according to Item 1, wherein the first block is encoded and decoded in a multiple hypothesis prediction (MHP) mode.
[0432] Item 8. The method according to Item 7, wherein affine motion compensation is applied to generate the MHP prediction of the first block.
[0433] Item 9. The method according to any one of items 1 to 8, wherein the affine model included in the affine information is utilized by at least one block that is encoded or decoded after the first block.
[0434] Item 10. The method according to Item 2, wherein the affine information is used by at least one block that is encoded or decoded after the first block, the first block being encoded or decoded using the sub-block-based CIIP.
[0435] Item 11. The method according to Item 10, wherein the affine information of the first block is included in the History Parameter Table (HPT).
[0436] Item 12. The method according to Item 11, wherein the HPT is updated after encoding or decoding at least one of the blocks encoded or decoded by sub-block-based CIIP encoding / decoding.
[0437] Item 13. The method according to Item 10, wherein the affine information of the first block is included in at least one of the affine merge list or the advanced motion vector prediction (AMVP) candidate list of blocks encoded and decoded after the first block.
[0438] Item 14. The method according to Item 10, wherein the affine information of the first block is included in the affine candidate list of the second block which is encoded and decoded in at least one of the sub-block-based CIIP mode or affine-GPM mode.
[0439] Item 15. The method according to Item 13 or Item 14 further comprises: determining, based on a comparison between the affine information and at least one affine candidate in the affine candidate list, to include the affine information in the affine candidate list.
[0440] Item 16. The method according to Item 15, wherein the affine information is excluded from the affine candidate list in response to the affine information being identical to at least one affine candidate in the affine candidate list.
[0441] Item 17. The method according to Item 15, wherein the affine information is excluded from the affine candidate list in response to a difference between the affine information and at least one affine candidate in the affine candidate list being less than or equal to a threshold.
[0442] Item 18. The method according to Item 1 further includes: constructing an affine list based on examining at least one of sub-blocks or blocks that cover a location related to the current video block.
[0443] Item 19. The method according to Item 18, wherein the affine candidate list includes at least one of the following: a sub-block merge list, an affine merge list, an affine AMVP list, an affine list for GPM, or an affine list for sub-block-based CIIP.
[0444] Item 20. The method according to Item 18 or Item 19 further includes: determining whether the encoding / decoding tool used for the at least one of the sub-blocks or blocks is an affine encoding / decoding tool.
[0445] Item 21. The method according to Item 20 further includes: if it is determined that the encoding / decoding tool is an affine-based encoding / decoding tool, then using affine information relating to the at least one of the sub-blocks or blocks to generate affine candidates for the affine candidate list.
[0446] Item 22. The method according to Item 18 or Item 19 further comprises: determining whether the encoding / decoding tool used for the at least one of the sub-blocks or blocks is at least one of a GPM affine encoding / decoding tool or a sub-block-based CIIP encoding / decoding tool.
[0447] Item 23. The method according to Item 22 further comprises: if it is determined that the codec tool is at least one of a GPM-based affine codec tool or a sub-block-based CIIP codec tool, then using affine information relating to the at least one of the sub-blocks or blocks to generate affine candidates for the affine candidate list.
[0448] Item 24. The method according to items 18 to 23, wherein at least one of the sub-blocks or blocks is examined to generate the affine candidate from the block encoded and decoded by sub-block-based CIIP.
[0449] Item 25. The method according to Item 24, wherein after examining the inherited affine candidates from the block adjacent to the current block, at least one of the sub-blocks or blocks is examined to immediately generate the affine candidates from the block encoded and decoded by the sub-block-based CIIP.
[0450] Item 26. The method according to Item 24, wherein after checking the affine candidates inherited from blocks that are not adjacent to the current block, at least one of the sub-blocks or blocks is checked to immediately generate the affine candidates from the block encoded and decoded by the sub-block-based CIIP.
[0451] Item 27. The method according to Item 24, wherein after checking regression-based motion vector field (RMVF) affine candidates, the at least one of the sub-blocks or blocks is checked to immediately generate the affine candidates from the blocks encoded and decoded by the sub-block-based CIIP.
[0452] Item 28. The method according to Item 24, wherein after examining the constructed affine candidate, at least one of the sub-blocks or blocks is examined to immediately generate the affine candidate from the block encoded and decoded by the sub-block-based CIIP.
[0453] Item 29. The method according to Item 24, wherein after checking the affine candidates of temporal inheritance, the at least one of the sub-blocks or blocks is checked to immediately generate the affine candidates from the blocks encoded and decoded by the sub-block-based CIIP.
[0454] Item 30. The method according to Item 24, wherein after checking the history-based affine candidates, at least one of the sub-blocks or blocks is checked to immediately generate the affine candidates from the blocks encoded and decoded by the sub-block-based CIIP.
[0455] Item 31. A method for video processing, comprising: a conversion between a current video block of a video and a bitstream of the video; encoding and decoding syntax elements related to intra-frame and inter-frame joint prediction (CIIP) based on more than one context; and performing the conversion based on the syntax elements.
[0456] Item 32. The method according to Item 31, wherein the syntax element is an index of at least one of CIIP Merge or Merge Pattern with Motion Vector Difference (MMVD).
[0457] Item 33. The method according to Item 31, wherein the syntax element indicates the use of at least one of the following: CIIP-Location-Related Intra-Prediction Combination (PDPC), CIIP-Template Matching (TM), Sub-Block-Based CIIP, or CIIP-MMVD.
[0458] Item 34. The method according to Item 31, wherein the syntax element is binarized into N binary bits, where N is an integer.
[0459] Item 35. The method according to Item 34, wherein a first number of binary bits of the N binary bits are encoded and decoded using a different context, and the remaining binary bits of the N binary bits are encoded and decoded using the same context, wherein the first number is one of the following: 1, 2, 3, 4, 5, 6 or 7.
[0460] Item 36. The method according to Item 34, wherein a first number of binary bits of the N binary bits are encoded and decoded using a different context, and the remaining binary bits of the N binary bits are bypassed and decoded, wherein the first number is one of the following: 1, 2, 3, 4, 5, 6 or 7.
[0461] Item 37. The method according to any one of items 31 to 33, wherein at least one context for encoding and decoding the syntax element is determined based on encoding and decoding information.
[0462] Item 38. The method according to Item 37, wherein the at least one context is determined based on whether a sub-block-based CIIP is used.
[0463] Item 39. The method according to Item 38, wherein the at least one context of the index of at least one of CIIP Merge or MMVD used for encoding and decoding CIIP blocks is determined based on whether CIIP based on sub-blocks is used.
[0464] Item 40. The method according to Item 38, wherein the at least one context for indicating the use of at least one of the following is determined based on whether sub-block-based CIIP is used: CIIP-PDPC, CIIP-TM, sub-block-based CIIP, or CIIP-MMVD.
[0465] Item 41. The method according to Item 38, wherein the at least one context is determined based on whether at least one of the following is used: CIIP-PDPC, CIIP-TM, sub-block based CIIP, or CIIP-MMVD.
[0466] Item 42. The method according to Item 38, wherein the at least one context is determined based on a quantization parameter (QP).
[0467] Item 43. The method according to Item 38, wherein the at least one context is determined based on the stripe type.
[0468] Item 44. The method according to Item 37, wherein the encoding / decoding information is associated with at least one neighboring block.
[0469] Item 45. The method according to Item 44, wherein the at least one context is selected from a candidate set including at least one candidate context.
[0470] Item 46. The method according to Item 45, wherein the index of the selected candidate context is determined based on at least one neighboring block.
[0471] Item 47. The method according to Item 46, wherein the index is determined by: initializing the index to 0; and incrementing the index by 1 in response to the availability of the at least one neighboring block and the satisfaction of a condition, wherein the at least one neighboring block includes at least one of a left neighboring block or an upper neighboring block.
[0472] Item 48. The method according to Item 47, wherein the condition includes that the at least one neighboring block is affine encoded.
[0473] Item 49. The method according to Item 48, wherein the at least one neighboring block encoded in an affine-advanced motion vector prediction (AMVP) mode is considered to be affine encoded.
[0474] Item 50. The method according to Item 48, wherein the at least one neighboring block encoded in affine-Merge mode is considered to be affine encoded.
[0475] Item 51. The method according to Item 48, wherein the at least one neighboring block for which CIIP based on the sub-block is true is considered to be affine encoded.
[0476] Item 52. The method according to Item 46, wherein the syntax element is a CIIP flag based on a sub-block.
[0477] Item 53. The method according to Item 31, wherein at least one context for the affine tag of a block is determined based on whether at least one neighboring block of the block is encoded or decoded in a sub-block-based CIIP mode.
[0478] Item 54. The method according to Item 53, wherein the at least one context is selected from a candidate set including at least one candidate context.
[0479] Item 55. The method according to Item 54, wherein the at least one context is selected by: incrementing the index of the selected candidate context by 1 in response to the availability of the at least one neighboring block and the condition being met; and selecting the at least one context based on the index, wherein the index is initialized to 0, and the at least one neighboring block includes at least one of the left neighboring block or the upper neighboring block.
[0480] Item 56. The method according to Item 55, wherein the condition includes that the at least one neighboring block is affine encoded.
[0481] Item 57. The method according to Item 56, wherein the at least one neighboring block encoded in affine-advanced motion vector prediction (AMVP) mode is considered to be affine encoded.
[0482] Item 58. The method according to Item 56, wherein the at least one neighboring block encoded in affine-Merge mode is considered to be affine encoded.
[0483] Item 59. The method according to Item 56, wherein the at least one neighboring block for which CIIP based on the sub-block is true is considered to be affine encoded.
[0484] Item 60. The method according to Item 54, wherein the at least one context is selected based on the number of available affine-coded neighboring blocks.
[0485] Item 61. The method according to Item 60, wherein a first context is used if the number of available affine-coded neighboring blocks is less than or equal to a threshold, and a second context is used if the number is greater than the threshold, wherein the threshold is an integer.
[0486] Item 62. The method according to Item 61, wherein the threshold is 0, 1, 2, 3, 4, 5 or 6.
[0487] Item 63. The method according to Item 60, wherein the available affine-coded neighbor blocks comprise five neighbor blocks or seven neighbor samples.
[0488] Item 64. The method according to Item 60, wherein the at least one neighboring block encoded in an affine-advanced motion vector prediction (AMVP) mode is considered to be affine encoded.
[0489] Item 65. The method according to Item 60, wherein the at least one neighboring block encoded in affine-Merge mode is considered to be affine encoded.
[0490] Item 66. The method according to Item 60, wherein the at least one neighboring block for which CIIP based on the sub-block is true is considered to be affine encoded / decoded.
[0491] Item 67. The method according to any one of items 31 to 33, wherein the maximum value of the syntax element is determined based on whether sub-block-based CIIP is used.
[0492] Item 68. The method according to Item 67, wherein if the syntax element is encoded or decoded using at least one of rounding unary code or rounding binary code, the last valid codeword of the syntax element is determined based on the maximum value.
[0493] Item 69. The method according to Item 68, wherein the maximum value of the syntax element is indicated individually for both affine and non-affine blocks in at least one of: Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture Header, or Strip Header.
[0494] Item 70. The method according to Item 67, wherein the maximum value is determined based on the number of available affine-coded neighboring blocks.
[0495] Item 71. The method according to Item 70, wherein a first context is used if the number of available affine-coded neighboring blocks is less than or equal to a threshold, and a second context is used if the number is greater than the threshold, wherein the threshold is an integer.
[0496] Item 72. The method according to Item 71, wherein the threshold is 0, 1, 2, 3, 4, 5 or 6.
[0497] Item 73. The method according to Item 70, wherein the available affine-coded neighbor blocks comprise five neighbor blocks or seven neighbor samples.
[0498] Item 74. The method according to Item 70, wherein the at least one neighboring block encoded in an affine-advanced motion vector prediction (AMVP) mode is considered to be affine encoded.
[0499] Item 75. The method according to Item 70, wherein the at least one neighboring block encoded in affine-Merge mode is considered to be affine encoded.
[0500] Item 76. The method according to Item 70, wherein the at least one neighboring block for which CIIP based on the sub-block is true is considered to be affine encoded / decoded.
[0501] Item 77. The method according to any one of Items 1 to 76, wherein the current video block or video unit includes at least one of the following: color component, sub-picture, strip, slice, codec tree unit (CTU), CTU row, CTU group, codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), prediction block (PB), transform block (TB), block, sub-block of block, sub-region within block, or region containing more than one sample point or pixel.
[0502] Item 78. The method according to any one of items 1 to 77, wherein an indication of whether and / or how the method is applied is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.
[0503] Item 79. The method according to any one of items 1 to 77, wherein an indication of whether and / or how to apply the method is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.
[0504] Item 80. The method according to any one of Items 1 to 77, wherein whether and / or how the method is applied is based on at least one of the following: a message including one of the following: Dependency Parameter Set (DPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Picture Header, Strip Header, Piece Group Header, Maximum Codec Unit (LCU), Codec Unit (CU), LCU Row, LCU Group, Transform Unit (TU), Prediction Unit (PU), Block or Video Codec Unit, Position of CU, PU, TU, Block or Video Codec Unit, Block Dimension of the current video block and / or Block Dimension of the current video block's neighboring blocks, Block Shape of the current video block and / or Block Shape of the current video block's neighboring blocks, Codec Mode of the block, Indicator of Color Format, Codec Tree Structure, Strip, Piece Group Type and / or Picture Type, Color Components, Temporal Layer Identifier (ID), or Standard Level, Grade or Layer.
[0505] Item 81. The method according to Item 80, wherein the encoding / decoding mode includes one of the following: intra-block copy (IBC), non-IBC inter-frame mode, or non-IBC sub-block mode, or wherein the color format includes one of the following: 4:2:0 or 4:4:4.
[0506] Item 82. The method according to any one of items 1 to 81, wherein the conversion includes encoding the current video block into the bitstream.
[0507] Item 83. The method according to any one of items 1 to 81, wherein the conversion includes decoding the current video block from the bitstream.
[0508] Item 84. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of items 1 to 83.
[0509] Item 85. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of items 1 to 83.
[0510] Item 86. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: determining affine information of a current video block based on first affine information of a first block of the video, the first block being encoded and decoded using a first encoding / decoding tool different from the affine encoding / decoding tool; determining a prediction of the current video block based on the affine information; and generating the bitstream based on the prediction.
[0511] Item 87. A method for storing a bitstream of video, comprising: determining affine information of a current video block based on first affine information of a first block of the video, the first block being encoded and decoded using a first encoding / decoding tool different from the affine encoding / decoding tool; determining a prediction of the current video block based on the affine information; generating the bitstream based on the prediction; and storing the bitstream in a non-transitory computer-readable recording medium.
[0512] Item 88. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: encoding and decoding syntax elements relating to intra-frame and inter-frame joint prediction (CIIP) based on more than one context; and generating the bitstream based on the syntax elements.
[0513] Item 89. A method for storing a bitstream of video, comprising: encoding and decoding syntax elements related to intra-frame and inter-frame joint prediction (CIIP) based on more than one context; generating the bitstream based on the syntax elements; and storing the bitstream in a non-transitory computer-readable recording medium.
[0514] Example device Figure 26 A block diagram of a computing device 2600 in which various embodiments of the present disclosure may be implemented is shown. The computing device 2600 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0515] It should be understood that, Figure 26 The computing device 2600 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.
[0516] like Figure 26 As shown, computing device 2600 includes general-purpose computing device 2600. Computing device 2600 may include at least one or more processors or processing units 2610, memory 2620, storage unit 2630, one or more communication units 2640, one or more input devices 2650, and one or more output devices 2660.
[0517] In some embodiments, the computing device 2600 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 2600 can support any type of interface to the user (such as "wearable" circuit systems, etc.).
[0518] Processing unit 2610 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 2620. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 2600. Processing unit 2610 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0519] Computing device 2600 typically includes various computer storage media. Such media can be any media accessible by computing device 2600, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 2620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 2630 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 2600.
[0520] The computing device 2600 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 26 Not shown, but may provide disk drives for reading from and / or writing to removable non-volatile disks, and optical disc drives for reading from and / or writing to removable non-volatile optical discs. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0521] Communication unit 2640 communicates with another computing device via a communication medium. Furthermore, the functionality of components in computing device 2600 can be implemented by a single computing cluster or by multiple computing machines communicating via communication connections. Therefore, computing device 2600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0522] Input device 2650 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 2660 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 2640, computing device 2600 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 2600 can also communicate with one or more devices that enable a user to interact with computing device 2600, or any device that enables computing device 2600 to communicate with one or more other computing devices (e.g., network card, modem, etc.), if needed. Such communication can be performed via an input / output (I / O) interface (not shown).
[0523] In some embodiments, some or all components of computing device 2600 may not be integrated into a single device, but may be deployed within a cloud computing architecture. In a cloud computing architecture, components may be provided remotely and may work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing is provided via a wide area network (WAN) such as the Internet using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed at locations in remote data centers. Cloud computing infrastructure may be provided through shared data centers, although they may appear as a single access point for users. Thus, cloud computing architectures can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, they may be provided from traditional servers, or directly installed or otherwise installed on client devices.
[0524] In embodiments of this disclosure, computing device 2600 can be used to implement video encoding / decoding. Memory 2620 may include one or more video codec modules 2625 having one or more program instructions. These modules can be accessed and executed by processing unit 2610 to perform the functions of the various embodiments described herein.
[0525] In an example embodiment of performing video encoding, input device 2650 may receive video data as input 2670 to be encoded. The video data may be processed, for example, by video codec module 2625 to generate an encoded bitstream. The encoded bitstream may be provided as output 2680 via output device 2660.
[0526] In an example embodiment of performing video decoding, input device 2650 may receive an encoded bitstream as input 2670. The encoded bitstream may be processed, for example, by a video codec module 2625 to generate decoded video data. The decoded video data may be provided as output 2680 via output device 2660.
[0527] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These changes are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.
Claims
1. A method for video processing, comprising: For the conversion between the current video block and the bitstream of the video, the affine information of the current video block is determined based on the first affine information of the first block of the video, and the first block is encoded and decoded using a first encoding and decoding tool that is different from the affine encoding and decoding tool; The prediction of the current video block is determined based on the affine information; as well as The transformation is performed based on the prediction.
2. The method of claim 1, wherein the first block is encoded or decoded in at least one of intra-frame joint prediction (CIIP) mode or sub-block-based CIIP mode.
3. The method of claim 2, wherein affine motion compensation is applied to generate the CIIP prediction for the first block.
4. The method of claim 1, wherein the affine information comprises at least one of the following: Inter-frame direction, Control point motion vector (CPMV) Affine parameters At least one reference index of at least one list of reference images Local Lighting Compensation (LIC) sign Overlapping Block Motion Compensation (OBMC) flag Utilizing codec unit-level weighted bidirectional prediction (BCW) indexing, Affine type Merge type, or Image index.
5. The method of claim 1, wherein the first block is encoded and decoded in geometric segmentation mode (GPM).
6. The method of claim 5, wherein affine motion compensation is applied to generate the GPM prediction for the first block.
7. The method of claim 1, wherein the first block is encoded and decoded in a multiple hypothesis prediction (MHP) mode.
8. The method of claim 7, wherein affine motion compensation is applied to generate the MHP prediction of the first block.
9. The method according to any one of claims 1 to 8, wherein the affine model included in the affine information is utilized by at least one block that is encoded or decoded after the first block.
10. The method of claim 2, wherein the affine information is used by at least one block that is encoded or decoded after the first block, the first block being encoded or decoded using the sub-block-based CIIP.
11. The method of claim 10, wherein the affine information of the first block is included in a history parameter table (HPT).
12. The method of claim 11, wherein the HPT is updated after encoding or decoding at least one of the blocks encoded or decoded using sub-block-based CIIP encoding / decoding.
13. The method of claim 10, wherein the affine information of the first block is included in at least one of an affine merge list or an advanced motion vector prediction (AMVP) candidate list of blocks encoded and decoded after the first block.
14. The method of claim 10, wherein the affine information of the first block is included in the affine candidate list of the second block which is encoded and decoded in at least one of the sub-block-based CIIP mode or affine-GPM mode.
15. The method according to claim 13 or claim 14, further comprising: Based on a comparison between the affine information and at least one affine candidate in the affine candidate list, it is determined that the affine information should be included in the affine candidate list.
16. The method of claim 15, wherein the affine information is excluded from the affine candidate list in response to the affine information being identical to at least one affine candidate in the affine candidate list.
17. The method of claim 15, wherein the affine information is excluded from the affine candidate list in response to a difference between the affine information and the at least one affine candidate in the affine candidate list being less than or equal to a threshold.
18. The method according to claim 1, further comprising: An affine list is constructed based on examining at least one of the sub-blocks or blocks that cover a location related to the current video block.
19. The method of claim 18, wherein the affine candidate list comprises at least one of the following: Sub-block Merge list Affine Merge list Affine AMVP List Affine list for GPM, or Affine list for sub-block-based CIIP.
20. The method according to claim 18 or claim 19, further comprising: Determine whether the encoding / decoding tool used for the at least one of the sub-blocks or blocks is an affine encoding / decoding tool.
21. The method of claim 20, further comprising: If the encoding / decoding tool is determined to be an affine-based encoding / decoding tool, then affine information relating to at least one of the sub-blocks or blocks is used to generate affine candidates for the affine candidate list.
22. The method according to claim 18 or claim 19, further comprising: Determine whether the encoding / decoding tool used for the at least one of the sub-blocks or blocks is at least one of a GPM affine encoding / decoding tool or a sub-block-based CIIP encoding / decoding tool.
23. The method of claim 22, further comprising: If it is determined that the encoding / decoding tool is at least one of a GPM-based affine encoding / decoding tool or a sub-block-based CIIP encoding / decoding tool, then affine information related to the sub-block or at least one of the blocks is used to generate affine candidates for the affine candidate list.
24. The method of claims 18 to 23, wherein at least one of the sub-blocks or blocks is examined to generate the affine candidate from the block encoded and decoded by sub-block-based CIIP.
25. The method of claim 24, wherein after examining the inherited affine candidates from the block adjacent to the current block, the at least one of the sub-blocks or blocks is examined to immediately generate the affine candidates from the block encoded and decoded by the sub-block-based CIIP.
26. The method of claim 24, wherein After examining affine candidates inherited from blocks that are not adjacent to the current block, at least one of the sub-blocks or blocks is examined to immediately generate the affine candidates from the block encoded and decoded by CIIP based on the sub-blocks.
27. The method of claim 24, wherein After examining regression-based motion vector field (RMVF) affine candidates, at least one of the sub-blocks or blocks is examined to immediately generate the affine candidates from the blocks encoded and decoded by the sub-blocks using CIIP.
28. The method of claim 24, wherein After examining the constructed affine candidates, at least one of the sub-blocks or blocks is examined to immediately generate the affine candidates from the blocks encoded and decoded using sub-block-based CIIP.
29. The method of claim 24, wherein After examining the affine candidates of temporal inheritance, at least one of the sub-blocks or blocks is examined to immediately generate the affine candidates from the blocks encoded and decoded based on the sub-blocks using CIIP.
30. The method of claim 24, wherein After examining the history-based affine candidates, at least one of the sub-blocks or blocks is examined to immediately generate the affine candidates from the blocks encoded and decoded by the sub-blocks using CIIP.
31. A method for video processing, comprising: For the conversion between the current video block and the bitstream of the video, encoding and decoding of syntax elements related to intra-frame and inter-frame joint prediction (CIIP) are performed based on more than one context; and The transformation is performed based on the syntax elements.
32. The method of claim 31, wherein the syntax element is an index of at least one of CIIP Merge or Merge Pattern with Motion Vector Difference (MMVD).
33. The method of claim 31, wherein the syntax element indicates the use of at least one of the following: CIIP - Location-Related Intra-Prediction Combination (PDPC) CIIP-Template Matching (TM) CIIP based on sub-blocks, or CIIP-MMVD.
34. The method of claim 31, wherein the syntax element is binarized into N binary bits, where N is an integer.
35. The method of claim 34, wherein a first number of binary bits of the N binary bits are encoded and decoded using different contexts, and the remaining binary bits of the N binary bits are encoded and decoded using the same context, wherein the first number is one of the following: 1, 2, 3, 4, 5, 6 or 7.
36. The method of claim 34, wherein a first number of binary bits of the N binary bits are encoded and decoded using a different context, and the remaining binary bits of the N binary bits are bypassed and encoded / decoded, wherein the first number is one of the following: 1, 2, 3, 4, 5, 6 or 7.
37. The method according to any one of claims 31 to 33, wherein at least one context for encoding and decoding the syntax element is determined based on encoding and decoding information.
38. The method of claim 37, wherein the at least one context is determined based on whether a sub-block-based CIIP is used.
39. The method of claim 38, wherein the at least one context of the index of at least one of CIIP Merge or MMVD for encoding / decoding CIIP blocks is determined based on whether CIIP based on sub-blocks is used.
40. The method of claim 38, wherein the at least one context for indicating the use of at least one of the following is determined based on whether sub-block-based CIIP is used: CIIP-PDPC, CIIP-TM, sub-block-based CIIP, or CIIP-MMVD.
41. The method of claim 38, wherein the at least one context is determined based on whether at least one of the following is used: CIIP-PDPC, CIIP-TM, sub-block based CIIP, or CIIP-MMVD.
42. The method of claim 38, wherein the at least one context is determined based on a quantization parameter (QP).
43. The method of claim 38, wherein the at least one context is determined based on the stripe type.
44. The method of claim 37, wherein the encoding / decoding information is associated with at least one neighboring block.
45. The method of claim 44, wherein the at least one context is selected from a candidate set including at least one candidate context.
46. The method of claim 45, wherein the index of the selected candidate context is determined based on at least one neighboring block.
47. The method of claim 46, wherein the index is determined by: Initialize the index to 0; and In response to the availability of at least one neighboring block and the fulfillment of the condition, the index is incremented by 1. The at least one neighboring block includes at least one of the left neighboring block or the upper neighboring block.
48. The method of claim 47, wherein the condition includes that the at least one neighboring block is affine encoded.
49. The method of claim 48, wherein the at least one neighboring block encoded in affine-advanced motion vector prediction (AMVP) mode is considered to be affine encoded.
50. The method of claim 48, wherein the at least one neighboring block encoded in affine-Merge mode is considered to be affine encoded.
51. The method of claim 48, wherein the at least one neighboring block for which CIIP based on the sub-block is true is considered to be affine-coded.
52. The method of claim 46, wherein the syntax element is a CIIP flag based on a sub-block.
53. The method of claim 31, wherein at least one context for the affine tag of the block is determined based on whether at least one neighboring block of the block is encoded or decoded in a sub-block-based CIIP mode.
54. The method of claim 53, wherein the at least one context is selected from a candidate set including at least one candidate context.
55. The method of claim 54, wherein the at least one context is selected by: In response to the availability of at least one neighboring block and the fulfillment of the condition, the index of the selected candidate context is incremented by 1; and The at least one context is selected based on the index. The index is initialized to 0, and The at least one neighboring block includes at least one of the left neighboring block or the upper neighboring block.
56. The method of claim 55, wherein the condition includes that the at least one neighboring block is affine encoded.
57. The method of claim 56, wherein the at least one neighboring block encoded in affine-advanced motion vector prediction (AMVP) mode is considered to be affine encoded.
58. The method of claim 56, wherein the at least one neighboring block encoded in affine-Merge mode is considered to be affine encoded.
59. The method of claim 56, wherein the at least one neighboring block for which CIIP based on the sub-block is true is considered to be affine-coded.
60. The method of claim 54, wherein the at least one context is selected based on the number of available affine-coded neighboring blocks.
61. The method of claim 60, wherein if the number of available affine-coded neighboring blocks is less than or equal to a threshold, the first context is used, and If the number is greater than the threshold, then the second context is used. The threshold mentioned therein is an integer.
62. The method of claim 61, wherein the threshold is 0, 1, 2, 3, 4, 5 or 6.
63. The method of claim 60, wherein the available affine-coded neighbor blocks comprise five neighbor blocks or seven neighbor samples.
64. The method of claim 60, wherein the at least one neighboring block encoded in affine-advanced motion vector prediction (AMVP) mode is considered to be affine encoded.
65. The method of claim 60, wherein the at least one neighboring block encoded in affine-Merge mode is considered to be affine encoded.
66. The method of claim 60, wherein the at least one neighboring block for which CIIP based on the sub-block is true is considered to be affine-coded.
67. The method according to any one of claims 31 to 33, wherein the maximum value of the syntax element is determined based on whether sub-block-based CIIP is used.
68. The method of claim 67, wherein if the syntax element is encoded or decoded using at least one of rounding unary code or rounding binary code, the last valid codeword of the syntax element is determined based on the maximum value.
69. The method of claim 68, wherein the maximum value of the syntax element is indicated separately for affine-coded blocks and non-affine-coded blocks in at least one of the following: Video Parameter Set (VPS) Sequence Parameter Set (SPS) Image Parameter Set (PPS) Image header, or Strip head.
70. The method of claim 67, wherein the maximum value is determined based on the number of available affine-coded neighboring blocks.
71. The method of claim 70, wherein if the number of available affine-coded neighboring blocks is less than or equal to a threshold, the first context is used, and If the number is greater than the threshold, then the second context is used. The threshold mentioned therein is an integer.
72. The method of claim 71, wherein the threshold is 0, 1, 2, 3, 4, 5 or 6.
73. The method of claim 70, wherein the available affine-coded neighbor blocks comprise five neighbor blocks or seven neighbor samples.
74. The method of claim 70, wherein the at least one neighboring block encoded in affine-advanced motion vector prediction (AMVP) mode is considered to be affine encoded.
75. The method of claim 70, wherein the at least one neighboring block encoded in affine-Merge mode is considered to be affine encoded.
76. The method of claim 70, wherein the at least one neighboring block for which CIIP based on the sub-block is true is considered to be affine-coded.
77. The method according to any one of claims 1 to 76, wherein the current video block or video unit comprises at least one of the following: Color components, sub-images, stripes, slices, codec tree units (CTUs), CTU rows, CTU groups, codec units (CUs), prediction units (PUs), transform units (TUs), codec tree blocks (CTBs), codec blocks (CBs), prediction blocks (PBs), transform blocks (TBs), blocks, sub-blocks of blocks, sub-regions within blocks, or regions containing more than one sample point or pixel.
78. The method according to any one of claims 1 to 77, wherein an indication of whether and / or how the method is applied is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.
79. The method according to any one of claims 1 to 77, wherein an indication of whether and / or how to apply the method is indicated in one of the following: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.
80. The method according to any one of claims 1 to 77, wherein whether and / or how the method is applied is based on at least one of the following: Messages that include one of the following: Dependency Parameter Set (DPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Picture Header, Strip Header, Slice Header, Maximum Codec Unit (LCU), Codec Unit (CU), LCU Row, LCU Group, Transform Unit (TU), Prediction Unit (PU) Block, or Video Codec Unit. The location of CU, PU, TU, block, or the video codec unit. The block dimension of the current video block and / or the block dimensions of the neighboring blocks of the current video block, The block shape of the current video block and / or the block shape of the neighboring blocks of the current video block, Block encoding / decoding modes, Indicators of color format, Encoder tree structure, Strip, slice group type and / or image type, Color components, Temporal layer identifier (ID), or Standard grade, level, or tier.
81. The method of claim 80, wherein the encoding / decoding mode includes one of the following: intra-block copy (IBC), non-IBC inter-frame mode, or non-IBC sub-block mode, or The color format mentioned includes one of the following: 4:2:0 or 4:4:
4.
82. The method according to any one of claims 1 to 81, wherein the conversion comprises encoding the current video block into the bitstream.
83. The method according to any one of claims 1 to 81, wherein the conversion comprises decoding the current video block from the bitstream.
84. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 83.
85. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 83.
86. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: The affine information of the current video block is determined based on the first affine information of the first block of the video, and the first block is encoded and decoded using a first encoding and decoding tool that is different from the affine encoding and decoding tool. The prediction of the current video block is determined based on the affine information; as well as The bitstream is generated based on the prediction.
87. A method for storing a bitstream of video, comprising: The affine information of the current video block is determined based on the first affine information of the first block of the video, wherein the first block is encoded and decoded using a first encoding and decoding tool that is different from the affine encoding and decoding tool; The prediction of the current video block is determined based on the affine information; The bitstream is generated based on the prediction; as well as The bitstream is stored in a non-transitory computer-readable recording medium.
88. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: Encode and decode syntax elements related to intra-frame and inter-frame joint prediction (CIIP) based on more than one context; as well as The bitstream is generated based on the syntax elements.
89. A method for storing a bitstream of video, comprising: Encode and decode syntax elements related to intra-frame and inter-frame joint prediction (CIIP) based on more than one context; The bitstream is generated based on the syntax elements; as well as The bitstream is stored in a non-transitory computer-readable recording medium.