Method and device for video processing and medium
By introducing a geometric segmentation mode into video encoding and decoding, the prediction and encoding process of video blocks is optimized, solving the problem of improving encoding and decoding efficiency and performance in existing technologies, and achieving more efficient video encoding and decoding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DOUYIN CO LTD
- Filing Date
- 2024-09-25
- Publication Date
- 2026-04-21
AI Technical Summary
Existing video coding and decoding technologies have room for improvement in terms of coding and decoding efficiency and performance. In particular, in multi-functional video coding and decoding standards, existing motion prediction modes are unable to effectively utilize the geometric segmentation information of video blocks.
The Geometric Partitioning Model (GPM) is used to predict video blocks. By determining the transformation between video units and bitstreams, including maximum value indexing, syntax element skipping, mixed candidate list and affine motion process, the video encoding and decoding process is optimized.
It improves the efficiency and performance of video encoding and decoding, reducing data transmission and improving decoding quality through more accurate motion prediction and more efficient encoding methods.
Smart Images

Figure CN121909651A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure generally relate to video processing techniques, and more specifically, to representations in geometric prediction patterns. Background Technology
[0002] Today, digital video capabilities are being applied to all aspects of people's lives. Various video compression technologies have been proposed for video encoding / decoding, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-TH.265 High Efficiency Video Codec (HEVC) standard, and Multifunctional Video Codec (VVC) standard. However, the overall expectation is to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the Invention
[0003] Embodiments of this disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method includes: for a conversion between video units and a video bitstream, performing a first determination regarding whether a sub-block-based prediction is applicable to the video unit encoded and decoded using a geometrically segmented mode (GPM) pattern, wherein a true result of the first determination indicates that the sub-block-based prediction is applicable to the video unit, and a false result indicates that the sub-block-based prediction is not applicable to the video unit; and performing a conversion based on the first determination. In this manner, it can improve encoding / decoding efficiency and performance.
[0005] In a second aspect, another method for video processing is proposed. This method includes: a conversion between video units and the video bitstream; determining, for a video unit encoded using Geometric Partitioning Mode (GPM), at least one of the following: the maximum value of the index transmitted via signaling in the GPM; whether syntax elements for the GPM are skipped; the manner of transmitting GPM merge candidate indices via signaling; a mixing candidate list; or an affine motion process for the components of the video unit; and performing the conversion based on the determination. In this way, it can improve encoding / decoding efficiency and performance.
[0006] In a third aspect, an apparatus for video processing is proposed. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform a method according to the first or second aspect of this disclosure.
[0007] In a fourth aspect, a non-transitory computer-readable storage medium is proposed. This non-transitory computer-readable storage medium stores instructions that cause a processor to execute a method according to the first or second aspect of this disclosure.
[0008] In a fifth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: performing a first determination regarding whether a sub-block-based prediction is applicable to video units of a video encoded and decoded using a geometrically segmented mode (GPM) pattern, wherein a true result of the first determination indicates that the sub-block-based prediction is applicable to the video units, and a false result indicates that the sub-block-based prediction is not applicable to the video units; and generating a bitstream based on the first determination.
[0009] In a sixth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining at least one of the following for a video unit encoded using Geometric Partitioning Mode (GPM): the maximum value of an index transmitted via signaling in the GPM, whether syntax elements for the GPM are skipped, the manner of transmitting GPM merge candidate indices via signaling, a mixing candidate list, or an affine motion process for components of the video unit; and generating a bitstream based on the determined value.
[0010] In a seventh aspect, a method for storing a bitstream of video is proposed. The method includes: performing a first determination regarding whether a sub-block-based prediction is applicable to video units of a video encoded and decoded using a geometrically segmented mode (GPM) pattern, wherein a true result of the first determination indicates that the sub-block-based prediction is applicable to the video units, and a false result indicates that the sub-block-based prediction is not applicable to the video units; generating a bitstream based on the first determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0011] In the eighth aspect, a method for storing a bitstream of video is proposed. The method includes: determining at least one of the following for a video unit encoded using Geometric Partitioning Mode (GPM): the maximum value of an index transmitted via signaling in the GPM, whether syntax elements for the GPM are skipped, the manner of transmitting GPM merge candidate indices via signaling, a mixing candidate list, or an affine motion process for the components of the video unit; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0012] This summary aims to present, in a simplified form, the concept choices further described below in the detailed embodiments. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0013] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0014] Figure 1 A block diagram illustrating an example video codec system according to some embodiments of the present disclosure is shown; Figure 2 A block diagram illustrating a first example video encoder according to some embodiments of the present disclosure is shown; Figure 3 A block diagram illustrating an example video decoder according to some embodiments of the present disclosure is shown; Figure 4 This shows the locations of spatial and temporal neighbor blocks used in the construction of the AMVP / Merge candidate list; Figure 5 The locations of non-adjacent candidates in the ECM are shown; Figure 6A and Figure 6B An affine motion model based on control points is shown; Figure 7 An example affine MVF for each sub-block is shown; Figure 8 The location of the inherited affine motion prediction value is shown; Figure 9 This demonstrates the inheritance of control point motion vectors; Figure 10 The locations of candidate positions for the constructed affine Merge pattern are shown; Figure 11A and Figure 11B The spatial nearest neighbor used to derive the affine Merge candidate is shown; Figure 12 The affine Merge candidates from non-nearest neighbors to the constructed ones are shown; Figure 13 An example of generating HAPC is shown; Figure 14 A diagram illustrating the regression-based affine Merge candidate derivation is shown. Figure 15 The template matching performance is shown over the search area around the initial MV; Figure 16 The template and the corresponding reference template are shown; Figure 17 A template and a reference template are shown for a block with sub-block motion that utilizes motion information of the current block's sub-blocks; Figure 18The derivation of the sub-CU motion field obtained by applying motion shift based on neighbor motion information is shown; Figure 19 An example of GPM partitioning grouped at the same angle is shown; Figure 20 The unidirectional prediction MV selection for geometric segmentation patterns is shown; Figure 21 An exemplary generation of the bending weight w_0 using a geometric segmentation pattern is shown; Figure 22 The ramp function for weighted GPM mixing is shown based on the displacement (d) from the predicted sample location to the GPM segmentation boundary and the mixing region size (τ). Figures 23A-23C The available IPM candidates are shown respectively; Figure 23D GPM with inter-frame and intra-frame prediction is shown; Figure 24 The edges on the template are shown; Figure 25 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown; Figure 26 A flowchart of a method for video processing according to embodiments of the present disclosure is shown; and Figure 27 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.
[0015] In all accompanying drawings, the same or similar reference numerals usually refer to the same or similar elements. Detailed Implementation
[0016] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.
[0017] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0018] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, it is claimed that, whether explicitly described or not, such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.
[0019] It should be understood that although the terms “first” and “second”, etc., may be used herein to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0020] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0021] Example Environment Figure 1 This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0022] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.
[0023] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded images and associated data. An encoded image is an encoded representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data can be directly transmitted to destination device 120 via network 130A through I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by destination device 120.
[0024] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.
[0025] The video encoder 114 and the video decoder 124 can operate according to video compression standards such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or future standards.
[0026] Figure 2 This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... Figure 1 An example of a video encoder 114 in system 100 is shown.
[0027] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. Figure 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0028] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.
[0029] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, in which at least one reference picture is the picture in which the current video block is located.
[0030] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... Figure 2 The examples are shown separately.
[0031] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0032] The mode selection unit 203 can, for example, select one of several coding modes (intra-coding or inter-coding) based on the error result, and provide the resulting intra-coded or inter-coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-inter-prediction joint prediction (CIIP) mode, in which prediction is based on inter-prediction signals and intra-prediction signals. In the case of inter-prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).
[0033] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.
[0034] Motion estimation unit 204 and motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks independent of macroblocks within the same image.
[0035] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0036] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search reference images in list 0 to find a reference video block for the current video block, and can also search reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference images containing multiple reference video blocks in lists 0 and 1, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. Motion estimation unit 204 can output the multiple reference indices and multiple motion vectors of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.
[0037] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0038] In one example, motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0039] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0040] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0041] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0042] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.
[0043] In other examples, such as in skip mode, residual data for the current video block may not exist, and residual generation unit 207 may not perform a subtraction operation.
[0044] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0045] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0046] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current video block for storage in the buffer 213.
[0047] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0048] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0049] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.
[0050] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0051] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.
[0052] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, which involves deriving several most likely candidates based on data from neighboring PBs and reference pictures. Motion information typically includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially or temporally neighboring blocks.
[0053] The motion compensation unit 302 can generate motion compensation blocks, possibly by performing interpolation based on an interpolation filter. Identifiers for interpolation filters used with sub-pixel precision can be included in the syntax elements.
[0054] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of a video block to calculate the interpolated values of sub-integer pixels for the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate a prediction block.
[0055] Motion compensation unit 302 may use at least some of the syntax information to determine the size of the blocks used to encode the encoded video sequence (multiple frames) and / or (multiple stripes), segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some respects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be an entire image or a region of an image.
[0056] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.
[0057] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0058] Some exemplary embodiments of this disclosure will be described in detail below. It should be noted that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section. Furthermore, although some embodiments are described with reference to multi-function video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Furthermore, although some embodiments describe video encoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. Additionally, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another or at different compression bitrates. 1. Summary of the Invention This disclosure relates to video coding and decoding techniques. Specifically, it concerns affine motion prediction methods in video coding and decoding. These ideas can be applied individually or in various combinations to any standard or non-standard video codec.
[0060] 2. Introduction The exponential growth of multimedia data has posed significant challenges to video encoding and decoding. To meet the ever-increasing demand for more efficient compression technologies, the ITU-T and ISO / IEC have developed a series of video encoding and decoding standards over the past few decades. Specifically, the ITU-T developed the H.261 and H.263 standards, and ISO / IEC developed the MPEG-1 and MPEG-4 visual standards. These two organizations have jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, the H.265 / HEVC standard, and the latest VVC standard. Starting with H.262 / MPEG-2, a hybrid video encoding and decoding framework has been adopted, in which intra / inter-frame prediction plus transform coding and decoding are used. Figure 4 The location of spatial and temporal neighbor blocks used in the construction of the AMVP / Merge candidate list is shown.
[0061] 2.1 MVP in Video Encoding and Decoding Inter-frame prediction aims to eliminate temporal redundancy between adjacent frames, an indispensable component in hybrid video codec frameworks. Specifically, inter-frame prediction utilizes the content specified by motion vectors (MVs) as the predicted version of the current block to be encoded / decoded, thus only the residual signal and motion information are transmitted in the bitstream. To reduce the cost of MV signaling, motion vector prediction (MVP) emerged as an efficient mechanism for conveying motion information. Early strategies simply used the MV of a specified neighboring block or the median MV of neighboring blocks as the MVP. H.265 / HEVC introduced a contention mechanism where rate-distortion optimization (RDO) selects the best MVP from multiple candidates. Specifically, Advanced MVP (AMVP) mode and Merge mode with different motion information signaling strategies were designed. Using AMVP mode, a reference index, an MVP candidate index referencing the AMVP candidate list, and motion vector difference (MVD) are transmitted via signaling. Regarding Merge mode, only the Merge index referencing the Merge candidate list is transmitted via signaling, and all motion information associated with the Merge candidate is inherited. Both the AMVP and Merge modes require building an MVP candidate list, and the details of the building process for these two modes are described below.
[0062] AMVP mode: AMVP utilizes the spatial-temporal correlation of motion vectors with neighboring blocks, which is used for explicit transfer of motion parameters. For each list of reference images, the motion vector candidate list is constructed by first checking the availability of temporally adjacent positions to the left and top, removing redundant candidates, and adding zero vectors to ensure the candidate list has a constant length. For the spatial motion vector candidate derivation, two motion vector candidates are ultimately based on those located in... Figure 4 The motion vectors of the five blocks at different locations shown are derived. The five neighboring blocks located at B0, B1, B2 and A0, A1 are classified into two groups, where group A includes the three spatial neighboring blocks above and group B includes the two spatial neighboring blocks to the left. Two MV candidates are derived respectively using the first available candidates from group A and group B in a predefined order. For the derivation of temporal motion vector candidates, as follows... Figure 4 As shown, a motion vector candidate is derived based on two distinct co-positions examined sequentially (bottom right (C0) and center (C1)). To avoid redundant MV candidates, duplicate motion vector candidates in the list are discarded. If the number of potential candidates is less than 2, additional zero motion vector candidates are added to the list. Figure 5 The positions of non-adjacent candidates in the ECM are shown.
[0063] Merge modeSimilar to the AMVP model, the MVP candidate list for the Merge model also includes spatial and temporal candidates. For spatial motion vector candidate derivation, after performing availability and redundancy checks, a maximum of four candidates are selected, in the order A1, B1, B0, A0, and B2. For temporal Merge candidate (TMVP) derivation, a candidate is selected from at most two temporally neighboring blocks (C0 and C1). When there are not enough Merge candidates using both spatial and temporal candidates, combined bidirectional prediction Merge candidates and zero MV candidates are added to the MVP candidate list. Once the number of available Merge candidates reaches the maximum allowed number for signal transmission, the Merge candidate list construction process is terminated.
[0064] In VVC, the process of constructing the Merge pattern is further improved by introducing a history-based MVP (HMVP), where the HMVP incorporates motion information from previously encoded / decoded blocks that can be far removed from the current block. In VVC, HMVP Merge candidates are appended to the Merge list, following the Spatial MVP and TMVP. In this method, motion information from previously encoded / decoded blocks is stored in a table and used as the MVP for the current CU. During the encoding / decoding process, the table with multiple HMVP candidates is maintained using a first-in, first-out (FIFO) strategy. Whenever a non-sub-block inter-encoding / decoding CU is present, the associated motion information is added to the last entry of the table as a new HMVP candidate.
[0065] During the standardization of VVC, non-adjacent MVPs were proposed to facilitate better motion information derivation by utilizing non-adjacent regions. In ECM software, non-adjacent MVPs are inserted between TMVPs and HMVPs, where the distance between the non-adjacent spatial candidate and the current codec block is based on the width and height of the current codec block, such as... Figure 5 As shown.
[0066] 2.2 Affine Motion Compensation Prediction In HEVC, only the translational motion model is applied for motion compensation prediction (MCP). In the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transformation motion compensation prediction is applied. Figure 6A and Figure 6B An affine motion model based on control points is shown. For example... Figure 6A and Figure 6B As shown, the affine motion field of a block is described by motion information from two control points (4 parameters) or three control point motion vectors (6 parameters).
[0067] For the 4-parameter affine motion model, the motion vector at the sample point position (x, y) in the block is derived as follows:
[0068] For the 6-parameter affine motion model, the motion vector at the sample point position (x, y) in the block is derived as follows:
[0069] in( mv 0x ,mv 0y ) is the motion vector of the upper left control point, ( mv 1x ,mv 1y ) is the motion vector of the upper right control point, and ( mv 2x ,mv 2y ) is the motion vector of the lower left control point.
[0070] To simplify motion compensation prediction, block-based affine transformation prediction is applied. To derive the motion vector for each 4×4 lumen sub-block, the motion vector of the center sample point of each sub-block is calculated according to the above equation, such as... Figure 7 As shown, the values are rounded to 1 / 16 fractional precision. A motion-compensated interpolation filter is then applied to generate a prediction for each sub-block with a derived motion vector. The sub-block size for the chroma components is also set to 4×4. The MV of the 4×4 chroma sub-block is calculated as the average of the MV of the upper-left luminance sub-block and the lower-right luminance sub-block in the corresponding 8×8 luminance region.
[0071] Similar to the inter-frame prediction for translational motion, there are two other inter-frame prediction modes for affine motion: affine Merge mode and affine AMVP mode.
[0072] 2.2.1 Affine Merge Prediction The Affine Merge pattern can be applied to CUs with a width and height greater than or equal to 8. In this pattern, the CPVM of the current CU is generated based on the motion information of spatially neighboring CUs. There can be up to five CPVM candidates, and a signal transmission index indicates which candidate should be used for the current CU. In VVC, the following three types of CPVM candidates are used to form the Affine Merge candidate list: Inherited affine Merge candidates inferred from the CPMV of neighboring CUs; Affine Merge candidate CPMVP constructed using translational MV derivation of neighboring CUs; Zero MV.
[0073] In VVC, there are at most two inherited affine candidates, which are derived from the affine motion model of neighboring blocks: one from the left neighboring CU and one from the upper neighboring CU. Candidate blocks are as follows: Figure 8 As shown. For the left-hand predicted value, the scan order is A0->A1, and for the upper-hand predicted value, the scan order is B0->B1->B2. Only candidates from the first inheritance on each side are selected. Deduplication checks between candidates from two inheritances are not performed. When a neighboring affine CU is identified, its control point motion vector is used to derive the CPMVP candidate in the affine Merge list of the current CU. Figure 9 As shown, if the adjacent lower-left block A is encoded and decoded using affine mode, the motion vectors of the upper-left, upper-right, and lower-left corners of the CU containing block A are obtained. When block A is encoded and decoded using a 4-parameter affine model, according to Calculate the two CPMVs of the current CU. When block A is encoded and decoded using a 6-parameter affine model, according to... Calculate the three CPMVs of the current CU.
[0074] Figure 8 The location of the inherited affine motion prediction value is shown. Figure 9 The inheritance of control point motion vectors is shown.
[0075] The constructed affine candidate means that the candidate is constructed by combining the translational motion information of the neighbors of each control point. The motion information for the control point is derived from... Figure 10 The derivation is shown among the specified spatial and temporal nearest neighbors. CPMVk (k=1,2,3,4) represents the k-th control point. For CPMV1, the B2->B3->A2 block is checked, and the MV of the first available block is used. For CPMV2, the B1->B0 block is checked, and for CPMV3, the A1->A0 block is checked. If the TMVP is available, it is used as CPMV4.
[0076] After the motion signatures (MVs) of the four control points are obtained, affine merge candidates are constructed based on this motion information. The following combinations of control point MVs are used for sequential construction: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3} Combinations of three CPMVs construct a 6-parameter affine merge candidate, and combinations of two CPMVs construct a 4-parameter affine merge candidate. To avoid motion scaling, combinations of control point MVs are discarded if the reference indices of the control points are different. Figure 10 The locations of candidate positions for the constructed affine Merge pattern are shown.
[0077] After the inherited affine Merge candidates and the constructed affine Merge candidates are checked, if the list is still not full, a zero MV is inserted at the end of the list.
[0078] 2.2.2 Affine AMVP Prediction The affine AMVP mode can be applied to CUs with a width and height greater than or equal to 16. An affine flag at the CU level is signaled in the bitstream to indicate whether the affine AMVP mode is used, and another flag is signaled to indicate whether it is a 4-parameter affine or a 6-parameter affine. In this mode, the difference between the current CU's CPVM and its predicted CPMVP is signaled in the bitstream. The affine AMVP candidate list is of size 2 and is generated by sequentially using the following four types of CPVM candidates: Inherited affine AMVP candidates inferred from the CPMV of neighboring CUs; A constructive affine AMVP candidate CPMVP derived using translational MV of neighboring CUs; Translation MV from the neighboring CU; Zero MV.
[0079] The checking order for inherited affine AMVP candidates is the same as that for inherited affine Merge candidates. The only difference is that, for AVMP candidates, only affine CUs with the same reference image as those in the current block are considered. Deduplication is not applied when inserting inherited affine motion predictions into the candidate list.
[0080] The constructed AMVP candidate is from Figure 10 The derivation is based on the specified spatial nearest neighbors. The same checking order as in the affine Merge candidate construction is used. Additionally, the reference picture index of neighboring blocks is also checked. The first block in the checking order to be inter-coded and has the same reference picture as in the current CU is used. The current CU uses 4-parameter affine mode encoding and decoding, and... mv0 and mv1 If all three CPMVs are available, they are added as candidates in the affine AMVP list. If the current CU uses 6-parameter affine mode encoding / decoding and all three CPMVs are available, they are added as candidates in the affine AMVP list. Otherwise, the constructed AMVP candidates are set to unavailable.
[0081] Figure 11A and Figure 11BThe spatial nearest neighbors used to derive affine Merge candidates are shown: Figure 11A Used to derive affine Merge candidates for inheritance, and Figure 11B Used to derive the constructed affine Merge candidate.
[0082] If the number of affine AMVP candidates is still less than 2 after the insertion of valid inherited and constructed AMVP candidates, then when available, mv 0、 mv 1 and mv 2 will be added sequentially as translation MVs to predict all control point MVs for the current CU. Finally, if the affine AMVP list is still not full, zero MVs are used to populate the affine AMVP list.
[0083] 2.2.3 New Affine Candidate Derivation Method In ECM-6.0, three additional affine Merge and AMVP candidate derivation methods are integrated: non-adjacent spatial domain candidates, historical parameter-based candidates, and regression-based affine candidates.
[0084] 2.2.3.1 Non-adjacent airspace candidates In ECM-6.0, non-adjacent airspace nearest neighbors are studied to provide candidates for both affine Merge and affine AMVP. The pattern for obtaining non-adjacent airspace candidates is... Figure 11A and Figure 11B As shown in the diagram, similar to the non-adjacent regular merge candidates, the distance between non-adjacent spatial candidates and the current codec block is also defined based on the width and height of the current CU.
[0085] Figure 11A and Figure 11B Motion information of non-adjacent spatial neighbors is used to generate additional inheritance and construct affine merge candidates. Specifically, to generate inheritance candidates, non-adjacent spatial neighbors are checked based on their distance from the current block (i.e., from nearest to farthest). At a specific distance, only the first available neighbors encoded in affine mode from each side of the current block (e.g., left and top) are included. Figure 11A As shown, the checks of the left and top nearest neighbors are performed from bottom to top and from right to left, respectively. For the constructed candidates, such as... Figure 11B As shown, firstly, the positions of non-adjacent spatial neighbors on the left and top are independently determined; then, the positions of the upper left neighbors can be determined accordingly to form a rectangular virtual block together with the non-adjacent neighbors on the left and top. Figure 12The diagram illustrates the affine Merge candidates from non-nearest neighbors to the build. Motion information from three non-nearest neighbors is used to form a CPMV at the top-left (A), top-right (B), and bottom-left (C) of the virtual block, which is projected onto the current CU to generate the corresponding build candidates, as shown below. Figure 12 As shown.
[0086] 2.2.3.2 Affine Candidates Based on Historical Parameters History-based Affine Model Inheritance (HAMI) allows affine models to be inherited from previously affine-encoded blocks that may not be adjacent to the current block. A History Parameter Table (HPT) is established. Each entry in the HPT stores a set of affine parameters: a, b, c, and d, each represented by a 16-bit signed integer. Entries in the HPT are categorized by reference lists and reference indices. Each reference list in the HPT supports five reference indices. The HPT category (denoted as HPTCat) is calculated in a formulaic manner.
[0087] Here, RefList and RefIdx represent the list of reference images (0 or 1) and the reference index, respectively. A maximum of seven entries can be stored for each category, resulting in a total of 70 entries in the HPT. At the beginning of each CTU line, the number of entries for each category is initialized to zero. After decoding the affine-encoded CU with reference lists RefListcur and RefIdxcur, the affine parameters are used to update the entries in category HPTCat(RefListcur, RefIdxcur) in a manner similar to HMVP table updates.
[0088] Candidates based on historical affine parameters (HAPC) from... Figure 10 The MV of the neighboring 4×4 blocks, denoted as A0, A1, B0, B1, or B2, and a set of affine parameters in the corresponding entries stored in the HPT are derived. The MV of the neighboring 4×4 blocks is used as the base MV. Formulatically, the MV of the current block at position (x, y) is calculated as: (4), in( , ) represents the MV of the nearest 4×4 block, (x base y base (x, y) represents the center position of the nearest 4×4 block. (x, y) can be the top left, top right, and bottom left corners of the current block to obtain the corner position MV (CPMV) for the current block, or it can be the center of the current block to obtain the regular MV for the current block.
[0089] Figure 13An example of how to derive an HAPC from block A0 is shown. The affine parameters {a0, b0, c0, d0} are directly obtained from an entry in the category HPTIdx(RefListA0, refIdx0A0) in the HPT. The affine parameters from the HPT (with the center position of A0 as the base position and the MV of block A0 as the base MV) are used together to derive the CPMV for either the affine MergeHAPC or the affine AMVP HAPC. They can also be used to derive the MV located at the center of the current block as a regular Merge candidate. The HAPC can be placed into the sub-block-based Merge candidate list, the affine AMVP candidate list, or the regular Merge candidate list. In response to the introduction of new HAPCs, the size of the sub-block-based Merge candidate list is increased from 5 to 10 and 12 for random access and low-latency B configurations, respectively. Furthermore, for the random access configuration, the size of the regular Merge candidate list is increased from 10 to 11 to accommodate the newly added regular Merge candidates.
[0090] 2.2.3.3 Regression-based Affine Candidates In ECM-6.0, regression-based affine merge candidates are derived and added to the affine merge list. The sub-block motion fields from previously encoded and decoded affine CUs and the motion information of neighboring sub-blocks from the current CU are used as inputs to the regression process to derive the proposed affine candidates.
[0091] Previously encoded and decoded affine CUs can be identified from scans of non-adjacent positions and the affine HMVP table. For example... Figure 14 As shown, the information of the adjacent sub-blocks of the current CU is obtained from the 4x4 sub-blocks represented by the gray area. For each sub-block, given a reference list, the corresponding motion vector and center coordinates of the sub-block can be used.
[0092] For each affine CU, at most two affine candidates can be derived. One has neighboring subblock information, and the other does not. All candidates generated by linear regression are deduplicated and collected into a candidate subgroup. When ARMC is enabled, the ARMC process based on TM cost is applied. Subsequently, when N affine CUs are found, at most N candidates generated by linear regression are added to the affine merge list. Figure 14 This is a schematic diagram of the affine Merge candidate derivation based on regression.
[0093] 2.3 Template Matching Merge / AMVP Pattern in ECM Template Matching (TM) Merge / AMVP mode is a decoder-side MV derivation method used to refine the motion information of the current CU by finding the closest match between a template in the current image (i.e., the top and / or left neighboring blocks of the current CU) and a block in the reference image (i.e., the block of the same size as the template). Figure 15 As shown, within the search range of [-8, +8] pixels, a better MV is searched around the initial motion of the current CU. Figure 15 This demonstrates template matching execution over the search area surrounding the initial MV.
[0094] In AMVP mode, an MVP candidate is determined based on the template matching error, selecting the one that minimizes the difference between the current block and the reference block template. Then, the TM process performs MV refinement only on that specific MVP candidate. The TM refines the MVP candidate using an iterative diamond search, starting with full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode) within a search range of [-8, +8] pixels. AMVP candidates can also be refined using a cross search with full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode), followed by half-pixels and quarter-pixels sequentially depending on the AMVR mode. This search process ensures that the MVP candidate maintains the same MV precision as indicated by the Adaptive Motion Vector Resolution (AMVR) mode after the TM process.
[0095] In Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. TMMerge can proceed up to 1 / 8 pixel MVD accuracy, or skip those accuracies beyond half-pixel MVD accuracy, depending on whether an alternative interpolation filter is used based on the merged motion information (i.e., used when AMVR is in half-pixel mode). Furthermore, when TM mode is enabled, template matching can operate as a standalone process, or as an additional MV refinement process between a block-based bilateral matching (BM) method and a sub-block-based bilateral matching method, depending on whether BM is enabled according to its enable condition check. When both BM and TM are enabled for CU, the TM search process stops at half-pixel MVD accuracy, and the resulting MV is further refined using the same model-based MVD derivation method as in DMVR.
[0096] 2.4 Adaptive Reordering of Merge Candidates (ARMC) Inspired by the spatial correlation between reconstructed neighboring pixels and the current codec block, we propose Adaptive Reordering of Merge Candidates (ARMC) to refine the order of candidates in a given candidate list. The basic assumption is that candidates with lower template matching costs have a higher probability of being selected through the RDO process and should therefore be placed earlier in the list to reduce signaling costs.
[0097] The reordering method is applied to the regular Merge pattern, Template Matching (TM) Merge pattern, and Affine Merge pattern (excluding SbTMVP candidates). For the TM Merge pattern, Merge candidates are reordered before the refinement process.
[0098] After the Merge candidate list is constructed, the Merge candidates are divided into several subgroups. The subgroup size is set to 5. The Merge candidates in each subgroup are reordered in ascending order based on the cost value of template matching. For simplicity, the Merge candidates in the last subgroup (not the first subgroup) are not reordered.
[0099] Template matching cost is measured by the sum of absolute differences (SAD) between the samples of the current block's template and the samples of its corresponding reference template. For example... Figure 16 As shown, the template includes a set of reconstructed samples adjacent to the current block, while the reference template is located using the same motion information as the current block. When the merge candidate utilizes bidirectional prediction, the reference samples of the merge candidate's template are also generated through bidirectional prediction.
[0100] For sub-block-based merge candidates with sub-block size equal to Wsub * Hsub, the upper template includes several sub-templates of size Wsub × K, and the left template includes several sub-templates of size K × Hsub. For example... Figure 17 As shown, the motion information of the sub-blocks in the first row and first column of the current block is used to derive the reference sample points of each sub-template.
[0101] 2.5 Sub-block-based temporal motion vector prediction (SbTMVP) VVC supports the Sub-Block-Based Temporal Motion Vector Prediction (SbTMVP) method. Similar to TMVP, SbTMVP leverages motion fields in co-located images to facilitate more accurate MVP derivation. The same co-located image used by TMVP is used for SbTVMP. SbTMVP differs from TMVP primarily in two ways. First, SbTMVP enables motion prediction at the sub-CU level, while TMVP predicts CU-level motion. Second, compared to TMVP, which obtains temporal MVs from co-located blocks in the co-located image (co-located blocks are the lower right or center blocks relative to the current CU), SbTMVP applies motion shifting before obtaining temporal motion information from the co-located image. This motion shifting is achieved by reusing the MV of one of the spatially neighboring blocks from the current CU. Figure 16 The template and the corresponding reference template are shown.
[0102] Figure 18 The derivation of the sub-block level motion field for SbTMVP is shown. Specifically, the motion information of the lower left sub-block A1 is acquired first. If any MV in reference list 0 and list 1 points to the same frame, the corresponding MV will be identified as a motion shift. Otherwise, zero MV will be used as a motion shift.
[0103] Once the motion shift is determined, a designated region within the same frame is used to derive the sub-block level motion field. Assuming... Figure 15 As shown, the motion of A1 is used as motion shift. Then, for each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the center sample point) in the co-location image is obtained to provide motion information, wherein the MV scaling operation is first performed to align the reference frame of the temporal motion vector with the reference frame of the current CU. Figure 17 Templates and reference templates are shown for blocks with sub-block motion that use motion information of the current block's sub-blocks. Figure 18 The derivation of the sub-CU motion field obtained by applying motion displacement based on neighbor motion information is shown.
[0104] In VVC and ECM, in addition to the CU-level MVP candidate list, a sub-CU-level MVP candidate list is also constructed to provide more accurate motion predictions for the current CU. This sub-CU-level MVP candidate list includes the motion field generated by both the SbTMVP and AFFINE methods. Specifically, only one SbTMVP candidate is included, and this SbTMVP candidate is always placed as the first entry in the constructed sub-CU-level MVP candidate list. After performing template matching-based reordering, multiple AFFINE candidates are included in the list, with those having lower costs placed earlier.
[0105] 2.6 Geometric Partitioning (GPM) In VVC, geometric segmentation modes are supported for inter-frame prediction. Geometric segmentation modes are transmitted via signaling using CU-level flags as a merge mode, where other merge modes include regular merge mode, MMVD mode, CIIP mode, and sub-block merge mode. This applies to each possible CU size. (in Excluding 8x64 and 64x8), the geometric segmentation mode supports a total of 64 segments.
[0106] When this mode is used, the CU is divided into two geometric segments by a straight line of geometric positioning. Figure 19 The position of the segmentation line is mathematically derived from the angle and offset parameters of a specific segmentation. Each part of the geometric segmentation in the CU is predicted inter-frame using its own motion; only unidirectional prediction is allowed for each segment, i.e., each part has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, as with traditional bidirectional prediction, only two motion-compensated predictions are needed for each CU.
[0107] If the geometric segmentation pattern is used for the current CU, the geometric segmentation index (angle and offset) and two merge indices (one for each segment) are further indicated via signal transmission. The number of maximum GPM candidate dimensions is explicitly transmitted in SPS, and the syntax binarization used for the GPM merge index is specified. After predicting each part of the geometric segment, a blending process with adaptive weights is used to adjust the sample values along the geometric segment edges. This is the prediction signal for the entire CU, and the transformation and quantization processes are applied to the entire CU as in other prediction patterns. Finally, the motion field of the CU predicted using the geometric segmentation pattern is stored.
[0108] 2.6.1 Construction of One-Way Prediction Candidate List The unidirectional prediction candidate list is directly derived from the Merge candidate list constructed according to the extended Merge prediction process. Let n denote the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector of the nth extended Merge candidate (where X equals the parity of n) is used as the nth unidirectional prediction motion vector for the geometric segmentation pattern. These motion vectors in... Figure 20 The vector is marked with "x". If the corresponding LX motion vector of the nth extended Merge candidate does not exist, the L(1-X) motion vector of the same candidate is used as the unidirectional predicted motion vector for the geometric segmentation pattern.
[0109] 2.6.2 Blending along geometric segmentation edges After predicting each segment of the geometric segment using its own motion, a blend is applied to the two predicted signals to derive samples around the segmentation edges. The blend weights for each location of the CU are derived based on the distance between the individual location and the segmentation edge.
[0110] For location The distance to the segmentation edge is derived as follows:
[0111] in It is an index used for the angle and offset of geometric segmentation, which depends on the geometric segmentation index transmitted via signal. The sign depends on the angle index. .
[0112] The weights for each part of the geometric segmentation are derived as follows:
[0113] partIdx depends on the angle index Weight An example in Figure 21 As shown in the image, the Figure 21 The mixed weights using geometric segmentation patterns are shown. An example of generation.
[0114] 2.6.3 Geometric Partitioning Pattern (GPM) with Merge Motion Vector Difference (MMVD) GPM in VVC is extended by applying motion vector refinement on top of the existing unidirectional MV of GPM. First, a flag is transmitted to the GPMCU to specify whether this mode is used. If the mode is used, each geometric segment of the GPM CU can further determine whether MVD is transmitted via signaling. If MVD is transmitted for each geometric segment, the segment's motion is further refined using the transmitted MVD information after selecting a GPM Merge candidate. All other processes remain the same as in GPM.
[0115] Similar to MMVD, MVD is transmitted as a pair of distance and direction via signals. In GPM with MMVD (GPM-MMVD), nine candidate distances (1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixel, 3 pixel, 4 pixel, 6 pixel, 8 pixel, 16 pixel) and eight candidate directions (four horizontal / vertical directions and four diagonal directions) are involved. Additionally, when pic_fpel_mmvd_enabled_flag equals 1, MVD is shifted left by 2, just as in MMVD.
[0116] 2.6.4 Geometric Partitioning Mode with Adaptive Blending (GPM) In VVC, the final predicted samples are generated by weighted averaging of the predictions from two predicted signals. The two integer mixing matrices ( W 0 and W 1) Used. The weights in the GPM mixing matrix are derived from the ramp function based on the displacement from the predicted sample location to the GPM segmentation boundary. The mixing region size is fixed at 2 (2 samples on each side of the GPM segmentation boundary).
[0117] The blending process in ECM is improved by adding four additional blending region sizes (one-quarter, half, twice, and four times the existing region size), such as Figure 22 As shown. The CU-level flags are encoded and decoded to accommodate the selected mixing region size via signal transmission. Furthermore, extended weighted precision is utilized, where the maximum weight value changes from 8 (in VVC) to 32 to accommodate the extended mixing region size. Figure 22 The ramp function for GPM mixing is shown, based on the displacement (d) from the predicted sample location to the GPM segmentation boundary and the size of the mixing region (τ).
[0118] 2.6.5 Geometric Segmentation Pattern (GPM) with Template Matching (TM) Template matching is applied to GPM. When GPM mode is enabled for CU, a CU-level flag is signaled to indicate whether TM is applied to the two geometric segments. Motion information for each geometric segment is refined using TM. When TM is selected, the template is constructed using left, top, or left and top neighbor samples based on the segmentation angle, as shown in Table 1. Then, with the half-pixel interpolation filter disabled, motion is refined by minimizing the difference between the current template and the template in the reference image using the same search pattern as the Merge mode.
[0119] Table 1. Templates for the first geometric segmentation and the second geometric segmentation, where A indicates the use of the top sample point, L indicates the use of the left sample point, and L+A indicates the use of both the left and top samples.
[0120]
[0121] The GPM candidate list is constructed as follows: The staggered List-0 and List-1 MV candidates are derived directly from the regular Merge candidate list, with List-0 MV candidates having a higher priority than List-1 MV candidates. A deduplication method with an adaptive threshold based on the current CU size is applied to remove redundant MV candidates.
[0122] The interleaved List-1 and List-0 MV candidates are also derived directly from the regular Merge candidate list, with List-1 MV candidates having a higher priority than List-0 MV candidates. The same deduplication method with an adaptive threshold is also applied to remove redundant MV candidates.
[0123] Zero MV candidates are filled until the GPM candidate list is full.
[0124] GPM-MMVD and GPM-TM are exclusively enabled to a single GPM CU. This is done first by signaling the GPM-MMVD syntax. When both GPM-MMVD control flags are false (i.e., GPM-MMVD is disabled for both GPM segments), the GPM-TM flag is signaled to indicate whether template matching is applied to both GPM segments. Otherwise (if at least one GPM-MMVD flag is true), the value of the GPM-TM flag is presumed to be false.
[0125] 2.6.6 GPM with inter-frame and intra-frame prediction In GPM with inter-frame and intra-frame prediction, the final prediction samples are generated by weighting inter-frame and intra-frame prediction samples for each GPM separately. Inter-frame prediction samples are derived from inter-frame GPMs, while intra-frame prediction samples are derived from the intra-frame prediction mode (IPM) candidate list and the index from the encoder through signal transmission. The IPM candidate list size is predefined as 3. Available IPM candidates are parallel angle mode (parallel mode) for GPM block boundaries, vertical angle mode (vertical mode) for GPM block boundaries, and planar mode, such as... Figures 23A to 23C As shown. Furthermore, as... Figure 23D The GPM with intra-frame and intra-frame prediction shown is limited to reduce signaling overhead for IPM and avoid increasing the size of intra-frame prediction circuitry on the hardware decoder. Additionally, direct motion vectors and IPM storage on the GPM mixing region are introduced to further improve encoding / decoding performance.
[0126] In IPM derivation based on DIMD and neighboring modes, parallel modes are registered first. Therefore, if no identical IPM candidates exist in the list, up to two IPM candidates derived from the decoder-side intra-frame pattern derivation (DIMD) method and / or neighboring block derivation can be registered. As for neighboring mode derivation, there are up to five locations for available neighboring blocks, but they are limited by the angle of the GPM block boundary, as shown in Table 2. These locations have already been used for GPM with template matching (GPM-TM).
[0127] Table 2. Positions of available neighboring blocks for IPM candidate derivation based on the angle of the GPM block boundary. A and L represent the top and left sides of the predicted block.
[0128]
[0129] GPM-intraframe can be combined with GPM with Merge and Motion Vector Difference (GPM-MMVD). TIMD is used as an IPM candidate for GPM-intraframe to further improve encoding and decoding performance. Parallel mode can be registered first, followed by TIMD, DIMD, and IPM candidates from neighboring blocks.
[0130] 2.6.7 Template Matching-Based Reordering for GPM Partitioning Patterns In template-match-based reordering of GPM partitioning patterns, given the motion information of the current GPM block, the corresponding TM value of the GPM partitioning pattern is calculated. Then, all GPM partitioning patterns are reordered in ascending order based on their TM values. Instead of sending the GPM partitioning patterns, an index indicating where the exact GPM partitioning pattern is located in the reordering list is transmitted via signaling, using Golomb-Rice codes.
[0131] The GPM partitioning pattern reordering method is a two-step process performed after the corresponding reference templates for the two GPM partitions in the encoding / decoding unit are generated, as shown below: The GPM segmentation edge is extended to the reference templates of the two GPM segments, resulting in 64 reference templates, and the corresponding TM cost of each of the 64 reference templates is calculated. The TM generation value based on the GPM partitioning pattern is reordered in ascending order for the GPM partitioning patterns, and the best 32 partitioning patterns are marked as available partitioning patterns.
[0132] like Figure 24 As shown, the edges on the template are extended from the edges of the current CU, but the GPM blending process is not applied to the template regions across the edges.
[0133] After ascending reordering using the TM cost, the index is transmitted via signaling.
[0134] 2.6.8 Motion field storage for geometric segmentation patterns Mv1 from the first part of the geometric segmentation, Mv2 from the second part of the geometric segmentation, and Mv, a combination of Mv1 and Mv2, are stored in the motion field of the CU encoded and decoded by the geometric segmentation pattern.
[0135] The type of motion vector stored for each individual location in the sports field is determined as follows:
[0136] Where motionIdx equals It is recalculated from equation (2-36). partIdx depends on the angle index. .
[0137] If sType equals 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field; otherwise, if sType equals 2, then the Mv resulting from the combination of Mv0 and Mv2 is stored. The combined Mv is generated using the following process: 1) If Mv1 and Mv2 come from different lists of reference images (one from L0 and the other from L1), then Mv1 and Mv2 are simply combined to form a bidirectional predicted motion vector.
[0138] 2) Otherwise, if Mv1 and Mv2 come from the same list, only the unidirectional predicted motion Mv2 is stored.
[0139] 2.7 Multiple Hypothesis Prediction (MHP) In the multi-hypothesis inter-frame prediction mode (JVET-M0425), in addition to the traditional bidirectional prediction signal, one or more additional motion-compensated prediction signals are transmitted via signal transmission. The resulting overall prediction signal is obtained by weighted superposition of samples. Utilizing the bidirectional prediction signal... and the first additional inter-frame prediction signal / hypothesis The generated prediction signal The following was obtained:
[0140] According to the following mapping, the weighting factor It is defined by the new syntax element add_hyp_weight_idx.
[0141]
[0142] Similar to the above, more than one additional prediction signal can be used. The resulting overall prediction signal is accumulated iteratively using each additional prediction signal.
[0143]
[0144] The resulting overall prediction signal is obtained as the last one. (i.e., has the largest index) Within this EE, a maximum of two additional prediction signals can be used (i.e., Limited to 2).
[0145] The motion parameters for each additional prediction hypothesis can be explicitly transmitted via signaling by specifying a reference index, motion vector prediction value index, and motion vector difference, or implicitly transmitted via signaling by specifying a merge index. A separate multi-hypothesis merge flag distinguishes between these two signaling modes.
[0146] For inter-frame AMVP mode, MHP is only applied if unequal weights are selected in BCW in bidirectional prediction mode.
[0147] Combining MHP and BDOF is possible; however, BDOF is only applied to the bidirectional prediction signal portion of the predicted signal (i.e., the ordinary first two assumptions).
[0148] 2.8 Affine Motion Compensation in Geometric Prediction Mode A sub-block-based motion compensation method was proposed for use in GPM mode.
[0149] a) In one example, sub-block-based motion compensation can be affine motion compensation.
[0150] b) In one example, sub-block-based motion compensation could be sbTMVP motion compensation.
[0151] c) In one example, a prediction of at least one geometric segmentation can be generated using sub-block-based motion compensation, such as affine motion compensation.
[0152] d) In one example, the final prediction can be generated by a weighted sum of two predictions, where at least one prediction is generated using sub-block-based motion compensation (such as affine motion compensation).
[0153] i. In one example, the weighted sum is performed using the weight values defined by GPM.
[0154] e) In one example, the two predictions used in the GPM pattern can be of type A and type B, where type A and type B can be (type A and type B can be the same type): i. Non-affine inter-frame prediction; ii. Affine inter-frame prediction; iii. Intra-frame prediction; iv. Intra-block copy (IBC) prediction; v. sb-TMVP inter-frame prediction; vi. Any combination or generated prediction.
[0155] abbreviation ACT Adaptive Color Transformation ALF Adaptive Loop Filter AMVR Adaptive Motion Vector Resolution APS Adaptive Parameter Set AU Access Unit AUD access unit separator AVC (Advanced Video Coding) (Recommendation ITU-T H.264 | ISO / IEC 14496-10) B Two-way prediction BCW features bidirectional prediction with CU-level weights. BDOF bidirectional optical flow BDPCM is based on block-based incremental pulse coding and decoding modulation. BP cache cycle CABAC Context-Based Adaptive Binary Arithmetic Encoding and Decoding CB codec block CBR constant bit rate CCALF Cross-Component Adaptive Loop Filter CPB encoded / decoded image cache CRA completely random access CRC Cyclic Redundancy Check CTB codec tree block CTU encoding / decoding tree unit CU encoding / decoding unit CVS encoded video sequence DPB decodes image cache DCI decoding capability information DRAP depends on random access points DU decoding unit DUI Decoding Unit Information EG Index Columbus EGk k-th exponent Columbus EOB bitstream ends EOS sequence ends FD Fill Data FIFO (First In First Out) FL fixed length GBR green, blue and red GCI General Constraints Information GDR is being gradually decoded and refreshed. GPM geometric segmentation mode HEVC High-Efficiency Video Codec (Recommendation ITU-T H.265 | ISO / IEC 23008-2) HRD Hypothetical Reference Decoder HSS Hypothesis Flow Scheduler Within I-frame IBC Intra-Block Copying IDR instant decoding refresh ILRP Inter-Frame Layer Reference Image IRAP Intra-Frame Random Access Point LFNST Low-Frequency Inseparable Transform LIC local lighting compensation LPS least likely symbol LSB least significant bit LTRP Long-Term Reference Image LMCS with chroma scaling luminance mapping MIP-based intra-frame prediction MPS most likely symbol MSB most significant bit MTS Multiple Transformation Selection MVP motion vector prediction NAL Network Abstraction Layer OBMC Overlap Block Motion Compensation OLS Output Layer Set OP operation point OPI Operation Point Information P prediction PH image header POC image sequential counting PPS Image Parameter Set PROF refines the prediction using optical flow. PT image timer PU image unit QP quantization parameters RADL random access decodeable front-end (image) RASL random access skipped prerequisites (image) RBSP raw byte sequence payload RGB red, green and blue RPL Reference Image List SAO Sample Adaptive Compensation SAR sample amplitude ratio SEI Supplemental Enhancement Information SH strip head SLI sub-picture level information SODB data bit string SP sequence parameter set STRP Short-Term Reference Image STSA Stepwise Temporal Sublayer Access TR truncated Rice VBR Variable Bit Rate VCL video codec layer VPS Video Parameter Set VSEI Multifunctional Supplemental Enhancement Information (Recommendation ITU-T H.274 | ISO / IEC 23002-7) VUI Video Availability Information VVC Multi-Functional Video Codec (Recommendation ITU-T H.266 | ISO / IEC 23090-3) 3. Problems to be solved In existing technologies, affine motion can also be used to generate predictions in GPM. However, it is currently unclear how to coordinate affine motion in GPM with other encoding / decoding tools.
[0156] 4. Detailed Solution In this disclosure, we propose to refine the affine CPMV using template matching. For a given affine candidate in the affine candidate list, the CPMV can be further refined using template matching, and the refined affine candidates are then used to derive sub-block or pixel-level affine motion information for the current block.
[0157] The detailed embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.
[0158] The terms “video unit” or “code-decoder unit” or “block” can refer to code-decoder tree block (CTB), code-decoder tree unit (CTU), code-decoder block (CB), CU, PU, TU, PB, TB.
[0159] The term "affine block" can refer to a block encoded or decoded using affine Merge, affine AMVP, or any other affine variant mode (i.e., affine MMVD, etc.), which can be described by motion information of two control points (4 parameters) or three control point motion vectors (6 parameters). The term "CPMV" can refer to the motion information of an affine block at its top-left, top-right, and / or bottom-left corners.
[0160] The term "template" can refer to a reconstructed region that can be used to refine a CPMV, and can mean either a "separate template" or a "uniform template." Here, a "separate template" can refer to a reconstructed region that can be used to refine a single CPMV (i.e., one of the top-left, top-right, and / or bottom-left corners), while a "uniform template" can refer to a reconstructed region that can be used to refine all or any (multiple) CPMVs for a block. The terms "template matching cost" or "TM cost" can refer to the matching cost of a separate template or the matching cost of a uniform template.
[0161] In this disclosure, regarding "a block encoded / decoded in mode N", "mode N" here can be a prediction mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or an encoding / decoding technique (e.g., DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-inter, GPM intra-intra, GPM inter-intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, ALF, deblocking, SAO, bilateral filter, LMCS, and corresponding variants, etc.).
[0162] It should be noted that the following terms are not limited to the specific terms defined in the existing standards. Any changes to the encoding / decoding tools are also applicable.
[0163] 1. It is proposed that a first determination can be made to decide whether sub-block based prediction (such as affine prediction) is applicable to a GPM block. For example, if the determination is true, then affine prediction is applicable to the GPM block, and if the determination is false, then affine prediction is not applicable to the GPM block.
[0164] a) In one example, the first determination can be made based on a syntax element (SE) signaled in advance.
[0165] i. For example, the SE can be signaled in the VPS / SPS / PPS / DPS / picture header / slice header / CTU row / CTU / CU, etc.
[0166] ii. In one example, the SE can indicate whether affine prediction is applicable to GPM.
[0167] iii. In one example, if the SE indicates that affine prediction is not applicable to GPM, the first determination can be set to false.
[0168] b) In one example, the first determination can be made based on the width (denoted as W) and / or height (denoted as H) of the block.
[0169] i. For example, if W < TW0, the first determination can be set to false.
[0170] ii. For example, if W > TW1, the first determination can be set to false.
[0171] iii. For example, if H < TH0, the first determination can be set to false.
[0172] iv. For example, if H > TH1, then the first determination can be set to false.
[0173] v. For example, TW0=TH0=8 and TW1=TH1=128.
[0174] c) In one example, the first determination may depend on the encoding / decoding information (such as stripe type, QP, image resolution, encoding / decoding mode, information about neighboring blocks or samples, and history table) being made.
[0175] i. In one example, the first determination may depend on the encoding / decoding modes (of the neighboring blocks) being made.
[0176] 1) In one example, the number of neighboring blocks (denoted as NN) encoded and decoded using affine mode can be counted.
[0177] a) For example, if a neighboring block is encoded and decoded using the affine AMVP mode, it can be considered affine encoded and decoded.
[0178] b) For example, if a neighboring block is encoded and decoded using the affine Merge mode, it can be considered affine encoded and decoded.
[0179] c) For example, if a neighboring block is encoded and decoded using a GPM mode with affine prediction, it can be considered affine encoded and decoded.
[0180] d) A neighboring block can be one of the spatial neighboring blocks A0, A1, B0, B1, and B2, such as... Figure 4 As shown.
[0181] e) A neighboring block can be one of the temporal neighboring blocks C0 and C1, such as Figure 4 As shown.
[0182] f) A neighboring block can be one of the spatial neighboring blocks A0, A1, A2, B0, B1, B2, and B3, such as... Figure 10 As shown.
[0183] g) Neighboring blocks can be adjacent to or not adjacent to the current block.
[0184] 2) In one example, if NN is less than the threshold Tn, then the first determination can be set to equal to false.
[0185] a) For example, Tn = 1, or 2, or 3, or 4, or 5.
[0186] 3) In one example, if NN is greater than the threshold Tn, then the first determination can be set to true.
[0187] a) For example, Tn = 0, or 1, or 2, or 3, or 4.
[0188] ii. In one example, the first determination can depend on a history table being made for a history-parameter table (HPT) for affine prediction.
[0189] 1) In one example, the number of valid candidates in the HPT (denoted as NT) can be counted.
[0190] 2) In one example, if NT is less than a threshold Tn, the first determination can be set to equal false.
[0191] a) For example, Tn = 1, or 2, or 3, or 4, or 5.
[0192] 3) In one example, if NT is greater than a threshold Tn, the first determination can be set to equal true.
[0193] a) For example, Tn = 0, or 1, or 2, or 3, or 4.
[0194] 2. It is proposed that a second determination can be made to decide whether a conventional prediction excluding sub-block-based prediction (such as affine prediction) is applicable to a GPM block. For example, if the determination is true, the conventional prediction is applicable to the GPM block, and if the determination is false, the conventional prediction is not applicable to the GPM block.
[0195] a) In one example, the second determination can be made based on a syntax element (SE) signaled in advance.
[0196] i. For example, SE can be signaled in a VPS / SPS / PPS / DPS / picture header / strip header / CTU row / CTU / CU, etc.
[0197] ii. In one example, SE can indicate whether the conventional prediction is applicable to GPM.
[0198] iii. In one example, if SE indicates that the conventional prediction is not applicable to GPM, the second determination can be set to false.
[0199] b) In one example, the second determination can be made based on the width of the block (denoted as W) and / or the height (denoted as H).
[0200] i. For example, if W < TW0, the first determination can be set to false.
[0201] ii. For example, if W > TW1, the first determination can be set to false.
[0202] iii. For example, if H < TH0, the first determination can be set to false.
[0203] iv. For example, if H > TH1, then the first determination can be set to false.
[0204] v. For example, TW0=TH0=8 and TW1=TH1=64.
[0205] c) In one example, the second determination may depend on the encoding / decoding information (such as stripe type, QP, image resolution, encoding / decoding mode, information about adjacent blocks or samples, and history table) being made.
[0206] 3. In one example, whether GPM applies to the current block can be determined by the first and second determinations. For example, if the third determination is true, GPM applies to the block; if the third determination is false, GPM does not apply to the block.
[0207] a) For example, if at least one of the first determination and the second determination is true, then the third determination is set to true.
[0208] b) For example, if both the first determination and the second determination are false, then the third determination is set to false.
[0209] c) If the third determination is false, the SE indicating the use of GPM can be skipped and presumed to be false.
[0210] 4. In one example, if the first determination is false, one or more of the following SEs can be skipped.
[0211] a) Indicates whether affine prediction is applied to the SE of GPM or GPM segmentation; b) Indicates whether intra-frame prediction is applied to the SE of GPM or GPM segmentation; c) Indicate whether the MMVD prediction is applied to the SE of the GPM or the GPM segmentation; d) Indicate whether the TM prediction is applied to the SE of the GPM or the GPM segmentation; e) Indicates the SE for GPM Merge candidate indexes for GPM partitions; f) An SE indicating the GPM MMVD index for a segment of GPM or GPM; g) Indicates the SE for intra-prediction mode for GPM or GPM segmentation; h) Instructs the SE that partitions the GPM index; i) Indicates the SE of the GPM hybrid index; j) The skipped SE can be presumed to be a value such as true, false, 0, or 1.
[0212] 5. In one example, if the second determination is false, one or more of the following SEs can be skipped.
[0213] a) Indicates whether affine prediction is applied to the SE of GPM or GPM segmentation; b) Indicates whether intra-frame prediction is applied to the SE of GPM or GPM segmentation; c) Indicate whether the MMVD prediction is applied to the SE of the GPM or the GPM segmentation; d) Indicate whether the TM prediction is applied to the SE of the GPM or the GPM segmentation; e) Indicates the SE for GPM Merge candidate indexes for GPM partitions; f) An SE indicating the GPM MMVD index for a segment of GPM or GPM; g) An SE indicating the intra-prediction mode index for a segment of GPM or GPM. h) Instructs the SE that partitions the GPM index; i) Indicates the SE of the GPM hybrid index; j) The skipped SE can be presumed to be a value such as true, false, 0, or 1.
[0214] 6. In one example, the maximum value of the index transmitted via signaling in the GPM can depend on whether affine prediction is applied in the segmentation of the GPM.
[0215] a) An index can be: i. GPM Merge candidate indexes for GPM segmentation; ii. GPM MMVD index for GPM or GPM segmentation; iii. Intra-prediction mode index for GPM or GPM segmentation; iv. GPM partitioning index; v. GPM hybrid index; b) Indexes can be binarized using encoding / decoding methods limited by maximum values (such as rounding unary codes or rounding binary codes).
[0216] c) In one example, the maximum value of the GPM segment index can be different when affine prediction is applied in the GPM segment or when affine prediction is not applied in the GPM segment.
[0217] d) In one example, the maximum value of the GPM index can be different when affine prediction is applied in at least one GPM segment or when affine prediction is not applied in at least one GPM segment.
[0218] e) In one example, the maximum value of the index for a GPM segment using affine prediction encoding and decoding can depend on factors such as stripe type, QP, image resolution, encoding / decoding mode, information about neighboring blocks or samples, and the history table.
[0219] i. In one example, the maximum value of the index for the GPM segmentation using affine prediction encoding and decoding can depend on the NN or NT as disclosed in Item 1.c.
[0220] 7. In one example, if the second SE for GPM meets the condition, the first SE for GPM can be skipped.
[0221] a) The first SE and / or the second SE may indicate whether affine prediction is applied to the GPM or the segmentation of the GPM; b) The first SE and / or the second SE may indicate whether intra-frame prediction is applied to GPM or GPM segmentation; c) The first SE and / or the second SE can indicate whether the MMVD prediction is applied to the GPM or the segmentation of the GPM; d) The first SE and / or the second SE may indicate whether the TM prediction is applied to the GPM or the segmentation of the GPM; e) The first SE and / or the second SE may indicate GPM Merge candidate indices for the segmentation of GPM; f) The first SE and / or the second SE may indicate the GPM MMVD index for the GPM or the GPM segment; g) The first SE and / or the second SE may indicate the intra-prediction mode index for the GPM or the segmentation of the GPM. h) The first SE and / or the second SE can instruct the GPM to partition the index; i) The first SE and / or the second SE can indicate a GPM mixed index; j) The first SE that is skipped can be presumed to be a value such as true, false, 0, or 1.
[0222] 8. In one example, how GPM Merge candidate indices are signaled can depend on the pattern of (multiple) GPM segmentation.
[0223] a) In one example, if intra-frame prediction is applied to another GPM segment, then for a GPM segment only one GPM Merge candidate index can be signaled.
[0224] b) In one example, two GPM Merge candidate indices can be transmitted via signaling in two schemes.
[0225] i. In the first scheme, the two GPM Merge candidate indices are transmitted via signaling independently.
[0226] ii. In the second scheme, the two GPM Merge candidate indices are transmitted via signaling in a correlated manner.
[0227] 1) The second GPM Merge candidate index is required to be not equal to the first GPM Merge candidate index.
[0228] 2) Assuming the valid values of the second GPM Merge candidate index are {0, 1,… Max_V}, and the first GPM Merge candidate index transmitted via signaling is C0, then the second GPM Merge candidate index C1 can depend on C0 being transmitted via signaling, parsed, and interpreted.
[0229] a) C0 can be binarized using encoding / decoding methods limited by the maximum value Max_V (such as rounding unary codes or rounding binary codes).
[0230] b) An SE represented as C1' can be parsed, where C1' is binarized using a codec limited by a maximum value Max_V-offset (such as rounding unary codec or rounding binary codec). The offset is an integer, such as 1.
[0231] c) C1 can be interpreted based on C1' and C0.
[0232] i. If C1' <= C0, then C1 = C1'; ii. Otherwise (C1'>C0), C1 = C1'+1; c) In one example, whether the GPM Merge candidate index is transmitted via signaling using the first scheme or the second scheme may depend on the encoding / decoding mode of the GPM segmentation.
[0233] i. In one example, if GPM splits one of the two by applying MMVD while the other does not apply MMVD, then the first scheme can be applied.
[0234] ii. In one example, if GPM splits one of the two by applying the TM mode while the other does not apply the TM mode, then the first scheme can be applied.
[0235] iii. In one example, if GPM splits one of the two by applying affine prediction while the other does not, then the first scheme can be applied.
[0236] iv. In one example, if both GPM partitions apply MMVD but have different MMVD indices, then the first scheme can be applied.
[0237] v. In one example, if GPM segmentation excludes both MMVD, TM, and affine prediction, then the first scheme can be applied.
[0238] vi. In one example, if GPM segmentation applies affine predictions excluding MMVD and TM, then the first approach can be applied.
[0239] vii. In one example, if both GPM segmentation methods apply a TM pattern with affine prediction, then the second approach can be applied.
[0240] viii. In one example, if both GPM segmentation methods apply a TM mode that excludes affine prediction, then the second approach can be applied.
[0241] ix. In one example, if both GPM partitions apply the MMVD that excludes affine predictions and have the same MMVD index, then the first scheme can be applied.
[0242] x. In one example, if both GPM splits apply an MMVD with affine prediction and have the same MMVD index, then the first scheme can be applied.
[0243] 9. In one example, a mixed candidate list can be built for GPM.
[0244] a) The mixed candidate list may include at least one affine prediction candidate.
[0245] b) The mixed candidate list may include at least one non-affine prediction candidate.
[0246] c) Whether affine or non-affine prediction is applied to GPM segmentation can depend on the candidate index of GPM segmentation, refer to the candidates in the mixed list.
[0247] d) No SE is transmitted via signal to indicate whether GPM segmentation is encoded or decoded using affine prediction.
[0248] e) The hybrid candidate list may be constructed based on at least a first regular candidate list and at least a second affine candidate list.
[0249] i. For example, a candidate in the mixed candidate list can be picked from the first candidate list and the second candidate list.
[0250] 1) For example, candidates in a mixed candidate list can be picked from two candidate lists in a fixed order, such as: a) Pick one from the first candidate list, then pick one from the second candidate list, and repeat until the mixed candidate list is full.
[0251] 2) For example, the candidates in the mixed candidate list may depend on the adaptive picking of codec information from both candidate lists, including: a) The direction of the candidate in the first or second candidate list.
[0252] b) Encoding and decoding information, such as stripe type, QP, image resolution, encoding and decoding mode, information on neighboring blocks or samples, and history table.
[0253] c) NN or NT as disclosed in Project 1.c.
[0254] 10. The affine motion compensation process for a specific component may depend on whether affine motion compensation is applied to blocks encoded and decoded using GPM mode.
[0255] a) For example, the sub-block size for the chroma component can be different for blocks encoded using GPM mode or blocks not encoded using GPM mode.
[0256] i. For example, for the chroma components of a block encoded using GPM, the sub-block size is 2×2.
[0257] ii. For example, for the chroma components of blocks that do not utilize GPM encoding / decoding, the sub-block size is 4×4.
[0258] b) For example, the motion vectors for sub-blocks of chroma components may differ for blocks encoded using GPM mode or blocks not encoded using GPM mode.
[0259] i. For example, the motion vectors of sub-blocks for chroma components of a block encoded using GPM can be derived from an affine model.
[0260] ii. For example, the motion vector for a sub-block of chroma component for a block encoded using GPM can be derived from the motion vector for a sub-block of luma component.
[0261] General aspects 11. Additional operations can be applied to the proposed method or applied together with the proposed method.
[0262] a) The syntax elements disclosed above can be binarized into flags, fixed-length encoding / decoding, EG(x) encoding / decoding, unary codes, rounded unary codes, rounded binary codes, etc. They can be signed or unsigned.
[0263] b) If a codec tool or codec method is deemed unsuitable or unusable, it means that the syntax elements of the codec tool or codec method may not be transmitted via signal and may be implicitly determined to be unused.
[0264] c) The syntactic elements disclosed above can be encoded or decoded using at least one context model. Alternatively, they can be encoded or decoded in a bypass manner.
[0265] d) The syntactic elements disclosed above can be transmitted conditionally via signals.
[0266] a. SE is transmitted via signal only when the corresponding function is applicable.
[0267] b. SE is transmitted via signal only if the dimensions of the block (width and / or height) meet the conditions.
[0268] e) The syntax elements disclosed above can be transmitted via signaling at the block level / sequence level / picture group level / picture level / strip level / piece group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0269] f) Whether and / or how the methods disclosed above can be applied to be transmitted via signal at the block level / sequence level / picture group level / picture level / strip level / piece group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0270] g) Whether and / or how the methods disclosed above are applied may depend on the encoded / decoded information, such as block size, color format, single / double tree segmentation, color components, and stripe / picture type.
[0271] h) The methods presented in this document can be used in other codec tools that require chroma blending.
[0272] Figure 25 A flowchart of a method 2500 for video processing according to an embodiment of the present disclosure is shown. Method 2500 is implemented during the conversion between video units of a video and a bitstream of a video.
[0273] At box 2510, for the conversion between video units and video bitstreams, a first determination is performed regarding whether sub-block-based predictions are applicable to video units encoded and decoded using Geometric Partition Mode (GPM). A true result for the first determination indicates that sub-block-based predictions are applicable to the video units. A false result for the first determination indicates that sub-block-based predictions are not applicable to the video units.
[0274] At box 2520, the conversion is performed based on a first determination. In some embodiments, the conversion includes encoding video units into a bitstream. In some embodiments, the conversion includes decoding video units from the bitstream. In this way, encoding / decoding efficiency and performance can be improved.
[0275] In some embodiments, the first determination is made based on a first syntax element (SE) transmitted in advance via signaling. For example, the first SE is transmitted via signaling in one of the following: video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), dependency parameter set (DPS), picture header, stripe header, codec tree unit (CTU), CTU line, or codec unit (CU).
[0276] In some embodiments, the first SE indicates whether affine prediction is applicable to the GPM. For example, if the first SE indicates that affine prediction is not applicable to the GPM, the first determination is set to false.
[0277] In some embodiments, the first determination is made based on at least one of the following: the width or height of the video unit. For example, if the width is less than TW0, the first determination is set to false, where TW0 represents a threshold. In some embodiments, if the width is greater than TW1, the first determination is set to false, where TW1 represents a threshold. In some other embodiments, if the height is less than TH0, the first determination is set to false, where TH0 represents a threshold. In some further embodiments, if the height is greater than TH1, the first determination is set to false, where TH1 represents a threshold. In this case, TW0 = TH0 = 8 and TW1 = TH1 = 128.
[0278] In some embodiments, the first determination is made based on codec information. For example, the codec information includes at least one of the following: stripe type, quantization parameter (QP), image resolution, codec mode, information on neighboring blocks or samples, or a history table.
[0279] In some embodiments, a first determination is made based on one or more encoding / decoding modes based on neighboring blocks. In some embodiments, the number of neighboring blocks encoded / decoded using affine modes is counted.
[0280] In some embodiments, if a neighboring block is encoded using an affine AMVP mode, then the neighboring block is considered affine-coded. In some other embodiments, if a neighboring block is encoded using an affine Merge mode, then the neighboring block is considered affine-coded. In some embodiments, if a neighboring block is encoded using a GPM mode with affine prediction, then the neighboring block is considered affine-coded.
[0281] In some embodiments, a neighboring block is one of the spatial neighboring blocks. The neighboring block can be one of the spatial neighboring blocks A0, A1, B0, B1, and B2, such as... Figure 4 As shown.
[0282] In some embodiments, a neighboring block is one of the temporal neighboring blocks. The neighboring block can be one of the temporal neighboring blocks C0 and C1, such as... Figure 4 As shown.
[0283] In some embodiments, a neighboring block may or may not be adjacent to a video unit. A neighboring block can be one of spatial neighboring blocks A0, A1, A2, B0, B1, B2, and B3, such as... Figure 10 As shown.
[0284] In some embodiments, if the number of neighboring blocks is less than a threshold Tn, the first determination is set to false. In some embodiments, the threshold Tn is equal to 1, or 2, or 3, or 4, or 5.
[0285] In some embodiments, if the number of neighboring blocks is greater than a threshold Tn, the first determination is set to true. In some embodiments, the threshold Tn is equal to 0, 1, 2, 3, or 4.
[0286] In some embodiments, the first determination is made based on a history table of the history-parameter table (HPT) for affine prediction. In some embodiments, the number of valid candidates in the HPT is counted. For example, if the number of valid candidates is less than a threshold Tn, the first determination is set to false. In some embodiments, Tn = 1, or 2, or 3, or 4, or 5.
[0287] In some embodiments, if the number of valid candidates is greater than a threshold Tn, the first determination is set to true. In some embodiments, Tn = 0, 1, 2, 3, or 4.
[0288] In some embodiments, a second determination is made to determine whether the general prediction excluding sub-block-based predictions applies to the GPM block. In some embodiments, a true result of the second determination indicates that the general prediction applies to the GPM block, and a false result indicates that the general prediction does not apply to the GPM block.
[0289] In some embodiments, the second determination is made based on a second SE transmitted in advance via signal transmission. In some embodiments, the second SE is transmitted via signal transmission in one of the following: VPS, SPS, PPS, DPS, image header, strip header, CTU line, CTU, or CU.
[0290] In some embodiments, the second SE indicates whether the regular forecast applies to GPM. For example, if the second SE indicates that the regular forecast does not apply to GPM, the second determination is set to false.
[0291] In some embodiments, the second determination is made based on at least one of the following: the width or height of the video unit. For example, if the width is less than TW0, the second determination is set to false, where TW0 represents a threshold. In some embodiments, if the width is greater than TW1, the second determination is set to false, where TW1 represents a threshold. In some other embodiments, if the height is less than TH0, the second determination is set to false, where TH0 represents a threshold. In some further embodiments, if the height is greater than TH1, the second determination is set to false, where TH1 represents a threshold. In some embodiments, TW0 = TH0 = 8 and TW1 = TH1 = 128.
[0292] In some embodiments, the second determination is made based on codec information. In some embodiments, the codec information includes at least one of the following: stripe type, quantization parameter (QP), image resolution, codec mode, information on neighboring blocks or samples, or a history table.
[0293] In some embodiments, a third determination of whether GPM applies to the video unit is based on a first determination and a second determination. In some embodiments, a true result of the third determination indicates that GPM applies to the video unit, and a false result indicates that GPM does not apply to the video unit.
[0294] In some embodiments, a third determination is set to true if at least one of a first determination or a second determination is true. In some other embodiments, a third determination is set to false if both the first determination and the second determination are false. In some embodiments, if the third determination is false, the SE indicating the use of GPM is skipped and presumed to be false.
[0295] In some embodiments, if the first determination is false, one or more of the following SEs are skipped: an SE indicating whether affine prediction is applied to GPM or GPM segmentation; an SE indicating whether intra-frame prediction is applied to GPM or GPM segmentation; an SE indicating whether Merge mode motion vector difference (MMVD) prediction is applied to GPM or GPM segmentation; an SE indicating whether template matching (TM) prediction is applied to GPM or GPM segmentation; an SE indicating the GPM Merge candidate index for GPM segmentation; an SE indicating the GPM MMVD index for GPM or GPM segmentation; an SE indicating the intra-prediction mode for GPM or GPM segmentation; an SE indicating the GPM partitioning index; and an SE indicating the GPM fusion index. In some other embodiments, if the second determination is false, one or more of the following SEs are skipped: an SE indicating whether affine prediction is applied to GPM or GPM segmentation; an SE indicating whether intra-frame prediction is applied to GPM or GPM segmentation; an SE indicating whether MMVD prediction is applied to GPM or GPM segmentation; an SE indicating whether TM prediction is applied to GPM or GPM segmentation; an SE indicating the GPM Merge candidate index for GPM segmentation; an SE indicating the GPM MMVD index for GPM or GPM segmentation; an SE indicating the intra-prediction mode index for GPM or GPM segmentation; an SE indicating the GPM partitioning index; or an SE indicating the GPM hybrid index.
[0296] In some embodiments, the skipped SE is presumed to be a value. For example, the value is one of the following: true, false, 0, or 1.
[0297] In some embodiments, whether and / or how the first determination is obtained is transmitted via signaling at one of the following: sequence level, picture group level, picture level, stripe level, or slice group level. In some other embodiments, whether and / or how the first determination is obtained is transmitted via signaling at one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), stripe header, or slice group header. In some further embodiments, whether and / or how the first determination is obtained is transmitted via signaling at one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), codec tree block (CTB), codec tree unit (CTU), CTU line, stripe, slice, sub-picture, or region comprising more than one sample or pixel.
[0298] In some embodiments, method 2500 further includes: determining whether and / or how to obtain a first determination based on the encoded / decoded information of the video unit. The encoded / decoded information may include at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.
[0299] In some embodiments, syntax elements are binarized into one of the following: a flag, a fixed-length codec, an EG(x) codec, a unary code, a rounded unary code, or a rounded binary code. For example, syntax elements can be signed or unsigned.
[0300] In some embodiments, syntax elements are encoded and decoded using at least one context model, or syntax elements are encoded and decoded in a bypass manner. In some embodiments, syntax elements are transmitted conditionally via signaling.
[0301] In some embodiments, syntax elements are transmitted via signals if the corresponding functionality applies. Alternatively, syntax elements are transmitted via signals if the dimensions of the video unit satisfy a condition. In some embodiments, dimensions include the width and / or height of the video unit.
[0302] In some embodiments, syntax elements are signaled at one of the following: sequence level, picture group level, picture level, stripe level, or slice group level. In some embodiments, syntax elements are signaled at one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), stripe header, or slice group header. In some embodiments, syntax elements are signaled at one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), codec tree block (CTB), codec tree unit (CTU), CTU line, stripe, slice, subpicture, or region comprising more than one sample or pixel. In some embodiments, video units are encoded and decoded using one or more other codec tools requiring chroma blending.
[0303] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: performing a first determination regarding whether a sub-block-based prediction is applicable to video units of a video encoded and decoded using a geometrically segmented mode (GPM) mode, wherein a true result of the first determination indicates that the sub-block-based prediction is applicable to the video units, and a false result indicates that the sub-block-based prediction is not applicable to the video units; and generating a bitstream based on the first determination.
[0304] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: performing a first determination regarding whether a sub-block-based prediction is applicable to video units of a video encoded and decoded using a geometrically segmented mode (GPM) pattern, wherein a true result of the first determination indicates that the sub-block-based prediction is applicable to the video units, and a false result indicates that the sub-block-based prediction is not applicable to the video units; generating a bitstream based on the first determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0305] Figure 26 A flowchart of a method 2600 for video processing according to an embodiment of the present disclosure is shown. Method 2600 is implemented during the conversion between video units of a video and a bitstream of a video.
[0306] At box 2610, for the conversion between video units and video bitstreams, determine at least one of the following for video units encoded using Geometric Partition Mode (GPM): the maximum value of the index transmitted via signaling in GPM, whether the syntax elements for GPM are skipped, the manner of transmitting GPM merge candidate indices via signaling, the mixing of candidate lists, or the affine motion process for the components of the video unit.
[0307] At box 2620, the conversion is performed based on determination. In some embodiments, the conversion includes encoding video units into a bitstream. In some embodiments, the conversion includes decoding video units from the bitstream. In this way, it can improve encoding / decoding efficiency and performance.
[0308] In some embodiments, the maximum value of the signaling index in the GPM depends on whether affine prediction is applied in the segmentation of the GPM. For example, the index is at least one of the following: a GPM Merge candidate index for a segmentation of the GPM; a GPM Merge mode motion vector difference (MMVD) index for a segmentation of the GPM or GPM; an intra-prediction mode index for a segmentation of the GPM or GPM; a GPM partitioning index; or a GPM hybrid index.
[0309] In some embodiments, the index is binarized using a maximum value-limited encoding / decoding. For example, the index is a rounded unary code or a rounded binary code.
[0310] In some embodiments, the maximum value of the index for a GPM segment is based on whether an affine prediction is applied to the GPM segment. In some other embodiments, the maximum value of the index for a GPM segment is based on whether an affine prediction is applied to at least one GPM segment.
[0311] In some embodiments, the maximum value of the index for a GPM segment using affine prediction encoding and decoding depends on encoding and decoding information. For example, the encoding and decoding information includes at least one of the following: stripe type, quantization parameter (QP), image resolution, encoding / decoding mode, information on neighboring blocks or samples, or a history table. In some embodiments, the maximum value of the index for a GPM segment using affine prediction encoding and decoding depends on the number of neighboring blocks or the number of valid candidates.
[0312] In some embodiments, if a second SE for the GPM meets a condition, the first SE for the GPM is skipped. In some embodiments, at least one of the first SE or the second SE indicates whether affine prediction is applied to the GPM or its segmentation. In some other embodiments, at least one of the first SE or the second SE indicates whether intra-frame prediction is applied to the GPM or its segmentation. In some still embodiments, at least one of the first SE or the second SE indicates whether MMVD prediction is applied to the GPM or its segmentation.
[0313] In some embodiments, at least one of the first SE or the second SE indicates whether the TM prediction is applied to the GPM or a GPM segmentation. In some other embodiments, at least one of the first SE or the second SE indicates a GPM Merge candidate index for a GPM segmentation. In some still embodiments, at least one of the first SE or the second SE indicates a GPM MMVD index for the GPM or a GPM segmentation.
[0314] In some embodiments, at least one of the first SE or the second SE indicates an intra-prediction mode index for GPM or GPM segmentation. In some other embodiments, at least one of the first SE or the second SE indicates a GPM segmentation index.
[0315] In some embodiments, at least one of the first SE or the second SE indicates the GPM mixed index. In some other embodiments, the skipped first SE is presumed to be a value, and the value includes one of the following: true, false, 0, or 1.
[0316] In some embodiments, the manner in which GPM Merge candidate indices are signaled depends on the mode of one or more GPM segments. In some embodiments, if intra-frame prediction is applied to another GPM segment, only one GPM Merge candidate index is signaled for a given GPM segment.
[0317] In some embodiments, the two GPM Merge candidate indices are transmitted via signaling in two schemes. In some embodiments, in a first scheme, the two GPM Merge candidate indices are transmitted via signaling independently. In some embodiments, in a second scheme, the two GPM Merge candidate indices are transmitted via signaling in a correlated manner.
[0318] In some embodiments, the second GPM Merge candidate index is required to be not equal to the first GPM Merge candidate index. In some embodiments, if the valid value of the second GPM Merge candidate index is {0, 1, ... Max_V}, and the first GPM Merge candidate index transmitted via signaling is C0, then the second GPM Merge candidate index C1 depends on C0 being transmitted via signaling, parsed, and interpreted. In some embodiments, C0 is binarized using a codec limited by the maximum value Max_V. In some embodiments, the SE represented as C1' is parsed, where C1' is binarized using a codec limited by the maximum value Max_V - offset, and the offset is an integer. For example, the offset is 1.
[0319] In some embodiments, C1 is interpreted based on C1' and C0. In some embodiments, C1' <= C0, C1 = C1'. Alternatively, C1' > C0, C1 = C1'+1.
[0320] In some embodiments, whether the GPM Merge candidate index is transmitted via signaling using a first scheme or a second scheme depends on the encoding / decoding mode of the GPM segment. In some embodiments, the first scheme is applied if MMVD is applied to one GPM segment but not the other. In some other embodiments, the first scheme is applied if TM mode is applied to one GPM segment but not the other. In some embodiments, the first scheme is applied if affine prediction is applied to one GPM segment but not the other. In some other embodiments, the first scheme is applied if MMVD is applied to both GPM segments but with different MMVD indices.
[0321] In some embodiments, a first scheme is applied if one of MMVD, TM, or affine prediction is not applied to GPM segmentation. In some other embodiments, the first scheme is applied if affine prediction excluding MMVD and TM is applied to GPM segmentation. In some embodiments, a second scheme is applied if GPM segmentation applies a TM pattern with affine prediction. In some other embodiments, the second scheme is applied if GPM segmentation applies a TM pattern excluding affine prediction.
[0322] In some embodiments, the first scheme is applied if the GPM segmentation application excludes affine-predicted MMVDs and has the same MMVD index. In some other embodiments, the first scheme is applied if the GPM segmentation application has affine-predicted MMVDs and has the same MMVD index.
[0323] In some embodiments, the hybrid candidate list is constructed for GPM. In some embodiments, the hybrid candidate list includes at least one affine prediction candidate. In some other embodiments, the hybrid candidate list includes at least one non-affine prediction candidate.
[0324] In some embodiments, whether affine prediction or non-affine prediction is applied to GPM segmentation depends on the candidate index of the GPM segmentation, and the candidate index references candidates in a mixed candidate list. In some embodiments, no SE is signaled to indicate whether GPM segmentation utilizes affine prediction encoding / decoding.
[0325] In some embodiments, the hybrid candidate list is constructed based on at least a first regular candidate list and at least a second affine candidate list. In some embodiments, a candidate in the hybrid candidate list is picked from the first candidate list and the second candidate list.
[0326] In some embodiments, candidates in the mixed candidate list are picked from the first candidate list and the second candidate list in a fixed order. For example, the fixed order includes repeating the following until the mixed candidate list is full: picking one from the first candidate list and then picking one from the second candidate list.
[0327] In some embodiments, the candidates in the hybrid candidate list are adaptively picked from a first candidate list and a second candidate list based on the encoding / decoding information. In some embodiments, the encoding / decoding information includes at least one of the following: the orientation, stripe type, QP, image resolution, encoding / decoding mode, information on neighboring blocks or samples, a history table, the number of neighboring blocks, or the number of valid candidates in the first or second candidate list.
[0328] In some embodiments, the affine motion compensation process for components depends on whether affine motion compensation is applied to video units encoded and decoded using GPM mode.
[0329] In some embodiments, the sub-block size for chroma components differs for blocks encoded using GPM mode versus blocks not encoded using GPM mode. For example, the sub-block size is 2×2 for the chroma components of a block encoded using GPM. As another example, the sub-block size is 4×4 for the chroma components of a block not encoded using GPM.
[0330] In some embodiments, the motion vectors for sub-blocks of the chroma component are different for blocks encoded using GPM mode versus blocks not encoded using GPM mode. For example, the motion vectors for sub-blocks of the chroma component of a block encoded using GPM are derived from an affine model. As another example, the motion vectors for sub-blocks of the chroma component of a block encoded using GPM are derived from the motion vectors for sub-blocks of the luma component.
[0331] In some embodiments, whether and / or how to determine for a video unit encoded using GPM whether at least one of the following is transmitted: the maximum value of an index transmitted via signaling in GPM, whether syntax elements for GPM are skipped, the manner in which GPM Merge candidate indices are transmitted via signaling, a mixed candidate list, or an affine motion process for the components of the video unit is transmitted via signaling at one of the following: sequence level, picture group level, picture level, stripe level, or slice group level. In some other embodiments, whether and / or how to determine for a video unit encoded using GPM whether at least one of the following is transmitted: the maximum value of an index transmitted via signaling in GPM, whether syntax elements for GPM are skipped, the manner in which GPM Merge candidate indices are transmitted via signaling, a mixed candidate list, or an affine motion process for the components of the video unit is transmitted via signaling at one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), stripe header, or slice group header. In some other embodiments, whether and / or how to determine for a video unit using GPM encoding / decoding at least one of the following: the maximum value of the index transmitted via signaling in GPM, whether the syntax elements for GPM are skipped, the manner of transmitting GPM merge candidate indices via signaling, the mixing of candidate lists, or the affine motion process of the components for the video unit being transmitted via signaling at one of the following locations: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), codec tree block (CTB), codec tree unit (CTU), CTU line, strip, slice, sub-picture, or region comprising more than one sample point or pixel.
[0332] In some embodiments, method 2600 further includes: determining, based on the encoded and decoded information of the video unit, whether and / or how to determine for the video unit encoded and decoded using GPM at least one of the following: the maximum value of the index transmitted via signaling in GPM, whether the syntax elements used for GPM are skipped, the manner of transmitting GPM merge candidate indices via signaling, the mixing of candidate lists, or the affine motion process for the components of the video unit. The encoded and decoded information may include at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.
[0333] In some embodiments, syntax elements are binarized into one of the following: flags, fixed-length encoding / decoding, EG(x) encoding / decoding, unary codes, rounded unary codes, or rounded binary codes. For example, syntax elements can be signed or unsigned.
[0334] In some embodiments, syntax elements are encoded or decoded using at least one context model, or the syntax elements are encoded or decoded in a bypass manner. In some embodiments, syntax elements are transmitted conditionally via signals.
[0335] In some embodiments, syntax elements are transmitted via signals if the corresponding functionality applies. Alternatively, syntax elements are transmitted via signals if the dimensions of the video unit satisfy a condition. In some embodiments, dimensions include the width and / or height of the video unit.
[0336] In some embodiments, syntax elements are signaled at one of the following: sequence level, picture group level, picture level, stripe level, or slice group level. In some embodiments, syntax elements are signaled at one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), stripe header, or slice group header. In some embodiments, syntax elements are signaled at one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), codec tree block (CTB), codec tree unit (CTU), CTU line, stripe, slice, subpicture, or region comprising more than one sample or pixel. In some embodiments, video units are encoded and decoded using one or more other codec tools requiring chroma blending.
[0337] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining, for a video unit encoded using a geometric segmentation pattern (GPM), at least one of the following: the maximum value of an index transmitted via signaling in the GPM, whether syntax elements for the GPM are skipped, the manner in which GPMMerge candidate indices are transmitted via signaling, a mixing of candidate lists, or an affine motion process for components of the video unit; and generating a bitstream based on the determination.
[0338] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: determining, for a video unit encoded using a Geometric Partition Mode (GPM), at least one of the following: the maximum value of an index transmitted via signaling in the GPM, whether syntax elements for the GPM are skipped, the manner of transmitting GPM Merge candidate indices via signaling, mixing candidate lists, or an affine motion process for components of the video unit; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0339] Embodiments of this disclosure can be described according to the following entries, and its features can be combined in any reasonable manner.
[0340] Item 1. For a conversion between a video unit and a bitstream of the video, performing a first determination regarding whether a sub-block-based prediction is applicable to the video unit encoded and decoded using a geometrically segmented mode (GPM) mode, wherein a true result of the first determination indicates that the sub-block-based prediction is applicable to the video unit, and a false result of the first determination indicates that the sub-block-based prediction is not applicable to the video unit; and performing the conversion based on the first determination.
[0341] Item 2. The method according to Item 1, wherein the first determination is made based on a first syntax element (SE) transmitted in advance via signaling.
[0342] Item 3. The method according to 2, wherein the first SE is transmitted via a signal in one of the following: video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), dependency parameter set (DPS), picture header, strip header, codec tree unit (CTU), CTU line, or codec unit (CU).
[0343] Item 4. The method according to Item 2, wherein the first SE indicates whether affine prediction is applicable to the GPM.
[0344] Item 5. The method according to Item 2, wherein if the first SE indicates that affine prediction is not applicable to the GPM, the first determination is set to false.
[0345] Item 6. The method according to Item 1, wherein the first determination is made based on at least one of the following: the width or height of the video unit.
[0346] Item 7. The method according to Item 6, wherein if the width is less than TW0, the first determination is set to false, where TW0 represents a threshold.
[0347] Item 8. The method according to Item 6, wherein if the width is greater than TW1, the first determination is set to false, where TW1 represents a threshold.
[0348] Item 9. The method according to Item 6, wherein if the height is less than TH0, the first determination is set to false, where TH0 represents a threshold.
[0349] Item 10. The method according to Item 6, wherein if the height is greater than TH1, the first determination is set to false, where TH1 represents a threshold.
[0350] Item 11. The method according to any one of items 7-10, wherein TW0=TH0=8 and TW1=TH1=128.
[0351] Item 12. The method according to Item 1, wherein the first determination is made based on encoding / decoding information.
[0352] Item 13. The method according to Item 12, wherein the encoding / decoding information includes at least one of the following: stripe type, quantization parameter (QP), image resolution, encoding / decoding mode, information on neighboring blocks or samples, or a history table.
[0353] Item 14. The method according to Item 12, wherein the first determination is made based on one or more encoding / decoding modes of neighboring blocks.
[0354] Item 15. The method according to Item 14, wherein the number of neighboring blocks encoded and decoded using affine mode is counted.
[0355] Item 16. The method according to Item 15, wherein if the neighboring block is encoded using an affine AMVP mode, then the neighboring block is considered affine encoded.
[0356] Item 17. The method according to Item 15, wherein if the neighboring block is encoded using an affine Merge mode, the neighboring block is considered affine encoded.
[0357] Item 18. The method according to Item 15, wherein if the neighboring block is encoded using a GPM mode with affine prediction, the neighboring block is considered affine encoded.
[0358] Item 19. The method according to Item 15, wherein the neighboring block is one of the spatial neighboring blocks.
[0359] Item 20. The method according to Item 15, wherein the neighboring block is one of the temporal neighboring blocks.
[0360] Item 21. The method according to Item 15, wherein the neighboring block is adjacent to or not adjacent to the video unit.
[0361] Item 22. The method according to Item 14, wherein if the number of neighboring blocks is less than a threshold Tn, the first determination is set to be equal to false.
[0362] Item 23. The method according to Item 22, wherein the threshold Tn is equal to 1, or 2, or 3, or 4, or 5.
[0363] Item 24. The method according to Item 14, wherein if the number of neighboring blocks is greater than a threshold Tn, the first determination is set to true.
[0364] Item 25. The method according to Item 24, wherein the threshold Tn is equal to 0, or 1, or 2, or 3, or 4.
[0365] Item 26. The method according to Item 12, wherein the first determination is made based on a history table of history-parameter-based tables (HPT) for affine prediction.
[0366] Item 27. The method according to Item 26, wherein the number of valid candidates in the HPT is counted.
[0367] Item 28. The method according to Item 27, wherein if the number of valid candidates is less than a threshold Tn, the first determination is set to be equal to false.
[0368] Item 29. The method according to Item 28, wherein Tn = 1, or 2, or 3, or 4, or 5.
[0369] Item 30. The method according to Item 27, wherein if the number of valid candidates is greater than a threshold Tn, the first determination is set to true.
[0370] Item 31. The method according to Item 30, wherein Tn = 0, or 1, or 2, or 3, or 4.
[0371] Item 32. The method according to any one of items 1-30, wherein a second determination is made to determine whether the general prediction excluding sub-block-based predictions is applicable to the GPM block.
[0372] Item 33. The method according to Item 32, wherein a true result of the second determination indicates that the conventional prediction applies to the GPM block, and a false result of the second determination indicates that the conventional prediction does not apply to the GPM block.
[0373] Item 34. The method according to Item 32, wherein the second determination is made based on a second SE transmitted in advance via signal transmission.
[0374] Item 35. The method according to Item 34, wherein the second SE is transmitted via a signal in one of the following ways: VPS, SPS, PPS, DPS, picture header, strip header, CTU line, CTU, or CU.
[0375] Item 36. The method according to Item 34, wherein the second SE indicates whether the conventional prediction is applicable to the GPM.
[0376] Item 37. The method according to Item 34, wherein if the second SE indicates that the conventional prediction is not applicable to the GPM, the second determination is set to false.
[0377] Item 38. The method according to Item 32, wherein the second determination is made based on at least one of the following: the width or height of the video unit.
[0378] Item 39. The method according to Item 38, wherein if the width is less than TW0, the second determination is set to false, where TW0 represents a threshold.
[0379] Item 40. The method according to Item 38, wherein if the width is greater than TW1, the second determination is set to false, where TW1 represents a threshold.
[0380] Item 41. The method according to Item 38, wherein if the height is less than TH0, the second determination is set to false, where TH0 represents a threshold.
[0381] Item 42. The method according to Item 38, wherein if the height is greater than TH1, the second determination is set to false, where TH1 represents a threshold.
[0382] Item 43. The method according to any one of items 39-42, wherein TW0=TH0=8 and TW1=TH1=128.
[0383] Item 44. The method according to Item 32, wherein the second determination is made based on encoding / decoding information.
[0384] Item 45. The method according to Item 44, wherein the encoding / decoding information includes at least one of the following: stripe type, quantization parameter (QP), image resolution, encoding / decoding mode, information on neighboring blocks or samples, or a history table.
[0385] Item 46. The method according to any one of items 1-45, wherein a third determination of whether GPM is applicable to the video unit is based on the first determination and the second determination.
[0386] Item 47. The method according to Item 46, wherein a true result of the third determination indicates that GPM applies to the video unit, and a false result of the third determination indicates that GPM does not apply to the video unit.
[0387] Item 48. The method according to Item 46, wherein the third determination is set to true if at least one of the first determination or the second determination is true.
[0388] Item 49. The method according to Item 46, wherein the third determination is set to false if both the first determination and the second determination are false.
[0389] Item 50. The method according to Item 46, wherein if the third is determined to be false, the SE indicating the use of the GPM is skipped and presumed to be false.
[0390] Item 51. The method according to any one of items 1-50, wherein if the first determination is equal to false, one or more of the following SEs are skipped: an SE indicating whether affine prediction is applied to GPM or GPM segmentation; an SE indicating whether intra-frame prediction is applied to GPM or GPM segmentation; an SE indicating whether Merge mode motion vector difference (MMVD) prediction is applied to GPM or GPM segmentation; an SE indicating whether template matching (TM) prediction is applied to GPM or GPM segmentation; an SE indicating the GPM Merge candidate index for GPM segmentation; an SE indicating the GPM MMVD index for GPM or GPM segmentation; an SE indicating the intra-frame prediction mode for GPM or GPM segmentation; an SE indicating the GPM partitioning index; an SE indicating the GPM mixing index.
[0391] Item 52. The method according to any one of items 1-51, wherein if the second determination is equal to false, one or more of the following SEs are skipped: an SE indicating whether affine prediction is applied to GPM or GPM segmentation; an SE indicating whether intra-frame prediction is applied to GPM or GPM segmentation; an SE indicating whether MMVD prediction is applied to GPM or GPM segmentation; an SE indicating whether TM prediction is applied to GPM or GPM segmentation; an SE indicating the GPM Merge candidate index for GPM segmentation; an SE indicating the GPM MMVD index for GPM or GPM segmentation; an SE indicating the intra-frame prediction mode index for GPM or GPM segmentation; an SE indicating the GPM partitioning index; or an SE indicating the GPM hybrid index.
[0392] Item 53. The method described according to Item 51 or 52, wherein the skipped SE is presumed to be a value.
[0393] Item 54. The method according to Item 53, wherein the value is one of the following: true, false, 0, or 1.
[0394] Item 55. The method according to any one of items 1-54, wherein whether and / or how the first determination is obtained is transmitted by signaling at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.
[0395] Item 56. The method according to any one of items 1-54, wherein whether and / or how the first determination is obtained is transmitted via signaling at one of the following locations: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.
[0396] Item 57. The method according to any one of items 1-54, wherein whether and / or how the first determination is obtained is transmitted by signal at one of the following locations: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), codec tree block (CTB), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region comprising more than one sample point or pixel.
[0397] Item 58. The method according to any one of items 1-54 further comprises: determining whether and / or how the first determination is obtained based on the encoded and decoded information of the video unit, said encoded and decoded information including at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.
[0398] Item 59. A method of video processing, comprising: a conversion between a video unit of a video and a bitstream of the video; determining, for the video unit encoded using a geometric segmentation mode (GPM), at least one of the following: a maximum value of an index transmitted via signaling in the GPM; whether syntax elements for the GPM are skipped; a manner of transmitting GPM Merge candidate indices via signaling; a mixing of candidate lists; or an affine motion process for components of the video unit; and performing the conversion based on the determination.
[0399] Item 60. The method according to Item 59, wherein the maximum value of the index transmitted via signaling in the GPM depends on whether affine prediction is applied in the segmentation of the GPM.
[0400] Item 61. The method according to Item 60, wherein the index is at least one of the following: a GPM Merge candidate index for GPM segmentation; a GPM Merge mode motion vector difference (MMVD) index for GPM or GPM segmentation; an intra-prediction mode index for GPM or GPM segmentation; a GPM partitioning index; or a GPM hybrid index.
[0401] Item 62. The method according to Item 60, wherein the index is binarized using a maximum value-limited encoding / decoding.
[0402] Item 63. The method according to Item 62, wherein the index is a rounding unary code or a rounding binary code.
[0403] Item 64. The method according to Item 60, wherein the maximum value of the index for the GPM segment is based on whether an affine prediction is applied to the GPM segment.
[0404] Item 65. The method according to Item 60, wherein the maximum value of the index for GPM is based on whether an affine prediction is applied in at least one GPM segment.
[0405] Item 66. The method according to Item 60, wherein the maximum value of the index for the GPM segmentation using affine prediction encoding and decoding depends on the encoding and decoding information.
[0406] Item 67. The method according to Item 66, wherein the encoding / decoding information includes at least one of the following: stripe type, quantization parameter (QP), image resolution, encoding / decoding mode, information on neighboring blocks or samples, or a history table.
[0407] Item 68. The method according to Item 67, wherein the maximum value of the index for the GPM segmentation using affine prediction coding depends on the number of neighboring blocks or the number of valid candidates.
[0408] Item 69. The method according to Item 59, wherein if the condition is met for the second SE for the GPM, the first SE for the GPM is skipped.
[0409] Item 70. The method according to Item 69, wherein at least one of the first SE or the second SE indicates whether affine prediction is applied to the GPM or the segmentation of the GPM.
[0410] Item 71. The method according to Item 69, wherein at least one of the first SE or the second SE indicates whether intra-frame prediction is applied to GPM or GPM segmentation.
[0411] Item 72. The method according to Item 69, wherein at least one of the first SE or the second SE indicates whether the MMVD prediction is applied to the GPM or the segmentation of the GPM.
[0412] Item 73. The method according to Item 69, wherein at least one of the first SE or the second SE indicates whether the TM prediction is applied to the GPM or the segmentation of the GPM.
[0413] Item 74. The method according to Item 69, wherein at least one of the first SE or the second SE indicates a GPM Merge candidate index for a segmentation of the GPM.
[0414] Item 75. The method according to Item 69, wherein at least one of the first SE or the second SE indicates a GPM MMVD index for a GPM or a segment of the GPM.
[0415] Item 76. The method according to Item 69, wherein at least one of the first SE or the second SE indicates an intra-prediction mode index for a segmentation of the GPM or GPM.
[0416] Item 77. The method according to Item 69, wherein at least one of the first SE or the second SE indicates the GPM partition index.
[0417] Item 78. The method according to Item 69, wherein at least one of the first SE or the second SE indicates a GPM hybrid index.
[0418] Item 79. The method according to Item 69, wherein the first SE that is skipped is presumed to be a value, and the value includes one of the following: true, false, 0, or 1.
[0419] Item 80. The method according to Item 59, wherein the manner in which the GPM Merge candidate index is transmitted via signaling depends on one or more GPM segmentation patterns.
[0420] Item 81. The method according to Item 80, wherein if intra-frame prediction is applied to another GPM segment, only one GPM Merge candidate index is transmitted via signaling for a GPM segment.
[0421] Item 82. The method according to Item 80, wherein two GPM Merge candidate indices are transmitted via signaling in two schemes.
[0422] Item 83. The method according to Item 82, wherein in the first scheme, the two GPM Merge candidate indices are transmitted by signaling independently.
[0423] Item 84. The method according to Item 83, wherein in the second scheme, the two GPM Merge candidate indices are transmitted in a correlated manner via signaling.
[0424] Item 85. The method according to Item 84, wherein the second GPM Merge candidate index is required to be not equal to the first GPM Merge candidate index.
[0425] Item 86. The method according to Item 84, wherein if the valid value of the second GPM Merge candidate index is {0, 1, … Max_V}, and the first GPM Merge candidate index transmitted by signal is C0, then the second GPM Merge candidate index C1 depends on C0 being transmitted by signal, parsed, and interpreted.
[0426] Item 87. The method according to Item 86, wherein C0 is binarized using encoding / decoding limited by the maximum value Max_V.
[0427] Item 88. The method according to Item 86, wherein the SE represented as C1' is parsed, wherein C1' is binarized using encoding and decoding limited by a maximum value Max_V-offset, and said offset is an integer.
[0428] Item 89. The method according to Item 88, wherein the offset is 1.
[0429] Item 90. The method according to Item 86, wherein C1 is interpreted based on C1' and C0.
[0430] Item 91. The method according to Item 90, wherein C1'<= C0, C1 = C1', or wherein C1'>C0, C1 =C1'+1.
[0431] Item 92. The method according to Item 80, wherein whether the GPM Merge candidate index is transmitted via signaling in a first scheme or a second scheme depends on the encoding / decoding mode of the GPM segmentation.
[0432] Item 93. The method according to Item 92, wherein if MMVD is applied to one of the GPM segments but not to the other GPM segment, then the first scheme is applied.
[0433] Item 94. The method according to Item 92, wherein the first scheme is applied if the TM mode is applied to one of the GPM segments but not to the other GPM segment.
[0434] Item 95. The method according to Item 92, wherein the first scheme is applied if affine prediction is applied to one of the GPM segments but not to the other GPM segment.
[0435] Item 96. The method according to Item 92, wherein if MMVD is applied to the GPM to split both, but with different MMVD indices, then the first scheme is applied.
[0436] Item 97. The method according to Item 92, wherein if no GPM segmentation is applied, then the first scheme is applied.
[0437] Item 98. The method according to Item 92, wherein the first scheme is applied if the GPM segmentation is applied with affine prediction excluding MMVD and TM.
[0438] Item 99. The method according to Item 92, wherein the second scheme is applied if the GPM segmentation is applied with a TM pattern having affine prediction.
[0439] Item 100. The method according to Item 92, wherein if the GPM segmentation is applied with a TM mode that excludes affine prediction, then the second scheme is applied.
[0440] Item 101. The method according to Item 92, wherein the first scheme is applied if the GPM segmentation is applied to exclude affine prediction MMVDs and has the same MMVD index.
[0441] Item 102. The method according to Item 92, wherein the first scheme is applied if the GPM segmentation is applied with an affine prediction MMVD and has the same MMVD index.
[0442] Item 103. The method according to Item 59, wherein the hybrid candidate list is constructed for GPM.
[0443] Item 104. The method according to Item 59, wherein the mixed candidate list includes at least one affine prediction candidate.
[0444] Item 105. The method according to Item 59, wherein the mixed candidate list includes at least one non-affine prediction candidate.
[0445] Item 106. The method according to Item 59, wherein whether affine prediction or non-affine prediction is applied to GPM segmentation depends on the candidate index of the GPM segmentation, and the candidate index refers to candidates in the mixed candidate list.
[0446] Item 107. The method according to Item 59, wherein no SE is transmitted via signaling to indicate whether GPM segmentation utilizes affine prediction encoding / decoding.
[0447] Item 108. The method according to Item 59, wherein the hybrid candidate list is constructed based on at least a first regular candidate list and at least a second affine candidate list.
[0448] Item 109. The method according to Item 108, wherein one candidate in the mixed candidate list is picked from the first candidate list and the second candidate list.
[0449] Item 110. The method according to Item 109, wherein the candidates in the mixed candidate list are picked from the first candidate list and the second candidate list in a fixed order.
[0450] Item 111. The method according to Item 110, wherein the fixed order comprises: until the mixed candidate list is filled, repeating the following: picking one from the first candidate list and then picking one from the second candidate list.
[0451] Item 112. The method according to Item 109, wherein the candidates in the hybrid candidate list are adaptively picked from the first candidate list and the second candidate list depending on the encoding / decoding information.
[0452] Item 113. The method according to Item 112, wherein the encoding / decoding information includes at least one of the following: the orientation of the candidates in the first candidate list or the second candidate list, the stripe type, the QP, the image resolution, the encoding / decoding mode, information on neighboring blocks or samples, the history table, the number of neighboring blocks, or the number of valid candidates.
[0453] Item 114. The method according to Item 59, wherein the affine motion compensation process for the component depends on whether the affine motion compensation is applied to the video unit encoded and decoded using GPM mode.
[0454] Item 115. The method according to Item 114, wherein the sub-block size for the chroma component is different for blocks encoded using GPM mode or blocks not encoded using GPM mode.
[0455] Item 116. The method according to Item 115, wherein for blocks encoded and decoded using GPM, the size of the sub-block for the chroma component is 2×2.
[0456] Item 117. The method according to Item 115, wherein for blocks that do not utilize GPM encoding / decoding, the sub-block size for the chroma components is 4×4.
[0457] Item 118. The method according to Item 114, wherein the motion vectors for sub-blocks of chroma components are different for blocks encoded using GPM mode or blocks not encoded using GPM mode.
[0458] Item 119. The method according to Item 118, wherein the motion vectors of sub-blocks for chroma components of blocks encoded and decoded using GPM are derived from an affine model.
[0459] Item 120. The method according to Item 118, wherein the motion vector of the sub-block for the chroma component of the block encoded using GPM is derived from the motion vector of the sub-block for the luminance component.
[0460] Item 121. The method according to any one of items 59-120, wherein whether and / or how to determine for the video unit encoded using the GPM: wherein whether and / or how to determine for the video unit encoded using the GPM: the maximum value of the index transmitted via signaling in the GPM, whether the syntax element for the GPM is skipped, the manner of transmitting the GPM Merge candidate index via signaling, the mixed candidate list, or the affine motion process of the component for the video unit is transmitted via signaling at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.
[0461] Item 122. The method according to any one of items 59-120, wherein whether and / or how to determine for the video unit encoded and decoded using GPM: the maximum value of the index transmitted via signaling in GPM, whether the syntax element for GPM is skipped, the manner in which the GPM Merge candidate index is transmitted via signaling, the mixed candidate list, or the affine motion process for the components of the video unit is transmitted via signaling at one of the following locations: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.
[0462] Item 123. The method according to any one of items 59-120, wherein at least one of the following is determined for the video unit encoded using GPM: the maximum value of the index transmitted via signaling in GPM, whether the syntax element for GPM is skipped, the manner in which the GPM Merge candidate index is transmitted via signaling, the mixed candidate list, or the affine motion process for the component of the video unit is transmitted via signaling at one of the following locations: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), codec tree block (CTB), codec tree unit (CTU), CTU line, strip, slice, sub-picture, or region comprising more than one sample or pixel.
[0463] Item 124. The method according to any one of items 59-120, further comprising: determining, based on the encoded and decoded information of the video unit, whether and / or how to determine at least one of the following for the video unit encoded and decoded using GPM: the maximum value of the index transmitted via signaling in GPM, whether the syntax element for GPM is skipped, the manner of transmitting the GPM Merge candidate index via signaling, the mixed candidate list, or the affine motion process for the components of the video unit, wherein the encoded and decoded information includes at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.
[0464] Item 125. The method according to any one of items 1-124, wherein the syntax element is binarized into one of the following: a flag, a fixed-length code, an EG(x) code, a unary code, a rounded unary code, or a rounded binary code.
[0465] Item 126. The method according to Item 125, wherein the syntax element is signed or unsigned.
[0466] Item 127. The method according to any one of items 1-124, wherein the syntax elements are encoded or decoded using at least one context model, or wherein the syntax elements are encoded or decoded in a bypass manner.
[0467] Item 128. The method according to any one of items 1-127, wherein the syntax elements are transmitted via signals in a conditional manner.
[0468] Item 129. The method according to Item 128, wherein the syntax element is transmitted via signaling if the corresponding function applies, or wherein the syntax element is transmitted via signaling if the dimension of the video unit satisfies the condition.
[0469] Item 130. The method according to Item 129, wherein the dimension includes the width and / or height of the video unit.
[0470] Item 131. The method according to any one of items 1-129, wherein the syntax element is transmitted by signaling at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.
[0471] Item 132. The method according to any one of Items 1-129, wherein the syntax element is transmitted via signaling at one of the following locations: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.
[0472] Item 133. The method according to any one of items 1-129, wherein the syntax element is transmitted by signal at one of the following locations: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), codec tree block (CTB), codec tree unit (CTU), CTU line, strip, slice, sub-picture, or region comprising more than one sample point or pixel.
[0473] Item 134. The method according to any one of items 1-133, wherein the video unit is encoded or decoded using one or more other encoding / decoding tools that require chroma blending.
[0474] Item 135. The method according to any one of items 1-134, wherein the conversion includes encoding the video unit into the bitstream.
[0475] Item 136. The method according to any one of items 1-134, wherein the conversion includes decoding the video unit from the bitstream.
[0476] Item 137. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-136.
[0477] Item 138. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of items 1-136.
[0478] Item 139. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of an apparatus for video processing, wherein the method comprises: performing a first determination, the first determination relating whether a sub-block-based prediction is applicable to video units of the video encoded and decoded using a geometrically segmented mode (GPM) mode, wherein a true result of the first determination indicates that the sub-block-based prediction is applicable to the video units, and a false result of the first determination indicates that the sub-block-based prediction is not applicable to the video units; and generating the bitstream based on the first determination.
[0479] Item 140. A method for storing a bitstream of video, comprising: performing a first determination, the first determination relating whether a sub-block-based prediction is applicable to video units of the video encoded and decoded using a geometrically segmented mode (GPM) mode, wherein a true result of the first determination indicates that the sub-block-based prediction is applicable to the video units, and a false result of the first determination indicates that the sub-block-based prediction is not applicable to the video units; generating the bitstream based on the first determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0480] Item 141. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of an apparatus for video processing, wherein the method comprises: determining for video units of the video encoded and decoded using a geometric segmentation pattern (GPM) at least one of the following: a maximum value of an index transmitted via signaling in the GPM, whether a syntax element for the GPM is skipped, a manner of merging candidate indices via signaling in the GPM, mixing a candidate list, or an affine motion process for components of the video unit; and generating the bitstream based on the determination.
[0481] Item 142. A method for storing a bitstream of video, comprising: determining for a video unit encoded using a geometric segmentation pattern (GPM) at least one of the following: a maximum value of an index transmitted via signaling in the GPM, whether a syntax element for the GPM is skipped, a manner of transmitting a GPM merge candidate index via signaling, a mixing candidate list, or an affine motion process for components of the video unit; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0482] Example device Figure 27A block diagram of a computing device 2700 in which various embodiments of the present disclosure may be implemented is shown. The computing device 2700 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0483] It should be understood that, Figure 27 The computing device 2700 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.
[0484] like Figure 27 As shown, computing device 2700 includes general-purpose computing device 2700. Computing device 2700 may include at least one or more processors or processing units 2710, memory 2720, storage unit 2730, one or more communication units 2740, one or more input devices 2750, and one or more output devices 2760.
[0485] In some embodiments, the computing device 2700 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 2700 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).
[0486] Processing unit 2710 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 2720. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 2700. Processing unit 2710 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0487] Computing device 2700 typically includes various computer storage media. Such media can be any media accessible by computing device 2700, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 2720 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 2730 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 2700.
[0488] The computing device 2700 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 27 Not shown, but a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disc drive for reading from and / or writing to a removable non-volatile optical disc may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0489] Communication unit 2740 communicates with another computing device via a communication medium. Furthermore, the functionality of components in computing device 2700 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, computing device 2700 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0490] Input device 2750 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 2760 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 2740, computing device 2700 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 2700 can also communicate with one or more devices that enable a user to interact with computing device 2700, or, if needed, with any device (e.g., network card, modem, etc.) that enables computing device 2700 to communicate with one or more other computing devices. Such communication can be performed via an input / output (I / O) interface (not shown).
[0491] In some embodiments, some or all of the components of computing device 2700 may be arranged in a cloud computing architecture, rather than integrated into a single device. In a cloud computing architecture, components may be remotely provided and work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed across remote data center locations. Cloud computing infrastructure may provide services through shared data centers, although they appear as a single access point to users. Therefore, a cloud computing architecture can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, the components and functionality described herein may be provided by conventional servers or installed directly or otherwise on client devices.
[0492] In embodiments of this disclosure, computing device 2700 may be used to implement video encoding / decoding. Memory 2720 may include one or more video encoding / decoding modules 2725 having one or more program instructions. These modules are accessible and executable by processing unit 2710 to perform the functions of the various embodiments described herein.
[0493] In an example embodiment of performing video encoding, input device 2750 may receive video data as input 2770 to be encoded. The video data may be processed, for example, by video codec module 2725 to generate an encoded bitstream. The encoded bitstream may be provided as output 2780 via output device 2760.
[0494] In an example embodiment of performing video decoding, input device 2750 may receive an encoded bitstream as input 2770. The encoded bitstream may be processed, for example, by video codec module 2725 to generate decoded video data. The decoded video data may be provided as output 2780 via output device 2760.
[0495] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These variations are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.
Claims
1. A video processing method, comprising: For the conversion between video units and the bitstream of the video, a first determination is performed, the first determination being whether a sub-block-based prediction is applicable to the video unit encoded and decoded using the Geometric Partition Mode (GPM) mode, wherein a true result of the first determination indicates that the sub-block-based prediction is applicable to the video unit, and a false result of the first determination indicates that the sub-block-based prediction is not applicable to the video unit. as well as The conversion is performed based on the first determination.
2. The method of claim 1, wherein the first determination is made based on a first syntax element (SE) transmitted in advance via signaling.
3. The method of claim 2, wherein the first SE is transmitted via a signal in one of the following ways: Video Parameter Set (VPS) Sequence Parameter Set (SPS) Image Parameter Set (PPS) Dependency Parameter Set (DPS) Image header, Strip head, Code-decode tree unit (CTU) CTU line, or Codec Unit (CU).
4. The method of claim 2, wherein the first SE indicates whether affine prediction is applicable to the GPM.
5. The method of claim 2, wherein if the first SE indicates that affine prediction is not applicable to the GPM, then the first determination is set to false.
6. The method of claim 1, wherein the first determination is made based on at least one of the following: the width or height of the video unit.
7. The method of claim 6, wherein if the width is less than TW0, the first determination is set to false, where TW0 represents a threshold.
8. The method of claim 6, wherein if the width is greater than TW1, the first determination is set to false, wherein TW1 represents a threshold.
9. The method of claim 6, wherein if the height is less than TH0, the first determination is set to false, where TH0 represents a threshold.
10. The method of claim 6, wherein if the height is greater than TH1, the first determination is set to false, where TH1 represents a threshold.
11. The method according to any one of claims 7-10, wherein TW0=TH0=8 and TW1=TH1=128.
12. The method of claim 1, wherein the first determination is made based on encoding / decoding information.
13. The method of claim 12, wherein the encoding / decoding information includes at least one of the following: stripe type, quantization parameter (QP), image resolution, encoding / decoding mode, information on neighboring blocks or samples, or a history table.
14. The method of claim 12, wherein the first determination is made based on one or more encoding / decoding modes of neighboring blocks.
15. The method of claim 14, wherein the number of neighboring blocks encoded and decoded using affine patterns is counted.
16. The method of claim 15, wherein if the neighboring block is encoded using an affine AMVP mode, then the neighboring block is considered affine encoded.
17. The method of claim 15, wherein if the neighboring block is encoded using an affine-Merge mode, the neighboring block is considered affine-encoded.
18. The method of claim 15, wherein if the neighboring block is encoded using a GPM mode with affine prediction, the neighboring block is considered affine-encoded.
19. The method of claim 15, wherein the neighboring block is one of the spatial neighboring blocks.
20. The method of claim 15, wherein the neighboring block is one of the temporal neighboring blocks.
21. The method of claim 15, wherein the neighboring block is adjacent to or not adjacent to the video unit.
22. The method of claim 14, wherein if the number of neighboring blocks is less than a threshold Tn, the first determination is set to false.
23. The method of claim 22, wherein the threshold Tn is equal to 1, or 2, or 3, or 4, or 5.
24. The method of claim 14, wherein if the number of neighboring blocks is greater than a threshold Tn, the first determination is set to true.
25. The method of claim 24, wherein the threshold Tn is equal to 0, or 1, or 2, or 3, or 4.
26. The method of claim 12, wherein the first determination is made based on a history table of the history-parameter table (HPT) for affine prediction.
27. The method of claim 26, wherein the number of valid candidates in the HPT is counted.
28. The method of claim 27, wherein if the number of valid candidates is less than a threshold Tn, the first determination is set to be equal to false.
29. The method according to claim 28, wherein Tn = 1, or 2, or 3, or 4, or 5.
30. The method of claim 27, wherein if the number of valid candidates is greater than a threshold Tn, the first determination is set to true.
31. The method of claim 30, wherein Tn = 0, or 1, or 2, or 3, or 4.
32. The method according to any one of claims 1-30, wherein a second determination is made to determine whether conventional predictions excluding sub-block-based predictions are applicable to GPM blocks.
33. The method of claim 32, wherein a true result of the second determination indicates that the conventional prediction applies to the GPM block, and a false result of the second determination indicates that the conventional prediction does not apply to the GPM block.
34. The method of claim 32, wherein the second determination is made based on a second SE transmitted in advance via signal transmission.
35. The method of claim 34, wherein the second SE is transmitted via a signal in one of the following: VPS, SPS, PPS, DPS, picture header, strip header, CTU line, CTU, or CU.
36. The method of claim 34, wherein the second SE indicates whether the conventional prediction is applicable to the GPM.
37. The method of claim 34, wherein if the second SE indicates that the conventional prediction is not applicable to the GPM, the second determination is set to false.
38. The method of claim 32, wherein the second determination is made based on at least one of the following: the width or height of the video unit.
39. The method of claim 38, wherein if the width is less than TW0, the second determination is set to false, where TW0 represents a threshold.
40. The method of claim 38, wherein if the width is greater than TW1, the second determination is set to false, where TW1 represents a threshold.
41. The method of claim 38, wherein if the height is less than TH0, the second determination is set to false, where TH0 represents a threshold.
42. The method of claim 38, wherein if the height is greater than TH1, the second determination is set to false, where TH1 represents a threshold.
43. The method according to any one of claims 39-42, wherein TW0=TH0=8 and TW1=TH1=128.
44. The method of claim 32, wherein the second determination is made based on encoding / decoding information.
45. The method of claim 44, wherein the encoding / decoding information includes at least one of the following: stripe type, quantization parameter (QP), image resolution, encoding / decoding mode, information on neighboring blocks or samples, or a history table.
46. The method according to any one of claims 1-45, wherein the third determination of whether GPM is applicable to the video unit is based on the first determination and the second determination.
47. The method of claim 46, wherein a true result of the third determination indicates that GPM applies to the video unit, and a false result of the third determination indicates that GPM does not apply to the video unit.
48. The method of claim 46, wherein the third determination is set to true if at least one of the first determination or the second determination is true.
49. The method of claim 46, wherein the third determination is set to false if both the first determination and the second determination are false.
50. The method of claim 46, wherein if the third determination is false, the SE indicating the use of the GPM is skipped and presumed to be false.
51. The method according to any one of claims 1-50, wherein if the first determination is equal to false, one or more of the following SEs are skipped: Indicates whether affine prediction is applied to the SE of GPM or GPM segmentation; Indicates whether intra-frame prediction is applied to the SE of GPM or GPM segmentation; Indicates whether the Merge Mode Motion Vector Difference (MMVD) prediction is applied to the SE of GPM or GPM segmentation; Indicate whether template matching (TM) prediction is applied to the SE of GPM or GPM segmentation; The SE indicating the GPM Merge candidate index for the GPM segmentation; The SE indicating the GPM MMVD index for a segment of GPM or GPM; SE indicating the intra-prediction mode for GPM or GPM segmentation; The SE that instructs GPM to partition the index; The SE that indicates the GPM mixed index.
52. The method according to any one of claims 1-51, wherein if the second determination is equal to false, one or more of the following SEs are skipped: Indicates whether affine prediction is applied to the SE of GPM or GPM segmentation; Indicates whether intra-frame prediction is applied to the SE of GPM or GPM segmentation; Indicates whether the MMVD prediction is applied to the SE of the GPM or the GPM segmentation; Indicates whether the TM prediction is applied to the SE of the GPM or the GPM segmentation; The SE indicating the GPM Merge candidate index for the GPM segmentation; The SE indicating the GPM MMVD index for a segment of GPM or GPM; The SE that indicates the intra-prediction mode index for GPM or GPM segmentation; The SE that instructs GPM to partition the index; or The SE that indicates the GPM mixed index.
53. The method of claim 51 or 52, wherein the skipped SE is presumed to be a value.
54. The method of claim 53, wherein the value is one of: true, false, 0, or 1.
55. The method according to any one of claims 1-54, wherein whether and / or how the first determination is obtained is transmitted via signal transmission in one of the following locations: sequence level, Image group level, Image quality, strip level, or Film series level.
56. The method according to any one of claims 1-54, wherein whether and / or how the first determination is obtained is transmitted via signal transmission in one of the following locations: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.
57. The method according to any one of claims 1-54, wherein whether and / or how the first determination is obtained is transmitted via signal transmission in one of the following locations: Predicted blocks (PB). Transform block (TB) Code block (CB) Prediction Unit (PU) Transformer Unit (TU) Codec Unit (CU) Code-decode tree block (CTB). Code-decode tree unit (CTU) CTU line, strip, piece, Sub-images, or This includes regions containing more than one sample point or pixel.
58. The method according to any one of claims 1-54, further comprising: Based on the encoded and decoded information of the video unit, it is determined whether and / or how the first determination is obtained, wherein the encoded and decoded information includes at least one of the following: Block size, Color format, Single-tree partitioning and / or dual-tree partitioning, Color components, Strip type, or Image type.
59. A video processing method, comprising: For the conversion between video units and the bitstream of the video, determine at least one of the following for the video unit encoded and decoded using Geometric Partition Mode (GPM): the maximum value of the index transmitted by signaling in GPM, whether the syntax elements for GPM are skipped, the manner of transmitting GPM Merge candidate indices by signaling, the mixing of candidate lists, or the affine motion process for the components of the video unit. as well as The conversion is performed based on the determination.
60. The method of claim 59, wherein the maximum value of the index transmitted via signaling in the GPM depends on whether affine prediction is applied in the segmentation of the GPM.
61. The method of claim 60, wherein the index is at least one of the following: GPM Merge candidate index for GPM segmentation; GPM Merge Mode Motion Vector Difference (MMVD) Index for GPM or GPM Segmentation; Intra-prediction mode index for GPM or GPM segmentation. GPM partitioning index; or GPM Hybrid Index.
62. The method of claim 60, wherein the index is binarized using a maximum value-limited encoding / decoding.
63. The method of claim 62, wherein the index is a rounding unary code or a rounding binary code.
64. The method of claim 60, wherein the maximum value of the index for the GPM segment is based on whether an affine prediction is applied to the GPM segment.
65. The method of claim 60, wherein the maximum value of the index for GPM is based on whether affine prediction is applied in at least one GPM segmentation.
66. The method of claim 60, wherein the maximum value of the index for the GPM segmentation using affine prediction encoding and decoding depends on the encoding and decoding information.
67. The method of claim 66, wherein the encoding / decoding information includes at least one of the following: stripe type, quantization parameter (QP), image resolution, encoding / decoding mode, information on neighboring blocks or samples, or a history table.
68. The method of claim 67, wherein the maximum value of the index for a GPM segment using affine prediction coding depends on the number of neighboring blocks or the number of valid candidates.
69. The method of claim 59, wherein if the condition is met for the second SE for the GPM, the first SE for the GPM is skipped.
70. The method of claim 69, wherein at least one of the first SE or the second SE indicates whether affine prediction is applied to the GPM or the segmentation of the GPM.
71. The method of claim 69, wherein at least one of the first SE or the second SE indicates whether intra-frame prediction is applied to GPM or GPM segmentation.
72. The method of claim 69, wherein at least one of the first SE or the second SE indicates whether the MMVD prediction is applied to the GPM or the segmentation of the GPM.
73. The method of claim 69, wherein at least one of the first SE or the second SE indicates whether the TM prediction is applied to the GPM or the segmentation of the GPM.
74. The method of claim 69, wherein at least one of the first SE or the second SE indicates a GPM Merge candidate index for a segmentation of the GPM.
75. The method of claim 69, wherein at least one of the first SE or the second SE indicates a GPM MMVD index for a segment of GPM or GPM.
76. The method of claim 69, wherein at least one of the first SE or the second SE indicates an intra-prediction mode index for a segmentation of the GPM or GPM.
77. The method of claim 69, wherein at least one of the first SE or the second SE indicates a GPM partition index.
78. The method of claim 69, wherein at least one of the first SE or the second SE indicates a GPM hybrid index.
79. The method of claim 69, wherein the skipped first SE is presumed to be a value, and the value includes one of: true, false, 0, or 1.
80. The method of claim 59, wherein the manner in which the GPM Merge candidate index is transmitted via signaling depends on one or more GPM segmentation patterns.
81. The method of claim 80, wherein if intra-frame prediction is applied to another GPM segmentation, only one GPM Merge candidate index is transmitted via signaling for a GPM segmentation.
82. The method of claim 80, wherein the two GPM Merge candidate indices are transmitted via signaling in two schemes.
83. The method of claim 82, wherein in the first scheme, the two GPM Merge candidate indices are transmitted by signaling independently.
84. The method of claim 83, wherein in the second scheme, the two GPM Merge candidate indices are transmitted in a correlated manner via signaling.
85. The method of claim 84, wherein the second GPM Merge candidate index is required to be not equal to the first GPM Merge candidate index.
86. The method of claim 84, wherein if the valid value of the second GPM Merge candidate index is {0,1, … Max_V}, and the first GPM Merge candidate index transmitted via signaling is C0, then the second GPM Merge candidate index C1 depends on C0 being transmitted via signaling, parsed, and interpreted.
87. The method of claim 86, wherein C0 is binarized using encoding / decoding limited by the maximum value Max_V.
88. The method of claim 86, wherein the SE represented as C1' is parsed, wherein C1' is binarized using encoding / decoding limited by a maximum value Max_V-offset, and the offset is an integer.
89. The method of claim 88, wherein the offset is 1.
90. The method of claim 86, wherein C1 is interpreted based on C1' and C0.
91. The method of claim 90, wherein C1' <= C0, C1 = C1', or Where C1'>C0, C1 = C1'+1.
92. The method of claim 80, wherein whether the GPM Merge candidate index is transmitted via signaling in a first scheme or a second scheme depends on the encoding / decoding mode of the GPM segmentation.
93. The method of claim 92, wherein the first scheme is applied if MMVD is applied to one of the GPM segments but not to the other GPM segment.
94. The method of claim 92, wherein the first scheme is applied if the TM mode is applied to one of the GPM segments but not to the other GPM segment.
95. The method of claim 92, wherein the first scheme is applied if affine prediction is applied to one of the GPM segments but not to the other GPM segment.
96. The method of claim 92, wherein if MMVD is applied to the GPM to split both, but they have different MMVD indices, then the first scheme is applied.
97. The method of claim 92, wherein if no GPM segmentation is applied, then the first scheme is applied.
98. The method of claim 92, wherein the first scheme is applied if the GPM segmentation is applied with affine prediction excluding MMVD and TM.
99. The method of claim 92, wherein the second scheme is applied if the GPM segmentation is applied with a TM pattern having affine prediction.
100. The method of claim 92, wherein the second scheme is applied if the GPM segmentation is applied with a TM mode that excludes affine prediction.
101. The method of claim 92, wherein the first scheme is applied if the GPM segmentation is applied to exclude affine prediction MMVDs and has the same MMVD index.
102. The method of claim 92, wherein the first scheme is applied if the GPM segmentation is applied with an affine prediction MMVD and has the same MMVD index.
103. The method of claim 59, wherein the hybrid candidate list is constructed for GPM.
104. The method of claim 59, wherein the mixed candidate list includes at least one affine prediction candidate.
105. The method of claim 59, wherein the mixed candidate list includes at least one non-affine prediction candidate.
106. The method of claim 59, wherein whether affine prediction or non-affine prediction is applied to GPM segmentation depends on the candidate index of the GPM segmentation, and the candidate index refers to candidates in the mixed candidate list.
107. The method of claim 59, wherein no SE is transmitted via signal to indicate whether the GPM segmentation is encoded or decoded using affine prediction.
108. The method of claim 59, wherein the hybrid candidate list is constructed based on at least a first regular candidate list and at least a second affine candidate list.
109. The method of claim 108, wherein one candidate in the mixed candidate list is picked from the first candidate list and the second candidate list.
110. The method of claim 109, wherein the candidates in the mixed candidate list are picked from the first candidate list and the second candidate list in a fixed order.
111. The method of claim 110, wherein the fixed sequence comprises: Until the mixed candidate list is full, repeat the following: pick one from the first candidate list, and then pick one from the second candidate list.
112. The method of claim 109, wherein the candidates in the hybrid candidate list are adaptively picked from the first candidate list and the second candidate list depending on the encoding / decoding information.
113. The method of claim 112, wherein the encoding / decoding information comprises at least one of the following: The direction of the candidates in the first candidate list or the second candidate list. Strip type, QP, Image resolution, Encoding / decoding modes Information about neighboring blocks or samples, Historical table, The number of neighboring blocks, or The number of valid candidates.
114. The method of claim 59, wherein the affine motion compensation process for the component depends on whether the affine motion compensation is applied to the video unit encoded and decoded using GPM mode.
115. The method of claim 114, wherein the sub-block size for the chroma component is different for blocks encoded using GPM mode or blocks not encoded using GPM mode.
116. The method of claim 115, wherein for blocks encoded and decoded using GPM, the sub-block size for the chroma components is 2×2.
117. The method of claim 115, wherein for blocks not using GPM encoding / decoding, the sub-block size for the chroma components is 4×4.
118. The method of claim 114, wherein the motion vectors for sub-blocks of chroma components are different for blocks encoded using GPM mode or blocks not encoded using GPM mode.
119. The method of claim 118, wherein the motion vectors of sub-blocks for chroma components of a block encoded using GPM are derived from an affine model.
120. The method of claim 118, wherein the motion vector of the sub-block for the chroma component of the block encoded using GPM is derived from the motion vector of the sub-block for the luminance component.
121. The method according to any one of claims 59-120, wherein whether and / or how to determine for the video unit encoded and decoded using GPM: the maximum value of the index transmitted via signaling in GPM, whether the syntax element for GPM is skipped, the manner of transmitting the GPM Merge candidate index via signaling, the mixed candidate list, or the affine motion process for the component of the video unit is transmitted via signaling at one of the following locations: sequence level, Image group level, Image quality, strip level, or Film series level.
122. The method according to any one of claims 59-120, wherein whether and / or how to determine for the video unit encoded using GPM: the maximum value of the index transmitted via signaling in GPM, whether the syntax element for GPM is skipped, the manner of transmitting the GPM Merge candidate index via signaling, the mixed candidate list, or the affine motion process for the component of the video unit is transmitted via signaling at one of the following locations: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.
123. The method according to any one of claims 59-120, wherein whether and / or how the video unit encoded and decoded using GPM is determined to have at least one of the following: the maximum value of the index transmitted via signaling in GPM, whether the syntax element for GPM is skipped, the manner in which the GPM Merge candidate index is transmitted via signaling, the mixed candidate list, or the affine motion process for the component of the video unit is transmitted via signaling at one of the following locations: Predicted blocks (PB). Transform block (TB) Code block (CB) Prediction Unit (PU) Transformer Unit (TU) Codec Unit (CU) Code-decode tree block (CTB). Code-decode tree unit (CTU) CTU line, strip, piece, Sub-images, or This includes regions containing more than one sample point or pixel.
124. The method according to any one of claims 59-120, further comprising: Based on the encoded and decoded information of the video unit, determine whether and / or how to determine at least one of the following for the video unit encoded and decoded using GPM: the maximum value of the index transmitted via signaling in GPM, whether the syntax element for GPM is skipped, the manner of transmitting the GPM Merge candidate index via signaling, the mixed candidate list, or the affine motion process for the component of the video unit, wherein the encoded and decoded information includes at least one of the following: Block size, Color format, Single-tree partitioning and / or dual-tree partitioning, Color components, Strip type, or Image type.
125. The method according to any one of claims 1-124, wherein the syntax element is binarized into one of the following: a flag, a fixed-length code, an EG(x) code, a unary code, a rounded unary code, or a rounded binary code.
126. The method of claim 125, wherein the syntax elements are signed or unsigned.
127. The method according to any one of claims 1-124, wherein the syntax elements are encoded and decoded using at least one context model, or Syntax elements are bypassed for encoding and decoding.
128. The method according to any one of claims 1-127, wherein the syntax elements are transmitted via signals in a conditional manner.
129. The method of claim 128, wherein the syntax element is transmitted via a signal if the corresponding function applies, or If the dimension of the video unit meets the condition, the syntax element is transmitted via signal.
130. The method of claim 129, wherein the dimension includes the width and / or height of the video unit.
131. The method according to any one of claims 1-129, wherein the syntax element is transmitted via a signal at one of the following locations: sequence level, Image group level, Image quality, strip level, or Film series level.
132. The method according to any one of claims 1-129, wherein the syntax element is transmitted via a signal at one of the following locations: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.
133. The method according to any one of claims 1-129, wherein the syntax element is transmitted via a signal at one of the following locations: Predicted blocks (PB). Transform block (TB) Code block (CB) Prediction Unit (PU) Transformer Unit (TU) Codec Unit (CU) Code-decode tree block (CTB). Code-decode tree unit (CTU) CTU line, strip, piece, Sub-images, or This includes regions containing more than one sample point or pixel.
134. The method according to any one of claims 1-133, wherein the video unit is encoded or decoded using one or more other encoding / decoding tools that require chroma blending.
135. The method according to any one of claims 1-134, wherein the conversion comprises encoding the video unit into the bitstream.
136. The method according to any one of claims 1-134, wherein the conversion comprises decoding the video unit from the bitstream.
137. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-136.
138. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of claims 1-136.
139. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: Perform a first determination regarding whether a sub-block-based prediction is applicable to a video unit of the video encoded and decoded using a geometric segmentation mode (GPM) mode, wherein a true result of the first determination indicates that the sub-block-based prediction is applicable to the video unit, and a false result of the first determination indicates that the sub-block-based prediction is not applicable to the video unit. as well as The bit stream is generated based on the first determination.
140. A method for storing a bitstream of video, comprising: Perform a first determination regarding whether a sub-block-based prediction is applicable to a video unit of the video encoded and decoded using a geometric segmentation mode (GPM) mode, wherein a true result of the first determination indicates that the sub-block-based prediction is applicable to the video unit, and a false result of the first determination indicates that the sub-block-based prediction is not applicable to the video unit. The bit stream is generated based on the first determination; as well as The bitstream is stored in a non-transitory computer-readable recording medium.
141. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: Determine at least one of the following for a video unit encoded and decoded using Geometric Partitioning Mode (GPM): the maximum value of the index transmitted via signaling in GPM, whether the syntax elements for GPM are skipped, the manner of transmitting GPM Merge candidate indices via signaling, mixing candidate lists, or affine motion processes for the components of the video unit. as well as The bit stream is generated based on the determination.
142. A method for storing a bitstream of video, comprising: Determine at least one of the following for a video unit encoded and decoded using Geometric Partitioning Mode (GPM): the maximum value of the index transmitted via signaling in GPM, whether the syntax elements for GPM are skipped, the manner of transmitting GPM Merge candidate indices via signaling, mixing candidate lists, or affine motion processes for the components of the video unit. The bit stream is generated based on the determination; as well as The bitstream is stored in a non-transitory computer-readable recording medium.