Method and device for video processing and medium
By introducing geometric segmentation mode and advanced motion vector prediction into video encoding and decoding, multiple candidate lists are constructed, and motion information transmission is optimized, solving the problem of insufficient encoding and decoding efficiency in existing technologies and achieving more efficient video processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DOUYIN CO LTD
- Filing Date
- 2024-10-01
- Publication Date
- 2026-05-01
AI Technical Summary
There is room for improvement in the efficiency of existing video coding and decoding technologies, especially in multi-functional video coding and decoding standards, where existing motion vector prediction methods are difficult to effectively reduce signaling costs and improve coding and decoding performance.
By employing geometric segmentation mode (GPM) related syntax elements (SE) to perform encoding and decoding transformations using multiple contexts, and combining advanced motion vector prediction (AMVP) and Merge mode signaling, the transmission and prediction of motion information are optimized by constructing multiple candidate lists.
It improves the performance and efficiency of video encoding and decoding, reduces the cost of motion vector signaling, and enhances the quality and efficiency of video processing.
Smart Images

Figure CN121970337A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure generally relate to video processing techniques, and more specifically, to signaling methods for geometric prediction patterns. Background Technology
[0002] Today, digital video capabilities are being applied to all aspects of people's lives. Various video compression technologies have been proposed for video encoding / decoding, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-TH.265 High Efficiency Video Codec (HEVC) standard, and Multi-Functional Video Codec (VVC) standard. However, the encoding and decoding efficiency of video encoding and decoding technologies is generally expected to be further improved. Summary of the Invention
[0003] Embodiments of this disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. This method includes: for the conversion between video units and the video bitstream, determining a first syntax element (SE) related to the geometric segmentation mode (GPM) and encoding / decoding it using multiple contexts; and performing the conversion based on the first SE. In this manner, encoding / decoding performance and efficiency can be improved.
[0005] In a second aspect, an apparatus for video processing is provided. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform the method according to the first aspect of this disclosure.
[0006] In a third aspect, a non-transitory computer-readable storage medium is proposed. This non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of this disclosure.
[0007] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining a first syntax element (SE) associated with a geometric segmentation mode (GPM) and encoding / decoding it using multiple contexts; and generating a bitstream based on the first SE.
[0008] In the fifth aspect, a method for storing video bitstreams is proposed. The method includes: determining a first syntax element (SE) associated with a geometric segmentation mode (GPM) and encoding / decoding it using multiple contexts; generating a bitstream based on the first SE; and storing the bitstream in a non-transitory computer-readable recording medium.
[0009] This summary aims to present, in a simplified form, the selected concepts further described below in the detailed embodiments. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0010] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become clearer from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0011] Figure 1 A block diagram of an example video codec system according to some embodiments of the present disclosure is shown; Figure 2 A block diagram of a first example video encoder according to some embodiments of the present disclosure is shown; Figure 3 A block diagram of an example video decoder according to some embodiments of the present disclosure is shown; Figure 4 This shows the locations of spatial and temporal neighbor blocks used in the construction of the AMVP / Merge candidate list; Figure 5 This shows the positions of non-adjacent candidates in the ECM; Figure 6A and Figure 6B An affine motion model based on control points is shown; Figure 7 An example affine MVF for each sub-block is shown; Figure 8 The location of the inherited affine motion prediction value is shown; Figure 9 This demonstrates the inheritance of control point motion vectors; Figure 10 The locations of candidate positions for the constructed affine Merge pattern are shown; Figure 11A and Figure 11B The spatial nearest neighbor used to derive the affine Merge candidate is shown; Figure 12 The diagram shows a range of non-nearest neighbors to constructed affine Merge candidates; Figure 13 An example of generating HAPC is shown; Figure 14 A diagram illustrating the regression-based affine Merge candidate derivation is shown. Figure 15 This demonstrates template matching execution over the search area surrounding the initial MV; Figure 16The template and the corresponding reference template are shown; Figure 17 The template and reference template of a block with sub-block motion information using the motion information of the current block are shown; Figure 18 The derivation of the sub-CU motion field obtained by applying motion displacement based on neighbor motion information is shown; Figure 19 An example of GPM partitioning grouped at the same angle is shown; Figure 20 The unidirectional prediction MV selection for geometric segmentation patterns is shown; Figure 21 An exemplary generation of the bending weight w_0 using a geometric segmentation pattern is shown; Figure 22 The ramp function for weighting GPM mixing is shown, based on the displacement (d) from the predicted sample location to the GPM segmentation boundary and the mixing region size (τ). Figures 23A to 23C The available IPM candidates are shown respectively; Figure 23D GPM with inter-frame and intra-frame prediction is shown; Figure 24 The edges on the template are shown; Figure 25 A flowchart of a method for video processing according to embodiments of the present disclosure is shown; and Figure 26 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.
[0012] In all accompanying drawings, the same or similar reference numerals usually refer to the same or similar elements. Detailed Implementation
[0013] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.
[0014] In the following description and claims, unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0015] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Additionally, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, whether explicitly described or not, it is believed that such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.
[0016] It should be understood that although the terms “first” and “second”, etc., can be used to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0017] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0018] Example Environment Figure 1 This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0019] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.
[0020] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded images and associated data. An encoded image is an encoded representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data can be directly transmitted to destination device 120 via network 130A through I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by destination device 120.
[0021] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.
[0022] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or further standards.
[0023] Figure 2 This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... Figure 1 An example of a video encoder 114 in system 100 is shown.
[0024] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. Figure 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0025] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.
[0026] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, in which at least one reference picture is the picture in which the current video block is located.
[0027] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... Figure 2 The examples are shown separately.
[0028] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0029] The mode selection unit 203 can select one of several encoding / decoding modes (intra-frame encoding / decoding or inter-frame encoding / decoding) based, for example, on the error result, and provide the resulting intra-coded or inter-coded blocks to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded blocks for use as reference images. In some examples, the mode selection unit 203 can select an intra-inter-frame joint prediction (CIIP) mode, in which prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).
[0030] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.
[0031] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks that do not depend on macroblocks within the same image.
[0032] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0033] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for reference images in list 0 to find a reference video block for the current video block, and can also search for reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference images in lists 0 and 1 containing multiple reference video blocks, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. Motion estimation unit 204 can output the multiple reference indices and multiple motion vectors of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.
[0034] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0035] In one example, the motion estimation unit 204 may indicate a value to the video decoder 300 in the syntax structure associated with the current video block, which indicates that the current video block has the same motion information as another video block.
[0036] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0037] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0038] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0039] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0040] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.
[0041] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0042] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0043] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current video block, which is then stored in the buffer 213.
[0044] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0045] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0046] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.
[0047] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0048] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.
[0049] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, which involves deriving several most likely candidates based on data from neighboring PBs and reference pictures. Motion information typically includes horizontal and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially or temporally neighboring blocks.
[0050] The motion compensation unit 302 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. Identifiers for interpolation filters used with sub-pixel precision can be included in the syntax elements.
[0051] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of the video block to calculate the interpolated values for sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate the prediction block.
[0052] Motion compensation unit 302 may use at least some of the syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be an entire image or a region of an image.
[0053] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.
[0054] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the response prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be used to filter the decoded block to eliminate block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0055] Some exemplary embodiments of this disclosure will be described in detail below. It should be understood that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section only. Furthermore, although some embodiments are described with reference to multi-function video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Furthermore, although some embodiments describe video encoding and decoding steps in detail, it will be understood that the corresponding decoding steps for decoding will be implemented by the decoder. Additionally, the term video processing includes video encoding / decoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another or at different compression bitrates.
[0056] 1. Brief Overview This disclosure relates to video encoding and decoding technologies. Specifically, this disclosure pertains to affine motion prediction methods in video encoding and decoding. This idea can be applied alone or in various combinations to any standard or non-standard video codec.
[0057] 2. Introduction The exponential growth of multimedia data poses a significant challenge to video encoding and decoding. To meet the ever-increasing demand for more efficient compression technologies, the ITU-T and ISO / IEC have developed a series of video encoding and decoding standards over the past few decades. Specifically, the ITU-T developed the H.261 and H.263 standards, and ISO / IEC developed MPEG-1 and MPEG-4 Vision. These two organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), H.265 / HEVC, and the latest VVC standard. Starting with H.262 / MPEG-2, a hybrid video encoding and decoding framework was adopted, utilizing intra / inter-frame prediction plus transform encoding and decoding. Figure 4 The location of spatial and temporal neighbor blocks used in the construction of the AMVP / Merge candidate list is shown.
[0058] 2.1 MVP in Video Encoding and Decoding Inter-frame prediction aims to eliminate temporal redundancy between adjacent frames and is an indispensable component in hybrid video codec frameworks. Specifically, inter-frame prediction utilizes the content specified by motion vectors (MVs) as the predicted version of the current block to be encoded / decoded, thus transmitting only residual signals and motion information in the bitstream. To reduce the cost of MV signaling, motion vector prediction (MVP) emerged as an efficient mechanism for conveying motion information. Early strategies simply used the MV of a specified neighboring block or the median MV of neighboring blocks as the MVP. H.265 / HEVC involves a contention mechanism where rate-distortion optimization (RDO) selects the best MVP from multiple candidates. Specifically, Advanced MVP (AMVP) mode and Merge mode were designed using different motion information signaling strategies. In AMVP mode, a reference index, an MVP candidate index referencing the AMVP candidate list, and motion vector difference (MVD) are transmitted via signaling. In Merge mode, only the Merge index referencing the Merge candidate list is transmitted via signaling, and all motion information associated with the Merge candidate is inherited. Both the AMVP and Merge modes require building an MVP candidate list, and the details of the building process for these two modes are described below.
[0059] AMVP mode: AMVP utilizes the spatial-temporal correlation of motion vectors with neighboring blocks for explicit transfer of motion parameters. For each list of reference images, a motion vector candidate list is constructed by first checking the availability of temporal neighbors to the left and top, removing redundant candidates, and adding zero vectors to ensure a constant candidate list length. For the derivation of spatial motion vector candidates, the final result is based on locations such as... Figure 4 The motion vectors of five blocks at different locations are shown to derive two motion vector candidates. The five neighboring blocks located at B0, B1, B2, and A0, A1 are classified into two groups: group A includes three upper spatial neighboring blocks, and group B includes two left spatial neighboring blocks. The two motion vector candidates are derived using the first available candidates from groups A and B in a predefined order, respectively. For temporal motion vector candidate derivation, a motion vector candidate is derived based on two distinct co-locations checked sequentially (lower right (C0) and center (C1)), as shown. Figure 4 As shown. To avoid redundant MV candidates, duplicate motion vector candidates in the list are discarded. If the number of potential candidates is less than 2, additional zero motion vector candidates are added to the list. Figure 5 The positions of non-adjacent candidates in the ECM are shown.
[0060] Merge modeSimilar to the AMVP mode, the MVP candidate list for the Merge mode also includes spatial and temporal candidates. For spatial motion vector candidate derivation, after performing availability and redundancy checks, a maximum of four candidates are selected, in the order A1, B1, B0, A0, and B2. For temporal Merge candidate (TMVP) derivation, a maximum of one candidate is selected from two temporally neighboring blocks (C0 and C1). When there are not enough Merge candidates using both spatial and temporal candidates, combined bidirectional prediction Merge candidates and zero MV candidates are added to the MVP candidate list. The Merge candidate list construction process terminates once the number of available Merge candidates reaches the maximum allowed number for signal transmission.
[0061] In VVC, the Merge pattern construction process is further improved by introducing a history-based MVP (HMVP), which incorporates motion information from previously encoded / decoded blocks that can be far removed from the current block. In VVC, HMVP merge candidates are appended to the Merge list, following the spatial MVP and TMVP. In this method, motion information from previously encoded / decoded blocks is stored in a table and used as the MVP for the current CU. During the encoding / decoding process, the table with multiple HMVP candidates is maintained using a first-in, first-out (FIFO) strategy. Whenever a non-sub-block inter-frame encoded / decoded CU is present, the associated motion information is added to the last entry of the table as a new HMVP candidate.
[0062] During the standardization of VVC, a non-adjacent MVP was proposed to facilitate better motion information derivation by utilizing non-adjacent regions. In ECM software, the non-adjacent MVP is inserted between the TMVP and HMVP, where the distance between the non-adjacent spatial candidate and the current codec block is based on, for example... Figure 5 The width and height of the current codec block are shown.
[0063] 2.2 Affine Motion Compensation Prediction In HEVC, only a translational motion model is applied for motion compensation prediction (MCP). In the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transformation motion compensation prediction is applied. Figure 6A and Figure 6B An affine motion model based on control points is shown. For example... Figure 6A and Figure 6B As shown, the affine motion field of a block is described by motion information from two control points (4 parameters) or three control point motion vectors (6 parameters).
[0064] For the 4-parameter affine motion model, the motion vector at the sample point position (x, y) in the block is derived as: (1); For the 6-parameter affine motion model, the motion vector at the sample point position (x, y) in the block is derived as: (2), in( mv0x, mv0y ) is the motion vector of the upper left control point, ( mv1x, mv1y ) is the motion vector of the upper right control point, and ( mv2x, mv2y ) is the motion vector of the lower left control point.
[0065] To simplify motion compensation prediction, a block-based affine transformation prediction is applied. To derive the motion vector for each 4×4 lumen sub-block, the motion vector of the center sample point of each sub-block is calculated according to the above equation (e.g., ...). Figure 7 (as shown), and rounded to 1 / 16 fractional precision. Then, a motion-compensated interpolation filter is applied to generate a prediction for each sub-block with a derived motion vector. The sub-block size for the chroma component is also set to 4×4. The MV of the 4×4 chroma sub-block is calculated as the average of the MV of the upper-left luminance sub-block and the lower-right luminance sub-block in the corresponding 8×8 luminance region.
[0066] Similar to translational motion inter-frame prediction, there are two affine motion inter-frame prediction modes: affine Merge mode and affine AMVP mode.
[0067] 2.2.1 Affine Merge Prediction The Affine Merge pattern can be applied to CUs with a width and height greater than or equal to 8. In this pattern, the CPVM of the current CU is generated based on the motion information of spatially neighboring CUs. There can be up to five CPVM candidates, and the one to be used for the current CU is indicated by a signal transmission index. In VVC, the following three types of CPVM candidates are used to form the Affine Merge candidate list: - Inherited affine Merge candidates inferred from the CPMV of neighboring CUs; - Affine Merge candidate CPMVP constructed using translational MV derivation of neighboring CUs; - Zero MV.
[0068] In VVC, there are at most two inherited affine candidates, which are derived from the affine motion model of neighboring blocks: one from the left neighboring CU and one from the upper neighboring CU. Candidate blocks are as follows: Figure 8As shown. For the predicted value on the left, the scan order is A0->A1, and for the predicted value above, the scan order is B0->B1->B2. Only the first inherited candidate from each side is selected. No deduplication check is performed between candidates from two inheritances. When a neighboring affine CU is identified, its control point motion vector is used to derive the CPMVP candidate in the affine Merge list of the current CU. Figure 9 As shown, if the adjacent lower-left block A is encoded and decoded in affine mode, then the motion vectors of the upper-left, upper-right, and lower-left corners of the CU containing block A are... , and Obtained. When block A is encoded and decoded using a 4-parameter affine model, the two CPMVs of the current CU are based on... and Calculation. When block A is encoded and decoded using a 6-parameter affine model, the three CPMVs of the current CU are calculated according to... , and calculate.
[0069] Figure 8 The location of the inherited affine motion prediction value is shown. Figure 9 The inheritance of control point motion vectors is shown.
[0070] The constructed affine candidate refers to the candidate built by combining the translational motion information of the neighbors of each control point. The motion information of the control points is derived from... Figure 10 The spatial and temporal nearest neighbors shown are derived. CPMVk (k=1, 2, 3, 4) represents the k-th control point. For CPMV1, check the B2->B3->A2 block and use the MV of the first available block. For CPMV2, check the B1->B0 block, and for CPMV3, check the A1->A0 block. If available, the TMVP is used as CPMV4.
[0071] After obtaining the motion signatures (MVs) of the four control points, the affine merge candidate is constructed based on this motion information. The following combinations of control point MVs are used for sequential construction: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}.
[0072] Combining three CPMVs constructs a 6-parameter affine merge candidate, and combining two CPMVs constructs a 4-parameter affine merge candidate. To avoid motion scaling, combinations of control point MVs are discarded if the reference indices of the control points are different. Figure 10The locations of candidate positions for constructing the affine Merge pattern are shown.
[0073] After the inherited affine Merge candidate and the constructed affine Merge candidate are checked, if the list is still not full, zero MV is inserted at the end of the list.
[0074] 2.2.2 Affine AMVP Prediction The affine AMVP mode can be applied to CUs with a width and height greater than or equal to 16. An affine flag at the CU level is signaled in the bitstream to indicate whether the affine AMVP mode is used, and another flag is signaled to indicate whether it is a 4-parameter affine or a 6-parameter affine. In this mode, the difference between the current CU's CPVM and its predicted CPMVP is signaled in the bitstream. The affine AMVP candidate list is of size 2 and is generated by sequentially using the following four types of CPVM candidates: – Inherited affine AMVP candidates inferred from the CPMV of neighboring CUs; – A constructive affine AMVP candidate CPMVP derived using translational MV of neighboring CUs; – Translation MV from the neighboring CU; – Zero MV.
[0075] The checking order for inherited affine AMVP candidates is the same as that for inherited affine Merge candidates. The only difference is that, for AVMP candidates, only affine CUs with the same reference picture as those in the current block are considered. No deduplication is applied when inserting inherited affine motion predictions into the candidate list.
[0076] The constructed AMVP candidate is from Figure 10 The specified spatial nearest neighbor derivation is shown. The same checking order as in the affine Merge candidate construction is used. Additionally, the reference picture index of neighboring blocks is checked. The first block in the checking order is used, which is inter-frame encoded and has the same reference picture as in the current CU. There is only one. Encoded and decoded in the current CU using a 4-parameter affine mode, and... mv0 and mv1 When all three CPMVs are available, they are added as candidates in the affine AMVP list. If the current CU is encoding / decoding in a 6-parameter affine mode and all three CPMVs are available, they are added as candidates in the affine AMVP list. Otherwise, the constructed AMVP candidates are set to unavailable.
[0077] Figure 11A and Figure 11B The spatial nearest neighbors used to derive affine Merge candidates are shown: Figure 11A Used to derive affine Merge candidates for inheritance, and Figure 11B Used to derive the constructed affine Merge candidate.
[0078] If the affine AMVP list still has fewer than 2 candidates after inserting valid inherited affine AMVP candidates and constructed AMVP candidates, then when available, mv0 , mv1 and mv2 The translation MVs will be added sequentially to predict the MVs of all control points in the current CU. Finally, if the affine AMVP list is still not full, it will be filled with zero MVs.
[0079] 2.2.3 A Novel Affine Candidate Derivation Method ECM-6.0 integrates three additional affine Merge and AMVP candidate derivation methods: non-adjacent spatial domain candidates, historical parameter-based candidates, and regression-based affine candidates.
[0080] 2.2.3.1 Non-adjacent airspace candidates In ECM-6.0, non-adjacent airspace neighbors were studied to provide candidates for both affine Merge and affine AMVP. The format for obtaining non-adjacent airspace candidates is as follows: Figure 11A and Figure 11B As shown. Similar to non-adjacent regular merge candidates, the distance between non-adjacent spatial candidates and the current codec block is also defined based on the width and height of the current CU.
[0081] Figure 11A and Figure 11B Motion information of non-adjacent spatial neighbors is used to generate additional inheritance and construct affine merge candidates. Specifically, to generate inheritance candidates, non-adjacent spatial neighbors are checked based on their distance from the current block (e.g., from nearest to farthest). At a specific distance, only the first available neighbors encoded in affine mode from each side (e.g., left and top) of the current block are included. Figure 11A As shown, the checks of the left and top nearest neighbors are performed from bottom to top and from right to left, respectively. For the constructed candidates, such as... Figure 11B As shown, the positions of non-adjacent spatial neighbors on the left and top are first determined independently; then, the positions of the upper left neighbors can be determined accordingly to form a rectangular virtual block together with the non-adjacent neighbors on the left and top. Figure 12 The diagram illustrates the affine Merge candidates from non-nearest neighbors to the build. Motion information from three non-nearest neighbors is used to form a CPMV at the top-left (A), top-right (B), and bottom-left (C) of the virtual block, which is projected onto the current CU to generate the corresponding build candidates, as shown below. Figure 12 As shown.
[0082] 2.2.3.2 Affine Candidates Based on Historical Parameters History-based Affine Model Inheritance (HAMI) allows affine models to inherit from previously affinely encoded / decoded blocks that may not be adjacent to the current block. A History-based Table (HPT) is created. Each entry in the HPT stores a set of affine parameters: a, b, c, and d, each represented by a 16-bit signed integer. Entry points in the HPT are categorized by reference lists and reference indices. Five reference indices are supported for each reference list in the HPT. The HPT category (denoted as HPTCat) is calculated in a formulaic manner. HPTCat (RefList, RefIdx) = 5×RefList + min (RefIdx, 4)(3) Here, RefList and RefIdx represent the list of reference images (0 or 1) and the reference index, respectively. A maximum of 7 entries can be stored for each category, resulting in a total of 70 entries in the HPT. At the beginning of each CTU row, the number of entries for each category is initialized to zero. After decoding the affine-encoded CU using the reference lists RefListcur and RefIdxcur, the affine parameters are used to update the entries in the category HPTCat(RefListcur, RefIdxcur) in a manner similar to HMVP table updates.
[0083] Candidates based on historical affine parameters (HAPC) from... Figure 10 The derivation of a set of affine parameters, represented as A0, A1, B0, B1, or B2, and their corresponding entries stored in the HPT, is as follows: The MV of the neighboring 4×4 blocks is used as the base MV. The MV of the current block at position (x, y) is calculated in a formulaic manner as follows: (4), Where (mvhbase, mvvbase) represents the MV of the nearest 4×4 blocks, and (xbase, ybase) represents the center position of the nearest 4×4 blocks. (x, y) can be the top left, top right, and bottom left corners of the current block to obtain the corner position MV (CPMV) for the current block, or it can be the center of the current block to obtain the regular MV for the current block.
[0084] Figure 13An example of how to derive the HAPC from block A0 is shown. The affine parameters {a0, b0, c0, d0} are obtained directly from an entry in the category HPTIdx(RefListA0, refIdx0A0) in the HPT. The affine parameters from the HPT (with the center position of A0 as the base position and the MV of block A0 as the base MV) are used together to derive the CPMV for either the affine MergeHAPC or the affine AMVP HAPC. They can also be used to derive the MV located at the center of the current block as a regular Merge candidate. The HAPC can be placed into the sub-block-based Merge candidate list, the affine AMVP candidate list, or the regular Merge candidate list. In response to the introduction of new HAPCs, the size of the sub-block-based Merge candidate list is increased from 5 to 10 and 12 for random access and low-latency B configurations, respectively. Furthermore, for the random access configuration, the size of the regular Merge candidate list is increased from 10 to 11 to accommodate the newly added regular Merge candidates.
[0085] 2.2.3.3 Regression-based Affine Candidates In ECM-6.0, regression-based affine merge candidates are derived and added to the affine merge list. The sub-block motion fields from previously encoded and decoded affine CUs and the motion information of neighboring sub-blocks from the current CU are used as inputs to the regression process to derive the proposed affine candidates.
[0086] Previously encoded and decoded affine CUs can be identified from scans of non-adjacent locations and the affine HMVP table. The neighboring sub-block information of the current CU is obtained from, for example... Figure 14 This is obtained from the 4x4 sub-blocks represented by the gray area shown. For each sub-block, given a list of references, the corresponding motion vector and center coordinates of the sub-block can be used.
[0087] For each affine CU, at most two affine candidates can be derived. One has neighboring subblock information, and the other does not. All candidates generated by linear regression are deduplicated and collected into a candidate subgroup. When ARMC is enabled, the ARMC process based on TM cost is applied. Subsequently, when N affine CUs are found, at most N candidates generated by linear regression are added to the affine merge list. Figure 14 This is a diagram illustrating the affine Merge candidate derivation based on regression.
[0088] 2.3 Template Matching Merge / AMVP Pattern in ECM Template Matching (TM) Merge / AMVP mode is a decoder-side MV derivation method used to refine the motion information of the current CU by finding the closest match between a template in the current image (i.e., the top and / or left neighboring blocks of the current CU) and a block in the reference image (i.e., of the same size as the template). Figure 15 As shown, within the search range of [-8, +8] pixels, a better MV is searched around the initial motion of the current CU. Figure 15 This illustrates template matching execution over the search area surrounding the initial MV.
[0089] In AMVP mode, an MVP candidate is determined based on the template matching error, selecting the one that minimizes the difference between the current block and the reference block template. Then, the TM process performs MV refinement only on that specific MVP candidate. The TM refines the MVP candidate using an iterative diamond search, starting with full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode) within a search range of [-8, +8] pixels. The AMVP candidate can be further refined using a cross search with full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode), followed by half-pixels and quarter-pixels sequentially depending on the AMVR mode. This search process ensures that the MVP candidate maintains the same MV precision as indicated by the Adaptive Motion Vector Resolution (AMVR) mode after the TM process.
[0090] In Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. TMMerge can proceed up to 1 / 8 pixel MVD accuracy, or skip those accuracies beyond half-pixel MVD accuracy, depending on whether an alternative interpolation filter is used based on the merged motion information (used when AMVR is in half-pixel mode). Furthermore, when TM mode is enabled, template matching can operate as a standalone process, or as an additional MV refinement process between block-based and sub-block-based bilateral matching (BM) methods, depending on whether BM is enabled according to its enable condition check. When both BM and TM are enabled for the CU, the TM search process stops at half-pixel MVD accuracy, and the resulting MV is further refined using the same model-based MVD derivation method as in DMVR.
[0091] 2.4 Adaptive Reordering of Merge Candidates (ARMC) Inspired by the spatial correlation between reconstructed neighboring pixels and the current codec block, we propose Adaptive Reordering of Merge Candidates (ARMC) to refine the order of candidates in a given candidate list. The basic assumption is that candidates with lower template matching costs have a higher probability of being selected through the RDO process and should therefore be placed earlier in the list to reduce signaling costs.
[0092] The reordering method is applied to the regular Merge pattern, the Template Matching (TM) Merge pattern, and the Affine Merge pattern (excluding SbTMVP candidates). For the TM Merge pattern, the Merge candidates are reordered before the refinement process.
[0093] After constructing the Merge candidate list, the Merge candidates are divided into several subgroups. The subgroup size is set to 5. The Merge candidates in each subgroup are reordered in ascending order based on the cost value of template matching. For simplicity, the Merge candidates in the last subgroup (not the first subgroup) are not reordered.
[0094] Template matching cost is measured by the sum of absolute differences (SAD) between the samples of the current block's template and its corresponding reference template. For example... Figure 16 As shown, the template includes a set of reconstructed samples adjacent to the current block, while the reference template is located using the same motion information of the current block. When the Merge candidate utilizes bidirectional prediction, the reference samples of the Merge candidate's template are also generated through bidirectional prediction.
[0095] For a sub-block-based merge candidate with a sub-block size equal to Wsub * Hsub, the upper template includes several sub-templates of size Wsub × K, and the left template includes several sub-templates of size K × Hsub. For example... Figure 17 As shown, the motion information of the sub-blocks in the first row and first column of the current block is used to derive the reference sample points of each sub-template.
[0096] 2.5 Sub-block-based temporal motion vector prediction (SbTMVP) VVC supports the Sub-Block-Based Temporal Motion Vector Prediction (SbTMVP) method. Similar to TMVP, SbTMVP leverages motion fields in co-located images to facilitate more accurate MVP derivation. The same co-located image used by TMVP is used for SbTVMP. SbTMVP differs from TMVP primarily in two ways. First, SbTMVP enables motion prediction at the sub-CU level, while TMVP predicts motion at the CU level. Second, compared to TMVP, which obtains temporal MVs from co-located blocks in a co-located image (where a co-located block is the lower right or center block relative to the current CU), SbTMVP applies motion shifting before obtaining temporal motion information from the co-located image. This motion shifting is achieved by reusing the MV of one of the spatially neighboring blocks from the current CU. Figure 16 The template and the corresponding reference template are shown.
[0097] Figure 18 The derivation process for the sub-block level motion field for SbTMVP is shown. Specifically, the motion information of the lower left sub-block A1 is first obtained. If any MV in reference list 0 and list 1 points to the same frame, the corresponding MV will be marked as a motion shift. Otherwise, zero MV will be used as a motion shift.
[0098] Once the motion shift is determined, a designated region in the same frame is used to derive the sub-block-level motion field. For example... Figure 15 As shown, assuming the motion of A1 is used for motion displacement, then for each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the center sample point) in the co-location image is obtained to provide motion information, wherein an MV scaling operation is first performed to align the reference frame of the temporal motion vector with the reference frame of the current CU. Figure 17 The template and reference template of a block with sub-block motion information are shown. Figure 18 The derivation of the sub-CU motion field obtained by applying motion displacement based on neighbor motion information is shown.
[0099] In VVC and ECM, in addition to the CU-level MVP candidate list, a sub-CU-level MVP candidate list is also constructed to provide more accurate motion predictions for the current CU. This list includes the motion field generated by both the SbTMVP and AFFINE methods. Specifically, only one SbTMVP candidate is included, and this SbTMVP candidate is always placed as the first entry in the constructed sub-CU-level MVP candidate list. Multiple affine candidates are included in the list after performing template matching-based reordering, with those affine candidates with lower costs placed earlier.
[0100] 2.6 Geometric Partitioning (GPM) In VVC, geometric segmentation modes are supported for inter-frame prediction. A CU-level flag is used as a merge mode to transmit the geometric segmentation mode via signaling. Other merge modes include regular merge mode, MMVD mode, CIIP mode, and sub-block merge mode. A total of 64 segments are supported for each possible CU size. ,in Excluding 8x64 and 64x8.
[0101] When using this mode, the CU is divided into two geometric segments by a straight line of geometric positioning. Figure 19 The position of the dividing line is mathematically derived from the angle and offset parameters of the specific segment. Each part of the geometric segment in the CU is predicted inter-frame using its own motion; only unidirectional prediction is allowed for each segment, i.e., each part has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, as with regular bidirectional prediction, only two motion-compensated predictions are required per CU.
[0102] If a geometric segmentation pattern is used for the current CU, the geometric segmentation index (angle and offset) and two merge indices (one for each segment) are further indicated via signal transmission. The number of maximum GPM candidate sizes is explicitly transmitted via signal transmission in SPS, and the syntax binarization used for the GPM merge indices is specified. After predicting each part of the geometric segmentation, a blending process with adaptive weights is used to adjust the sample values along the geometric segmentation edges. This is the prediction signal for the entire CU, and the transformation and quantization processes are applied to the entire CU as in other prediction patterns. Finally, the motion field of the CU predicted using the geometric segmentation pattern is stored.
[0103] 2.6.1 Construction of One-Way Prediction Candidate List The unidirectional prediction candidate list is directly derived from the Merge candidate list constructed according to the extended Merge prediction process. Let n denote the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector of the nth extended Merge candidate (where X equals the parity of n) is used as the nth unidirectional prediction motion vector for the geometric segmentation pattern. These motion vectors in... Figure 20 The value is marked with "x". If the corresponding LX motion vector of the nth extended Merge candidate does not exist, then the L(1) of the same candidate... The X motion vector is used as a unidirectional predictive motion vector for the geometric segmentation pattern.
[0104] 2.6.2 Blending along geometric segmentation edges After each part of the geometric segmentation using its own motion prediction, blending is applied to the two prediction signals to derive samples around the geometric segmentation edges. The blending weights at each location of the CU are derived based on the distance between the individual location and the segmentation edge.
[0105] Location The distance to the segmentation edge is derived as follows: (2-1) (2-2) (2-3) (2-4) in It is an index for the angle and offset of the geometric segmentation, which depends on the geometric segmentation index transmitted via signal. and The sign depends on the angle index. .
[0106] The weights of each part of the geometric segment are derived as follows: (2-5) (2-6) (2-7)
[0107] partIdx depends on the angle index Weight An example in Figure 21 As shown in the image, the Figure 21 The bending weights using the geometric segmentation pattern are shown. An example of generation.
[0108] 2.6.3 Geometric Partitioning Pattern (GPM) with Merge Motion Vector Difference (MMVD) GPM in VVC extends the existing GPM unidirectional MV by applying motion vector refinement. First, a flag is transmitted to the GPMCU to specify whether to use this mode. If this mode is used, each geometric segment of the GPM CU can further determine whether to transmit MVD via signal transmission. If MVD is transmitted via signal transmission for a geometric segment, the segment's motion is further refined using the transmitted MVD information after selecting a GPM Merge candidate. All other processes remain the same as in GPM.
[0109] Similar to MMVD, MVD is transmitted as a pair of distance and direction via signals. GPM with MMVD (GPM-MMVD) involves nine candidate distances (1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixel, 3 pixel, 4 pixel, 6 pixel, 8 pixel, 16 pixel) and eight candidate directions (four horizontal / vertical directions and four diagonal directions). Additionally, when pic_fpel_mmvd_enabled_flag equals 1, MVD is shifted left by 2, just as in MMVD.
[0110] 2.6.4 Geometric Partitioning Mode with Adaptive Blending (GPM) In VVC, the final predicted samples are generated by weighted averaging of the predictions from the two predicted signals. This is achieved using two integer mixing matrices (...). W 0 and W 1). The weights in the GPM blending matrix are derived from the ramp function based on the displacement from the predicted sample location to the GPM segmentation boundary. The blending region size is fixed at 2 (2 samples on each side of the GPM segmentation boundary).
[0111] By adding four additional blending region sizes (one-quarter, half, twice, and four times the existing region size), such as Figure 22 As shown, the mixing process in the ECM is improved. The CU-level flags are encoded and decoded to select the mixing region size via signal transmission. Furthermore, extended weighting precision is utilized, where the maximum weight value is changed from 8 (in VVC) to 32 to accommodate the extended mixing region size. Figure 22 The ramp function for GPM mixing is shown, based on the displacement (d) from the predicted sample location to the GPM segmentation boundary and the size of the mixing region (τ).
[0112] 2.6.5 Geometric Segmentation Pattern (GPM) with Template Matching (TM) Template matching is applied to GPM. When GPM mode is enabled for CU, a CU-level flag is transmitted via signaling to indicate whether TM is applied to the two geometric segments. Motion information for each geometric segment is refined using TM. When TM is selected, a template is constructed using neighboring samples from the left and top, or left and top, depending on the segmentation angle, as shown in Table 1. Then, with the half-pixel interpolation filter disabled, motion is refined by minimizing the difference between the current template and the template in the reference image using the same search style of Merge mode.
[0113] Table 1. Templates for the first and second geometric segmentations, where A indicates the use of the top sample point, L indicates the use of the left sample point, and L+A indicates the use of both the left and top sample points.
[0114]
[0115] The GPM candidate list is constructed as follows: The interleaved list 0 MV candidates and list 1 MV candidates are directly derived from the regular merge candidate list, where list 0 MV candidates have a higher priority than list 1 MV candidates. A deduplication method with an adaptive threshold based on the current CU size is applied to remove redundant MV candidates.
[0116] The interleaved list 1 MV candidates and list 0 MV candidates are further derived directly from the regular merge candidate list, where list 1 MV candidates have a higher priority than list 0 MV candidates. The same deduplication method with an adaptive threshold is also applied to remove redundant MV candidates.
[0117] Zero MV candidates are filled until the GPM candidate list is full.
[0118] GPM-MMVD and GPM-TM are exclusively enabled for a single GPM CU. This is done first by signaling the GPM-MMVD syntax. When both GPM-MMVD control flags are false (i.e., GPM-MMVD is disabled for both GPM segments), the GPM-TM flag is signaled to indicate whether template matching is applied to both GPM segments. Otherwise (if at least one GPM-MMVD flag is true), the value of the GPM-TM flag is presumed to be false.
[0119] 2.6.6 GPM with inter-frame and intra-frame prediction In a GPM with inter-frame and intra-frame prediction, the final prediction samples are generated by weighting the inter-frame and intra-frame prediction samples for each GPM region. Inter-frame prediction samples are derived from the inter-frame GPM, while intra-frame prediction samples are derived from the intra-frame prediction mode (IPM) candidate list and the index from the encoder transmitted through the signal. The IPM candidate list size is predefined as 3. The available IPM candidates are the parallel angle mode (parallel mode) opposite to the GPM block boundary, the vertical angle mode (vertical mode) opposite to the GPM block boundary, and the planar mode, such as... Figures 23A to 23C As shown. Furthermore, as... Figure 23D The GPM with intra-frame and intra-frame prediction shown is limited to reduce signaling overhead for IPM and avoid increasing the size of intra-frame prediction circuitry on the hardware decoder. Additionally, direct motion vectors and IPM storage are introduced over the GPM mixing region to further improve encoding / decoding performance.
[0120] In IPM derivation based on DIMD and neighboring modes, parallel modes are registered first. Therefore, if no identical IPM candidates exist in the list, up to two IPM candidates can be registered from the decoder-side intra-frame mode derivation (DIMD) method and / or neighboring block derivation. As for neighboring mode derivation, up to five locations are available for neighboring blocks, but they are limited by the angle of the GPM block boundary, as shown in Table 2, and have already been used for GPM with template matching (GPM-TM).
[0121] Table 2. Positions of available neighboring blocks for IPM candidate derivation based on the angle of the GPM block boundary. A and L represent the top and left sides of the predicted block.
[0122]
[0123] GPM-intraframe can be combined with GPM with Merge Motion Vector Difference (GPM-MMVD). TIMD is used as an IPM candidate for GPM-intraframe to further improve encoding and decoding performance. Parallel modes can be registered first, followed by TIMD, DIMD, and IPM candidates for neighboring blocks.
[0124] 2.6.7 Template Matching-Based Reordering for GPM Partitioning Patterns In template-match-based reordering for GPM partitioning patterns, given the motion information of the current GPM block, the corresponding TM generation value for each GPM partitioning pattern is calculated. Then, all GPM partitioning patterns are reordered in ascending order based on their TM generation values. Instead of sending the GPM partitioning patterns, an index using Golomb-Rice codes is transmitted via signaling, indicating the exact location of the GPM partitioning pattern in the reordering list.
[0125] The reordering method for GPM partitioning patterns is a two-step process performed after the corresponding reference templates for the two GPM partitions in the encoding / decoding unit are generated, as shown below: ● Extend the GPM segmentation edge to the reference templates of the two GPM segments to obtain 64 reference templates, and calculate the corresponding TM cost for each of the 64 reference templates; ● The TM generation values based on the GPM partitioning pattern are reordered in ascending order, and the 32 best partitioning patterns are marked as available partitioning patterns.
[0126] The edges on the template extend from the edges of the current CU, such as Figure 24 As shown, however, the GPM blending process is not used in template regions across edges.
[0127] After the index is reordered in ascending order using the TM cost, it is transmitted via signaling.
[0128] 2.6.8 Motion field storage for geometric segmentation patterns Mv1 from the first part of the geometric segmentation, Mv2 from the second part of the geometric segmentation, and the combination Mv of Mv1 and Mv2 are stored in the motion field of the CU encoded and decoded by the geometric segmentation pattern.
[0129] The type of motion vector stored for each individual location in the sports field is determined as follows: (2-43) Where motionIdx equals It is recalculated from equation (2-36). partIdx depends on the angle index. .
[0130] If sType equals 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field; otherwise, if sType equals 2, then the combined Mv from Mv0 and Mv2 is stored. The combined Mv is generated using the following process: 1) If Mv1 and Mv2 come from different lists of reference images (one from L0 and the other from L1), then Mv1 and Mv2 are simply combined to form a bidirectional predicted motion vector.
[0131] 2) Otherwise, if Mv1 and Mv2 come from the same list, only the unidirectional predicted motion Mv2 is stored.
[0132] 2.7 Multiple Hypothesis Prediction (MHP) In the multi-hypothesis inter-frame prediction mode (JVET-M0425), in addition to the regular bidirectional prediction signal, one or more additional motion-compensated prediction signals are transmitted via signal transmission. The resulting overall prediction signal is obtained by weighted superposition of samples. The bidirectional prediction signal is utilized... and the first additional inter-frame prediction signal / hypothesis The generated prediction signal It is obtained as follows: .
[0133] According to the following mapping, the weighting factor It is specified by the new syntax element add_hyp_weight_idx.
[0134]
[0135] Similar to the above, more than one additional prediction signal can be used. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal.
[0136]
[0137] The resulting overall prediction signal serves as the final (i.e., has the largest index) of ) is obtained. Within this EE, a maximum of two additional prediction signals can be used (i.e., Limited to 2).
[0138] The motion parameters for each additional prediction hypothesis can be explicitly transmitted via signaling by specifying a reference index, motion vector prediction index, and motion vector difference, or implicitly transmitted via signaling by specifying a merge index. A separate multi-hypothesis merge flag distinguishes between these two signaling modes.
[0139] For inter-frame AMVP mode, MHP is applied only when unequal weights are selected in BCW in bidirectional prediction mode.
[0140] Combining MHP and BDOF is possible; however, BDOF is only applied to the bidirectional prediction signal portion of the predicted signal (i.e., the ordinary first two assumptions).
[0141] 2.8 Affine Motion Compensation in Geometric Prediction Mode A sub-block-based motion compensation method was proposed for use in GPM mode.
[0142] a) In one example, sub-block-based motion compensation can be affine motion compensation.
[0143] b) In one example, sub-block-based motion compensation could be sbTMVP motion compensation.
[0144] c) In one example, a prediction of at least one geometric segmentation can be generated using sub-block-based motion compensation, such as affine motion compensation.
[0145] d) In one example, the final prediction can be generated by a weighted sum of two predictions, at least one of which is generated using sub-block-based motion compensation such as affine motion compensation.
[0146] i. In one example, a weighted sum is performed using weight values defined by GPM.
[0147] e) In one example, the two predictions used in the GPM pattern can be type A and type B, where type A and type B can be (type A and type B can be the same type): i. Non-affine inter-frame prediction; ii. Affine inter-frame prediction; iii. Intra-frame prediction; iv. Intra-block copy (IBC) prediction; v. sb-TMVP inter-frame prediction; vi. Any combination or generated prediction.
[0148] abbreviation ACT Adaptive Color Transformation ALF Adaptive Loop Filter AMVR Adaptive Motion Vector Resolution APS Adaptive Parameter Set AU Access Unit AUD access unit delimiter AVC (Advanced Video Coding) (Recommendation ITU-T H.264 | ISO / IEC 14496-10) B Two-way prediction BCW features bidirectional prediction with CU-level weights. BDOF bidirectional optical flow BDPCM is based on block-based incremental pulse code modulation. BP buffer cycle CABAC Context-Based Adaptive Binary Arithmetic Encoding and Decoding CB codec block CBR constant bit rate CCALF Cross-Component Adaptive Loop Filter CPB encoded / decoded image cache CRA Pure Random Access CRC Cyclic Redundancy Check CTB codec tree block CTU encoding / decoding tree unit CU encoding / decoding unit CVS encoded video sequence DPB decodes image cache DCI decoding capability information DRAP depends on random access point DU decoding unit DUI Decoding Unit Information EG Index - Columbus EGk k-th exponent - Columbus End of EOB bitstream End of EOS sequence FD filtered data FIFO (First In First Out) FL fixed length GBR green, blue and red GCI General Constraints Information GDR is being gradually decoded and refreshed. GPM geometric segmentation mode HEVC High-Efficiency Video Codec (Recommendation ITU-T H.265 | ISO / IEC 23008-2) HRD Hypothetical Reference Decoder HSS Hypothesis Flow Scheduler Within I-frame IBC Intra-Block Copying IDR instant decoding refresh ILRP Inter-Frame Layer Reference Image IRAP Intra-Frame Random Access Point LFNST Low-Frequency Inseparable Transform LIC local lighting compensation LPS Lowest Probability Symbol LSB Least Significant Bit LTRP Long-Term Reference Image LMCS with chroma scaling luminance mapping MIP-based intra-frame prediction MPS Maximum Probability Symbol MSB most significant bit MTS Multiple Transformation Selection MVP motion vector prediction NAL Network Abstraction Layer OBMC Overlap Block Motion Compensation OLS Output Layer Set OP operation point OPI Operation Point Information P prediction PH image header POC image sequential counting PPS Image Parameter Set PROF refines the prediction using optical flow. PT image timer PU image unit QP quantization parameters RADL random access decodeable front-end (image) RASL random access skipped prerequisites (image) RBSP raw byte sequence payload RGB red, green and blue RPL Reference Image List SAO Sample Adaptive Compensation SAR sample amplitude ratio SEI Supplemental Enhancement Information SH strip head SLI sub-picture level information SODB data bit string SPS sequence parameter set STRP Short-Term Reference Image STSA Stepwise Temporal Sublayer Access TR discard Rice VBR Variable Bit Rate VCL video codec layer VPS Video Parameter Set VSEI Multifunctional Supplemental Enhancement Information (Recommendation ITU-T H.274 | ISO / IEC 23002-7) VUI Video Availability Information VVC Multi-Functional Video Codec (Recommendation ITU-T H.266 | ISO / IEC 23090-3) SE syntax elements 3. Problems to be solved In the following statements, “context” may refer to the context model used in arithmetic encoding and decoding.
[0149] 1) In VVC and ECM, SEs such as Merge indexes in GPM are encoded and decoded using only one context, which can be inefficient.
[0150] 2) In VVC and ECM, SEs such as Merge indexes in different cases, such as GPM and regular Merge, are encoded and decoded using the same context, which can be inefficient.
[0151] 3) For GPM, SEs such as Merge indexes are encoded and decoded using the same context with or without affine motion compensation, which can be inefficient.
[0152] 4) It is unclear how the SE is indicated by signal transmission whether affine motion compensation is being effectively applied.
[0153] 4. Detailed Solution The detailed embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.
[0154] The terms “video unit” or “code-decoder unit” or “block” can refer to code-decoder tree block (CTB), code-decoder tree unit (CTU), code-decoder block (CB), CU, PU, TU, PB, TB.
[0155] The term "affine block" can refer to a block encoded using affine Merge, affine AMVP, or any other affine variant mode (i.e., affine MMVD, etc.), which can be described by motion information of two control points (4 parameters) or three control point motion vectors (6 parameters). The term "CPMV" can refer to the motion information of an affine block at the top left, top right, and / or bottom left corners.
[0156] The term "template" can refer to a reconstructed region that can be used to refine a CPMV, and it can mean either a "separate template" or a "uniform template." Here, a "separate template" can refer to a reconstructed region that can be used to refine a single CPMV (i.e., a specific(s) CPMV(s) at the top left, top right, and / or bottom left corner), while a "uniform template" can refer to a reconstructed region that can be used to refine all or any(s) CPMV(s) for a block. The terms "template matching cost" or "TM cost" can refer to the matching cost of a separate template or the matching cost of a uniform template.
[0157] In this disclosure, with respect to "blocks encoded and decoded in mode N", "mode N" can be a predictive mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a encoding / decoding technique (e.g., DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-inter, GPM intra-intra, GPM inter-intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, ALF, deblocking, SAO, bilateral filter, LMCS and corresponding variants, etc.).
[0158] It should be noted that the following terms are not limited to the specific terms defined in existing standards. Any changes to encoding / decoding tools also apply.
[0159] In the following discussion, SE can be binarized into fixed-length code, EG(x) code, unary code, rounded unary code, rounded binary code, etc. It can be signed or unsigned.
[0160] 1. A first SE related to GPM was proposed that can be encoded and decoded using more than one context.
[0161] a) In one example, the first SE could be the GPM Merge index.
[0162] b) In one example, the first SE can be a GPM MMVD index.
[0163] c) In one example, the first SE could be the GPM bending weight index.
[0164] d) In one example, the first SE can be the GPM partitioning mode index.
[0165] e) In one example, the first SE can be the GPM intra-prediction mode.
[0166] f) In one example, SE can be binarized to represent B0, B1, ..., B N-1 N binary bits. In one example, if n != m, then B n and B m It can be encoded and decoded using different contexts.
[0167] i. In one example, B0, B1, ..., B W They can be represented as C0, C1, ..., C W Different contexts are encoded and decoded, while B W B W+1 B N-1 Shared can be represented as C W+1 The same context. W is an integer, such as 1, or 2, or 3, or 4, or 5, or 6, or 7.
[0168] ii. In one example, B0, B1, ..., B W They can be represented as C0, C1, ..., C W Different contexts are encoded and decoded, while B W B W+1 B N-1 It can be bypassed for encoding and decoding. W is an integer, such as 1, 2, 3, 4, 5, 6, or 7.
[0169] 2. It is proposed that at least one context for encoding and decoding a first SE related to GPM can be selected based on encoding and decoding information.
[0170] a) In one example, the first SE could be the GPM Merge index.
[0171] b) In one example, the first SE can be a GPM MMVD index.
[0172] c) In one example, the first SE could be the GPM bending weight index.
[0173] d) In one example, the first SE can be the GPM partitioning mode index.
[0174] e) In one example, the first SE can be the GPM intra-prediction mode.
[0175] f) In one example, the first SE can be the GPM intra-frame codec flag.
[0176] g) In one example, the first SE can be a GPM MMVD encoding / decoding flag.
[0177] h) In one example, the first SE can be a GPM™ codec flag.
[0178] i) In one example, the first SE can be a GPM affine codec flag.
[0179] j) In one example, context selection may depend on whether at least one GPM segment corresponding to the first SE is encoded or decoded using affine motion compensation.
[0180] i. For example, when the GPM segment is affine-coded or not, the context(s) used to encode and decode the GPM Merge index can be different.
[0181] ii. For example, when the GPM segment is affine-coded or not, the GPM MMVD(s) used to encode and decode the GPM segment can have different contexts.
[0182] iii. For example, when at least one GPM segment is affine-coded or not affine-coded, the context(s) used to encode the GPM bending weight index can be different.
[0183] k) In one example, context selection may depend on whether at least one GPM segment corresponding to the first SE is encoded or decoded using intra-frame prediction.
[0184] l) In one example, context selection may depend on whether at least one GPM segment corresponding to the first SE is encoded or decoded using the GPM MMVD mode.
[0185] m) In one example, context selection may depend on whether at least one GPM segment corresponding to the first SE is encoded or decoded using GPM TM mode.
[0186] n) In one example, the context selection for the first GPM segment may depend on the encoding / decoding information of the second GPM segment.
[0187] o) In one example, the context selection may depend on the GPM partitioning mode.
[0188] p) In one example, context selection may depend on mixed weights.
[0189] q) In one example, the context selection may depend on QP.
[0190] r) In one example, the context selection can depend on the stripe type.
[0191] 3. It is proposed that at least one context for encoding and decoding a first SE related to GPM can be selected based on encoding and decoding information of at least one neighboring block.
[0192] a) In one example, the first SE could be the GPM Merge index.
[0193] b) In one example, the first SE can be a GPM MMVD index.
[0194] c) In one example, the first SE could be the GPM bending weight index.
[0195] d) In one example, the first SE can be the GPM partitioning mode index.
[0196] e) In one example, the first SE can be the GPM intra-prediction mode.
[0197] f) In one example, the first SE can be the GPM intra-frame codec flag.
[0198] g) In one example, the first SE can be a GPM MMVD encoding / decoding flag.
[0199] h) In one example, the first SE can be a GPM™ codec flag.
[0200] i) In one example, the first SE can be a GPM affine codec flag.
[0201] j) In one example, the context used to encode and decode the SE can be selected from a candidate set that includes more than one candidate context.
[0202] i. In one example, the candidate set may include three members represented as {C[0], C[1], C[2]}, and the selected context is C[S].
[0203] 1) S may depend on the encoding / decoding mode of at least one neighboring block.
[0204] 2) For example, S is initialized to 0. If the left neighboring block is available and a certain condition is met, S is incremented by 1; if the upper neighboring block is available and a certain condition is met, S is incremented by 1.
[0205] 3) In one example, SE can be a GPM affine codec flag.
[0206] 4) In one example, a specific condition could be that neighboring blocks are affine encoded or decoded.
[0207] a) For example, if a block is encoded or decoded in affine-AMVP mode, it can be considered to be affine encoded or decoded.
[0208] b) For example, if a block is encoded or decoded in affine-Merge mode, it can be considered to be affine encoded or decoded.
[0209] c) In one example, if at least one of the GPM affine flags of a block is true, then the block can be decoded as affine encoded.
[0210] d) In one example, if the GPM affine flag of a block is true, then the block can be decoded as affine encoded.
[0211] 4. In one example, the context of the affine flag used to encode / decode a block may depend on whether neighboring blocks are encoded / decoded in GPM affine mode.
[0212] a) In one example, the context used to encode and decode affine flags can be selected from a candidate set that includes more than one candidate context.
[0213] b) In one example, the candidate set may include three members represented as {C[0], C[1], C[2]}, and the selected context is C[S].
[0214] i. For example, S is initialized to 0. If the left neighboring block is available and a certain condition is met, S is incremented by 1; if the upper neighboring block is available and a certain condition is met, S is incremented by 1.
[0215] ii. In one example, a specific condition could be that neighboring blocks are affine encoded or decoded.
[0216] 1) For example, if a block is encoded or decoded in affine-AMVP mode, it can be considered to be affine encoded or decoded.
[0217] 2) For example, if a block is encoded or decoded in affine-Merge mode, it can be considered to be affine encoded or decoded.
[0218] 3) In one example, if at least one of the GPM affine flags of a block is true, then the block can be decoded as affine encoded.
[0219] 4) In one example, if the GPM affine flag of a block is true, then the block can be decoded as affine encoded.
[0220] c) For example, the context can be selected based on the number of available affine-coded neighboring blocks (denoted as N).
[0221] i. For example, if N <= T, the first context can be used; if N > T, the second context can be used. T is an integer, such as 0, 1, 2, 3, 4, 5, or 6.
[0222] ii. May include, for example Figure 1 The five neighboring blocks shown, or as... Figure 7 The seven neighboring sample points are shown.
[0223] iii. For example, if a block is encoded or decoded in affine-AMVP mode, it can be considered as affine encoded or decoded.
[0224] iv. For example, if a block is encoded or decoded in affine-Merge mode, it can be considered to be affine encoded or decoded.
[0225] v. In one example, if at least one of the GPM affine flags of a block is true, then the block can be decoded as affine-encoded.
[0226] vi. In one example, if the GPM affine flag of a block is true, then the block can be decoded as affine encoded / decoded.
[0227] 5. The maximum permissible value of the first SE may depend on whether at least one GPM segment corresponding to the first SE is encoded or decoded using affine motion compensation.
[0228] a) In one example, the first SE could be the GPM Merge index.
[0229] b) In one example, the first SE can be a GPM MMVD index.
[0230] c) In one example, the first SE could be the GPM bending weight index.
[0231] d) In one example, the first SE can be the GPM partitioning mode index.
[0232] e) In one example, the first SE can be the GPM intra-prediction mode.
[0233] f) If the SE is encoded or decoded using a rounding unary code (e.g., a rounding unary code or a rounding binary code), the maximum allowed value can determine the last valid codeword of the SE.
[0234] g) For example, the allowed value of the first SE can be transmitted via signaling for both affine codec and non-affine codec cases, such as in VPS / SPS / PPS / image header / strip header.
[0235] 6. In one example, the maximum allowed value of the first SE can be determined based on the number of available affine-coded neighboring blocks (denoted as N).
[0236] a) In one example, the first SE could be the GPM Merge index.
[0237] b) In one example, the first SE can be a GPM MMVD index.
[0238] c) In one example, the first SE could be the GPM bending weight index.
[0239] d) In one example, the first SE can be the GPM partitioning mode index.
[0240] e) In one example, the first SE can be the GPM intra-prediction mode.
[0241] f) For example, if N <= T, the first maximum allowed value can be used; if N > T, the maximum allowed value can be used. T is an integer, such as 0, 1, 2, 3, 4, 5, or 6.
[0242] g) may include, for example Figure 1 The five neighboring blocks shown, or as... Figure 7 The seven neighboring sample points are shown.
[0243] h) For example, if a block is encoded or decoded in affine-AMVP mode, it can be considered as affine encoded or decoded.
[0244] i) For example, if a block is encoded or decoded in affine-Merge mode, it can be considered to be affine encoded or decoded.
[0245] j) In one example, if at least one of the GPM affine flags of a block is true, then the block can be decoded as affine encoded / decoded.
[0246] k) In one example, if the GPM affine flag of a block is true, then the block can be decoded as affine encoded.
[0247] 7. The weight values used in GPM may depend on whether at least one GPM segment is affine encoded or decoded.
[0248] a) In one example, the weight values used in GPM can differ in different situations, for example: i. Both GPM segments are affine encoded and decoded.
[0249] ii. Both GPM segments are non-affine encoded / decoded.
[0250] iii. One GPM segment is encoded and decoded affinely, and the other GPM segment is encoded and decoded non-affinely.
[0251] General aspects 8. Additional operations can be applied to the proposed method or applied together with the proposed method.
[0252] a) The syntax elements disclosed above can be binarized into flags, fixed-length codes, EG(x) codes, unary codes, rounded unary codes, rounded binary codes, etc. These can be signed or unsigned.
[0253] b) If a codec tool or codec method is deemed unsuitable or unusable, it means that the syntax elements of the codec tool or codec method may not be transmitted via signal and are implicitly determined to be unused.
[0254] c) The syntax elements disclosed above can be encoded or decoded using at least one context model. Alternatively, they can be encoded or decoded using a bypass method.
[0255] d) The syntax elements disclosed above can be transmitted conditionally via signals.
[0256] a. SE is transmitted via signal only when the corresponding function is applicable.
[0257] b. SE is transmitted via signal only when the dimensions of the block (width and / or height) meet the conditions.
[0258] e) The syntax elements disclosed above can be transmitted via signaling at the block level / sequence level / picture group level / picture level / strip level / piece group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0259] f) Whether and / or how the methods disclosed above can be applied to signal transmission at the block level / sequence level / picture group level / picture level / strip level / piece group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0260] g) Whether and / or how the methods disclosed above are applied may depend on the encoded / decoded information, such as block size, color format, single-tree / double-tree partitioning, color components, and stripe / picture type.
[0261] h) The methods presented in this document can be used in other codec tools that require chroma blending.
[0262] Figure 25 A flowchart of a method 2500 for video processing according to an embodiment of the present disclosure is shown. Method 2500 is implemented during the conversion between video units of a video and a bitstream of a video.
[0263] At box 2510, for the conversion between video units and video bitstreams, the first syntax element (SE) related to the geometric segmentation mode (GPM) is determined to be encoded and decoded using multiple contexts.
[0264] At box 2520, the conversion is performed based on the first SE. In some embodiments, the conversion includes encoding video units into a bitstream. In some embodiments, the conversion includes decoding video units from the bitstream. In this way, encoding / decoding efficiency and performance can be improved.
[0265] In some embodiments, the first SE is a GPM Merge index, or a GPM Merge Motion Vector Difference (MMVD) index. In some other embodiments, the first SE is a GPM Bending Weight index. In some further embodiments, the first SE is a GPM Partition Mode index. In some embodiments, the first SE is a GPM Intra-Prediction Mode, where the first SE is a GPM Intra-Codec Mode. In some other embodiments, the first SE is a GPM MMVD Codec Flag. In some further embodiments, the first SE is a GPM Template Matching (TM) Codec Flag. In some embodiments, the first SE is a GPM Affine Codec Flag.
[0266] In some embodiments, the first SE is binarized to be represented as B0, B1, ..., B N-1 It consists of N binary bits, where N is an integer. In some embodiments, if n is not equal to m, Bn and Bm are encoded and decoded using different contexts, where n and m are integers not greater than N-1.
[0267] In some embodiments, B0, B1, ..., B W Using the denoted C0, C1, ..., C respectively W Different contexts are encoded and decoded, and B W B W+1 B N-1 Shared is represented as C W+1In some embodiments, B0, B1, ..., B W Using the denoted C0, C1, ..., C respectively W Different contexts are encoded and decoded, and B W B W+1 B N-1 The code is bypassed and encoded / decoded. In this case, W is an integer. For example, W equals 1, 2, 3, 4, 5, 6, or 7.
[0268] In some embodiments, at least one context for encoding and decoding a first SE associated with a GPM is selected based on encoding and decoding information. In some embodiments, the selection of the context depends on whether at least one GPM segment corresponding to the first SE is encoded and decoded using affine motion compensation. For example, if the GPM segment is affinely encoded, a first context is used to encode the GPM Merge index of the GPM segment; if the GPM segment is not affinely encoded, a second context is used to encode the GPM Merge index of the GPM segment, wherein the first context is different from the second context. In some other embodiments, if the GPM segment is affinely encoded, a third context is used to encode the GPM MMVD index of the GPM segment; if the GPM segment is not affinely encoded, a fourth context is used to encode the GPM MMVD index of the GPM segment, wherein the third context is different from the fourth context. In some further embodiments, if the GPM segment is affinely encoded, a fifth context is used to encode the GPM curvature weight index of the GPM segment; if the GPM segment is not affinely encoded, a sixth context is used to encode the GPM curvature weight index of the GPM segment, wherein the fifth context is different from the sixth context.
[0269] In some embodiments, the selection of the context depends on whether at least one GPM segment corresponding to the first SE is encoded / decoded using intra-frame prediction. In some other embodiments, the selection of the context depends on whether at least one GPM segment corresponding to the first SE is encoded / decoded using GPM MMVD mode. In some further embodiments, the selection of the context depends on whether at least one GPM segment corresponding to the first SE is encoded / decoded using GPM TM mode. In some embodiments, the selection of the context for the first GPM segment depends on the encoding / decoding information of the second GPM segment. In some other embodiments, the selection of the context depends on the GPM partitioning mode. In some further embodiments, the selection of the context depends on the mixing weights. In some embodiments, the selection of the context depends on the quantization parameter (QP). In some other embodiments, the selection of the context depends on the stripe type.
[0270] In some embodiments, at least one context for encoding and decoding a first SE related to the GPM is selected based on encoding and decoding information of at least one neighboring block. In some embodiments, the context for encoding and decoding the first SE is selected from a candidate set that includes more than one candidate context. For example, the candidate set includes three context candidates denoted as {C[0], C[1], C[2]}, and the selected context is C[S], where S is an integer.
[0271] In some embodiments, S depends on the encoding / decoding mode of at least one neighboring block. In some other embodiments, S is initialized to 0, and incremented by 1 if the left neighboring block is available and a certain condition is met, and if the upper neighboring block is available and a certain condition is met.
[0272] In some embodiments, a specific condition is that neighboring blocks are affine-coded. For example, if a block is encoded in affine-AMVP mode, then the block is considered affine-coded. In some embodiments, if a block is encoded in affine-Merge mode, then the block is considered affine-coded. In some other embodiments, a block is decoded as affine-coded if at least one GPM affine flag of the block is true, or if the block's GPM affine flag is true. In some embodiments, the first SE is a GPM affine codec flag.
[0273] In some embodiments, the context for the affine flag used to encode / decode a block depends on whether neighboring blocks are encoded / decoded in GPM affine mode. For example, the context for the affine flag used to encode / decode is selected from a candidate set that includes more than one candidate context.
[0274] In some embodiments, the candidate set includes three context candidates denoted as {C[0], C[1], C[2]}, and the selected context is C[S]. In some embodiments, S is initialized to 0, and incremented by 1 if the left neighboring block is available and meets a certain condition, and incremented by 1 if the upper neighboring block is available and meets a certain condition.
[0275] In some embodiments, a specific condition is that neighboring blocks are affine-coded. In some embodiments, a block is considered affine-coded if it is encoded in affine-AMVP mode. In some other embodiments, a block is considered affine-coded if it is encoded in affine-Merge mode. In some further embodiments, a block is decoded as affine-coded if at least one GPM affine flag of the block is true. Alternatively, a block is decoded as affine-coded if its GPM affine flag is true.
[0276] In some embodiments, the context is selected based on the number of available affine-coded neighbor blocks. In some embodiments, a first context is used if N <= T, and a second context is used if N > T, where N represents the number of available affine-coded neighbor blocks, and T is an integer. For example, T is equal to 0, or 1, or 2, or 3, or 4, or 5, or 6.
[0277] In some embodiments, five neighboring blocks (e.g., such as...) Figure 4 One or more neighboring blocks (as shown) are included in the available affine-coded neighboring blocks. Alternatively, seven neighboring samples (e.g., as shown) are used. Figure 10 One or more neighboring samples (as shown) are included in the available affine-coded neighbor blocks, such as Figure 10 As shown.
[0278] In some embodiments, if a block is encoded / decoded in affine-AMVP mode, the block is considered affine-encoded. Alternatively, if a block is encoded / decoded in affine-Merge mode, the block is considered affine-encoded. In some other embodiments, a block is decoded as affine-encoded if at least one of its GPM affine flags is true. In some further embodiments, a block is decoded as affine-encoded if its GPM affine flag is true.
[0279] In some embodiments, the maximum permissible value of the first SE depends on whether at least one GPM segment corresponding to the first SE is encoded or decoded using affine motion compensation. In some embodiments, if the first SE is encoded or decoded using rounded unary or rounded binary codes, the maximum permissible value determines the last valid codeword of the first SE. In some other embodiments, the maximum permissible value of the first SE is transmitted via signaling for both affine and non-affine encoding / decoding cases.
[0280] In some embodiments, the maximum allowed value of the first SE is determined based on the number of available affine-coded neighbor blocks. For example, if N <= T, the first maximum allowed value is used, and if N > T, the second maximum allowed value is used. In this case, N represents the number of available affine-coded neighbor blocks, and T is an integer. In some embodiments, T is equal to 0, or 1, or 2, or 3, or 4, or 5, or 6.
[0281] In some embodiments, five neighboring blocks (e.g., such as...) Figure 4 One or more neighboring blocks (as shown) are included in the available affine-coded neighboring blocks. Alternatively, seven neighboring samples (e.g., as shown) are used. Figure 10 One or more neighboring samples (as shown) are included in the available affine-coded neighboring blocks.
[0282] In some embodiments, if a block is encoded / decoded in affine-AMVP mode, the block is considered affine-encoded. Alternatively, if a block is encoded / decoded in affine-Merge mode, the block is considered affine-encoded. In some other embodiments, a block is decoded as affine-encoded if at least one of its GPM affine flags is true. In some further embodiments, a block is decoded as affine-encoded if its GPM affine flag is true.
[0283] In some embodiments, the weight values used in the GPM depend on whether at least one GPM segment is affine-coded. Alternatively, the weight values used in the GPM may vary based on different conditions. For example, different conditions include at least one of the following: both GPM segments are affine-coded, both GPM segments are non-affine-coded, or one GPM segment is affine-coded and the other GPM segment is non-affine-coded.
[0284] In some embodiments, whether and / or how to determine whether the first SE is encoded and decoded using multiple contexts is transmitted via signaling at one of the following: sequence level, picture group level, picture level, strip level, or slice group level. In some embodiments, whether and / or how to determine whether the first SE is encoded and decoded using multiple contexts is transmitted via signaling at one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header. In some embodiments, whether and / or how to determine whether the first SE is encoded and decoded using multiple contexts is transmitted via signaling at one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), codec tree block (CTB), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region comprising more than one sample or pixel.
[0285] In some embodiments, method 2500 further includes: determining, based on the encoded / decoded information of the video unit, whether and / or how to determine that the first SE is encoded / decoded using multiple contexts. The encoded / decoded information may include at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.
[0286] In some embodiments, syntax elements are binarized into one of the following: a flag, a fixed-length code, an EG(x) code, a unary code, a rounded unary code, or a rounded binary code. In some embodiments, syntax elements are signed or unsigned.
[0287] In some embodiments, syntax elements are encoded and decoded using at least one context model, or syntax elements are encoded and decoded in a bypass manner. In some embodiments, syntax elements are conditionally transmitted via signaling. In some embodiments, syntax elements are transmitted via signaling if the corresponding functionality applies, or if the dimensions of the video unit satisfy a condition. In some embodiments, dimensions include the width and / or height of the video unit.
[0288] In some embodiments, syntax elements are signaled at one of the following levels: sequence level, picture group level, picture level, strip level, or slice group level. In some embodiments, syntax elements are signaled at one of the following levels: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header. In some embodiments, syntax elements are signaled at one of the following levels: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), codec tree block (CTB), codec tree unit (CTU), CTU row, strip, slice, subpicture, or region comprising more than one sample or pixel. In some embodiments, video units are encoded and decoded using one or more other codec tools requiring chroma blending.
[0289] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining a first syntax element (SE) associated with a geometric segmentation mode (GPM) that is encoded and decoded using multiple contexts; and generating a bitstream based on the first SE.
[0290] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: determining a first syntax element (SE) associated with a geometric segmentation mode (GPM) that is encoded and decoded using multiple contexts; generating a bitstream based on the first SE; and storing the bitstream in a non-transitory computer-readable recording medium.
[0291] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.
[0292] Item 1. A video processing method, comprising: a conversion between video units of a video and a bitstream of the video; determining a first syntax element (SE) associated with a geometric segmentation mode (GPM) that is encoded and decoded using multiple contexts; and performing the conversion based on the first SE.
[0293] Item 2. The method according to Item 1, wherein the first SE is a GPM Merge index, or wherein the first SE is a GPM Merge Motion Vector Difference (MMVD) index, or wherein the first SE is a GPM Bending Weight index, or wherein the first SE is a GPM Partition Mode index, or wherein the first SE is a GPM Intra Prediction Mode, or wherein the first SE is a GPM Intra Coding Mode, or wherein the first SE is a GPM MMVD Coding Flag, or wherein the first SE is a GPM Template Matching (TM) Coding Flag, or wherein the first SE is a GPM Affine Coding Flag.
[0294] Item 3. The method according to Item 1, wherein the first SE is binarized to be represented as B0, B1, ..., B N-1 It consists of N binary bits, where N is an integer.
[0295] Item 4. The method described in Item 3, wherein if n is not equal to m, Bn and Bm are encoded and decoded using different contexts, where n and m are integers not greater than N-1.
[0296] Item 5. The method described in Item 3, wherein B0, B1, ..., B W Using the denoted C0, C1, ..., C respectively W Different contexts are encoded and decoded, and B W B W+1 B N-1 Shared is represented as C W+1 The same context, where W is an integer.
[0297] Item 6. The method according to Item 3, wherein B0, B1, ..., B W Using the denoted C0, C1, ..., C respectively W Different contexts are encoded and decoded, and B W B W+1 B N-1 It is bypassed encoding / decoding, where W is an integer.
[0298] Item 7. The method according to Item 5 or 6, wherein W is equal to 1, or 2, or 3, or 4, or 5, or 6, or 7.
[0299] Item 8. The method according to Item 1, wherein at least one context for encoding and decoding the first SE in relation to GPM is selected based on encoding and decoding information.
[0300] Item 9. The method according to Item 8, wherein the choice of context depends on whether at least one GPM segment corresponding to the first SE is encoded or decoded using affine motion compensation.
[0301] Item 10. The method according to Item 9, wherein if the GPM segment is affine encoded, a first context is used to encode the GPM Merge index of the GPM segment, and if the GPM segment is not affine encoded, a second context is used to encode the GPM Merge index of the GPM segment, wherein the first context is different from the second context.
[0302] Item 11. The method according to Item 9, wherein if the GPM segment is affine encoded, a third context is used to encode the GPM MMVD index of the GPM segment, and if the GPM segment is not affine encoded, a fourth context is used to encode the GPM MMVD index of the GPM segment, wherein the third context is different from the fourth context.
[0303] Item 12. The method according to Item 9, wherein if the GPM segment is affine encoded, a fifth context is used to encode the GPM bend weight index of the GPM segment, and if the GPM segment is not affine encoded, a sixth context is used to encode the GPM bend weight index of the GPM segment, wherein the fifth context is different from the sixth context.
[0304] Item 13. The method according to Item 8, wherein the selection of the context depends on whether at least one GPM segment corresponding to the first SE is encoded or decoded using intra-frame prediction.
[0305] Item 14. The method according to Item 8, wherein the selection of the context depends on whether at least one GPM segment corresponding to the first SE is encoded or decoded using GPM MMVD mode.
[0306] Item 15. The method according to Item 8, wherein the selection of the context depends on whether at least one GPM segment corresponding to the first SE is encoded or decoded using GPM TM mode.
[0307] Item 16. The method according to Item 8, wherein the selection of the context for the first GPM segment depends on the encoding / decoding information of the second GPM segment.
[0308] Item 17. The method described in Item 8, wherein the choice of context depends on the GPM partitioning mode.
[0309] Item 18. The method described in Item 8, wherein the choice of context depends on the mixed weights.
[0310] Item 19. The method according to Item 8, wherein the choice of context depends on the quantization parameter (QP).
[0311] Item 20. The method described in Item 8, wherein the choice of context depends on the stripe type.
[0312] Item 21. The method according to Item 1, wherein at least one context for encoding and decoding the first SE in relation to the GPM is selected based on encoding and decoding information of at least one neighboring block.
[0313] Item 22. The method according to Item 21, wherein the context for encoding and decoding the first SE is selected from a candidate set including more than one candidate context.
[0314] Item 23. The method according to Item 22, wherein the candidate set comprises three context candidates denoted as {C[0], C[1], C[2]}, and the selected context is C[S], where S is an integer.
[0315] Item 24. The method according to Item 23, wherein S depends on the encoding / decoding mode of at least one neighboring block.
[0316] Item 25. The method according to Item 23, wherein S is initialized to 0, incremented by 1 if the left neighboring block is available and a certain condition is met, and incremented by 1 if the upper neighboring block is available and a certain condition is met.
[0317] Item 26. The method according to Item 25, wherein the specific condition is that the neighboring block is affine encoded or decoded.
[0318] Item 27. The method according to Item 26, wherein if the block is encoded in affine-AMVP mode, the block is considered affine-encoded; or wherein if the block is encoded in affine-Merge mode, the block is considered affine-encoded; or wherein if at least one of the GPM affine flags of the block is true, the block is decoded as affine-encoded; or wherein if the GPM affine flag of the block is true, the block is decoded as affine-encoded.
[0319] Item 28. The method according to Item 23, wherein the first SE is the GPM affine codec flag.
[0320] Item 29. The method according to Item 1, wherein the context of the affine flag used for encoding / decoding a block depends on whether the neighboring block is encoded / decoded in GPM affine mode.
[0321] Item 30. The method according to Item 29, wherein the context for encoding and decoding the affine sign is selected from a candidate set including more than one candidate context.
[0322] Item 31. The method according to Item 29, wherein the candidate set comprises three context candidates denoted as {C[0], C[1], C[2]}, and the selected context is C[S].
[0323] Item 32. The method according to Item 31, wherein S is initialized to 0, and S is incremented by 1 if the left neighboring block is available and a certain condition is met, and S is incremented by 1 if the upper neighboring block is available and a certain condition is met.
[0324] Item 33. The method according to Item 32, wherein the specific condition is that the neighboring block is affine encoded or decoded.
[0325] Item 34. The method according to Item 33, wherein if the block is encoded in affine-AMVP mode, the block is considered affine-encoded; or if the block is encoded in affine-Merge mode, the block is considered affine-encoded; or if at least one of the GPM affine flags of the block is true, the block is decoded as affine-encoded; or if the GPM affine flag of the block is true, the block is decoded as affine-encoded.
[0326] Item 35. The method according to Item 29, wherein the context is selected based on the number of available affine-coded neighboring blocks.
[0327] Item 36. The method according to Item 35, wherein a first context is used if N <= T, and a second context is used if N > T, and wherein N represents the number of available affine-coded neighbor blocks, and T is an integer.
[0328] Item 37. The method according to Item 36, wherein T is equal to 0, or 1, or 2, or 3, or 4, or 5, or 6.
[0329] Item 38. The method according to Item 35, wherein one or more of the five neighboring blocks are included in the available affine-coded neighboring blocks, or wherein one or more of the seven neighboring samples are included in the available affine-coded neighboring blocks.
[0330] Item 39. The method according to Item 35, wherein if the block is encoded in affine-AMVP mode, the block is considered affine-encoded; or if the block is encoded in affine-Merge mode, the block is considered affine-encoded; or if at least one of the GPM affine flags of the block is true, the block is decoded as affine-encoded; or if the GPM affine flag of the block is true, the block is decoded as affine-encoded.
[0331] Item 40. The method according to Item 1, wherein the maximum permissible value of the first SE depends on whether at least one GPM segment corresponding to the first SE is encoded or decoded using affine motion compensation.
[0332] Item 41. The method according to Item 40, wherein if the first SE is encoded or decoded using a rounding unary code or a rounding binary code, the maximum allowable value determines the last valid codeword of the first SE.
[0333] Item 42. The method according to Item 40, wherein the maximum permissible value of the first SE is transmitted via signaling for both affine and non-affine codec cases.
[0334] Item 43. The method according to Item 1, wherein the maximum allowed value of the first SE is determined based on the number of available affine-coded neighboring blocks.
[0335] Item 44. The method according to Item 43, wherein a first maximum allowed value is used if N <= T, and a second maximum allowed value is used if N > T, and wherein N represents the number of available affine-coded neighboring blocks, and T is an integer.
[0336] Item 45. The method according to Item 44, wherein T is equal to 0, or 1, or 2, or 3, or 4, or 5, or 6.
[0337] Item 46. The method according to Item 43, wherein one or more of the five neighboring blocks are included in the available affine-coded neighboring blocks, or wherein one or more of the seven neighboring samples are included in the available affine-coded neighboring blocks.
[0338] Item 47. The method according to Item 43, wherein if the block is encoded in affine-AMVP mode, the block is considered affine-encoded, or if the block is encoded in affine-Merge mode, the block is considered affine-encoded, or if at least one of the GPM affine flags of the block is true, the block is decoded as affine-encoded, or if the GPM affine flag of the block is true, the block is decoded as affine-encoded.
[0339] Item 48. The method according to Item 1, wherein the weight values used in the GPM depend on whether at least one GPM segment is affine encoded or decoded.
[0340] Item 49. The method described in Item 48, wherein the weight values used in GPM vary based on different conditions.
[0341] Item 50. The method according to Item 49, wherein the different conditions include at least one of the following: both GPM segments are affine encoded, both GPM segments are non-affine encoded, or one GPM segment is affine encoded while the other GPM segment is non-affine encoded.
[0342] Item 51. The method according to any one of items 1 to 50, wherein whether and / or how the first SE is determined to be encoded and decoded using multiple contexts at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.
[0343] Item 52. The method according to any one of Items 1 to 50, wherein whether and / or how it is determined that the first SE is encoded and decoded using multiple contexts at one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header or slice header.
[0344] Item 53. The method according to any one of items 1 to 50, wherein whether and / or how the first SE is determined to be encoded and decoded using multiple contexts at one of the following locations for signal transmission: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), codec tree block (CTB), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region comprising more than one sample point or pixel.
[0345] Item 54. The method according to any one of items 1 to 50 further comprises: determining, based on the encoded information of the video unit, whether and / or how to determine that the first SE is encoded using multiple contexts, said encoded information including at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.
[0346] Item 55. The method according to any one of items 1 to 54, wherein the syntax element is binarized into one of the following: a flag, a fixed-length code, an EG(x) code, a unary code, a rounded unary code, or a rounded binary code.
[0347] Item 56. The method according to Item 55, wherein the syntax element is signed or unsigned.
[0348] Item 57. The method according to any one of items 1 to 54, wherein the syntax elements are encoded or decoded using at least one context model, or wherein the syntax elements are encoded or decoded in a bypass manner.
[0349] Item 58. The method according to any one of items 1 to 57, wherein the syntax element is transmitted via signal in a conditional manner.
[0350] Item 59. The method according to Item 58, wherein the syntax element is transmitted via signaling if the corresponding function applies, or wherein the syntax element is transmitted via signaling if the dimension of the video unit satisfies a condition.
[0351] Item 60. The method according to Item 59, wherein the dimension includes the width and / or height of the video unit.
[0352] Item 61. The method according to any one of items 1 to 60, wherein the syntax element is transmitted by signaling at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.
[0353] Item 62. The method according to any one of Items 1 to 60, wherein the syntax element is transmitted via signaling at one of the following locations: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header or slice header.
[0354] Item 63. The method according to any one of items 1 to 60, wherein the syntax element is transmitted by signal at one of the following locations: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), codec tree block (CTB), codec tree unit (CTU), CTU line, strip, slice, sub-picture, or region comprising more than one sample point or pixel.
[0355] Item 64. The method according to any one of items 1 to 63, wherein the video unit is encoded or decoded using one or more other encoding / decoding tools that require chroma blending.
[0356] Item 65. The method according to any one of items 1 to 64, wherein the conversion includes encoding the video unit into the bitstream.
[0357] Item 66. The method according to any one of items 1 to 64, wherein the conversion includes decoding the video unit from the bitstream.
[0358] Item 67. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1 to 66.
[0359] Item 68. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of items 1 to 66.
[0360] Item 69. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of an apparatus for video processing, wherein the method includes: determining a first syntax element (SE) associated with a geometric segmentation mode (GPM) and encoding / decoding it using a plurality of contexts; and generating the bitstream based on the first SE.
[0361] Item 70. A method for storing a bitstream of video, comprising: determining a first syntax element (SE) associated with a geometric segmentation mode (GPM) and encoding / decoding it using a plurality of contexts; generating the bitstream based on the first SE; and storing the bitstream in a non-transitory computer-readable recording medium.
[0362] Example device Figure 26A block diagram of a computing device 2600 in which various embodiments of the present disclosure may be implemented is shown. The computing device 2600 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0363] It should be understood that, Figure 26 The computing device 2600 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.
[0364] like Figure 26 As shown, computing device 2600 includes general-purpose computing device 2600. Computing device 2600 may include at least one or more processors or processing units 2610, memory 2620, storage unit 2630, one or more communication units 2640, one or more input devices 2650, and one or more output devices 2660.
[0365] In some embodiments, the computing device 2600 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 2600 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).
[0366] Processing unit 2610 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 2620. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 2600. Processing unit 2610 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0367] Computing device 2600 typically includes various computer storage media. Such media can be any media accessible by computing device 2600, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 2620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 2630 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 2600.
[0368] The computing device 2600 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 26 Not shown, but may provide disk drives for reading from and / or writing to removable non-volatile disks, and optical disc drives for reading from and / or writing to removable non-volatile optical discs. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0369] Communication unit 2640 communicates with another computing device via a communication medium. Furthermore, the functionality of components in computing device 2600 can be implemented by a single computing cluster or by multiple computing machines communicating via communication connections. Therefore, computing device 2600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0370] Input device 2650 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 2660 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 2640, computing device 2600 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 2600 can also communicate with one or more devices that enable a user to interact with computing device 2600, or any device that enables computing device 2600 to communicate with one or more other computing devices (e.g., network card, modem, etc.), if needed. Such communication can be performed via an input / output (I / O) interface (not shown).
[0371] In some embodiments, some or all components of computing device 2600 may not be integrated into a single device, but may be deployed in a cloud computing architecture. In a cloud computing architecture, components may be provided remotely and may work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN) such as the Internet using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at a remote location. Computing resources in a cloud computing environment may be consolidated or distributed at locations in remote data centers. Cloud computing infrastructure may provide services through shared data centers, although they may appear as a single access point for users. Therefore, cloud computing architectures can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, they may be provided from conventional servers or installed directly or otherwise on client devices.
[0372] In embodiments of this disclosure, computing device 2600 can be used to implement video encoding / decoding. Memory 2620 may include one or more video codec modules 2625 having one or more program instructions. These modules can be accessed and executed by processing unit 2610 to perform the functions of the various embodiments described herein.
[0373] In an example embodiment of performing video encoding, input device 2650 may receive video data as input 2670 to be encoded. The video data may be processed, for example, by video codec module 2625 to generate an encoded bitstream. The encoded bitstream may be provided as output 2680 via output device 2660.
[0374] In an example embodiment of performing video decoding, input device 2650 may receive an encoded bitstream as input 2670. The encoded bitstream may be processed, for example, by a video codec module 2625 to generate decoded video data. The decoded video data may be provided as output 2680 via output device 2660.
[0375] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These changes are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.
Claims
1. A video processing method, comprising: For the conversion between video units and the bitstream of the video, the first syntax element (SE) related to the geometric segmentation mode (GPM) is determined to be encoded and decoded using multiple contexts; and The transformation is performed based on the first SE.
2. The method of claim 1, wherein the first SE is a GPM Merge index, or Where the first SE is the GPM Merge motion vector difference (MMVD) index, or Where the first SE is the GPM bending weight index, or Where the first SE is the GPM partitioning pattern index, or The first SE is the GPM intra-frame prediction mode. Wherein the first SE is the GPM intra-frame encoding / decoding mode, or Wherein the first SE is the GPM MMVD encoding / decoding flag, or Wherein the first SE is a GPM template matching (TM) encoding / decoding flag, or The first SE is the GPM affine codec flag.
3. The method according to claim 1, wherein the first SE is binarized to be represented as B0, B1, ..., B N-1 It consists of N binary bits, where N is an integer.
4. The method of claim 3, wherein if n is not equal to m, then Bn and Bm are encoded and decoded using different contexts, wherein n and m are integers not greater than N-1.
5. The method according to claim 3, wherein B0, B1, ..., B W Using the denoted C0, C1, ..., C respectively W Different contexts are encoded and decoded, and B W B W+1 B N-1 Shared is represented as C W+1 The same context, where W is an integer.
6. The method according to claim 3, wherein B0, B1, ..., B W Using the denoted C0, C1, ..., C respectively W Different contexts are encoded and decoded, and B W B W+1 B N-1 It is bypassed encoding / decoding, where W is an integer.
7. The method according to claim 5 or 6, wherein W is equal to 1, or 2, or 3, or 4, or 5, or 6, or 7.
8. The method of claim 1, wherein at least one context for encoding and decoding the first SE in relation to GPM is selected based on encoding and decoding information.
9. The method of claim 8, wherein the selection of the context depends on whether at least one GPM segment corresponding to the first SE is encoded or decoded using affine motion compensation.
10. The method of claim 9, wherein if the GPM segment is affine-coded, a first context is used to encode the GPM Merge index of the GPM segment, and if the GPM segment is not affine-coded, a second context is used to encode the GPM Merge index of the GPM segment, wherein the first context is different from the second context.
11. The method of claim 9, wherein if the GPM segment is affine-coded, a third context is used to encode the GPM MMVD index of the GPM segment, and if the GPM segment is not affine-coded, a fourth context is used to encode the GPM MMVD index of the GPM segment, wherein the third context is different from the fourth context.
12. The method of claim 9, wherein if the GPM segment is affine encoded, a fifth context is used to encode the GPM bend weight index of the GPM segment, and if the GPM segment is not affine encoded, a sixth context is used to encode the GPM bend weight index of the GPM segment, wherein the fifth context is different from the sixth context.
13. The method of claim 8, wherein the selection of the context depends on whether at least one GPM segment corresponding to the first SE is encoded or decoded using intra-frame prediction.
14. The method of claim 8, wherein the selection of the context depends on whether at least one GPM segment corresponding to the first SE is encoded or decoded using GPM MMVD mode.
15. The method of claim 8, wherein the selection of the context depends on whether at least one GPM segment corresponding to the first SE is encoded or decoded using GPM TM mode.
16. The method of claim 8, wherein the selection of the context for the first GPM segment depends on the encoding / decoding information of the second GPM segment.
17. The method of claim 8, wherein the choice of context depends on the GPM partitioning mode.
18. The method of claim 8, wherein the choice of context depends on the mixed weights.
19. The method of claim 8, wherein the choice of context depends on the quantization parameter (QP).
20. The method of claim 8, wherein the choice of context depends on the stripe type.
21. The method of claim 1, wherein at least one context for encoding and decoding the first SE in relation to the GPM is selected based on encoding and decoding information of at least one neighboring block.
22. The method of claim 21, wherein the context for encoding and decoding the first SE is selected from a candidate set including more than one candidate context.
23. The method of claim 22, wherein the candidate set comprises three context candidates denoted as {C[0], C[1], C[2]}, and the selected context is C[S], where S is an integer.
24. The method of claim 23, wherein S depends on the encoding / decoding mode of at least one neighboring block.
25. The method of claim 23, wherein S is initialized to 0, incremented by 1 if the left neighboring block is available and a certain condition is met, and incremented by 1 if the upper neighboring block is available and a certain condition is met.
26. The method of claim 25, wherein the specific condition is that the neighboring block is affine encoded or decoded.
27. The method of claim 26, wherein if the block is encoded / decoded in affine-AMVP mode, then the block is considered to be affine encoded / decoded, or Where a block is encoded or decoded in affine-Merge mode, then the block is considered to be affine encoded or decoded. Where at least one of the GPM affine flags of the block is true, the block is decoded as affine encoded or decoded. If the GPM affine flag of a block is true, then the block is decoded as affine encoded / decoded.
28. The method of claim 23, wherein the first SE is the GPM affine codec flag.
29. The method of claim 1, wherein the context of the affine flag used for encoding / decoding blocks depends on whether neighboring blocks are encoded / decoded in GPM affine mode.
30. The method of claim 29, wherein the context for encoding and decoding the affine sign is selected from a candidate set including more than one candidate context.
31. The method of claim 29, wherein the candidate set comprises three context candidates denoted as {C[0], C[1], C[2]}, and the selected context is C[S].
32. The method of claim 31, wherein S is initialized to 0, and S is incremented by 1 if the left neighboring block is available and a certain condition is met, and S is incremented by 1 if the upper neighboring block is available and a certain condition is met.
33. The method of claim 32, wherein the specific condition is that the neighboring block is affine encoded or decoded.
34. The method of claim 33, wherein if the block is encoded / decoded in affine-AMVP mode, then the block is considered to be affine-encoded / decoded, or Where a block is encoded or decoded in affine-Merge mode, then the block is considered to be affine encoded or decoded. Where at least one of the GPM affine flags of the block is true, the block is decoded as affine encoded or decoded. If the GPM affine flag of a block is true, then the block is decoded as affine encoded / decoded.
35. The method of claim 29, wherein the context is selected based on the number of available affine-coded neighbor blocks.
36. The method of claim 35, wherein if N <= T, then the first context is used, and If N > T, then the second context is used, and Where N represents the number of available affine-coded neighbor blocks, and T is an integer.
37. The method of claim 36, wherein T is equal to 0, or 1, or 2, or 3, or 4, or 5, or 6.
38. The method of claim 35, wherein one or more of the five neighboring blocks are included in the available affine-coded neighboring blocks, or One or more of the seven neighboring samples are included in the available affine-coded neighbor block.
39. The method of claim 35, wherein if the block is encoded or decoded in affine-AMVP mode, then the block is considered to be affine encoded or decoded. Where a block is encoded or decoded in affine-Merge mode, then the block is considered to be affine encoded or decoded. Where at least one of the GPM affine flags of the block is true, the block is decoded as affine encoded or decoded. If the GPM affine flag of a block is true, then the block is decoded as affine encoded / decoded.
40. The method of claim 1, wherein the maximum permissible value of the first SE depends on whether at least one GPM segment corresponding to the first SE is encoded or decoded using affine motion compensation.
41. The method of claim 40, wherein if the first SE is encoded or decoded using a rounding unary code or a rounding binary code, the maximum allowable value determines the last valid codeword of the first SE.
42. The method of claim 40, wherein the maximum permissible value of the first SE is transmitted via signaling for both affine codec and non-affine codec cases.
43. The method of claim 1, wherein the maximum allowable value of the first SE is determined based on the number of available affine-coded neighbor blocks.
44. According to the method described in 43, wherein if N <= T, then the first maximum allowed value is used, and If N > T, then the second maximum allowed value is used, and Where N represents the number of available affine-coded neighbor blocks, and T is an integer.
45. The method of claim 44, wherein T is equal to 0, or 1, or 2, or 3, or 4, or 5, or 6.
46. The method of claim 43, wherein one or more of the five neighboring blocks are included in the available affine-coded neighboring blocks, or One or more of the seven neighboring samples are included in the available affine-coded neighbor block.
47. The method of claim 43, wherein if the block is encoded or decoded in affine-AMVP mode, then the block is considered to be affine encoded or decoded. Where a block is encoded or decoded in affine-Merge mode, then the block is considered to be affine encoded or decoded. Where at least one of the GPM affine flags of the block is true, the block is decoded as affine encoded or decoded. If the GPM affine flag of a block is true, then the block is decoded as affine encoded / decoded.
48. The method of claim 1, wherein the weight values used in the GPM depend on whether at least one GPM segment is affine encoded or decoded.
49. The method of claim 48, wherein the weight values used in GPM vary based on different conditions.
50. The method of claim 49, wherein the different conditions include at least one of the following: Both GPM segments are affine encoded and decoded. Both GPM segments are non-affine encoded or decoded, or One GPM segment is encoded and decoded affinely, while the other GPM segment is encoded and decoded non-affinely.
51. The method according to any one of claims 1 to 50, wherein it is determined whether and / or how the first SE is encoded / decoded using the plurality of contexts at one of the following locations: sequence level, Image group level, Image quality, strip level, or Film series level.
52. The method according to any one of claims 1 to 50, wherein it is determined whether and / or how the first SE is encoded / decoded using the plurality of contexts at one of the following: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.
53. The method according to any one of claims 1 to 50, wherein it is determined whether and / or how the first SE is encoded / decoded using the plurality of contexts at one of the following locations: Predicted blocks (PB). Transform block (TB) Code block (CB) Prediction Unit (PU) Transformer Unit (TU) Codec Unit (CU) Code-decode tree block (CTB). Code-decode tree unit (CTU) CTU line, strip, piece, Sub-images, or This includes regions containing more than one sample point or pixel.
54. The method according to any one of claims 1 to 50, further comprising: Based on the encoded and decoded information of the video unit, determine whether and / or how to determine that the first SE is encoded and decoded using the plurality of contexts, wherein the encoded and decoded information includes at least one of the following: Block size, Color format, Single-tree partitioning and / or dual-tree partitioning, Color components, Strip type, or Image type.
55. The method according to any one of claims 1 to 54, wherein the syntax element is binarized into one of the following: a flag, a fixed-length code, an EG(x) code, a unary code, a rounded unary code, or a rounded binary code.
56. The method of claim 55, wherein the syntax element is signed or unsigned.
57. The method according to any one of claims 1 to 54, wherein the syntax elements are encoded and decoded using at least one context model, or Syntax elements are bypassed for encoding and decoding.
58. The method according to any one of claims 1 to 57, wherein the syntax element is transmitted via signal in a conditional manner.
59. The method of claim 58, wherein the syntax element is transmitted via a signal if the corresponding function applies, or If the dimension of the video unit meets the condition, the syntax element is transmitted via signal.
60. The method of claim 59, wherein the dimension includes the width and / or height of the video unit.
61. The method according to any one of claims 1 to 60, wherein the syntax element is transmitted via a signal at one of the following locations: sequence level, Image group level, Image quality, strip level, or Film series level.
62. The method according to any one of claims 1 to 60, wherein the syntax element is transmitted via a signal at one of the following locations: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.
63. The method according to any one of claims 1 to 60, wherein the syntax element is transmitted via a signal at one of the following locations: Predicted blocks (PB). Transform block (TB) Code block (CB) Prediction Unit (PU) Transformer Unit (TU) Codec Unit (CU) Code-decode tree block (CTB). Code-decode tree unit (CTU) CTU line, strip, piece, Sub-images, or This includes regions containing more than one sample point or pixel.
64. The method according to any one of claims 1 to 63, wherein the video unit is encoded or decoded using one or more other encoding / decoding tools that require chroma blending.
65. The method according to any one of claims 1 to 64, wherein the conversion comprises encoding the video unit into the bitstream.
66. The method according to any one of claims 1 to 64, wherein the conversion comprises decoding the video unit from the bitstream.
67. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 66.
68. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 66.
69. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: The first syntax element (SE) associated with the geometric segmentation pattern (GPM) is determined using multiple contexts for encoding and decoding; as well as The bit stream is generated based on the first SE.
70. A method for storing a bitstream of video, comprising: The first syntax element (SE) associated with the geometric segmentation pattern (GPM) is determined using multiple contexts for encoding and decoding; The bit stream is generated based on the first SE; as well as The bitstream is stored in a non-transitory computer-readable recording medium.