Method and device for geometric partitioning mode with motion vector refinement - Patents.com
The method improves video encoding/decoding efficiency by refining motion vectors in GPM using predefined MVD magnitudes and directions, addressing inefficiencies in existing standards like VVC and AVS3, particularly in non-merge inter modes.
Patent Information
- Application Number
- JP2023577272
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-14
- Filing Date
- 2022-06-14
- Publication Date
- 2025-12-11
- Estimated Expiration
- 2042-06-14
AI Technical Summary
Existing video encoding and decoding standards, such as VVC and AVS3, face inefficiencies in the geometric partitioning mode (GPM) due to suboptimal motion vector refinement, particularly when applied to non-merge inter modes, leading to inaccurate motion representation and increased signaling overhead.
A method is proposed to enhance GPM by applying motion vector refinement (MVR) on top of existing unidirectional MVD, using predefined MVD magnitudes and directions, and extending GPM mode to explicit inter modes with individual motion signaling, minimizing signaling costs while improving accuracy.
The proposed method enhances coding/decoding efficiency by providing more accurate motion vectors for geometric partitions, reducing signaling overhead and improving overall performance in video encoding/decoding processes.
Smart Images

Figure 0007784452000028 
Figure 0007784452000029 
Figure 0007784452000030
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 210,484, filed June 14, 2021, the disclosure of which is incorporated herein by reference in its entirety for all purposes. [Technical Field]
[0002] This disclosure relates to video encoding / decoding and compression, and more particularly to methods and apparatus for improving encoding / decoding efficiency of geometric partitioning mode (GPM), also known as angle weighted prediction (AWP) mode. [Background technology]
[0003] Various video encoding and decoding techniques can be used to compress video data. Video encoding and decoding is performed according to one or more video encoding and decoding standards. For example, some well-known video encoding and decoding standards today include the Universal Video Coding and Decoding (VVC), jointly developed by ISO / IEC MPEG and ITU-T VECG, High Efficiency Video Coding and Decoding (HEVC, also known as H.265 or MPEG-H Part 2), and Advanced Video Coding and Decoding (AVC, also known as H.264 or MPEG-4 Part 10). AOMedia Video 1 (AV1) was developed by the Alliance for Open Media (AOM) as a successor to the predecessor standard VP9. Audio-Video Coding and Decoding (AVS), referring to digital audio and digital video compression standards, is another video compression standard series developed by the China Audio and Video Coding and Decoding Standards Workgroup. Many of the existing video coding and decoding standards are built on a well-known hybrid video coding and decoding framework, such as using block-based prediction methods (e.g., inter-prediction, intra-prediction) to reduce redundancy present in a video or sequence, and using transform coding to compress the energy of prediction errors, etc. One important goal of video coding and decoding technology is to compress video data into a format with a lower bit rate while avoiding or minimizing degradation of video quality. Summary of the Invention
[0004] The present disclosure provides a video encoding / decoding method and apparatus, and a non-transitory computer-readable storage medium.
[0005] In a first aspect of the present disclosure, a method for decoding a video block using a global motion vector refinement (GPM) is provided. The method may include receiving a control variable associated with the video block, the control variable enabling adaptive switching between a set of motion vector refinement (MVR) offsets and applied at a coding level. The method may also include receiving an indicator variable associated with the video block, the control variable enabling adaptive switching between a set of codewords for binarizing magnitudes of offsets in the set of MVR offsets under the coding level. The method may also include partitioning the video block into a first geometric partition and a second geometric partition. The method may also include selecting a set of MVR offsets from the set of MVR offsets based on the control variable. The method may also include receiving one or more syntax elements and determining first and second MVR offsets to be applied to the first and second geometric partitions from the selected set of MVR offsets. The method may also include obtaining a first motion vector (MV) and a second MV from candidate lists for the first and second geometric partitions. The method may also include calculating a first refined MV and a second refined MV based on the first and second MVs and the first and second MVR offsets. The method may also include obtaining a prediction sample for the video block based on the first and second refined MVs.
[0006] In a second aspect of the present disclosure, there is provided a video decoding apparatus. The apparatus may include one or more processors and a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium is configured to store instructions executable by the one or more processors. The one or more processors are configured, upon execution of the instructions, to perform a method according to the first aspect.
[0007] In a third aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium, the non-transitory computer-readable storage medium may store computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform a method according to the first aspect.
[0008] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. [Brief explanation of the drawings]
[0009] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure.
[0010] [Figure 1] FIG. 2 is a block diagram of an encoder according to one embodiment of the present disclosure. [Figure 2A] FIG. 1 illustrates block division of a multi-type tree structure according to one embodiment of the present disclosure. [Figure 2B] FIG. 1 illustrates block division of a multi-type tree structure according to one embodiment of the present disclosure. [Figure 2C] FIG. 1 illustrates block division of a multi-type tree structure according to one embodiment of the present disclosure. [Figure 2D] FIG. 1 illustrates block division of a multi-type tree structure according to one embodiment of the present disclosure. [Figure 2E] FIG. 1 illustrates block division of a multi-type tree structure according to one embodiment of the present disclosure. [Figure 3] FIG. 2 is a block diagram of a decoder according to one embodiment of the present disclosure. [Figure 4] 1 is a diagrammatic view of allowed geometric partition mode (GPM) partitions according to one embodiment of the present disclosure. [Figure 5] 10 is a table illustrating selection of a unidirectional predictive motion vector according to one embodiment of the present disclosure. [Figure 6A] 1 is a diagram illustrating a merge mode with a motion vector differential (MMVD) mode according to one embodiment of the present disclosure. [Figure 6B] 1 is a diagram illustrating an MMVD mode according to one embodiment of the present disclosure. [Figure 7] 1 is a diagram of a template matching (TM) algorithm according to one embodiment of the present disclosure. [Figure 8] 1 is a method for decoding a video block in a GPM according to one embodiment of the present disclosure. [Figure 9] FIG. 1 illustrates a computing environment coupled to a user interface according to one embodiment of the present disclosure. [Figure 10] FIG. 1 is a block diagram illustrating a system for encoding and decoding video blocks in accordance with some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0011] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which like numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following description of the embodiments do not represent all implementations consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with aspects related to the present disclosure, as set forth in the appended claims.
[0012] The terms used in this disclosure are for the purpose of describing particular embodiments only and are not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an," and the like are intended to include the plural forms as well, unless the context clearly indicates otherwise. Additionally, as used herein, the term "and / or" is intended to mean and include any and all possible combinations of one or more of the associated listed items.
[0013] It should be noted that, although terms such as "first," "second," and "third" may be used herein to describe various pieces of information, the information should not be limited by these terms. These terms are used only to distinguish one category of information from another. For example, first information could be referred to as second information, and similarly, second information could be referred to as first information, without departing from the scope of this disclosure. As used herein, the term "if" may be understood to mean "when," "when," or "depending on the determination," depending on the context.
[0014] The first generation of AVS standards includes the Chinese national standards "Information Technology, Advanced Audio-Video Coding / Decoding, Part 2: Video" (known as AVS1) and "Information Technology, Advanced Audio-Video Coding / Decoding, Part 16: Radio-Television Video" (known as AVS+). Compared to the MPEG-2 standard, it can save approximately 50% of the bitrate while maintaining the same perceptual quality. The video portion of the AVS1 standard was promulgated as a Chinese national standard in February 2006. The second generation of AVS standards includes a series of Chinese national standards, "Information Technology, Efficient Multimedia Coding / Decoding" (known as AVS2), primarily aimed at transmitting additional HD TV programs. The coding / decoding efficiency of AVS2 is twice that of AVS+. In May 2016, AVS2 was published as a Chinese national standard. Meanwhile, the video portion of the AVS2 standard was submitted by the Institute of Electrical and Electronics Engineers (IEEE) as one of the international application standards. The AVS3 standard is one of a new generation of video encoding / decoding standards for UHD video applications, aiming to exceed the encoding / decoding efficiency of the latest international standard, HEVC. In March 2019, the 68th AVS Conference finalized the AVS3-P2 baseline, which offers approximately 30% bitrate savings over the HEVC standard. Currently, there is one reference software, called the High Performance Model (HPM), which is maintained by the AVS Group and demonstrates a reference implementation of the AVS3 standard.
[0015] The AVS3 standard, like HEVC, is built on a block-based hybrid video encoding / decoding framework.
[0016] 10 is a block diagram illustrating an example system 10 for encoding and decoding video blocks in parallel, according to some implementations of this disclosure. 0 As shown, system 10 includes a source device 12 that generates and encodes video data that is subsequently decoded by a target device 14. Source device 12 and target device 14 may include any of a wide variety of electronic devices, including desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some implementations, source device 12 and target device 14 are equipped with wireless communication capabilities.
[0017] In some implementations, target device 14 may receive the encoded video data to be decoded via link 16. Link 16 may include any type of communication medium or device capable of moving encoded video data from source device 12 to target device 14. In one embodiment, link 16 may include a communication medium that enables source device 12 to transmit encoded video data directly to target device 14 in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to target device 14. The communication medium may include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful in facilitating communication from source device 12 to target device 14.
[0018] In some other implementations, the encoded video data may be transmitted from output interface 22 to storage device 32. The encoded video data in storage device 32 may then be accessed by target device 14 via input interface 28. Storage device 32 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In further embodiments, storage device 32 may correspond to a file server or other intermediate storage device that may hold encoded video data generated by source device 12. Target device 14 may access the stored video data from storage device 32 via streaming or download. The file server may be any type of computer capable of storing encoded video data and transmitting the encoded video data to target device 14. Exemplary file servers include a web server (e.g., for a website), a file transfer protocol (FTP) server, a network-attached storage (NAS) device, or a local disk drive. Target device 14 may access the encoded video data over any standard data connection suitable for accessing encoded video data stored on a file server, including a wireless channel (e.g., a Wi-Fi (Wireless Fidelity) connection), a wired connection (e.g., a DSL (Digital Subscriber Line), a cable modem, etc.), or a combination of both. Transmission of the encoded video data from storage device 32 may be a streaming transmission, a download transmission, or a combination of both.
[0019] 10 , source device 12 includes video source 18, video encoder 20, and output interface 22. Video source 18 may include a video capture device, such as a video camera, a video archive containing previously captured video, a video feed interface receiving video from a video content provider, and / or a computer graphics system generating computer graphics data as source video, or a combination of these sources. As one example, if video source 18 is a video camera in a security surveillance system, source device 12 and target device 14 may form a camera phone or video phone. However, implementations described herein are applicable to video encoding and decoding generally, and may also be applicable to wireless and / or wired applications.
[0020] The captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video data may be transmitted directly to target device 14 via output interface 22 of source device 12. The encoded video data may also (or alternatively) be stored in storage device 32 for later access, decoding, and / or playback by target device 14 or other devices. Output interface 22 may further include a modem and / or a transmitter.
[0021] Target device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 may include a receiver and / or modem to receive encoded video data over link 16. The encoded video data communicated over link 16 or provided to storage device 32 may include various syntax elements generated by video encoder 20 that video decoder 30 uses to decode the video data. Such syntax elements may be included within the encoded video data that is transmitted over a communications medium, stored on a storage medium, or stored on a file server.
[0022] In some implementations, target device 14 may include a display device 34, which may be an integrated display device and an external display device configured to communicate with target device 14. Display device 34 displays decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.
[0023] Video encoder 20 and video decoder 30 may operate in accordance with proprietary or industry standards, such as, for example, VVC, HEVC, MPEG-4 Part 10, AVC, or extensions of these standards. It should be understood that the present application is not limited to a particular video encoding / decoding standard and may be applicable to other video encoding / decoding standards. It is generally contemplated that video encoder 20 of source device 12 may be configured to encode video data in accordance with any of these current or future standards. Similarly, it is generally contemplated that video decoder 30 of target device 14 may be configured to decode video data in accordance with any of these current or future standards.
[0024] Video encoder 20 and video decoder 30 may each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or combinations thereof. If the electronic device is implemented partially in software, the software instructions may be stored on a suitable non-transitory computer-readable medium and executed in hardware by one or more processors to perform the video encoding / decoding operations disclosed in this disclosure. Each of video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, or either may be integrated as part of a combined encoder / decoder (CODEC) within the respective device.
[0025] Figure 1 shows a general diagram of a VVC block-based video encoder. Da Show. Encode Da is The video encoder 20 may be the one shown in FIG. Da is , video input force, Motion Compensation 101 ,motion estimation 102 , Intra / Inter mode decision 103 , block prediction child, Adder 128, conversion 108 , quantization 109 , prediction related information Information, Intra prediction 104 , image buffer 106 , inverse quantization 111 , inverse transformation 112 , adder 126, memory 105 , in-loop filter 107 , entropy coding 110 , and Bitstream M include.
[0026] Encode DaIn general, a video frame is divided into multiple video blocks for processing. For each given video block, a prediction is formed based on an inter-prediction or intra-prediction approach. Note that the term "frame" may be used synonymously with the terms "picture" or "image" in the field of video encoding and decoding.
[0027] Video Included of power The current video block that it is part of and the block prediction Child The prediction residual, which represents the difference between the predictor and the 108 The transform coefficients are then sent to the transform 108 quantization from 109 The quantized coefficients are then entropy coded. 110 1, the video block partition information, the motion vector (MV), the reference image index, and the intra / inter mode decision such as the intra prediction mode are provided to generate the compressed video bitstream. 103 Prediction related information from News Entropy Coding 110 The compressed bitstream is provided by Mu The compressed bitstream is saved. M is , including the video bitstream.
[0028] Encode Da Also requires decoder-related circuitry to reconstruct pixels for prediction purposes. 111 and the inverse transformation 112 The prediction residual is reconstructed by the following. This reconstructed prediction residual is used as the block prediction With children are combined to produce unfiltered reconstructed pixels for the current video block.
[0029] Spatial prediction (or "intra prediction") predicts the current video block using pixels from samples (called reference samples) of already-encoded neighboring blocks in the same video frame as the current video block.
[0030] Temporal prediction (also called "inter-prediction") uses reconstructed pixels from already coded video images to predict a current video block. Temporal prediction reduces the temporal redundancy inherent in video signals. The temporal prediction signal for a given coding unit (CU) or coding block is typically signaled by one or more motion vectors (MVs) that indicate the amount and direction of motion between the current CU and its temporal reference. Additionally, if multiple reference images are supported, a reference image index is also sent and is used to identify which reference image in a reference image store the temporal prediction signal comes from.
[0031] Motion Estimation 102 is a video input influence and image buffer 106 The signal from the motion estimation signal is taken in and motion compensated. 101 Output to Motion Compensation 101 is a video input force, Image Buffer 106 Signals from, and motion estimation 102 The motion estimation signal is taken from the input signal, and the motion compensation signal is determined as intra / inter mode. 103 Output to.
[0032] After spatial and / or temporal prediction is performed, the encoding Da's Intra / Inter mode decision 103 selects the best prediction mode based on, for example, a rate-distortion optimization method. The child , is subtracted from the current video block, and the resulting prediction residual is transformed 108 and quantization 109 The resulting quantized residual coefficients are dequantized by 111 It is dequantized by 112 , and the reconstructed residual is then added to the prediction block to form a reconstructed signal for the CU. Further in-loop filtering, such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive in-loop filter (ALF),107 The reconstructed CU is then stored in the image buffer. 106 The reference picture is then stored in the reference picture store and can be applied to the reconstructed CU before being used to encode / decode future video blocks. M To form the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients, all are sent to the entropy coding unit 110 , where it is further compressed and packed to form a bitstream.
[0033] Figure 1 shows a block diagram of a typical block-based hybrid video coding system. An input video signal is processed block by block (called a coding unit (CU)). Unlike HEVC, which divides blocks based only on a quadtree, AVS3 divides a single coding tree unit (CTU) into CUs based on a quadtree, binary tree, or ternary tree to adapt to varying local characteristics. Furthermore, the concept of multiple division unit types in HEVC is eliminated; that is, the separation of CUs, prediction units (PUs), and transform units (TUs) does not exist in AVS3. Instead, each CU is always used as a basic unit for both prediction and transformation without further division. In the AVS3 tree division structure, a CTU is first divided based on a quadtree structure. Then, each quadtree leaf node can be further divided based on a binary tree structure and an extended quadtree structure.
[0034] As shown in Figures 2A, 2B, 2C, 2D, and 2E, there are five partition types: 4-partition, horizontal 2-partition, vertical 2-partition, horizontal extended quadtree partition, and vertical extended quadtree partition.
[0035] FIG. 2A shows a diagram illustrating a quadrant of blocks in a multi-type tree structure according to the present disclosure.
[0036] FIG. 2B shows a diagram illustrating vertical bisection of a block in a multi-type tree structure according to the present disclosure.
[0037] FIG. 2C shows a diagram illustrating horizontal bisection of a block in a multi-type tree structure according to the present disclosure.
[0038] FIG. 2D shows a diagram illustrating a vertical division of blocks into thirds in a multi-type tree structure according to the present disclosure.
[0039] FIG. 2E shows a diagram illustrating a horizontal division of blocks into thirds in a multi-type tree structure according to the present disclosure.
[0040] In FIG. 1, spatial prediction and / or temporal prediction may be performed. Spatial prediction (or "intra prediction") predicts a current video block using pixels from samples (called reference samples) of already coded neighboring blocks of the same video image / slice. Spatial prediction reduces spatial redundancy inherent in video signals. Temporal prediction (also called "inter prediction" or "motion-compensated prediction") predicts a current video block using reconstructed pixels from already coded video images. Temporal prediction reduces temporal redundancy inherent in video signals. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and its temporal references. If multiple reference images are supported, a reference image index is also sent, which is used to identify which reference image in the reference image store the temporal prediction signal comes from. After spatial and / or temporal prediction, a mode decision block of the encoder selects the best prediction mode, for example, based on a rate-distortion optimization method. The prediction block is then subtracted from the current video block, and the prediction residual is de-correlated using a transform and quantized. The quantized residual coefficients are inverse quantized and inverse transformed to form a reconstructed residual, which is then added back to the prediction block to form a reconstructed signal for the CU. Furthermore, in-loop filtering, such as a deblocking filter, sample adaptive offset (SAO), and adaptive in-loop filter (ALF), can be applied to the reconstructed CU before it is placed in a reference picture store and used as a reference for encoding and decoding future video blocks. The coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to an entropy coding unit for further compression and packing to form an output video bitstream.
[0041] FIG. 3 is a block diagram illustrating a block-based video decoder according to some implementations of this disclosure. The block-based video decoder may be, for example, the video decoder 30 illustrated in FIG. 10. A video bitstream is first entropy decoded in an entropy decoding unit (e.g., entropy decoding 301). Coding mode and prediction information are sent to a spatial prediction unit (if intra-coded) (e.g., intra prediction 308) or a temporal prediction unit (if inter-coded) (e.g., motion compensation 307) to form a prediction block. Residual transform coefficients are sent to an inverse quantization unit (e.g., inverse quantization 302) and an inverse transform unit (e.g., inverse transform 303) to reconstruct a residual block. The prediction block and residual block are then summed (e.g., via intra / inter mode selection 309 and / or stored in memory 304). The reconstructed block undergoes in-loop filtering before being stored in a reference image store (e.g., image buffer 306). (e.g., in-loop filter 305) The reconstructed video in the reference picture store can then be used to predict future video blocks as well as sent to drive a display device.
[0042] The focus of this disclosure is to improve the encoding / decoding performance of the geometric partitioning mode (GPM) used in both the VVC standard and the AVS3 standard. In AVS3, this tool is also known as angle weighted prediction (AWP). AWP follows the same design philosophy as GPM, but with some subtle differences in specific design details. To facilitate the explanation of this disclosure, the existing GPM design in the VVC standard will be used as an example below to explain the main aspects of the GPM / AWP tool. Meanwhile, another existing inter-prediction technique called merge mode with motion vector differential (MMVD), which applies to both the VVC standard and the AVS3 standard, will also be briefly outlined because it is closely related to the technology proposed in this disclosure. After that, several drawbacks of the current GPM / AWP design will be identified. Finally, the proposed method will be described in detail. Note that while the existing GPM design in the VVC standard will be used as an example throughout this disclosure, those skilled in the art of modern video encoding / decoding technology will understand that the proposed technology can also be applied to other GPM / AWP designs or other encoding tools with the same or similar design philosophy.
[0043] Geometric Partitioning Mode (GPM) In VVC, geometric partitioning mode is supported for inter prediction. Geometric partitioning mode is signaled by a CU-level flag as a special merge mode. In the current GPM design, a total of 64 partitions are supported by GPM modes for each CU with possible sizes of both width and height between 8 and 64, except for 8x64 and 64x8.
[0044] When this mode is used, a CU is divided into two parts by a geometrically positioned line, as shown in FIG. 4 (an explanation is provided below). The position of the division line is mathematically derived from the angle and offset parameters of a particular partition. Each part of the geometric partition within a CU is inter-predicted using its own motion, and only unidirectional prediction is allowed per partition; that is, each part has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that only two motion-compensated predictions are required for each CU, similar to traditional bidirectional prediction. When the geometric partition mode is used for the current CU, a geometric partition index (angle and offset) indicating the partition mode of the geometric partition and two merge indices (one for each partition) are further signaled. The number of maximum GPM candidate sizes is explicitly signaled at the sequence level.
[0045] FIG. 4 shows the allowed GPM partitions, where each picture partition has one and the same partition direction.
[0046] Building a list of unidirectional prediction candidates To derive a unidirectional prediction motion vector for one geometric partition, one unidirectional prediction candidate list is first derived directly from the normal merge candidate list generation process. n is ,single It is represented as the index of the unidirectionally predicted motion in the directional prediction candidate list. The LX motion vector (X equals the parity of n) of the nth merge candidate is used as the nth unidirectionally predicted motion vector in the geometric partitioning mode.
[0047] These motion vectors are marked with an "x" in Figure 5 (described below). If the corresponding LX motion vector of the nth extended merge candidate does not exist, the L(1-X) motion vector of the same candidate is used instead as the unidirectional predicted motion vector for the geometric partitioning mode.
[0048] FIG. 5 illustrates the selection of a unidirectional predictive motion vector from the motion vectors in the merge candidate list of the GPM.
[0049] Blending along geometric partition edges After each geometric partition is obtained using its respective motion, blending is applied to the two unidirectionally predicted signals to derive samples around the geometric partition edges. The blending weights for each position of the CU are derived based on the distance from each individual sample position to the corresponding partition edge.
[0050] GPM signaling design According to the current GPM design, the use of GPM is indicated by signaling one flag at the CU level. This flag is signaled only if the current CU is coded in merge mode or skip mode. Specifically, if the flag is equal to 1, it indicates that the current CU is predicted by GPM. Otherwise (if the flag is equal to 0), the CU is coded by other merge modes, such as normal merge mode, merge mode with motion vector differential, or combined inter and intra prediction. If GPM is enabled for the current CU, one syntax element, merge_gpm_partition_idx, indicating the applied geometric partition mode (specifying the direction and offset of the line from the CU center that divides the CU into two partitions, as shown in FIG. 4) is further signaled. Then, two syntax elements, merge_gpm_idx0 and merge_gpm_idx1, indicating the indexes of the unidirectionally predicted merge candidates used for the first and second GPM partitions are signaled. More specifically, these two syntax elements are used to determine the unidirectional MVs of two GPM partitions from the unidirectional prediction merge list, as described in the section "Constructing a Unidirectional Prediction Merge List." According to the current GPM design, the two indices cannot be the same to make the two unidirectional MVs more different. Based on this prior knowledge, the unidirectional prediction merge index of the first GPM partition is signaled first and used as a predictor to reduce the signaling overhead of the unidirectional prediction merge index of the second GPM partition. Specifically, if the second unidirectional prediction merge index is smaller than the first unidirectional prediction merge index, its original value is signaled as is. Otherwise (if the second unidirectional prediction merge index is larger than the first unidirectional prediction merge index), its value is decremented by one before being signaled to the bitstream. At the decoder side, the first unidirectional prediction merge index is decoded first.Next, for decoding the second unidirectional prediction merge index, if the parsed value is less than the first unidirectional prediction merge index, the second unidirectional prediction merge index is set equal to the parsed value; otherwise (if the parsed value is greater than or equal to the first unidirectional prediction merge index), the second unidirectional prediction merge index is set equal to the parsed value plus 1. Table 1 shows the existing syntax elements used for GPM mode in the current VVC specification. Table 1 Existing GPM syntax elements in the merge data syntax table of the VVC specification [Table 1]
[0051] On the other hand, in the current GPM design, a truncated unary code is used to binarize the two unidirectional prediction merge indices, namely, merge_gpm_idx0 and merge_gpm_idx1. Also, since the two unidirectional prediction merge indices cannot be the same, different maximum values are used to truncate the codewords of the two unidirectional prediction merge indices, which are set equal to MaxGPMMergeCand-1 and MaxGPMMergeCand-2 for merge_gpm_idx0 and merge_gpm_idx1, respectively. MaxGPMMergeCand is the number of candidates in the unidirectional prediction merge list.
[0052] When the GPM / AWP mode is applied, two different binarization methods are applied to convert the syntax element merge_gpm_partition_idx into a string of binary bits. Specifically, the syntax element is binarized by a fixed-length code and a truncated binary code in the VVC and AVS3 standards, respectively. On the other hand, for the AWP mode in AVS3, a different maximum value is used to binarize the value of the syntax element. Specifically, in AVS3, the number of allowed GPM / AWP partition modes is 56 (i.e., the maximum value of merge_gpm_partition_idx is 55), while in VVC, the number increases to 64 (i.e., the maximum value of merge_gpm_partition_idx is 63).
[0053] Merge Mode with Motion Vector Difference (MMVD) In addition to the conventional merge mode that derives the motion information of a current block from its spatial / temporal neighboring blocks, the MMVD / UMVE mode is introduced as a special merge mode in both the VVC and AVS standards. Specifically, in both VVC and AVS3, the mode is signaled by an MMVD flag at the coding block level. In the MMVD mode, the first two candidates in the merge list of the conventional merge mode are selected as two base merge candidates for the MMVD. After a base merge candidate is selected and signaled, an additional syntax element is signaled indicating the motion vector differential (MVD) to be added to the motion of the selected merge candidate. The MMVD syntax element includes a merge candidate flag for selecting the base merge candidate, a distance index specifying the magnitude of the MVD, and a direction index indicating the MVD direction.
[0054] In existing MMVD designs, the distance index specifies the size of the MVD, which is defined based on a predetermined set of offsets from the starting point. As shown in Figures 6A and 6B, the offsets are added to the horizontal or vertical components of the starting MV (i.e., the MV of the selected base merge candidate).
[0055] Figure 6A shows the MMVD mode with L0 reference, and Figure 6B shows the MMVD mode with L1 reference.
[0056] Table 2 shows the MVD offsets applied in AVS3. Size are shown respectively. Table 2 MVD offsets used in AVS3 Size [Table 2]
[0057] As shown in Table 3, the direction index is used to identify the code of the signaled MVD. Note that the meaning of the MVD code may vary depending on the starting MV. If the starting MV is a unidirectionally predicted MV or a bidirectionally predicted MV with MVs pointing to two reference pictures whose POCs are both greater than or less than the POC of the current picture, the signaled code is the code of the MVD added to the starting MV. If the starting MV is a bidirectionally predicted MV pointing to two reference pictures whose POCs are one greater than the POC of the current picture and the other less than the POC of the current picture, the signaled code is applied to the L0 MVD, and the opposite value of the signaled code is applied to the L1 MVD. Table 3 MVD Codes Specified by Direction Index [Table 3]
[0058] Normal inter-mode movement signaling Similar to the HEVC standard, in addition to merge / skip mode, both VVC and AVS3 allow one inter CU to explicitly specify its motion information in the bitstream. Overall, motion information signaling in both VVC and AVS3 is kept the same as that in the HEVC standard. Specifically, one inter prediction syntax, inter_pred_idc, indicating whether the prediction signal is from list L0, L1, or both, is signaled first. For each used reference list, the corresponding reference picture is identified by signaling one reference picture index ref_idx_lx (x=0, 1) of the corresponding reference list, and the corresponding MV is represented by one MVP index mvp_lx_flag (x=0, 1) used to select an MV predictor (MVP), followed by the motion vector difference (MVD) between the target MV and the selected MVP. Also, in the VVC standard, one control flag mvd_l1_zero_flag is signaled at the slice level. If mvd_l1_zero_flag is equal to 0, then L1MVD is signaled in the bitstream; otherwise (mvd_l1_zero_flag is equal to 1), then L1MVD is not signaled and its value is always assumed to be 0 in the encoder and decoder.
[0059] Bidirectional prediction with CU-level weights In standards prior to VVC and AVS3, if weighted prediction (WP) is not applied, the bidirectional prediction signal is generated by averaging the unidirectional prediction signals obtained from two reference pictures. In VVC, a single tool coding, namely, bidirectional prediction with CU-level weights (BCW), is introduced to improve the efficiency of bidirectional prediction. Specifically, bidirectional prediction in BCW is extended by allowing weighted averaging of two prediction signals instead of simple averaging, as follows:
number
[0060] JPEG0007784452000005.jpg50167
[0061] Template Matching Template matching (TM) is a decoder-side motion vector (MV) derivation method that refines the motion information of the current CU by finding the best match between a template consisting of the reconstructed samples above and to the left of the current CU and a reference block (i.e., the same size as the template) in a reference image. As shown in Figure 7, a MV is searched for within a pixel search range of [-8, +8], centered on the initial motion vector of the current CU. The best match may be defined as the MV that achieves the minimum matching cost, such as the sum of absolute differences (SAD) or the sum of absolute transformed differences (SATD) between the current template and the reference template. TM mode is applied to inter-coding / decoding in two different ways:
[0062] In AMVP mode, an MVP candidate is determined based on the template matching difference to pick the one with the smallest difference between the current block template and the reference block template. Then, TM is performed only on this specific MVP candidate for MV refinement. TM refines this MVP candidate by using an iterative diamond search, starting with full-pixel MVD accuracy (or 4 pixels in the case of 4-pixel AMVR mode) within a search range of [-8, +8] pixels. The AMVP candidate can be further refined by a full-pixel MVD accuracy (4 pixels in the case of 4-pixel AMVR mode) cross search, followed by a half-pixel cross search and a quarter-pixel cross search depending on the AMVR mode, as shown in Table 14 below. This search process ensures that the MVP candidate has the same MVD accuracy as that shown in AMVR mode after the TM process. Table 14 [Table 4]
[0063] In merge mode, a similar search method is applied to the merge candidates indicated by the merge index. As shown in the table above, the TM may go all the way up to 1 / 8-pixel MVD precision or skip beyond half-pixel MVD precision, depending on whether or not to use an alternative interpolation filter (used when AMVR is in half-pixel mode) depending on the merged motion information.
[0064] As mentioned above, the unidirectional motion used to generate prediction samples for two GPM partitions is obtained directly from the merge candidate. If there is no strong correlation between the MVs of spatially / temporally adjacent blocks, the unidirectional MVs derived from the merge candidate may not be accurate enough to capture the true motion of the GPM partitions. Motion estimation can provide more accurate motion, but this comes at the expense of significant signaling overhead, since any motion refinement can be applied on top of the existing unidirectional MVs. On the other hand, MVD mode, utilized in both the VVC and AVS3 standards, has proven to be an efficient signaling mechanism for reducing MVD signaling overhead. Therefore, combining GPM mode with MMVD mode may also be advantageous. Such a combination may improve the overall coding / decoding efficiency of the GPM tool by providing more accurate MVs to capture the individual motion of each GPM partition.
[0065] As mentioned above, in both the VVC and AVS3 standards, the GPM mode is only applicable to the merge / skip mode. Such a design may not be optimal in terms of coding / decoding efficiency, given that not all non-merge inter CUs can benefit from the flexible non-rectangular partitions of the GPM. Meanwhile, for similar reasons as those mentioned above, unidirectionally predicted motion candidates derived from the regular merge / skip mode may not be able to accurately capture the true motion of two geometric partitions. Based on this analysis, additional coding / decoding gains can be expected by reasonably extending the GPM mode to non-merge inter modes (i.e., CUs that explicitly signal their motion information in the bitstream). However, improved MV accuracy comes at the expense of increased signaling overhead. Therefore, to efficiently apply the GPM mode to explicit inter modes, it is important to identify an effective signaling scheme that can minimize signaling costs while providing more accurate MVs for the two geometric partitions.
[0066] Proposed method This disclosure proposes a method to further improve the coding / decoding efficiency of GPM by applying additional motion refinement on top of the existing unidirectional MVD applied to each GPM partition. The proposed method is called Geometric Partitioning Mode with Motion Vector Refinement (GPM-MVR). In the proposed scheme, motion refinement is signaled in a manner similar to existing MMVD designs, i.e., based on a set of predefined MVD magnitudes and motion refinement directions.
[0067] In another aspect of this disclosure, a solution is provided for extending GPM mode to explicit inter mode. For simplicity, these schemes are referred to as geometric partition mode with explicit motion signaling (GPM-EMS). Specifically, to achieve better integration with regular inter mode, the proposed GPM-EMS scheme utilizes the existing motion signaling mechanism, i.e., MVP+MVD, to specify corresponding unidirectional MVs of two geometric partitions.
[0068] Geometric partitioning mode with individual motion vector refinement To improve the coding / decoding efficiency of GPM, this section proposes an improved geometric partitioning mode with individual motion vector refinement. Specifically, given a GPM partition, the proposed method first uses the existing syntax merge_gpm_idx0 and merge_gpm_idx1 to identify unidirectional motion vectors for two GPM partitions from the existing unidirectional prediction merge candidate list and uses them as base motion vectors. After determining the two base motion vectors, two sets of new syntax elements are introduced to individually specify the motion refinement values to be applied on the base motion vectors of the two GPM partitions. Specifically, two flags, namely, gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag, which indicate whether GPM-MVR is applied to the first and second GPM partitions, respectively, are first signaled. If the flag of one GPM partition is equal to 1, the corresponding value of the MVR applied to the base MV of this partition, i.e., one distance index (indicated by the syntax elements gpm_mvr_partIdx0_distance_idx and gpm_mvr_partIdx1_distance_idx) to specify the size of the MVR and one direction index (indicated by the syntax elements gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx1_direction_idx) to specify the direction of the MVR, are signaled in MMVD style. Table 4 shows the syntax elements introduced in the proposed GPM-MVR method. Table 4. Syntax elements of the proposed GPM-MVR method with separate MVRs for the two GPM partitions (Method 1) [Table 5]
[0069] Based on the proposed syntax elements shown in Table 4, at the decoder, the final MV used to generate unidirectionally predicted samples for each GPM partition is equal to the sum of the signaled motion vector refinement and the corresponding base MV. In practice, different sets of MVR magnitudes and directions may be predefined and applied to the proposed GPM-MVR scheme, which can provide different trade-offs between motion vector accuracy and signaling overhead. In one specific example, the eight MVDs used in the VVC standard are: Size It is proposed to reuse four MVD dimensions (i.e., 1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 4 pixels, 8 pixels, 16 pixels, and 32 pixels) and four MVD directions (i.e., + / -x axis and + / -y axis) in the proposed GPM-MVR scheme. In another embodiment, the existing five MVD dimensions used in the AVS3 standard are Size {1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels and 4 pixels} and four MVD directions (i.e., + / -x axis and + / -y axis) are applied in the proposed GPM-MVR scheme.
[0070] As discussed in the "GPM Signaling Design" section, the existing GPM design enforces a constraint that the two unidirectional prediction merge indices must be different because the unidirectional MVs used for two GPM partitions cannot be the same. However, in the proposed GPM-MVR scheme, an additional motion refinement is applied on top of the existing GPM unidirectional MV. Therefore, even if the base MVs of the two GPM partitions are the same, if the two motion vector refinement values are not the same, the final unidirectional MVs used to predict the two partitions may be different. Based on the above considerations, when the proposed GPM-MVR scheme is applied, the constraint (which restricts the two unidirectional prediction merge indices to be different) is removed. Also, because the two unidirectional prediction merge indices can be the same, the same maximum value, MaxGPMMergeCand-1, is used for binarization of both merge_gpm_idx0 and merge_gpm_idx1, where MaxGPMMergeCand is the number of candidates in the unidirectional prediction merge list.
[0071] As analyzed above, when the unidirectional prediction merge indexes (i.e., merge_gpm_idx0 and merge_gpm_idx1) of two GPM partitions are the same, the values of the two motion vector refinements cannot be the same to ensure that the final MVs of the two partitions are different. Based on this condition, in one embodiment of the present disclosure, when the unidirectional prediction merge indexes of two GPM partitions are the same (i.e., merge_gpm_idx0 is equal to merge_gpm_idx1), a signaling redundancy elimination method is proposed that uses the MVR of the first GPM partition to reduce the signaling overhead of the MVR of the second GPM partition. In one embodiment, the following signaling condition applies:
[0072] First, if the flag gpm_mvr_partIdx0_enable_flag is equal to 0 (i.e., GPM-MVR does not apply to the first GPM partition), the flag gpm_mvr_partIdx1_enable_flag is not signaled but is inferred to be 1 (i.e., GPM-MVR applies to the second GPM partition).
[0073] Second, if the flags gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are both equal to 1 (i.e., GPM_MVR applies to two GPM partitions) and gpm_mvr_partIdx0_direction_idx is equal to gpm_mvr_partIdx1_direction_idx (i.e., the MVRs of the two GPM partitions have the same direction), the MVR size of the first GPM partition (i.e., gpm_mvr_partIdx0_distance_idx) is used to predict the MVR size of the second GPM partition (i.e., gpm_mvr_partIdx1_distance_idx). Specifically, if gpm_mvr_partIdx1_distance_idx is smaller than gpm_mvr_partIdx0_distance_idx, its original value is signaled directly, otherwise (gpm_mvr_partIdx1_distance_idx is larger than gpm_mvr_partIdx0_distance_idx), its value is decremented by 1 before being signaled in the bitstream. On the decoder side, to decode the value of gpm_mvr_partIdx1_distance_idx, if the parsed value is less than gpm_mvr_partIdx0_distance_idx, then gpm_mvr_partIdx1_distance_idx is set equal to the parsed value, otherwise (if the parsed value is greater than or equal to gpm_mvr_partIdx0_distance_idx), then gpm_mvr_partIdx1_distance_idx is set equal to the parsed value plus one. In such cases, to further reduce overhead, different maximum values MaxGPMMVRDistance-1 and MaxGPMMVRDistance-2 can be used for binarization of gpm_mvr_partIdx0_distance_idx and gpm_mvr_partIdx1_distance_idx, where MaxGPMMVRDistance is the number of allowed magnitudes of motion vector refinement.
[0074] In another embodiment, the MVR direction It is proposed to switch the signaling order of gpm_mvr_partIdx0_direction_idx / gpm_mvr_partIdx1_direction_idx and gpm_mvr_partIdx0_distance_idx / gpm_mvr_partIdx1_distance_idx so that the MVR direction of the first GPM partition is signaled before the MVR magnitude. In this way, following a logic similar to that described above, the encoder / decoder can adjust the signaling of the MVR direction of the second GPM partition depending on the MVR direction of the first GPM partition. In other embodiments, the MVR magnitude and direction of the second GPM partition can be signaled first, and then they can be signaled after the second GPM partition. 1 It is proposed to condition the signaling of the MVR size and direction of the GPM partition.
[0075] In other embodiments, it is proposed to signal syntax elements related to GPM-MVR before signaling existing GPM syntax elements. Specifically, in such a design, two flags, namely, gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag, indicating whether GPM-MVR applies to the first and second GPM partitions, respectively, are signaled first. If the flag for one GPM partition is equal to 1, the distance index (indicated by syntax elements gpm_mvr_partIdx0_distance_idx and gpm_mvr_partIdx1_distance_idx) and the direction index (indicated by syntax elements gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx1_distance_idx) are Signaled Specifies the direction of MVR. Then, the existing syntax merge_gpm_idx0 and merge_gpm_idx1 are signaled to identify the unidirectional MVs of the two GPM partitions, i.e., the base MVs. Table 5 shows the proposed GPM-MVR signaling scheme. Table 5. Syntax elements of the proposed GPM-MVR method with separate MVRs for the two GPM partitions (Method 2) [Table 6]
[0076] Similar to the signaling method of Table 4, certain conditions may be applied to ensure that the resulting MVs used for predicting the two GPM partitions are not the same when the GPM-MVR signaling method of Table 5 is applied. Specifically, the following conditions are proposed to constrain the signaling of the unidirectional prediction merge indices merge_gpm_idx0 and merge_gpm_idx1 depending on the values of MVR applied to the first and second GPM partitions:
[0077] First, if the values of both gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 0 (ie, GPM-MVR is disabled in the two GPM partitions), then the values of merge_gpm_idx0 and merge_gpm_idx1 cannot be the same.
[0078] Second, if gpm_mvr_partIdx0_enable_flag is equal to 1 (i.e., GPM-MVR is enabled in the first GPM partition) and gpm_mvr_partIdx1_enable_flag is equal to 0 (i.e., GPM-MVR is disabled in the second GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 can be the same.
[0079] Third, if gpm_mvr_partIdx0_enable_flag is equal to 0 (i.e., GPM-MVR is disabled in the first GPM partition) and gpm_mvr_partIdx1_enable_flag is equal to 1 (i.e., GPM-MVR is enabled in the second GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 can be the same.
[0080] Fourth, if the values of both gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 1 (i.e., GPM_MVR is enabled for both two GPM partitions), the determination of whether the values of merge_gpm_idx0 and merge_gpm_idx1 can be the same depends on the values of the MVR (indicated by gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx0_distance_idx, and gpm_mvr_partIdx1_direction_idx and gpm_mvr_partIdx1_distance_idx) that apply to the two GPM partitions. If the values of the two MVRs are equal, merge_gpm_idx0 and merge_gpm_idx1 cannot be the same. Otherwise (when the two MVR values are not equal), the values of merge_gpm_idx0 and merge_gpm_idx1 can be the same.
[0081] In the above four cases, when the values of merge_gpm_idx0 and merge_gpm_idx1 cannot be the same, the index value of one partition can be used as a predictor of the index value of the other partition. One method proposes to signal merge_gpm_idx0 first and use its value to predict merge_gpm_idx1. Specifically, in the encoder, if merge_gpm_idx1 is greater than merge_gpm_idx0, the value of merge_gpm_idx1 sent to the decoder is decreased by 1. In the decoder, if the received value of merge_gpm_idx1 is greater than or equal to the received value of merge_gpm_idx0, the value of merge_gpm_idx1 is increased by 1. Another method proposes to signal merge_gpm_idx1 first and use its value to predict merge_gpm_idx0. Therefore, in such a case, if merge_gpm_idx0 is greater than merge_gpm_idx1 in the encoder, the value of merge_gpm_idx0 sent to the decoder is decreased by 1. If the received value of merge_gpm_idx0 is greater than or equal to the received value of merge_gpm_idx1 in the decoder, the value of merge_gpm_idx0 is increased by 1. Also, similar to existing GPM signaling designs, different maximum values MaxGPMMergeCand-1 and MaxGPMMergeCand-2 can be used to binarize the first and second index values, respectively, depending on the signaling order. On the other hand, if the values of merge_gpm_idx0 and merge_gpm_idx1 can be the same because there is no correlation between the two index values, the same maximum value MaxGPMMergeCand-1 can be used to binarize both of the two index values.
[0082] In the above method, to reduce signaling costs, different maximum values may be applied to the binarization of merge_gpm_idx0 and merge_gpm_idx1. The selection of the corresponding maximum value depends on the decoded value of MVR (indicated by gpm_mvr_partIdx0_enable, gpm_mvr_partIdx1_enable, gpm_mvr_partIdx0_direction_idx, gpm_mvr_partIdx1_direction_idx, gpm_mvr_partIdx0_distance_idx, and gpm_mvr_partIdx1_distance_idx). Such a design may cause undesired parsing dependencies between different GPM syntax elements, which may affect the overall parsing. To solve this problem, in one embodiment, it is proposed that one and the same maximum value (e.g., MaxGPMMergeCand-1) is always used to parse the values of merge_gpm_idx0 and merge_gpm_idx1. When using such a method, one bitstream compatibility constraint may be used to prevent two decoded MVs of two GPM partitions from being the same. Alternatively, such a non-identity constraint may be removed so that the decoded MVs of two GPM partitions can be made the same. On the other hand, when such a method is applied (i.e., using the same maximum value for merge_gpm_idx0 and merge_gpm_idx1), there is no parsing dependency between merge_gpm_idx0 / merge_gpm_idx1 and other GPM-MVR syntax elements. Therefore, the order in which these syntax elements are signaled is no longer important.In one embodiment, it is proposed to move the merge_gpm_idx0 / merge_gpm_idx1 signaling before the gpm_mvr_partIdx0_enable, gpm_mvr_partIdx1_enable, gpm_mvr_partIdx0_direction_idx, gpm_mvr_partIdx1_direction_idx, gpm_mvr_partIdx0_distance_idx, and gpm_mvr_partIdx1_distance_idx signaling.
[0083] Geometric partitioning mode with symmetric motion vector refinement. In the above GPM-MVR method, two separate MVR values are signaled, one of which is applied to improve the base MV of only one GPM partition. This method is effective in improving prediction accuracy by enabling independent motion refinement for each GPM partition. However, such flexible motion refinement comes at the expense of increased signaling overhead, considering that two different sets of GMP-MVR syntax elements must be sent from the encoder to the decoder. To reduce signaling overhead, this section proposes a geometric partitioning mode with symmetric motion vector refinement. Specifically, in this method, one MVR value is signaled to one GPM CU and is used for both GPM partitions depending on the symmetry relationship of the picture order count (POC) values between the current picture and the reference pictures associated with the two GPM partitions. Table 6 shows the syntax elements when the proposed method is applied. Table 6. Syntax elements of the proposed GPM-MVR method with symmetric MVR of two GPM partitions (Method 1) [Table 7]
[0084] As shown in Table 6, after the base MVs of the two GPM partitions are selected (based on merge_gpm_idx0 and merge_gpm_idx1), one flag, gpm_mvr_enable_flag, is signaled to indicate whether the GPM-MVR mode is applied to the current GPM CU. If the flag is equal to 1, it indicates that motion refinement is applied to enhance the base MVs of the two GPM partitions. Otherwise (if the flag is equal to 0), it indicates that no motion refinement is applied to either of the two partitions. If the GPM-MVR mode is enabled, additional syntax elements are further signaled to specify the values of the applied MVR by the direction index, gpm_mvr_direction_idx, and the magnitude index, gpm_mvr_distance_idx. Also, similar to the MMVD mode, the meaning of the MVR code may change depending on the relationship between the POC of the current picture and the two reference pictures of the GPM partition. Specifically, if both POCs of two reference images are larger or smaller than the POC of the current image, the signaled code is the code of the MVR added to both two base MVs. Otherwise (if the POC of one reference image is larger than the current image and the POC of the other reference image is smaller than the current image), the signaled code is applied to the MVR of the first GPM partition, and the inverse code is applied to the MVR of the second GPM partition. In Table 6, the values of merge_gpm_idx0 and merge_gpm_idx1 can be the same.
[0085] In another embodiment, it is proposed to signal two different flags to separately control the enable / disable of GPM-MVR mode for the two GPM partitions. However, if GPM-MVR mode is enabled, only one MVR is signaled based on the syntax elements gpm_mvr_direction_idx and gpm_mvr_distance_idx. The corresponding syntax table for such a signaling method is shown in Table 7. Table 7. Syntax elements of the proposed GPM-MVR method with symmetric MVR of two GPM partitions (Method 2) [Table 8]
[0086] The values of merge_gpm_idx0 and merge_gpm_idx1 can be the same when the signaling method of Table 7 is applied. However, to ensure that the resulting MVs applied to the two GPM partitions are not redundant, if the flag gpm_mvr_partIdx0_enable_flag is equal to 0 (i.e., GPM-MVR is not applied to the first GPM partition), the flag gpm_mvr_partIdx1_enable_flag is not signaled but is inferred to be 1 (i.e., GPM-MVR is applied to the second GPM partition).
[0087] Accepted MVR indications for GPM-MVR In the above GPM-MVR method, a group of fixed MVR values is used for the GPM CU in both the encoder and decoder for a single video sequence. Such a design is suboptimal for high-resolution video content or video content with intense motion. In these cases, the MV tends to be very large, so the fixed MVR values may not be optimal for capturing the actual motion of those blocks. To further improve the encoding / decoding performance of the GPM-MVR mode, this disclosure proposes supporting adaptation of selectable MVR values by the GPM-MVR mode at various coding levels, such as the sequence level, picture / slice picture level, and coding block group level. For example, multiple MVR sets and corresponding codewords may be derived offline according to the specific motion characteristics of different video sequences. The encoder can select the best MVR set and signal the corresponding index of the selected MVR set to the decoder.
[0088] In some embodiments of the present disclosure, for GPM-MVR mode, eight offset magnitudes (i.e., 1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 4 pixels, 8 pixels, 16 pixels, and 32 pixels) and four offset magnitudes (i.e., 1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 4 pixels, 8 pixels, 16 pixels, and 32 pixels) are available. offset In addition to the default MVR offsets, including directions (+ / -x and + / -y axes), other MVR offsets are proposed, as defined in the following tables: Table 15 shows the magnitude of the offsets in the proposed set of secondary MVR offsets. Table 16 shows the magnitude of the offsets in the proposed set of secondary MVR offsets. offset Indicates direction. Table 15 [Table 9] Table 16 [Table 10]
[0089] In Tables 15 and 16, the values +½ and −½ on the x-axis and y-axis indicate the horizontal and vertical diagonal directions (+45° and −45°). As shown in Tables 15 and 16, compared to the existing set of MVR offsets, the second set of MVR offsets introduces two new offset magnitudes (i.e., 3 pixels and 6 pixels) and four new offset directions (45°, 135°, 225°, and 315°). The newly added MVR offsets make the second set of MVR offsets more suitable for encoding video blocks with complex motion. In addition, to enable adaptive switching between the two sets of MVR offsets, it is proposed to signal one control flag at a specific coding level (e.g., sequence, picture, slice, CTU, coding block, etc.) to indicate which set of MVR offsets is selected for the GPM-MVR mode applied at that coding level. Assuming that the proposed adaptation is performed at the picture level, Table 17 below shows the corresponding syntax elements signaled in the picture header. Table 17 [Table 11]
[0090] In Table 17 above, a new flag, ph_gpm_mvr_offset_set_flag, is used to indicate the selection of the corresponding GPM_MVR offset to be used for the image. If the flag is equal to 0, the default MVR offset (i.e., 1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 4 pixels, 8 pixels, 16 pixels, and 32 pixels) is used. Offset Size and four offset This means that the direction (+ / -x axis and + / -y axis) is applied to the GPM-MVR mode in the image. Otherwise (flag equals 1), the second MVR offset (i.e., 1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 3 pixels, 4 pixels, 6 pixels, 8 pixels, and 16 pixels) is applied. Offset Size and 8 offset This means that the directions (+ / -x axis, + / -y axis, 45°, 135°, 225°, and 315°) are applied to the GPM-MVR mode in the image.
[0091] The MVR offset may be signaled in different ways. First, considering that the MVR direction is usually statistically uniformly distributed, it is proposed to binarize the MVR direction using a fixed-length codeword. Taking the default MVR offset as an example, there are a total of four directions, and codewords 00, 01, 10, and 11 can be used to represent the four directions. On the other hand, since the magnitude of the MVR offset may have various distributions adapted to the specific motion characteristics of the video content, it is proposed to use a variable-length codeword to binarize the MVR direction. offset It is proposed to binarize the magnitude of the MVR offsets. Table 18 below shows the MVR offsets for the default MVR offset set and the second MVR offset set. offset Here is one particular codeword table that can be used to binarize the magnitude of Table 18 [Table 12]
[0092] In other embodiments, different fixed-length variable code words may be applied to binarize the magnitude of the MVR offsets in the set of default MVR offsets and the set of second MVR offsets, for example, bin “0” and bin “1” in the above code word table may be swapped to accommodate different 0 / 1 statistics of a context-adaptive binary arithmetic coding and decoding (CABAC) engine.
[0093] In one embodiment, two different codeword tables are provided to allow for MVR. offset The magnitude values of MVR offsets are binarized. The following tables show the corresponding codewords for the default and second MVR offset sets applied in the first and second codeword tables. Table 19 shows the codewords for the magnitude of MVR offsets in the first codeword table. Table 20 shows the codewords for the magnitude of MVR offsets in the second codeword table. Table 19 [Table 13] Table 20 [Table 14]
[0094] To enable adaptive switching between two codeword tables, one indication flag is signaled at one coding level (e.g., sequence, picture, slice, CTU, coded block, etc.) to enable MVR at that coding level. offset It is proposed to specify which codeword table is used to binarize the magnitude of . Assuming the proposed adaptation is performed at the picture level, Table 21 below shows the corresponding syntax elements signaled in the picture header, with the newly added syntax elements shown in italic and bold. Table 21 [Table 15]
[0095] In the above syntax table, the new flag ph_gpm_mvr_step_codeword_flag specifies the MVR offset is used to indicate the selection of the corresponding code table to be used for binarization of the magnitude of . If the flag is equal to 0, it indicates that the first code table is applied to the image, otherwise (i.e., the flag is equal to 1), it indicates that the second code table is applied to the image.
[0096] In another embodiment, it is proposed to always use one codeword table to binarize the magnitude of the MVR offset during encoding / decoding of the entire video sequence. offset It is proposed to always use the first codeword table for binarizing the magnitude of MVR. offset It is proposed to always use the second codeword table for binarization of the magnitude of .
[0097] In another method, a statistics-based binarization method may be applied to adaptively design an optimal codeword for the MVR offset magnitude on the fly without signaling. The statistics used to determine the optimal codeword may be, but are not limited to, a probability distribution of the MVR offset magnitude collected for a number of previously coded images, slices, and / or coding blocks. The codeword may be redetermined / updated at various frequency levels. For example, the update may be performed every time a CU is coded / decoded in GPM-MVR mode. In another embodiment, the update may be redetermined and / or updated every time there are multiple (e.g., 8 or 16) CUs coded / decoded in GPM-MVR mode.
[0098] Instead of redesigning one new set of codewords in the other method, the proposed statistical-based method redesigns the MVR based on the same set of codewords to assign shorter codewords to more frequently used magnitudes and longer codewords to less frequently used magnitudes. offsetTaking the table below as an example, assuming that statistics are collected at the image level, the column labeled "Used" indicates the corresponding proportion of different MVR offset magnitudes used by the coding blocks of GPM-MVR in the previously coded image. Depending on the value of the "Used" column, with the same binarization method (i.e., truncated unary codewords), the encoder / decoder will offset The encoder / decoder then ranks the most frequently used MVRs. offset (i.e., 1 pixel), and assigns the shortest codeword (i.e., "1") to the second most frequently used MVR offset , the size (i.e., 1 / 2 pixel) of the second shortest codeword (i.e., “01”), …, the two most rarely used MVRs. offset Therefore, with this reordering scheme, the MVR offset The same set of codewords can be freely reordered to accommodate dynamic changes in the statistical distribution of the magnitudes of . [Table 16]
[0099] Encoder speedup logic for GPM-MVR rate-distortion optimization For the proposed GPM-MVR scheme, to determine the optimal MVR for each GPM partition, the encoder needs to test the rate-distortion cost of each GPM partition multiple times, with different MVR values applied each time. This may significantly increase the encoding / decoding complexity of the GPM mode. To solve the encoding / decoding complexity problem, this section proposes the following fast encoding / decoding logic:
[0100] First, due to the quadtree / binary / ternary tree block partitioning structure applied in VVC and AVS3, the same coding block can be checked during the rate-distortion optimization (RDO) process, and each time, the same coding block is partitioned by different partitioning passes. In the current implementation of the VTM / HPM encoder, GPM, GPM-MVR mode, and other inter and intra coding / decoding modes are always tested each time the same CU is obtained through a different block partition combination. Generally speaking, only the neighboring blocks of a CU may differ for different partitioning passes, but this should have a relatively small impact on the optimal coding / decoding mode selected by a CU. Based on this consideration, to reduce the total number of applied GPM RDOs, it is proposed to store the decision on whether GPM mode is selected when the RD cost of a CU is checked for the first time. If the same CU is subsequently checked again by the RDO process (by another partitioning pass), the RD cost of GPM (including GPM-MVR) is checked only when GPM is selected for the CU for the first time. If a GPM is not selected for the first RD check of a CU, only the GPM (without GPM-MVR) is tested when the same CU is obtained by another split pass. Alternatively, if a GPM is not selected for the first RD check of a CU, both the GPM and GPM-MVR are not tested when the same CU is obtained by another split pass.
[0101] Second, to reduce the number of GPM partitions in GPM-MVR mode, the first time the RD cost of a CU is checked, the one with the smallest RD cost is selected. is It is proposed to keep the first M GPM split modes. Then, if the same CU is checked again by the RDO process (by another split pass), only those m GPM split modes are tested for the GPM-MVR mode.
[0102] Third, to reduce the number of GPM partitions tested in the initial RDO process, it is proposed to first calculate the sum of absolute differences (SAD) for each GPM partition when using different unidirectional predictive merge candidates for two GPM partitions. Then, for each GPM partition under one specific partition mode, select the best unidirectional predictive merge candidate with the smallest SAD value, and calculate the corresponding SAD value for that partition mode, which is equal to the sum of the SAD values of the best unidirectional predictive merge candidates for the two GPM partitions. Then, in the following RD process, for the GPM-MVR mode, only the first N partition modes with the best SAD values in the previous step are tested.
[0103] Geometric partitioning with explicit motion signaling - Patents.com In this section, several methods are proposed to extend the GPM mode, in which two unidirectional MVs are explicitly signaled from the encoder to the decoder, to regular inter-mode bi-prediction.
[0104] In the first solution (Solution 1), it is proposed to fully reuse the existing bi-prediction motion signaling to signal two unidirectional MVs in GPM mode. Table 8 shows a modified syntax table of the proposed scheme, with newly added syntax elements shown in bold italics. As shown in Table 8, in this solution, all existing syntax elements signaling L0 and L1 motion information are fully reused to indicate the unidirectional MVs of the two GPM partitions, respectively. It is also assumed that the L0 MV is always associated with the first GPM partition, and the L1 MV is always associated with the second GPM partition. Meanwhile, in Table 8, the inter-prediction syntax, i.e., inter_pred_idc, is signaled before the GPM flag (i.e., gpm_flag), so that the value of inter_pred_idc can be used to condition the presence of the GPM flag. Specifically, the flag gpm_flag needs to be signaled only if inter_pred_idc is equal to PRED_BI (i.e., bidirectional prediction) and both inter_affine_flag and sym_mvd_flag are equal to 0 (i.e., the CU is not encoded or decoded in either affine or SMVD mode). If the flag gpm_flag is not signaled, its value is always inferred to be 0 (i.e., GPM mode is disabled). If gpm_flag is 1, another syntax element gpm_partition_idx is further signaled, indicating the selected GPM mode (out of a total of 64 GPM partitions) for the current CU. Table 8. Revised syntax table for motion signaling for Solution 1 (Option 1) [Table 17] JPEG0007784452000020.jpg79169
[0105] Another method proposes signaling the flag gpm_flag before other inter-signaling syntax elements so that the value of gpm_flag can be used to determine whether other inter syntax elements need to be present. Table 9 shows the corresponding syntax table when such a method is applied, with newly added syntax elements shown in bold italics. As can be seen, gpm_flag is signaled first in Table 9. If gpm_flag is equal to 1, the signaling of the corresponding inter_pred_idc, inter_affine_flag, and sym_mvd_flag can be bypassed. Instead, the corresponding values of the three syntax elements can be inferred as PRED_BI, 0, and 0, respectively. Table 9. Revised syntax table for motion signaling for Solution 1 (Option 2) [Table 18] JPEG0007784452000022.jpg85170
[0106] In both Table 8 and Table 9, the SMVD mode cannot be combined with the GPM mode. In another embodiment, it is proposed to allow the SMVD mode if the current CU is encoded or decoded by the GPM mode. If such a combination is allowed, by following the same design of SMVD, the MVDs of the two GPM partitions are assumed to be symmetric, so that only the MVD of the first GPM partition needs to be signaled, and the MVD of the second GPM partition is always symmetric with the first MVD. If such a method is applied, the corresponding signaling condition of the sym_mvd_flag on the gpm_flag can be deleted.
[0107] As shown above, in Solution 1, it is always assumed that L0 MVs are used for the first GPM partition and L1 MVs are used for the second GPM partition. Such a design may not be optimal in the sense that it prohibits MVs of two GPM partitions from being obtained from the same prediction list (L0 or L1). To solve this problem, Solution 2, an alternative GPM-EMS scheme, is proposed, with a signaling design as shown in Table 10. In Table 10, newly added syntax elements are shown in bold italics. As shown in Table 10, the flag gpm_flag is signaled first. If the flag is equal to 1 (i.e., GPM is enabled), the syntax gpm_partition_idx for specifying the selected GPM mode is signaled. Next, an additional flag gpm_pred_dir_flag0 is signaled, indicating the corresponding prediction list from which the MVs of the first GPM partition originate. If the flag gpm_pred_dir_flag0 is equal to 1, it indicates that the MVs of the first GPM partition come from L1; otherwise (if the flag is equal to 0), it indicates that the MVs of the first GPM partition come from L0. Then, the existing syntax elements ref_idx_l0, mvp_l0_flag, and mvd_coding() are utilized to signal the values of the reference picture index, mvp index, and MVD of the first GPM partition. Meanwhile, similar to the first partition, another syntax element gpm_pred_dir_flag1 is introduced to select the corresponding prediction list of the second GPM partition, followed by the existing syntax elements ref_idx_l1, mvp_l1_flag, and mvd_coding() that are used to derive the MVs of the second GPM partition. Table 10. Revised syntax table for motion signaling in Solution 2 [Table 19] JPEG0007784452000024.jpg123170
[0108] Finally, considering that the GPM mode consists of two unidirectional prediction partitions (excluding blending samples on the split edge), when the proposed GPM-EMS scheme is enabled for one inter-CU, some existing coding tools in VVC and AVS3 that are specifically designed for bidirectional prediction, such as bidirectional optical flow, decoder-side motion vector refinement (DMVR), and bidirectional prediction with CU weights (BCW), can be automatically bypassed. For example, when one of the proposed GPM-EMS schemes is enabled for one CU, considering that BCW is not applicable to the GPM mode, there is no need to further signal the corresponding BCW weights to the CU to reduce signaling overhead.
[0109] Combination of GPM-MVR and GPM-EMS In this section, we propose combining GPM-MVR and GPM-EMS for one CU with geometric partitions. Specifically, unlike GPM-MVR or GPM-EMS, in which only one of merge-based motion signaling and explicit signaling can be applied to signal unidirectionally predicted MVs of two GPM partitions, the proposed scheme allows 1) one partition to use GPM-MVR-based motion signaling while the other partition uses GPM-EMS-based motion signaling, or 2) two partitions to use GPM-MVR-based motion signaling, or 3) two partitions to use GPM-EMS-based motion signaling. With the GPM-MVR signaling in Table 4 and the GPM-EMS in Table 10, the corresponding syntax table after the proposed combination of GPM-MVR and GPM-EMS is shown in Table 11. In Table 11, newly added syntax elements are shown in bold italics. As shown in Table 11, two additional syntax elements gpm_merge_flag0 and gpm_merge_flag1 are introduced for partitions #1 and #2, respectively, which specify that the corresponding partition uses GPM-MVR-based merge signaling or GPM-EMS-based explicit signaling. If the flag is 1, it means that GPM-MVR-based signaling is enabled for the partition where GPM unidirectional prediction motion is signaled by merge_gpm_idxX, gpm_mvr_partIdxX_enabled_flag, gpm_mvr_partIdxX_direction_idx, and gpm_mvr_partIdxX_distance_idx (X=0, 1). Otherwise, if the flag is 0, it means that the unidirectional predicted motion of the partition is explicitly signaled in the GPM-EMS manner by the syntax elements gpm_pred_dir_flagX, ref_idx_lX, mvp_lX_flag, and mvd_lX (X=0, 1). Table 11. Syntax table for GPM mode with proposed combination of GPM-MVR and GPM-EMS [Table 20]
[0110] Combining GPM-MVR with template matching In this section, different solutions are presented to combine GPM-MVR and template matching.
[0111] In Method 1, when one CU is encoded and decoded in GPM mode, it is proposed to signal two separate flags for two GPM partitions, respectively, indicating whether the unidirectional motion of the corresponding partition is further refined by template matching. If the flag is enabled, a template is generated using the reconstructed samples to the left and above the current CU, and then the unidirectional motion of the partition is refined by minimizing the difference between the template and its reference samples, following the same procedure as described in the "Template Matching" section. Otherwise (if the flag is disabled), template matching is not applied to the partition, and GPM-MVR may be further applied. Taking the GPM-MVR signaling method in Table 5 as an example, Table 12 shows the corresponding syntax table when GPM-MVR and template matching are combined. In Table 12, newly added syntax elements are shown in bold italics. Table 12. Syntax elements of the proposed method combining GPM-MVR and template matching (Method 1) [Table 21]
[0112] As shown in Table 12, in the proposed scheme, two additional flags, gpm_tm_enable_flag0 and gpm_tm_enable_flag1, respectively indicating whether the motion is refined for two GPM partitions, are signaled first. If the flag is 1, it indicates that TM is applied to refine unidirectional MV of one partition. If the flag is 0, one flag (gpm_mvr_partIdx0_enable_flag or gpm_mvr_partIdx1_enable_flag) is further signaled indicating whether GPM_MVR is applied to the GPM partition. If the flag for one GPM partition is equal to 1, the distance index (indicated by the syntax elements gpm_mvr_partIdx0_distance_idx and gpm_mvr_partIdx1_distance_idx) and the direction index (syntax elements gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx1_distance_idx) specify the size of the MVR. Then, the existing syntaxes merge_gpm_idx0 and merge_gpm_idx1 for specifying the unidirectional MVs of the two GPM partitions are signaled. Meanwhile, similar to the signaling conditions that apply in Table 5, the following conditions may apply to ensure that the resulting MVs used for prediction of the two GPM partitions are not the same:
[0113] First, if the values of both gpm_tm_enable_flag0 and gpm_tm_enable_flag1 are equal to 1 (i.e., TM is enabled for both GPM partitions), the values of merge_gpm_idx0 and merge_gpm_idx1 cannot be the same.
[0114] Second, if one of gpm_tm_enable_flag0 and gpm_tm_enable_flag1 is 1 and the other is 0, the values of merge_gpm_idx0 and merge_gpm_idx1 can be the same.
[0115] Otherwise, i.e., if both gpm_tm_enable_flag0 and gpm_tm_enable_flag1 are equal to 1, first, when both gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 0 (i.e., GPM-MVR is disabled for both GPM partitions), the values of merge_gpm_idx0 and merge_gpm_idx1 cannot be the same, and second, when gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 0 (i.e., GPM-MVR is disabled for both GPM partitions), the values of merge_gpm_idx0 and merge_gpm_idx1 cannot be the same. The values of merge_gpm_idx0 and merge_gpm_idx1 can be the same when partIdx0_enable_flag is equal to 1 (i.e., GPM-MVR is enabled for the first GPM partition) and gpm_mvr_partIdx1_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for two GPM partitions); third, when gpm_mvr_partIdx0_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for the first GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 can be the same when partIdx0_enable_flag is equal to 1 (i.e., GPM-MVR is enabled for the first GPM partition) and gpm_mvr_partIdx1_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for two GPM partitions). The values of merge_gpm_idx0 and merge_gpm_idx1 can be the same when gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are both equal to 1 (i.e., GPM-MVR is enabled for both of the two GPM partitions), and gpm_mvr_partIdx1_enable_flag is equal to 1 (i.e., GPM-MVR is enabled for the second GPM partition). When two GPM partitions have the same MVR (which is valid for both GPM partitions), the determination of whether the values of merge_gpm_idx0 and merge_gpm_idx1 can be the same depends on the values of the MVRs (indicated by gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx0_distance_idx, and gpm_mvr_partIdx1_direction_idx and gpm_mvr_partIdx1_distance_idx) that apply to the two GPM partitions. If the values of the two MVRs are equal, then merge_gpm_idx0 and merge_gpm_idx1 cannot be the same.Otherwise (when the two MVR values are not equal), the values of merge_gpm_idx0 and merge_gpm_idx1 can be the same.
[0116] In the above method 1, TM and MVR are applied exclusively to GPM. In such a scheme, further application of MVR on the refined MV in TM mode is prohibited. Therefore, to provide more MV candidates for GPM, Method 2 is proposed, which allows applying MVR offset on the refined MV in TM. Table 13 shows the corresponding syntax table when GPM-MVR and template matching are combined. In Table 13, newly added syntax elements are shown in bold italics. Table 13. Syntax elements of the proposed method combining GPM-MVR and template matching (Method 2) [Table 22]
[0117] As shown in Table 13, unlike Table 12, the signaling conditions of gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag on gpm_tm_enable_flag0 and gpm_tm_enable_flag1 are deleted. Therefore, regardless of whether the TM refines a unidirectional movement of one GPM partition or not, MV refinement can always be applied to the MV of a GPM partition. As before, to ensure that the resulting MVs of two GPM partitions are not the same, the following condition must be applied:
[0118] First, if one of gpm_tm_enable_flag0 and gpm_tm_enable_flag1 is 1 and the other is 0, the values of merge_gpm_idx0 and merge_gpm_idx1 can be the same.
[0119] Otherwise, i.e., if both gpm_enable_flag0 and gpm_enable_flag1 are equal to 1 or both flags are equal to 0, firstly, when both gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 0 (i.e., GPM-MVR is disabled for both GPM partitions), the values of merge_gpm_idx0 and merge_gpm_idx1 cannot be the same; and secondly, When gpm_mvr_partIdx0_enable_flag is equal to 1 (i.e., GPM-MVR is enabled for the first GPM partition) and gpm_mvr_partIdx1_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for the second GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 can be the same; third, when gpm_mvr_partIdx0_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for the first GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 can be the same; Fourth, when the values of gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are both equal to 1 (i.e., GPM-MVR is enabled for the second GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 can be the same; and fourth, when the values of gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are both equal to 1 (i.e., GPM-MVR is enabled for the second GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 can be the same. When two MVRs are equal (valid for both), the determination of whether the values of merge_gpm_idx0 and merge_gpm_idx1 can be the same depends on the values of the MVRs (indicated by gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx0_distance_idx, and gpm_mvr_partIdx1_direction_idx and gpm_mvr_partIdx1_distance_idx) that apply to the two GPM partitions. If the values of the two MVRs are equal, then merge_gpm_idx0 and merge_gpm_idx1 cannot be the same.Otherwise (when the two MVR values are not equal), the values of merge_gpm_idx0 and merge_gpm_idx1 can be the same.
[0120] In the above two methods, two separate flags need to be signaled to indicate whether TM is applied to each GPM partition. The additional signaling may reduce overall encoding / decoding efficiency due to additional overhead, especially at low bitrates. To reduce the signaling overhead, Method 3 is proposed, which inserts TM-based unidirectional MVs into the unidirectional MV candidate list for GPM mode instead of introducing additional signaling. TM-based unidirectional MVs are generated according to the same TM process as described in the "Template Matching" section, with the original unidirectional MV of the GPM as the first MV. This scheme eliminates the need to further signal an additional control flag from the encoder to the decoder. Instead, the decoder can identify whether an MV is refined by TM through the corresponding merge index (i.e., merge_gpm_idx0 and merge_gpm_idx1) received from the bitstream. Regular GPM MV candidates (i.e., non-TM) and TM-based MV candidates may be arranged in different ways. One approach proposes placing TM-based MV candidates first at the beginning of the MV candidate list, followed by non-TM-based MV candidates. Another approach proposes placing non-TM-based MV candidates first, followed by TM-based MV candidates. Another approach proposes placing TM-based MV candidates and non-TM-based MV candidates in an interleaved manner. For example, the first N non-TM-based candidates can be placed, followed by all TM-based candidates, and finally the remaining non-TM-based candidates. In another embodiment, the first N TM-based candidates can be placed, followed by all non-TM-based candidates, and finally the remaining TM-based candidates. In another embodiment, it is proposed to alternate between non-TM-based candidates and TM-based candidates, i.e., one non-TM-based candidate, one TM-based candidate, etc.
[0121] The above methods may be implemented employing an apparatus including one or more circuits, including an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic components. The apparatus may employ circuits in combination with other hardware or software components to perform the above-described methods. Each module, sub-module, unit, or sub-unit disclosed above may be at least partially implemented with one or more circuits.
[0122] 9 illustrates a computing environment (or computing device) 910 coupled to a user interface 960. The computing environment 910 may be part of a data processing server. In some embodiments, the computing device 910 may perform any of the various methods or processes (e.g., encoding / decoding methods or processes) according to various embodiments of the present disclosure, as described above. The computing environment 910 may include a processor 920, a memory 940, and an I / O interface 950.
[0123] The processor 920 typically controls the overall operation of the computing environment 910, such as operations related to display, data acquisition, data communication, and image processing. The processor 920 may include one or more processors for executing instructions to perform all or some of the steps of the methods described above. The processor 920 may also include one or more modules that facilitate interaction between the processor 920 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip machine, a GPU, etc.
[0124] Memory 940 is configured to store various types of data to support the operation of computing environment 910. Memory 940 may include predefined software 942. Examples of such data include instructions for any applications or methods operating on computing environment 910, video data sets, image data, etc. Memory 940 may be implemented using any type of volatile or non-volatile memory device, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic or optical disk, or a combination thereof.
[0125] The I / O interface 950 provides an interface between the processor 920 and a peripheral interface module, such as a keyboard, a click wheel, buttons, etc. The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. The I / O interface 950 can be coupled to an encoder and a decoder.
[0126] In some embodiments, a non-transitory computer-readable storage medium is also provided that includes a plurality of programs, such as those contained in memory 940, executable by processor 920 in computing environment 910 to perform the methods described above. For example, the non-transitory computer-readable storage medium may be a ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0127] A non-transitory computer-readable storage medium stores a plurality of programs that are executed by a computing device having one or more processors, and the plurality of programs, when executed by the one or more processors, cause the computing device to perform the motion estimation method described above.
[0128] In some embodiments, the computing environment 910 may be implemented using one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), graphical processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0129] FIG. 8 is a flowchart illustrating a method for decoding a video block with a GPM according to one embodiment of the present disclosure.
[0130] In step 801, processor 920 may receive a control flag associated with a video block. The control flag may be a control variable or any other variable including one or more flags, such as a binary flag, a non-binary flag, etc. In one or more embodiments, the control variable may be the flag "ph_gpm_mvr_offset_set_flag" as shown in Table 17 or Table 21.
[0131] In some embodiments, a control variable allows adaptive switching between multiple sets of MVR offsets and is applied at the coding level.
[0132] In some embodiments, the coding level may be the sequence level, the picture / slice level, the CTU level, or the coded block level. For example, if a control variable is signaled at the encoder side at the picture level, the decoder side correspondingly receives a control variable at the picture level to indicate which MVR offset set to select for the purpose of selecting the corresponding MVR offset associated with the current video block.
[0133] In step 802, the processor 920 may receive an indicator flag associated with the video block. The indicator flag may be an indicator variable or any other variable including one or more flags, such as a binary flag, a non-binary flag, etc. In one or more embodiments, instructions The variable may be the flag "ph_gpm_mvr_step_codeword_flag" as shown in Table 21. The value of the flag "ph_gpm_mvr_step_codeword_flag" allows more flexibility by switching binarization tables to binarize different sets of offsets with different codeword tables.
[0134] In some embodiments, the indicator variable enables adaptive switching between multiple codeword tables that binarize the magnitudes of multiple offsets in a set of multiple MVR offsets under a coding level.
[0135] In step 803, processor 920 may divide the video block into a first geometric partition and a second geometric partition.
[0136] In step 804, processor 920 may select a set of MVR offsets from the sets of MVR offsets based on the control variable.
[0137] In step 805, processor 920 may receive one or more syntax elements and determine first and second MVR offsets to apply to the first and second geometric partitions from a set of selected MVR offsets. Set of is one MVR offset selected by the control variable Set of may be.
[0138] In some embodiments, the set of MVR offsets may include a first set of MVR offsets and a second set of MVR offsets. In some embodiments, the first set of MVR offsets may include a plurality of default offset sizes and a plurality of default offset sizes. offset In some embodiments, the set of second MVR offsets may include multiple default MVR offsets, including multiple alternate offset magnitudes and multiple alternate offsets. offset In some embodiments, the second set of MVR offsets may include multiple alternative MVR offsets, including multiple offsets with different magnitudes and different orientations than the first set of MVR offsets. offset For example, multiple default offset magnitudes and multiple default offset The direction is selected from eight offset magnitudes (e.g., 1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 4 pixels, 8 pixels, 16 pixels, and 32 pixels) and four offset May include direction (+ / -x axis and + / -y axis). Multiple alternative offset magnitudes and multiple alternatives. offset The directions are as shown in Tables 15, 16, 19 and 20. Size and direction.
[0139] As shown in Tables 15, 16, 19, and 20, the set of alternate MVR offsets may include more offset sizes in addition to multiple default offset sizes. offset In addition to directions, many more offset It may also include a direction.
[0140] In some embodiments, the processor 920 may determine to apply a first set of MVR offsets in response to determining that the control variable is equal to 0, and may determine to apply a second set of MVR offsets in response to determining that the control variable is equal to 1.
[0141] In some embodiments, the plurality of code word tables includes a first code word table and a second code word table; The first code word table and the second code word table are 19 and table As shown in 20, The magnitude of the offsets in the first and second MVR offset sets is Binarization It is intended to .
[0142] As shown in Tables 19 and 20, the first default offset magnitude (e.g., 1 / 4 pixel) indicates a distance of 1 / 4 pixel from the video block, the second default offset magnitude (e.g., 1 / 2 pixel) indicates a distance of 1 / 2 pixel from the video block, the third default offset magnitude (e.g., 1 pixel) indicates a distance of 1 pixel from the video block, the fourth default offset magnitude (e.g., 2 pixels) indicates a distance of 2 pixels from the video block, the fifth default offset magnitude (e.g., 4 pixels) indicates a distance of 4 pixels from the video block, the sixth default offset magnitude (e.g., 8 pixels) indicates a distance of 8 pixels from the video block, the seventh default offset magnitude (e.g., 16 pixels) indicates a distance of 16 pixels from the video block, and the eighth default offset magnitude (e.g., 32 pixels) indicates a distance of 32 pixels from the video block.
[0143] Furthermore, as shown in Tables 19 and 20, the magnitude of the first alternate offset (e.g., 1 / 4 pixel) indicates a distance of 1 / 4 pixel from the video block, the magnitude of the second alternate offset (e.g., 1 / 2 pixel) indicates a distance of 1 / 2 pixel from the video block, the magnitude of the third alternate offset (e.g., 1 pixel) indicates a distance of 1 pixel from the video block, the magnitude of the fourth alternate offset (e.g., 2 pixels) indicates a distance of 2 pixels from the video block, the magnitude of the fifth alternate offset (e.g., 3 pixels) indicates a distance of 3 pixels from the video block, the magnitude of the sixth alternate offset (e.g., 4 pixels) indicates a distance of 4 pixels from the video block, the magnitude of the seventh alternate offset (e.g., 6 pixels) indicates a distance of 6 pixels from the video block, the magnitude of the eighth alternate offset (e.g., 8 pixels) indicates a distance of 8 pixels from the video block, and the magnitude of the ninth alternate offset (e.g., 16 pixels) indicates a distance of 16 pixels from the video block.
[0144] In some embodiments, processor 920 may further determine, in response to determining that the control variable is equal to 0 and the indicator variable is equal to 0, to apply a first set of MVR offsets and binarize the magnitudes of the plurality of default offsets using a first code word table. As shown in Table 19, the magnitude of the first default offset is binarized as 1, the magnitude of the second default offset is binarized as 10, the magnitude of the third default offset is binarized as 110, the magnitude of the fourth default offset is binarized as 1110, the magnitude of the fifth default offset is binarized as 11110, the magnitude of the sixth default offset is binarized as 111110, the magnitude of the seventh default offset is binarized as 1111110, and the magnitude of the eighth default offset is binarized as 1111111.
[0145] In some embodiments, processor 920 may further determine, in response to determining that the control variable is equal to 0 and the indicator variable is equal to 1, to apply a first set of MVR offsets and binarize the magnitudes of the plurality of default offsets using a second code word table. As shown in Table 20, the magnitude of the first default offset is binarized as 111110, the magnitude of the second default offset is binarized as 1, the magnitude of the third default offset is binarized as 10, the magnitude of the fourth default offset is binarized as 110, the magnitude of the fifth default offset is binarized as 1110, the magnitude of the sixth default offset is binarized as 11110, the magnitude of the seventh default offset is binarized as 1111110, and the magnitude of the eighth default offset is binarized as 1111111.
[0146] In some embodiments, processor 920 may further determine, in response to determining that the control variable is equal to 1 and the indicator variable is equal to 0, to apply a second set of MVR offsets and binarize the magnitudes of the plurality of alternate offsets using a first code word table. As shown in Table 19, the magnitude of the first alternate offset is binarized as 1, the magnitude of the second alternate offset is binarized as 10, the magnitude of the third alternate offset is binarized as 110, the magnitude of the fourth alternate offset is binarized as 1110, the magnitude of the fifth alternate offset is binarized as 11110, the magnitude of the sixth alternate offset is binarized as 111110, the magnitude of the seventh alternate offset is binarized as 1111110, the magnitude of the eighth alternate offset is binarized as 11111110, and the magnitude of the ninth alternate offset is binarized as 1111111.
[0147] In some embodiments, processor 920 may further determine, in response to determining that the control variable is equal to 1 and the indicator variable is equal to 1, to apply a second set of MVR offsets and binarize the magnitudes of the plurality of alternate offsets using a second code word table. As shown in Table 20, the magnitude of the first alternate offset is binarized as 111110, the magnitude of the second alternate offset is binarized as 1, the magnitude of the third alternate offset is binarized as 10, the magnitude of the fourth alternate offset is binarized as 110, the magnitude of the fifth alternate offset is binarized as 1110, the magnitude of the sixth alternate offset is binarized as 11110, the magnitude of the seventh alternate offset is binarized as 1111110, the magnitude of the eighth alternate offset is binarized as 11111110, and the magnitude of the ninth alternate offset is binarized as 11111111.
[0148] In some embodiments, the processor may further binarize the directions of the offsets in the first and second sets of MVR offsets using fixed length codewords, respectively.
[0149] In some embodiments, processor 920 further receives a first geometric partition enable syntax element (e.g., gpm_mvr_partIdx0_enable_flag) indicating whether MVR is applied to the first geometric partition, and in response to determining that the geometric partition enable syntax element is equal to 1, determines a first MVR offset of the first geometric partition determined based on the set of selected MVR offsets. offset Direction and offset receiving a first direction syntax element (e.g., gpm_mvr_partIdx0_direction_idx) and a first magnitude syntax element (e.g., gpm_mvr_partIdx0_distance_idx) indicating a magnitude, receiving a second geometric partition enable syntax element (e.g., gpm_mvr_partIdx1_enable_flag) indicating whether MVR is applied to the second geometric partition, and determining that the second geometric partition enable syntax element is equal to 1, offset Direction and offset A second direction syntax element (eg, gpm_mvr_partIdx1_direction_idx) and a second magnitude syntax element (eg, gpm_mvr_partIdx1_distance_idx) indicating the magnitude may be received.
[0150] In step 806, processor 920 may obtain a first MV and a second MV from the candidate lists for the first and second geometric partitions.
[0151] In step 807, processor 920 may calculate a first refine MV and a second refine MV based on the first and second MVs and the first and second MVR offsets.
[0152] At step 808, processor 920 may obtain a predicted sample for the video block based on the first and second refined MVs.
[0153] In some embodiments, an apparatus for decoding video blocks in a GPM is provided, the apparatus including a processor 920 and a memory 940 storing instructions executable by the processor, the processor configured, upon execution of the instructions, to perform the method illustrated in FIG.
[0154] In some other embodiments, a non-transitory computer-readable storage medium is provided that stores instructions that, when executed by the processor 920, cause the processor to perform the method illustrated in FIG.
[0155] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the present disclosure disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure in accordance with its general principles, including departures from the present disclosure within known or customary practice in the art. It is intended that the specification and examples be considered as illustrative only.
[0156] It will be understood that the present disclosure is not limited to the exact embodiments described above and illustrated by the accompanying drawings, and that various modifications and changes can be made without departing from the scope of the present disclosure.
Claims
1. receiving a control variable to be applied at a coding level that enables adaptive switching between a set of multiple motion vector refinement (MVR) offsets associated with the video block; receiving an indicator variable associated with the video block that enables adaptive switching between a plurality of codeword tables that binarize magnitudes of offsets in the set of MVR offsets under the coding level; Partitioning the video block into a first geometric partition and a second geometric partition; selecting a set of MVR offsets from the plurality of sets of MVR offsets based on the control variable; receiving one or more syntax elements to determine first and second MVR offsets from the selected set of MVR offsets to be applied to the first and second geometric partitions; obtaining a first motion vector (MV) and a second motion vector (MV) from the candidate list of the first geometric partition and the candidate list of the second geometric partition; calculating a first refine MV and a second refine MV based on the first and second MVs and the first and second MVR offsets; obtaining a prediction sample for the video block based on the first and second refined MVs; and 1. A method for decoding a video block in geometric partition mode (GPM), comprising:
2. The method of claim 1 , wherein the coding level comprises a sequence level, a picture level, a coding tree unit level, or a coding block level.
3. the plurality of sets of MVR offsets includes a first set of MVR offsets and a second set of MVR offsets; the second set of MVR offsets includes a magnitude of at least one of the offsets in the first set of MVR offsets; The method of claim 1 , wherein the plurality of code word tables includes a first code word table and a second code word table.
4. determining to apply the first set of MVR offsets in response to determining that the control variable is equal to zero; determining to apply the second set of MVR offsets in response to determining that the control variable is equal to one; The method of claim 3 further comprising:
5. determining to apply the first code table in response to determining that the indicator variable is equal to 0; determining to apply the second code table in response to determining that the indicator variable is equal to 1; The method of claim 4 further comprising:
6. The first set of MVR offsets includes a plurality of default offset magnitudes, the plurality of default offset magnitudes being: a first default offset magnitude indicating a quarter pixel distance from the video block; a second default offset magnitude indicating a half pixel distance from the video block; a third default offset magnitude indicating a distance of one pixel from the video block; a fourth default offset magnitude indicating a distance of two pixels from the video block; a fifth default offset magnitude indicating a distance of four pixels from the video block; a sixth default offset magnitude indicating a distance of 8 pixels from the video block; a seventh default offset magnitude indicating a distance of 16 pixels from the video block; an eighth default offset magnitude indicating a distance of 32 pixels from the video block; The second set of MVR offsets includes a plurality of alternative offset magnitudes, the plurality of alternative offset magnitudes being: a first alternate offset magnitude indicating a quarter pixel distance from the video block; a second alternate offset magnitude indicating a half pixel distance from the video block; a third alternative offset magnitude indicating a distance of one pixel from the video block; a fourth alternative offset magnitude indicating a distance of two pixels from the video block; a fifth alternate offset magnitude indicating a distance of three pixels from said video block; a sixth alternate offset magnitude indicating a distance of four pixels from the video block; a seventh alternate offset magnitude indicating a distance of 6 pixels from said video block; an eighth alternate offset magnitude indicating a distance of eight pixels from said video block; a ninth alternative offset magnitude indicating a distance of 16 pixels from said video block; The method of claim 5 , comprising:
7. determining, in response to determining that the control variable is equal to 0 and the indicator variable is equal to 0, to apply the first set of MVR offsets and to binarize magnitudes of the plurality of default offsets using the first code table; binarizing the magnitudes of the plurality of default offsets using the first code table binarizing the first default offset with a magnitude of 1; binarizing the second default offset with a magnitude of 10; binarizing the third default offset to a magnitude of 110; binarizing the magnitude of the fourth default offset as 1110; binarizing the magnitude of the fifth default offset as 11110; binarizing the magnitude of the sixth default offset as 111110; binarizing the magnitude of the seventh default offset as 1111110; binarizing the magnitude of the eighth default offset as 1111111; The method of claim 6, comprising:
8. and determining, in response to determining that the control variable is equal to 0 and the indicator variable is equal to 1, to apply the first set of MVR offsets and to binarize magnitudes of the plurality of default offsets using the second code word table; binarizing the magnitudes of the plurality of default offsets using the second code table binarizing the magnitude of the first default offset as 111110; binarizing the second default offset with a magnitude of 1; binarizing the third default offset with a magnitude of 10; binarizing the fourth default offset to a magnitude of 110; binarizing the magnitude of the fifth default offset as 1110; binarizing the magnitude of the sixth default offset as 11110; binarizing the magnitude of the seventh default offset as 1111110; binarizing the magnitude of the eighth default offset as 11111111; The method of claim 6, comprising:
9. and applying the second set of MVR offsets and binarizing magnitudes of the plurality of alternative offsets using the first code word table in response to determining that the control variable is equal to one and the indicator variable is equal to zero. binarizing the magnitudes of the plurality of alternative offsets using the first code table binarizing the first alternative offset with a magnitude of 1; binarizing the second alternative offset to a magnitude of 10; binarizing the magnitude of the third alternative offset as 110; binarizing the magnitude of the fourth alternative offset as 1110; binarizing the magnitude of the fifth alternative offset as 11110; binarizing the magnitude of the sixth alternative offset as 111110; binarizing the magnitude of the seventh alternative offset as 1111110; binarizing the magnitude of the eighth alternative offset as 11111110; binarizing the magnitude of the ninth alternative offset as 11111111; The method of claim 6, comprising:
10. and applying the second set of MVR offsets and binarizing magnitudes of the plurality of alternative offsets using the second code table in response to determining that the control variable is equal to one and the indicator variable is equal to one. binarizing the magnitudes of the plurality of alternative offsets using the second code table binarizing the magnitude of the first alternative offset as 111110; binarizing the second alternative offset with a magnitude of 1; binarizing the third alternative offset to a magnitude of 10; binarizing the magnitude of the fourth alternative offset as 110; binarizing the magnitude of the fifth alternative offset as 1110; binarizing the magnitude of the sixth alternative offset as 11110; binarizing the magnitude of the seventh alternative offset as 1111110; binarizing the magnitude of the eighth alternative offset as 11111110; binarizing the magnitude of the ninth alternative offset as 11111111; The method of claim 6, comprising:
11. the plurality of sets of MVR offsets includes a first set of MVR offsets and a second set of MVR offsets; the second set of MVR offsets includes a direction of at least one offset of the first set of MVR offsets; The method of claim 1 , wherein the directions of the offsets in the first and second sets of MVR offsets are each binarized using a fixed-length codeword.
12. Receiving one or more syntax elements and determining first and second MVR offsets to apply to the first and second geometric partitions from the selected set of MVR offsets includes: receiving a first geometric partition validity syntax element indicating whether the MVR applies to the first geometric partition; receiving a first direction syntax element and a first magnitude syntax element indicating an offset direction and an offset magnitude of the first MVR offset of the first geometric partition determined based on the selected set of MVR offsets in response to determining that the first geometric partition valid syntax element is equal to 1; receiving a second geometric partition validity syntax element indicating whether the MVR applies to the second geometric partition; receiving a second direction syntax element and a second magnitude syntax element indicating an offset direction and an offset magnitude of the second MVR offset of the second geometric partition determined based on the selected set of MVR offsets in response to determining that the second geometric partition valid syntax element is equal to 1; The method of claim 1 , comprising:
13. The first geometric partition enable syntax element includes: gpm_mvr_partIdx0_enable_flag; The first direction syntax element and the first magnitude syntax element include gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx0_distance_idx, respectively; The second geometric partition enable syntax element includes: gpm_mvr_partIdx1_enable_flag; The method of claim 12 , wherein the second direction syntax element and the second magnitude syntax element comprise gpm_mvr_partIdx1_direction_idx and gpm_mvr_partIdx1_distance_idx, respectively.
14. one or more processors; a memory configured to store instructions executable by the one or more processors; Video decoding apparatus, wherein the one or more processors are configured to perform the method of any one of claims 1 to 13 upon execution of the instructions.
15. A non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method of any one of claims 1 to 13.
16. performing an encoding method to generate a bitstream; storing the bitstream; 1. A method for storing a bitstream, comprising: The encoding method comprises: determining a control variable to be applied at a coding level that enables adaptive switching between a set of multiple motion vector refinement (MVR) offsets associated with the video block; determining an indicator variable associated with the video block that enables adaptive switching between a plurality of codeword tables that binarize magnitudes of offsets in the set of MVR offsets under the coding level; Partitioning the video block into a first geometric partition and a second geometric partition; selecting a set of MVR offsets from the plurality of sets of MVR offsets based on the control variable; determining a first MVR offset and a second MVR offset to be applied to the first and second geometric partitions from the selected set of MVR offsets; obtaining a first motion vector (MV) and a second motion vector (MV) from the candidate list of the first geometric partition and the candidate list of the second geometric partition; calculating a first refine MV and a second refine MV based on the first and second MVs and the first and second MVR offsets; obtaining a prediction sample for the video block based on the first and second refined MVs; and A method comprising:
17. A computer program comprising instructions stored therein, which, when executed by a processor, cause the processor to perform a method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Extensions of inter prediction with geometric partitioning
WO2020094049A1
Method for encoding / decoding image signal, and device for same
WO2020096426A1
An encoder, a decoder and corresponding methods for merge mode
WO2020106189A1