Method and device for geometric partition mode with motion vector refinement - Patents.com
Patent Information
- Application Number
- JP2023580626
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-06-28
- Filing Date
- 2022-06-28
- Publication Date
- 2025-07-08
AI Technical Summary
Existing video encoding techniques, such as those in VVC and AVS3, face inefficiencies in geometric partition mode (GPM) due to inaccurate unidirectional motion vectors and increased signaling overhead, particularly when applied to non-merge inter-CUs, which can hinder optimal coding efficiency.
The proposed method, Geometric Partition Mode with Motion Vector Refinement (GPM-MVR), enhances GPM by applying motion refinement on top of existing unidirectional motion vectors using predefined motion vector differences and directions, allowing for more accurate motion estimation while minimizing signaling overhead through optimized syntax element signaling.
GPM-MVR improves coding efficiency by providing more accurate motion vectors for geometric partitions, reducing signaling overhead and enhancing compression performance in video encoding, particularly for non-merge inter-CUs.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is based on and claims priority to U.S. Provisional Patent Application No. 63 / 215,957, filed June 28, 2021, the disclosure of which is incorporated by reference in its entirety and for all purposes.
[0002] FIELD OF THE DISCLOSURE This disclosure relates to video encoding and compression. More particularly, this disclosure relates to a method and apparatus for improving coding efficiency of Geometric Partitioning (GPM) mode, also known as Angle Weighted Prediction (AWP) mode. [Background technology]
[0003] Various video encoding techniques may be used to compress video data. Video encoding is performed according to one or more video encoding standards. For example, today, some well-known video encoding standards include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC, also known as H.265 or MPEG-H Part2), and Advanced Video Coding (AVC, also known as H.264 or MPEG-4 Part 10), which are jointly developed by ISO / IEC MPEG and ITU-T VECG. AOMedia Video 1 (AV1) was developed by the Alliance for Open Media (AOM) as a successor to the predecessor standard VP9. Audio Video Coding (AVS), which refers to digital audio and digital video compression standards, is another series of video compression standards developed by the Audio and Video Coding Standard Workgroup of China. Most of the existing video coding standards are built on the well-known hybrid video coding framework, i.e., using block-based prediction methods (e.g., inter-prediction, intra-prediction) to reduce the redundancy present in a video image or sequence, and using transform coding to compress the energy of the prediction error. An important goal of video coding techniques is to compress video data into a format that uses a lower bit rate while avoiding or minimizing the degradation of video quality. Summary of the Invention [Problem to be solved by the invention]
[0004] The present disclosure provides methods and apparatus, and non-transitory computer-readable storage media, for video encoding. [Means for solving the problem]
[0005] According to a first aspect of the present disclosure, a method for decoding a video block with a GPM is provided. The method can include partitioning the video block into first and second geometric partitions. The method can include constructing a unidirectional motion vector (MV) candidate list of the GPM by adding a plurality of regular merge candidates. The method can include constructing a first updated unidirectional MV candidate list by adding one or more additional unidirectional MVs derived from one or more bi-predictive MVs of the regular merge candidate list to the unidirectional MV candidate list in response to determining that the unidirectional MV candidate list is not complete.
[0006] The method may include constructing a second updated unidirectional MV candidate list by adding one or more pair average candidates to the first updated unidirectional MV candidate list in response to determining that the first updated unidirectional MV candidate list is not complete. The method may also include periodically adding zero unidirectional MVs to the second updated unidirectional MV candidate list until a maximum length is reached in response to determining that the second updated unidirectional MV candidate list is not complete. The method may further include generating a unidirectional MV for the first geometric partition and a unidirectional MV for the second geometric partition.
[0007] According to a second aspect of the present disclosure, an apparatus for video decoding is provided. The apparatus may include one or more processors and a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium is configured to store instructions executable by the one or more processors. The one or more processors are configured to perform the method of the first aspect when executing the instructions.
[0008] According to a third aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium, which can store computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method of the first aspect.
[0009] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure. [Brief description of the drawings]
[0010] [Figure 1] FIG. 2 is a block diagram of an encoder according to an example of the present disclosure. [Diagram 2] FIG. 2 is a block diagram of a decoder according to an example of the present disclosure. [Figure 3A] FIG. 13 is a diagram showing block partitions in a multi-type tree structure according to an example of the present disclosure. [Figure 3B] FIG. 13 is a diagram showing block partitions in a multi-type tree structure according to an example of the present disclosure. [Figure 3C] FIG. 13 is a diagram showing block partitions in a multi-type tree structure according to an example of the present disclosure. [Figure 3D] FIG. 13 is a diagram showing block partitions in a multi-type tree structure according to an example of the present disclosure. [Figure 3E] FIG. 13 is a diagram showing block partitions in a multi-type tree structure according to an example of the present disclosure. [Figure 4] FIG. 1 is a diagram of a permitted geometric partition (GPM) according to an example of the present disclosure. [Diagram 5] 11 is a table illustrating motion vector selection by unidirectional prediction according to an example of the present disclosure. [Figure 6A] FIG. 2 is a diagram of a motion vector differential (MMVD) mode according to an example of the present disclosure. [Figure 6B] FIG. 2 is a diagram of an MMVD mode according to an example of the present disclosure. [Figure 7] FIG. 1 is a diagram of a template matching (TM) algorithm according to an example of the present disclosure. [Figure 8] FIG. 2 is a diagram of a method for decoding a video block with a GPM according to an example of the present disclosure. [Figure 9] FIG. 1 illustrates a computing environment coupled to a user interface according to an example of the present disclosure. [Figure 10] FIG. 1 is a block diagram illustrating a system for encoding and decoding video blocks according to some examples of this disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0011] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which the same numerals in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following description of the embodiments do not represent all implementations consistent with the present disclosure. Instead, these implementations are merely examples of apparatus and methods consistent with aspects related to the present disclosure as set forth in the appended claims.
[0012] The terms used in this disclosure are intended only to describe certain embodiments and are not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. It will also be understood that the term "and / or" as used herein is intended to mean and encompass any or all possible combinations of one or more of the associated listed items.
[0013] While terms such as "first," "second," and "third" may be used herein to describe various pieces of information, it will be understood that the information should not be limited by these terms. These terms are used only to distinguish one category of information from another category of information. For example, a first piece of information may be referred to as a second piece of information, and similarly, a second piece of information may be referred to as a first piece of information, without departing from the scope of this disclosure. As used herein, the term "if" will be understood to mean "when" or "upon," or "in response to a judgment," depending on the context.
[0014] The first generation of AVS standards includes the Chinese national standard "Information Technology, Advanced Audio Video Coding, Part 2: Video" (also known as AVS1) and "Information Technology, Advanced Audio Video Coding Part 16: Radio Television Video" (also known as AVS+). It can provide about 50% bitrate savings with the same perceptual quality compared to the MPEG-2 standard. The video part of the AVS1 standard was promulgated as a Chinese national standard in February 2006. The second generation of AVS standards includes a series of Chinese national standards "Information Technology, Efficient Multimedia Coding" (also known as AVS2), which is mainly targeted at the transmission of additional HD TV programs. The coding efficiency of AVS2 is twice that of AVS+. In May 2016, AVS2 was issued as a Chinese national standard. Meanwhile, the video part of the AVS2 standard was submitted by the Institute of Electrical and Electronics Engineers (IEEE) as an international standard for application. The AVS3 standard is a new generation video coding standard for UHD video applications that aims to exceed the coding efficiency of the latest international standard HEVC. In March 2019, at the 68th AVS Conference, the AVS3-P2 baseline was completed, which provides about 30% bitrate savings compared to the HEVC standard. Currently, a reference software called the High Performance Model (HPM) is maintained by the AVS group to demonstrate the reference implementation of the AVS3 standard.
[0015] Like HEVC, the AVS3 standard is built on a block-based hybrid video coding framework.
[0016] 10 is a block diagram illustrating an example system 10 for encoding and decoding video blocks in parallel according to some implementations of the present disclosure. As shown in FIG. 1, the system 10 includes a source device 12 that generates video data and encodes it to be later decoded by a destination device 14. The source device 12 and the destination device 14 can comprise any of a wide variety of electronic devices, including desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, and the like. In some implementations, the source device 12 and the destination device 14 include wireless communication capabilities.
[0017] In some implementations, the destination device 14 may receive the encoded video data to be decoded via a link 16. The link 16 may comprise any type of communication medium or device capable of moving the encoded video data from the source device 12 to the destination device 14. In one example, the link 16 may comprise a communication medium to enable the source device 12 to transmit the encoded video data directly to the destination device 14 in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to the destination device 14. The communication medium may include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful in facilitating communication from the source device 12 to the destination device 14.
[0018] In some other implementations, the encoded video data may be transmitted from the output interface 22 to the storage device 32. The encoded video data in the storage device 32 may then be accessed by the destination device 14 via the input interface 28. The storage device 32 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data. In a further example, the storage device 32 may correspond to a file server or another intermediate storage device that may hold the encoded video data generated by the source device 12. The destination device 14 may access the stored video data from the storage device 32 via streaming or download. The file server may be any type of computer capable of storing the encoded video data and transmitting the encoded video data to the destination device 14. Exemplary file servers include web servers (e.g., for websites), file transfer protocol (FTP) servers, network attached storage (NAS) devices, or local disk drives. Destination device 14 can access the encoded video data over any standard data connection, including a wireless channel (e.g., a Wireless Fidelity (Wi-Fi) connection), a wired connection (e.g., a Digital Subscriber Line (DSL), cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from storage device 32 can be a streaming transmission, a download transmission, or a combination of both.
[0019] As shown in FIG. 10, source device 12 includes a video source 18, a video encoder 20, and an output interface 22. Video source 18 may include sources such as a video capture device, e.g., a video camera, a video archive containing previously captured video, a video supply interface receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as the source video, or a combination of such sources. As an example, if video source 18 is a video camera of a security surveillance system, source device 12 and destination device 14 may form a camera phone or a video phone. However, implementations described herein may be applicable to video encoding in general, and may be applied to wireless and / or wired applications.
[0020] Captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video data may be transmitted directly to destination device 14 via output interface 22 of source device 12. The encoded video data may also (or alternatively) be stored on storage device 32 for later access for decoding and / or playback by destination device 14 or other devices. Output interface 22 may further include a modem and / or a transmitter.
[0021] Destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 may include a receiver and / or modem and may receive encoded video data over link 16. The encoded video data communicated over link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements may be included within the encoded video data that is transmitted over a communication medium, stored on a storage medium, or stored on a file server.
[0022] In some implementations, destination device 14 may include a display device 34, which may be an integrated display device and an external display device configured to communicate with destination device 14. Display device 34 displays the decoded video data to a user and may comprise any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0023] Video encoder 20 and video decoder 30 may operate according to a proprietary or industry standard, such as VVC, HEVC, MPEG-4, Part 10, AVC, or an extension of such a standard. It should be understood that the present application is not limited to a particular video encoding / decoding standard and may be applicable to other video encoding / decoding standards. It is generally contemplated that video encoder 20 of source device 12 may be configured to encode video data according to any of these current or future standards. Similarly, it is generally contemplated that video decoder 30 of destination device 14 may be configured to decode video data according to any of these current or future standards.
[0024] Video encoder 20 and video decoder 30 may each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When implemented partially in software, an electronic device may store instructions for the software in a suitable non-transitory computer-readable medium and execute those instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in this disclosure. Each of video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) in the respective device.
[0025] FIG 1 shows a schematic diagram of a block-based video encoder for VVC. Specifically, FIG 1 shows an exemplary encoder 100. The encoder 100 may be the video encoder 20 shown in FIG 10. The encoder 100 includes a video input 110, a motion compensation 112, a motion estimation 114, an intra / inter mode decision 116, a block predictor 140, a summing unit 128, a transform 130, a quantization 132, a prediction related information 142, an intra prediction 118, a picture buffer 120, an inverse quantization 134, an inverse transform 136, a summing unit 126, a memory 124, an in-loop filter 122, an entropy coding 138, and a bitstream 144.
[0026] Within encoder 100, a video frame is partitioned into video blocks for processing. For each given video block, a prediction is formed based on either inter-prediction or intra-prediction techniques.
[0027] A prediction residual, which represents the difference between the current video block and a portion of the video input 110 and its predictor and a portion of the block predictor 140, is sent from the adder 128 to the transform 130. The transform coefficients are then sent from the transform 130 to the quantizer 132 for entropy reduction. The quantized coefficients are then sent to the entropy encoder 138 to generate a compressed video bitstream. As shown in FIG. 1, prediction related information 142 from the intra / inter mode decision 116, such as video block partition information, motion vectors (MVs), reference picture indexes, and intra prediction modes, are also sent by the entropy encoder 138 and stored in the compressed bitstream 144. The compressed bitstream 144 comprises the video bitstream.
[0028] Within the encoder 100, decoder-related circuitry is also required to reconstruct pixels for prediction purposes. First, a prediction residual is reconstructed by inverse quantization 134 and inverse transform 136. This reconstructed prediction residual is combined with a block predictor 140 to generate an unfiltered reconstructed pixel for the current video block.
[0029] Spatial prediction (or "intra prediction") predicts the current video block using pixels from samples of already coded neighboring blocks (called reference samples) in the same video frame as the current video block.
[0030] Temporal prediction (also called "inter prediction") uses reconstructed pixels from already coded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in video signals. Typically, a temporal prediction signal for a given coding unit (CU) or coding block is signaled by one or more MVs that indicate the amount and direction of motion between the current CU and its temporal reference. Additionally, if multiple reference pictures are supported, one reference picture index is additionally transmitted and used to identify which reference picture in the reference picture store the temporal prediction signal originated from.
[0031] Motion estimation 114 takes signals from video input 110 and picture buffer 120 and outputs a motion estimation signal to motion compensation 112. Motion compensation 112 takes signals from video input 110, picture buffer 120, and a motion estimation signal from motion estimation 114 and outputs a motion compensation signal to intra / inter mode decision 116.
[0032] After spatial and / or temporal prediction is performed, an intra / inter mode decision 116 in the encoder 100 chooses the best prediction mode, for example based on a rate-distortion optimization method. A block predictor 140 is then subtracted from the current video block, and the resulting prediction residual is decorrelated using a transform 130 and a quantization 132. The resulting quantized residual coefficients are inverse quantized by an inverse quantization 134 and inverse transformed by an inverse transform 136 to form a reconstructed residual, which is then added back to the predictive block to form a reconstructed signal for the CU. Further, in-loop filtering 122, such as a deblocking filter, sample adaptive offset (SAO), and / or an adaptive in-loop filter (ALF), may be applied to the reconstructed CU, which is then placed in a reference picture store of a picture buffer 120 and used to code future video blocks. To form the output video bitstream 144, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy encoding unit 138, which further compresses and packs them to form the bitstream.
[0033] Figure 1 gives a block diagram of a generic block-based hybrid video coding system. The input video signal is processed by blocks (called coding units (CUs)). Unlike HEVC, which partitions blocks based only on quadtrees, in AVS3, one coding tree unit (CTU) is partitioned into CUs to adapt to varying local characteristics based on quadtrees / binary trees / extended quadtrees. In addition, the concept of multiple partition unit types in HEVC is removed, i.e., there is no separation of CUs, prediction units (PUs), and transform units (TUs) in AVS3, instead each CU is always used as a basic unit for both prediction and transformation without further partitioning. In the tree partition structure of AVS3, one CTU is first partitioned based on a quadtree structure. Then each quadtree leaf node may be further partitioned based on binary tree and extended quadtree structures.
[0034] As shown in Figures 3A, 3B, 3C, 3D, and 3E, there are five partition types, namely, 4-partition, horizontal 2-partition, vertical 2-partition, horizontal extended quadtree partition, and vertical extended quadtree partition.
[0035] FIG. 3A shows a diagram illustrating four block partitions in a multi-type tree structure according to the present disclosure.
[0036] FIG. 3B shows a diagram illustrating two vertical partitions of blocks in a multi-type tree structure according to the present disclosure.
[0037] FIG. 3C shows a diagram illustrating two horizontal partitions of blocks in a multi-type tree structure according to the present disclosure.
[0038] FIG. 3D shows a diagram illustrating three vertical partitions of blocks in a multi-type tree structure according to the present disclosure.
[0039] FIG. 3E illustrates a diagram showing three horizontal partitions of blocks in a multi-type tree structure according to the present disclosure.
[0040] In FIG. 1, spatial prediction and / or temporal prediction may be performed. Spatial prediction (or "intra prediction") predicts a current video block using pixels from samples of already coded neighboring blocks (called reference samples) in the same video picture / slice. Spatial prediction reduces spatial redundancy inherent in video signals. Temporal prediction (also called "inter prediction" or "motion compensated prediction") predicts a current video block using reconstructed pixels from already coded video pictures. Temporal prediction reduces temporal redundancy inherent in video signals. Typically, a temporal prediction signal for a given CU is signaled by one or more motion vectors (MVs) that indicate the amount and direction of motion between the current CU and its temporal reference. Also, if multiple reference pictures are supported, one reference picture index is further transmitted and used to identify which reference picture in the reference picture store the temporal prediction signal originated from. After spatial and / or temporal prediction, a mode decision block in the encoder chooses the best prediction mode, for example based on a rate-distortion optimization method. The prediction block is then subtracted from the current video block, and the prediction residual is decorrelated using a transform and then quantized. The quantized residual coefficients are inverse quantized and inverse transformed to form a reconstructed residual, which is then added back to the prediction block to form a reconstructed signal for the CU. Furthermore, in-loop filtering, such as a deblocking filter, sample adaptive offset (SAO), and adaptive in-loop filter (ALF), may be applied to the reconstructed CU, which is then placed in a reference picture store and used as a reference for coding future video blocks. To form an output video bitstream, the coding mode (inter or intra), prediction mode information, motion information, and the quantized residual coefficients are all sent to an entropy coding unit for further compression and packing.
[0041] FIG 2 shows a schematic block diagram of a video decoder for VVC. Specifically, FIG 2 shows a block diagram of an exemplary decoder 200. The block-based video decoder 200 may be the video decoder 30 shown in FIG 10. The decoder 200 includes a bitstream 210, an entropy decoding 212, an inverse quantization 214, an inverse transform 216, an adder 218, an intra / inter mode selection 220, an intra prediction 222, a memory 230, an in-loop filter 228, a motion compensation 224, a picture buffer 226, prediction related information 234, and a video output 232.
[0042] The decoder 200 is similar to the reconstruction-related parts present in the encoder 100 of FIG. 1. First, in the decoder 200, an incoming video bitstream 210 is decoded by an entropy decoding 212 to derive quantized coefficient levels and prediction-related information. The quantized coefficient levels are then processed by an inverse quantization 214 and an inverse transform 216 to obtain a reconstructed prediction residual. A block prediction mechanism is implemented in an intra / inter mode selection unit 220 and configured to perform intra prediction 222 or motion compensation 224 based on the decoded prediction information. A set of unfiltered reconstructed pixels is obtained by summing the reconstructed prediction residual from the inverse transform 216 and the prediction output generated by the block prediction mechanism using a summation unit 218.
[0043] The reconstructed blocks may further pass through an in-loop filter 228 and then be stored in a picture buffer 226, which serves as a reference picture store. The reconstructed video in the picture buffer 226 may be transmitted to drive a display device, as well as used to predict future video blocks. In situations where the in-loop filter 228 is turned on, a filtering operation is performed on these reconstructed pixels to derive the final reconstructed video output 232.
[0044] FIG. 2 gives a schematic block diagram of a block-based video decoder. First, the video bitstream is entropy decoded in an entropy decoding unit. The coding mode and prediction information are sent to a spatial prediction unit (if intra-coded) or a temporal prediction unit (if inter-coded) to form a prediction block. The residual transform coefficients are sent to an inverse quantization unit and an inverse transform unit to reconstruct the residual block. The prediction block and the residual block are then added together. The reconstructed block may further pass through an in-loop filter and then stored in a reference picture store. The reconstructed video in the reference picture store is then sent for display as well as used to predict future video blocks.
[0045] The focus of this disclosure is to improve the coding performance of the Geometric Partition Mode (GPM) used in both the VVC and AVS3 standards. In AVS3, the tool is also known as Angle Weighted Prediction (AWP), which follows the same design spirit as GPM, but with some slight differences in certain design details. To facilitate the description of this disclosure, hereinafter, the existing GPM design in the VVC standard is used as an example to explain the main aspects of the GPM / AWP tool. Meanwhile, another existing inter-prediction technique called merge mode with motion vector differential (MMVD), which is applied in both the VVC and AVS3 standards, is also briefly considered, provided that it is closely related to the technology proposed in this disclosure. After that, some shortcomings of the current GPM / AWP design are identified. Finally, the proposed method is provided in detail. Throughout this disclosure, the existing GPM design in the VVC standard is used as an example, but those skilled in the art of modern video coding technology should note that the proposed technology may also be applied to other GPM / AWP designs or other coding tools with the same or similar design spirit.
[0046] Geometric Partition Mode (GPM) In VVC, geometric partitioning modes are supported for inter prediction. The geometric partitioning modes are signaled by one CU level flag as one special merge mode. In the current GPM design, a total of 64 partitions are supported by GPM modes for each possible CU size where both width and height are 8 or more and 64 or less, except for 8x64 and 64x8.
[0047] When this mode is used, the CU is divided into two parts by a straight line that is geometrically located as shown in Figure 4 (explained below). The location of the dividing line is mathematically derived from the angle and offset parameters of the specific partition. Each part of the geometric partition within the CU is inter predicted using its own motion, and only unidirectional prediction is allowed for each partition, i.e., each part has one motion vector and one reference index. As with conventional bidirectional prediction, the motion constraint with unidirectional prediction is applied to ensure that only two motion compensated predictions are required for each CU. If the geometric partitioning mode is used for the current CU, the geometric partition index indicating the partition mode (angle and offset) of the geometric partition, and two merge indices (one for each partition) are further signaled. The number of maximum GPM candidate sizes is explicitly signaled at the sequence level.
[0048] FIG. 4 shows the allowed GPM partitions, where the partitions within each picture have one and the same partition direction.
[0049] Unidirectional prediction candidate list structure To derive a unidirectional prediction motion vector for one geometric partition, first, one unidirectional prediction candidate list is directly derived from the regular merge candidate list generation process. Let n be the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector of the nth merge candidate is used as the nth unidirectional prediction motion vector for the geometric partition mode, where X is equal to the parity of n.
[0050] These motion vectors are denoted by "x" in Figure 5 (described below). If the corresponding LX motion vector of the nth extended merge candidate does not exist, the L(1-X) motion vector of the same candidate is used instead as the unidirectional predictive motion vector for the geometric partitioning mode.
[0051] FIG. 5 shows unidirectional predictive motion vector selection from the motion vectors in the merge candidate list for the GPM.
[0052] Blending along geometry edges After each geometric partition is obtained using its unique motion, blending is applied to the two unidirectional prediction signals to derive samples around the geometric partition edges. The blending weights for each position of the CU are derived based on the distance from each individual sample position to the corresponding partition edge.
[0053] GPM Signaling Design According to the current GPM design, the use of GPM is indicated by signaling one flag at the CU level. This flag is signaled only when the current CU is coded in merge mode or skip mode. Specifically, when this flag is equal to 1, it indicates that the current CU is predicted by GPM. Otherwise (flag is equal to 0), the CU is coded by another merge mode, such as regular merge mode, merge mode with motion vector differential, a combination of inter and intra prediction, etc. When GPM is enabled for the current CU, one syntax element, namely merge_gpm_partition_idx, is further signaled to indicate the applied geometric partition mode (specifying the direction and offset of the line from the CU center that divides the CU into two partitions, as shown in Figure 4). Then, two syntax elements merge_gpm_idx0 and merge_gpm_idx1 are signaled to indicate the indexes of the unidirectional prediction merge candidates used for the first and second GPM partitions. More specifically, those two syntax elements are used to determine the unidirectional MVs of the two GPM partitions from the unidirectional prediction merge list described in the chapter "Unidirectional Prediction Merge List Structure". According to the current GPM design, the two indices cannot be the same to make the two unidirectional MVs more different. Based on such prior knowledge, first, the unidirectional prediction merge index of the first GPM partition is signaled and used as a predictor to reduce the signaling overhead of the unidirectional prediction merge index of the second GPM partition. In detail, if the second unidirectional prediction merge index is smaller than the first unidirectional prediction merge index, its original value is directly signaled. Otherwise (the second unidirectional prediction merge index is larger than the first unidirectional prediction merge index), its value is subtracted by 1 and then signaled to the bitstream. At the decoder side, the first unidirectional prediction merge index is the decoder.Then, for decoding of the second unidirectional prediction merge index, if the parsed value is less than the first unidirectional prediction merge index, the second unidirectional prediction merge index is set equal to the parsed value, otherwise (the parsed value is equal to or greater than the first unidirectional prediction merge index), the second unidirectional prediction merge index is set equal to the parsed value plus 1. Table 1 shows the existing syntax elements used for GPM mode in the current VVC specification.
[0054] [Table 1]
[0055] On the other hand, in the current GPM design, shortened unary codes are used for binarization of two unidirectional prediction merge indexes, i.e., merge_gpm_idx0 and merge_gpm_idx1. In addition, since the two unidirectional prediction merge indexes cannot be the same, different maximum values are used to shorten the codewords of the two unidirectional prediction merge indexes, and the two unidirectional prediction merge indexes are set equal to MaxGPMMergeCand-1 and MaxGPMMergeCand-2 for merge_gpm_idx0 and merge_gpm_idx1, respectively. MaxGPMMergeCand is the number of candidates in the unidirectional prediction merge list. On the decoder side, when the received value of merge_gpm_idx1 is equal to or greater than the value of merge_gpm_idx0, its value is increased by one, provided that the values of merge_gpm_idx0 and merge_gpm_idx1 cannot be the same.
[0056] When the GPM / AWP mode is applied, two different binarization methods are applied to translate the syntax merge_gpm_partition_idx into a 2-bit string. Specifically, the syntax element is binarized by a fixed-length code and a shortened binary code in the VVC and AVS3 standards, respectively. Meanwhile, for the AWP mode in AVS3, a different maximum value is used for binarizing the value of the syntax element. Specifically, in AVS3, the number of allowed GPM / AWP partition modes is 56 (i.e., the maximum value of merge_gpm_partition_idx is 55), and in VVC, the number is increased to 64 (i.e., the maximum value of merge_gpm_partition_idx is 63).
[0057] Merge Mode with Motion Vector Difference (MMVD) In addition to the conventional merge modes that derive the motion information of one current block from its spatial / temporal neighbors, the MMVD / UMVE mode is introduced as one special merge mode in both VVC and AVS standards. Specifically, in both VVC and AVS3, the mode is signaled by one MMVD flag at the coding block level. In the MMVD mode, the first two candidates in the merge list for the regular merge mode are selected as two base merge candidates for the MMVD. After one base merge candidate is selected and signaled, an additional syntax element is signaled to indicate the motion vector differential (MVD) to be added to the motion of the selected merge candidate. The MMVD syntax element includes a merge candidate flag to select the base merge candidate, a distance index to specify the magnitude of the MVD, and a direction index to indicate the direction of the MVD.
[0058] In existing MMVD designs, the distance index specifies the size of the MVD, which is defined based on a set of predefined offsets from the starting point. As shown in Figures 6A and 6B, the offsets are added to the horizontal or vertical components of the starting MV (i.e., the MV of the selected base merge candidate).
[0059] Figure 6A shows the MMVD mode for the L0 standard, and Figure 6B shows the MMVD mode for the L1 standard.
[0060] Table 2 shows the MVD offsets applied in AVS3, respectively.
[0061] [Table 2]
[0062] As shown in Table 3, the direction index is used to specify the sign of the signaled MVD. Note that the meaning of the MVD code may vary according to the starting MV. When the starting MV is a unidirectionally predicted MV or a bidirectionally predicted MV, and the MV points to two reference pictures whose POCs are both greater than or less than the current picture's POC, the signaled code is the code of the MVD added to the starting MV. When the starting MV is a bidirectionally predicted MV that points to two reference pictures, and one picture's POC is greater than the current picture and the other picture's POC is less than the current picture, the signaled code is applied to the L0 MVD, and the inverse value of the signaled code is applied to the L1 MVD.
[0063] [Table 3]
[0064] Motion Signaling for Regular Inter-Modes Similar to the HEVC standard, in addition to merge / skip modes, both VVC and AVS3 allow one inter CU to explicitly specify its motion information in the bitstream. Overall, the signaling of motion information in both VVC and AVS3 remains the same as that of the HEVC standard. Specifically, first, one inter prediction syntax, namely inter_pred_idc, is signaled to indicate whether the prediction signal is from list L0, L1, or both. For each reference list used, the corresponding reference picture is identified by signaling one reference picture index ref_idx_lx(x=0,1) for the corresponding reference list, and the corresponding MV is represented by one MVP index mvp_lx_flag(x=0,1) that is used to select the MV predictor (MVP), followed by its motion vector difference (MVD) between the target MV and the selected MVP. In addition, in the VVC standard, one control flag mvd_l1_zero_flag is signaled at the slice level. When mvd_l1_zero_flag is equal to 0, the L1 MVD is signaled in the bitstream; otherwise (mvd_l1_zero_flag flag is equal to 1), the L1 MVD is not signaled and its value is always inferred to 0 in the encoder and decoder.
[0065] Bidirectional prediction with CU-level weights In standards prior to VVC and AVS3, when weighted prediction (WP) is not applied, the bidirectional prediction signal is generated by averaging the unidirectional prediction signals obtained from two reference pictures. In VVC, one tool coding, namely Bidirectional Prediction with CU-level Weights (BCW), is introduced to improve the efficiency of bidirectional prediction. Specifically, instead of simple averaging, the bidirectional prediction of BCW is extended by allowing a weighted average of two prediction signals, as shown below. P'(i,j)=((8-w)·P0(i,j)+w·P1(i,j)+4)≫3
[0066] In VVC, when the current picture is a low-delay picture, the weight of one BCW coding block is allowed to be selected from a set of predefined weight values w ∈ {-2,3,4,5,10}, with weight 4 representing the traditional bi-prediction case where two uni-directional prediction signals are weighted equally. For low-delay, only three weights w ∈ {3,4,5} are allowed. In general, some design similarities exist between WP and BCW, but the two coding tools target solving the illumination change problem at different granularity. However, the interaction between WP and BCW may complicate the VVC design in some cases, so the two tools are not allowed to be enabled at the same time. Specifically, when WP is enabled for a slice, the BCW weights for all bi-predicted CUs in the slice are not signaled and are inferred to be 4 (i.e., equal weights are applied).
[0067] Template Matching Template Matching (TM) is a decoder-side MV derivation method for improving the motion information of the current CU by finding the best match between one template consisting of the neighboring reconstructed samples above and to the left of the current CU and a reference block (i.e., the same size as the template) in the reference picture. As shown in Fig. 7, one MV is searched around the initial motion vector of the current CU within a search range of [-8, +8] pels. The best match can be defined as the MV that achieves the lowest matching cost between the current template and the reference template, e.g., sum of absolute differences (SAD), sum of absolute transformed differences (SATD), etc. There are two different methods for applying the TM mode to inter-coding.
[0068] In AMVP mode, the MVP candidate is determined to choose the one that reaches the minimum difference between the template of the current block and the template of the reference block based on the template matching difference, and then TM is performed only on this particular MVP candidate for MV refinement. TM refines this MVP candidate from 1-pel MVD accuracy (or 4-pel for 4-pel AMVR mode) in the [-8, +8]-pel search range by using an iterative diamond search. The AMVP candidate can be further refined by using a cross search with 1-pel MVD accuracy (or 4-pel for 4-pel AMVR mode), followed by sequential 1 / 2-pel and 1 / 4-pel ones, depending on the AMVR mode specified in Table 13 below. This search process ensures that the MVP candidate still maintains the same MV accuracy as indicated by the AMVR mode after the TM process.
[0069] [Table 4]
[0070] In merge mode, a similar search method is applied to the merge candidates indicated by the merge index. As shown in the table above, the TM can go all the way up to 1 / 8-pel MVD precision, or omit after 1 / 2-pel MVD precision, depending on whether an alternative interpolation filter (used when AMVR is in 1 / 2-pel mode) is used according to the merged motion information.
[0071] As described above, the unidirectional motions used to generate the prediction samples of the two GPM partitions are obtained directly from the regular merge candidates. In the absence of strong correlation between the MVs of spatial / temporal neighboring blocks, the derived unidirectional MVs from the merge candidates may not be accurate enough to capture the true motion of each GPM partition. Motion estimation can provide more accurate motions, but at the cost of non-negligible signaling overhead due to any motion refinement that may be applied on top of the existing unidirectional MVs. On the other hand, the MVMD mode is utilized in both the VVC and AVS3 standards and has proven to be one efficient signaling mechanism to reduce the MVD signaling overhead. Therefore, it may also be beneficial to combine GPM with the MMVD mode. Such a combination may potentially improve the overall coding efficiency of the GPM tool by providing more accurate MVs to capture the individual motion of each GPM partition.
[0072] As discussed earlier, in both the VVC and AVS3 standards, the GPM mode is applied only to the merge / skip mode. Such a design may be suboptimal in terms of coding efficiency, considering that not all non-merge inter CUs can benefit from the flexible non-rectangular partitions of GPM. On the other hand, for the same reasons as mentioned above, the unidirectional predictive motion candidates derived from the regular merge / skip mode are not always accurate in capturing the true motion of the two geometric partitions. Based on such an analysis, additional coding gains can be expected by a reasonable extension of the GPM mode to non-merge inter modes (i.e., CUs that explicitly signal motion information in the bitstream). However, the improvement in MV accuracy comes at the cost of increased signaling overhead. Therefore, to efficiently apply the GPM mode to the explicit inter mode, it will be important to identify one effective signaling scheme that can minimize the signaling cost while providing more accurate MV for the two geometric partitions.
[0073] Proposed method In this disclosure, a method is proposed to further improve the coding efficiency of GPM by applying additional motion refinement on top of the existing unidirectional MV applied to each GPM partition. The proposed method is called Geometric Partition Mode with Motion Vector Refinement (GPM-MVR). In addition, in the proposed scheme, motion refinement is signaled in a similar manner to one of the existing MMVD designs, i.e., based on a set of predefined MVD magnitudes and directions of the motion refinement.
[0074] In another aspect of the present disclosure, a solution is provided to extend GPM modes to explicit inter modes. For ease of explanation, these schemes are called Geometry Partition Mode with Explicit Motion Signaling (GPM-EMS). Specifically, to achieve better consistency with regular inter modes, the proposed GPM-EMS scheme utilizes existing motion signaling mechanisms, namely MVP and MVD, to specify corresponding unidirectional MVs of two geometry partitions.
[0075] Geometry partition mode with separate motion vector refinement In order to improve the coding efficiency of GPM, one improved geometric partition mode with separate motion vector refinement is proposed in this chapter. Specifically, considering a GPM partition, the proposed method first uses the existing syntaxes merge_gpm_idx0 and merge_gpm_idx1 to identify unidirectional MVs for two GPM partitions from the existing unidirectional prediction merge candidate list, and uses them as base MVs. After the two base MVs are determined, two sets of new syntax elements are introduced to separately specify the values of motion refinement applied on top of the base MVs of the two GPM partitions. Specifically, first, two flags, namely gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag, are signaled to indicate whether GPM-MVR has been applied to the first and second GPM partitions, respectively. When the flag of one GPM partition is equal to 1, the corresponding value of the MVR applied to the base MV of that partition is signaled in MMVD format, i.e., one distance index (indicated by syntax elements gpm_mvr_partIdx0_distance_idx and gpm_mvr_partIdx1_distance_idx) specifies the magnitude of the MVR, and one direction index (indicated by syntax elements gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx1_distance_idx) specifies the direction of the MVR. Table 4 shows the syntax elements introduced by the proposed GPM-MVR method.
[0076] [Table 5]
[0077] Based on the proposed syntax elements shown in Table 4, at the decoder, the final MV used to generate unidirectional predicted samples for each GPM partition is equal to the sum of the signaled motion vector refinement and the corresponding base MV. In practice, different sets of MVR magnitudes and directions may be predefined and applied to the proposed GPM-MVR scheme, providing various trade-offs between motion vector accuracy and signaling overhead. In one particular example, it is proposed to re-use the eight MVD offsets (i.e., 1 / 4, 1 / 2, 1, 2, 4, 8, 16, and 32 pels) and four MVD directions (i.e., ±x and y axes) used in the VVC standard for the proposed GPM-MVR scheme. In another example, the existing five MVD offsets {1 / 4, 1 / 2, 1, 2, and 4 pels} and four MVD directions (i.e., ±x and y axes) used in the AVS3 standard are applied in the proposed GPM-MVR scheme.
[0078] As discussed in the chapter "GPM Signaling Design", the unidirectional MVs used for two GPM partitions cannot be identical, so in the existing GPM design, one constraint is applied that the two unidirectional prediction merge indices are different. However, in the proposed GPM-MVR scheme, an additional motion refinement is applied on top of the existing GPM unidirectional MV. Thus, even when the base MVs of the two GPM partitions are identical, the final unidirectional MVs used to predict the two partitions will still be different unless the values of the two motion vector refinements are the same. Based on the above consideration, this constraint (which restricts the two unidirectional prediction merge indices to be different) is removed when the proposed GPM-MVR scheme is applied. In addition, since the two unidirectional prediction merge indices are allowed to be identical, the same maximum value MaxGPMMergeCand-1 is used for binarization of both merg_gpm_idx0 and merge_gpm_idx1, where MaxGPMMergeCand is the number of candidates in the unidirectional prediction merge list.
[0079] As analyzed above, when the unidirectional prediction merge indexes (i.e., merge_gpm_idx0 and merge_gpm_idx1) of two GPM partitions are identical, the values of the two motion vector refinements cannot be the same to ensure that the final MV used for the two partitions is different. Based on such conditions, in one embodiment of the present disclosure, one signaling redundancy elimination method is proposed to reduce the signaling overhead of the MVR of the second GPM partition using the MVR of the first GPM partition when the unidirectional prediction merge indexes of two GPM partitions are the same (i.e., merge_gpm_idx0 is equal to merge_gpm_idx1). In one example, the following signaling conditions are applied:
[0080] First, when the flag gpm_mvr_partIdx0_enable_flag is equal to 0 (i.e., GPM-MVR does not apply to the first GPM partition), the flag gpm_mvr_partIdx1_enable_flag is not signaled but is inferred to be 1 (i.e., GPM-MVR applies to the second GPM partition).
[0081] Second, when both flags gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 1 (i.e., GPM-MVR applies to two GPM partitions) and gpm_mvr_partIdx0_direction_idx is equal to gpm_mvr_partIdx1_direction_idx (i.e., the MVRs of the two GPM partitions have the same direction), the MVR magnitude of the first GPM partition (i.e., gpm_mvr_partIdx0_distance_idx) is used to predict the MVR magnitude of the second GPM partition (i.e., gpm_mvr_partIdx1_distance_idx). Specifically, if gpm_mvr_partIdx1_distance_idx is smaller than gpm_mvr_partIdx0_distance_idx, its original value is signaled directly, otherwise (gpm_mvr_partIdx1_distance_idx is larger than gpm_mvr_partIdx0_distance_idx), its value is subtracted by 1 and signaled in the bitstream. On the decoder side, to decode the value of gpm_mvr_partIdx1_distance_idx, if the parsed value is less than gpm_mvr_partIdx0_distance_idx, then gpm_mvr_partIdx1_distance_idx is set equal to the parsed value, otherwise (the parsed value is equal to or greater than gpm_mvr_partIdx0_distance_idx), then gpm_mvr_partIdx1_distance_idx is set equal to the parsed value plus one. In such a case, to further reduce overhead, different maximum values MaxGPMMVRDistance-1 and MaxGPMMVRDistance-2 may be used for binarization of gpm_mvr_partIdx0_distance_idx and gpm_mvr_partIdx1_distance_idx, where MaxGPMMVRDistance is the number of allowed magnitudes for motion vector refinement.
[0082] In another embodiment, it is proposed to switch the signaling order to gpm_mvr_partIdx0_direction_idx / gpm_mvr_partIdx1_direction_idx and gpm_mvr_partIdx0_distance_idx / gpm_mvr_partIdx1_distance_idx, so that the MVR size is signaled before the MVR size. This allows the encoder / decoder to use the MVR direction of the first GPM partition to adjust the signaling of the MVR direction of the second GPM partition, following the same logic above. In another embodiment, it is proposed to signal the MVR size and direction of the second GPM partition first, and use them to adjust the signaling of the MVR size and direction of the second GPM partition.
[0083] In another embodiment, it is proposed to signal syntax elements related to GPM-MVR before signaling existing GPM syntax elements. Specifically, in such a design, first, two flags gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are signaled to indicate whether GPM-MVR is applied to the first and second GPM partitions, respectively. When the flag of one GPM partition is equal to 1, a distance index (indicated by syntax elements gpm_mvr_partIdx0_distance_idx and gpm_mvr_partIdx1_distance_idx) and a direction index (indicated by syntax elements gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx1_distance_idx) specify the direction of MVR. Then, the existing syntax merge_gpm_idx0 and merge_gpm_idx1 are signaled to identify the unidirectional MVs, i.e., the base MVs, for the two GPM partitions. Table 5 shows the proposed GPM-MVR signaling scheme.
[0084] [Table 6]
[0085] Similar to the signaling method of Table 4, when the GPM-MVR signaling method of Table 5 is applied, certain conditions may be applied to ensure that the resulting MVs used for prediction of the two GPM partitions are not identical. Specifically, the following conditions are proposed for suppressing the signaling of the unidirectional prediction merge indexes merge_gpm_idx0 and merge_gpm_idx1 depending on the values of MVR applied to the first and second GPM partitions:
[0086] First, when the values of both gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 0 (ie, GPM-MVR is disabled for both GPM partitions), the values of merge_gpm_idx0 and merge_gpm_idx1 cannot be the same.
[0087] Second, when gpm_mvr_partIdx0_enable_flag is equal to 1 (i.e., GPM-MVR is enabled for the first GPM partition) and gpm_mvr_partIdx1_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for the second GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical.
[0088] Third, when gpm_mvr_partIdx0_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for the first GPM partition) and gpm_mvr_partIdx1_enable_flag is equal to 1 (i.e., GPM-MVR is enabled for the second GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical.
[0089] Fourth, when the values of gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are both equal to 1 (i.e., GPM-MVR is enabled for both two GPM partitions), the determination of whether the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical depends on the value of MVR applied to the two GPM partitions (indicated by gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx0_distance_idx, and gpm_mvr_partIdx1_direction_idx and gpm_mvr_partIdx1_distance_idx). If the values of the two MVR are equal, merge_gpm_idx0 and merge_gpm_idx1 are not allowed to be identical. Otherwise (the two MVR values are not equal), the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be the same.
[0090] In the above four cases, when the values of merge_gpm_idx0 and merge_gpm_idx1 are not allowed to be the same, the index value of one partition may be used as a predictor of the index value of the other partition. One method is proposed to signal merge_gpm_idx0 first and use its value to predict merge_gpm_idx1. Specifically, in the encoder, when merge_gpm_idx1 is greater than merge_gpm_idx0, the value of merge_gpm_idx1 transmitted to the decoder is reduced by 1. In the decoder, when the received value of merge_gpm_idx1 is equal to or greater than the received value of merge_gpm_idx0, the value of merge_gpm_idx1 is increased by 1. Another method is proposed to signal merge_gpm_idx1 first and use its value to predict merge_gpm_idx0. Thus, in such a case, at the encoder, when merge_gpm_idx0 is greater than merge_gpm_idx1, the value of merge_gpm_idx0 transmitted to the decoder is reduced by 1. At the decoder, when the received value of merge_gpm_idx0 is equal to or greater than the received value of merge_gpm_idx1, the value of merge_gpm_idx0 is increased by 1. In addition, similar to existing GPM signaling designs, different maximum values MaxGPMMergeCand-1 and MaxGPMMergeCand-2 may be used for binarization of the first and second index values, respectively, according to the signaling order. On the other hand, when the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical because there is no correlation between the two index values, the same maximum value MaxGPMMergeCand-1 is used for binarization of both of the two index values.
[0091] In the above method, different maximum values may be applied to the binarization of merge_gpm_idx0 and merge_gpm_idx1 to reduce the signaling cost. The selection of the corresponding maximum value depends on the decoded value of MVR (indicated by gpm_mvr_partIdx0_enable, gpm_mvr_partIdx1_enable, gpm_mvr_partIdx0_direction_idx, gpm_mvr_partIdx1_direction_idx, gpm_mvr_partIdx0_distance_idx, and gpm_mvr_partIdx1_distance_idx). Such a design may result in undesired parsing dependency between different GPM syntax elements and affect the overall parsing. To solve such a problem, in one embodiment, one and the same maximum value (e.g., MaxGPMMergeCand-1) is always proposed for the values of the parsing merge_gpm_idx0 and merge_gpm_idx1. When such a method is used, one bitstream conformance constraint may be used to prevent the two decoded MVs of the two GPM partitions from being the same. Alternatively, such a non-matching constraint may be removed so that the decoded MVs of the two GPM partitions are allowed to be the same. On the other hand, when such a method is applied (i.e., using the same maximum value for merge_gpm_idx0 and merge_gpm_idx1), there is no parsing dependency between merge_gpm_idx0 / merge_gpm_idx1 and other GPM-MVR syntax elements. Therefore, the order of signaling of these syntax elements becomes irrelevant. In one example, it is proposed to move merge_gpm_idx0 / merge_gpm_idx1 signaling before gpm_mvr_partIdx0_enable, gpm_mvr_partIdx1_enable, gpm_mvr_partIdx0_direction_idx, gpm_mvr_partIdx1_direction_idx, gpm_mvr_partIdx0_distance_idx, and gpm_mvr_partIdx1_distance_idx signaling.
[0092] Geometry partition mode with symmetric motion vector refinement For the GPM-MVR method discussed above, two separate MVR values are signaled, one applied to improve the base MV of only one GPM partition. Such a method can be efficient in terms of improving prediction accuracy by allowing independent motion refinement for each GPM partition. However, such flexible motion refinement comes at the cost of increasing the signaling overhead, provided that two different sets of GMP-MVR syntax elements need to be transmitted from the encoder to the decoder. To reduce the signaling overhead, this chapter proposes one geometric partition mode with symmetric motion vector refinement. Specifically, in this method, one single MVR value is signaled for one GPM CU and used for both two GPM partitions according to the symmetric relationship between the Picture Order Count (POC) values of the current picture and the reference picture associated with the two GPM partitions. Table 6 shows the syntax elements when the proposed method is applied.
[0093] [Table 7]
[0094] As shown in Table 6, after the base MVs of the two GPM partitions are selected (based on merge_gpm_idx0 and merge_gpm_idx1), one flag gpm_mvr_enable_flag is signaled to indicate whether the GPM-MVR mode is applied to the current GPM CU. When this flag is equal to 1, it indicates that motion refinement is applied to enhance the base MVs of the two GPM partitions. Otherwise (when the flag is equal to 0), it indicates that motion refinement is not applied to either of the two partitions. If the GPM-MVR mode is enabled, additional syntax elements are further signaled to specify the value of the applied MVR by the direction index gpm_mvr_direction_idx and the magnitude index gpm_mvr_distance_idx. In addition, similar to the MMVD mode, the meaning of the MVR code may vary according to the relationship between the POCs of the current picture and the two reference pictures of the GPM partition. Specifically, when the POCs of the two reference pictures are both greater than or less than the POC of the current picture, the signaled code is the code of the MVR added to both of the two base MVs. Otherwise (when the POC of one reference picture is greater than the current picture and the POC of the other reference picture is less than the current picture), the signaled code is applied to the MVR of the first GPM partition and the inverse code is applied to the MVR of the second GPM partition. In Table 6, the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical.
[0095] In another embodiment, it is proposed to signal two different flags to separately control the enabling / disabling of GPM-MVR mode for two separate GPM partitions. However, when GPM-MVR mode is enabled, only one MVR is signaled based on syntax elements gpm_mvr_direction_idx and gpm_mvr_distance_idx. The corresponding syntax table of such a signaling method is shown in Table 7.
[0096] [Table 8]
[0097] When the signaling method of Table 7 is applied, the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical. However, to ensure that the resulting MVs applied to the two GPM partitions are not redundant, when the flag gpm_mvr_partIdx0_enable_flag is equal to 0 (i.e., GPM-MVR is not applied to the first GPM partition), the flag gpm_mvr_partIdx1_enable_flag is not signaled but is inferred to be 1 (i.e., GPM-MVR is applied to the second GPM partition).
[0098] Adaptation of the Permitted MVR to the GPM-MVR In the GPM-MVR method discussed above, a set of fixed MVR values is used for GPM CUs at both the encoder and the decoder in one video sequence. Such a design is suboptimal for video content with high resolution or high motion. In those cases, the MV tends to be much larger, so that the fixed MVR values may not be optimal for capturing the true motion of those blocks. To further improve the coding performance of the GPM-MVR mode, this disclosure proposes to support the adaptation of the MVR values allowed to be selected by the GPM-MVR mode at various coding levels, such as sequence level, picture / slice picture, coding block group level, etc. For example, multiple MVR sets as well as corresponding codewords may be derived offline according to the specific motion characteristics of different video sequences. The encoder can select the best MVR set and signal the corresponding index of the selected set to the decoder.
[0099] In a specific embodiment of the present disclosure, in addition to the default MVR offsets, which include eight offset magnitudes (i.e., 1 / 4, 1 / 2, 1, 2, 4, 8, 16, and 32 pels) and four MVR directions (i.e., ±x and y axes), another MVR offset defined in the table below is proposed for GPM-MVR mode.
[0100] [Table 9]
[0101] [Table 10]
[0102] In the above Tables 15 and 16, the values +1 / 2 and -1 / 2 in the x-axis and y-axis indicate the diagonal directions (+45° and -45°) in the horizontal and vertical directions. As shown in Tables 15 and 16, compared with the existing MVR offset set, the second MVR offset set introduces two new offset magnitudes (i.e., 3-pel and 6-pel) and four offset directions (45°, 135°, 225°, and 315°). The newly added MVR offsets make the second MVR offset set more suitable for coding video blocks with high motion. In addition, to enable adaptive switching between the two MVR offset sets, one control flag is proposed to be signaled at one specific coding level (e.g., sequence, picture, slice, CTU, and coding block, etc.) to indicate which set of MVR offsets is selected for the GPM-MVR mode applied under the coding level. Assuming that the proposed adaptation is performed at the picture level, Table 17 below shows the corresponding syntax elements signaled in the picture header.
[0103] [Table 11]
[0104] In Table 17 above, a new flag ph_gpm_mvr_offset_set_flag is used to indicate the selection of the corresponding GPM MVR offset used for that picture. When this flag is equal to 0, it means that the default MVR offset (i.e., ¼, ½, 1, 2, 4, 8, 16, and 32 pel magnitudes, and the four MVR orientations ±x and y axis) is applied in this picture for GPM-MVR mode. Otherwise, when this flag is equal to 1, it means that the second MVR offset (i.e., ¼, ½, 1, 2, 3, 4, 6, 8, 16 pel magnitudes, and the eight MVR orientations ±x, y axis, and 45°, 135°, 225°, and 315°) is applied in this picture for GPM-MVR mode.
[0105] Different methods may be applied to signal the MVR offset. First, it is proposed to binarize the MVR direction using a fixed length codeword, provided that the MVR direction is usually statistically uniformly distributed. Taking the default MVR offset as an example, there are a total of four directions, and codewords 00, 01, 10, and 11 may be used to represent these four directions. On the other hand, since the MVR offset magnitude may have a varying distribution adapted to the specific motion characteristics of the video content, it is proposed to binarize the MVR magnitude using a variable length codeword. Table 18 below shows one specific codeword table that may be used for binarizing the MVR magnitude for the default MVR offset set and the second MVR offset set.
[0106] [Table 12]
[0107] In other embodiments, different fixed length variable codewords may be applied to binarize the MVR offset magnitudes of the default and second MVR offset sets, for example, bins "0" and "1" in the codeword table above may be swapped to accommodate different 0 / 1 statistics of a context-adaptive binary arithmetic coding (CABAC) engine.
[0108] In one specific example, two different codeword tables are provided to binarize the MVR magnitude values. The tables below show the corresponding codewords for the default and secondary MVR offset sets applied in the first and second codeword tables. Table 19 shows the MVR offset magnitude codewords in the first codeword table. Table 20 shows the MVR offset magnitude codewords in the second codeword table.
[0109] [Table 13]
[0110] [Table 14]
[0111] To enable the adaptive switch between the two codeword tables, one indication flag is proposed to be signaled at one particular coding level (e.g., sequence, picture, slice, CTU, and coding block, etc.) to specify which codeword table is used to binarize the MVR magnitude under that coding level. Assuming that the proposed adaptation is performed at the picture level, Table 21 below shows the corresponding syntax elements signaled in the picture header, with the newly added syntax elements in italic and bold.
[0112] [Table 15]
[0113] In the above syntax table, a new flag ph_gpm_mvr_step_codeword_flag is used to indicate the selection of the corresponding codeword table used for binarization of the MVR magnitude of a picture. When this flag is equal to 0, it indicates that the first codeword table is applied to the picture, otherwise (i.e., the flag is equal to 1), it indicates that the second codeword table is applied to the picture.
[0114] In another embodiment, it is proposed to always use one codeword table for binarizing the MVR offset magnitudes during encoding / decoding of the entire video sequence. In one example, it is proposed to always use the first codeword table for binarizing the MVR magnitudes. In another example, it is proposed to always use the second codeword table for binarizing the MVR magnitudes. In another method, it is proposed to use one fixed codeword table (e.g., the second codeword table) for binarizing all MVR magnitudes.
[0115] In another method, a statistically based binarization method may be applied to adaptively design the optimal codeword for the MVR offset magnitude on the fly rather than signaling. The statistical information used to determine the optimal codeword may be, but is not limited to, a probability distribution of the MVR offset magnitude collected over multiple previously coded pictures, slices, and / or coding blocks. The codeword may be re-determined / updated at various frequency levels. For example, the update may be performed every time a CU is coded in GPM-MVR mode. In another example, the update may be re-determined and / or updated every time multiple, e.g., 8 or 16, CUs are coded in GPM-MVR mode.
[0116] Alternatively, instead of redesigning a new set of codewords, the proposed statistics-based method can also be used to reorder the MVR magnitude values based on the same set of codewords, such that shorter codewords are assigned to the more used magnitudes and longer codewords are assigned to the less used magnitudes.Taking the table below as an example, assuming that the statistics are collected at the picture level, the "Used" row shows the corresponding proportion of different MVR offset magnitudes used by the GPM-MVR coding block in the previously coded picture. According to the values in the "usage" row (i.e., shortened unary codewords) using the same binarization method, the encoder / decoder can order the MVR magnitude values based on their usage, and then the encoder / decoder can assign the shortest codeword (i.e., "1") to the most frequently used MVR magnitude (i.e., 1 pel), the next shortest codeword (i.e., "01") to the next most frequently used MVR magnitude (i.e., ½ pel), and the longest codewords (i.e., "0000001" and "0000000") to the two least used MVR magnitudes (i.e., 16 pels and 32 pels). Thus, by such a reordering scheme, the same set of codewords can be freely reordered to accommodate dynamic changes in the statistical distribution of the MVR magnitudes.
[0117] [Table 16]
[0118] Encoder acceleration algorithm for GPM-MVR rate-distortion optimization For the proposed GPM-MVR scheme, to determine the optimal MVR for each GPM partition, the encoder may need to test the rate-distortion cost of each GPM partition multiple times, varying the applied MVR value each time. This may significantly increase the encoding complexity of the GPM mode. To address the encoding complexity issue, the following fast encoding logic is proposed in this section.
[0119] First, due to the quadtree / binary / ternary block partition structure applied in VVC and AVS3, one and the same coding block may be identified during the rate-distortion optimization (RDO) process and split by one different partition path each. In the current VTM / HPM encoder implementation, GPM and GPM-MVR modes are always tested whenever one and the same CU is obtained by different block partition combinations, along with other inter- and intra-coding modes. In general, for different partition paths, only adjacent blocks of one CU may differ, which should have a relatively minor impact on the optimal coding mode that one CU selects. Based on such consideration, in order to reduce the total number of applied GPM RDOs, it is proposed to memorize the decision of whether GPM mode is selected when the RD cost of one CU is identified for the first time. After that, when the same CU is identified again by the RDO process (by another partition path), the RD cost of GPM (including GPM-MVR) is identified only if GPM is selected for that CU for the first time. If a GPM is not selected for the initial RD check of a CU, then when the same CU is realized by another partition path, only the GPM (without GPM-MVR) is tested. Alternatively, if a GPM is not selected for the initial RD check of a CU, then when the same CU is realized by another partition path, then neither the GPM nor the GPM-MVR is tested.
[0120] Second, to reduce the number of GPM partitions for GPM-MVR mode, when the RD cost of one CU is confirmed for the first time, it is proposed to keep the first M GPM partition modes without the minimum RD cost. Then, when the same CU is confirmed again by the RDO process (by another partition path), only those M GPM partition modes are tested for GPM-MVR mode.
[0121] Thirdly, in order to reduce the number of GPM partitions tested for one initial RDO process for each GPM partition, it is proposed to first calculate the sum of absolute difference (SAD) value when using different unidirectional predictive merge candidates for two GPM partitions.Then, for each GPM partition under one specific partition mode, select the best unidirectional predictive merge candidate with the smallest SAD value, and calculate the corresponding SAD value of partition mode that is equal to the sum of the SAD value of the best unidirectional predictive merge candidate for two GPM partitions.Then, for the subsequent RD process, only the first N partition modes with the best SAD value for the previous step are tested against GPM-MVR mode.
[0122] Geometric partitioning with explicit motion signaling In this chapter, several methods are proposed to extend the GPM mode to regular inter-mode bi-prediction, where the two unidirectional MVs of the GPM mode are explicitly signaled from the encoder to the decoder.
[0123] In the first solution (Solution 1), it is proposed to fully re-use the existing motion signaling of bi-prediction to signal two unidirectional MVs of GPM mode. Table 8 shows the modified syntax table of the proposed scheme, with the newly added syntax elements in bold italics. As shown in Table 8, in this solution, all existing syntax elements signaling L0 and L1 motion information are fully re-used to indicate the unidirectional MVs of the two GPM partitions, respectively. In addition, it is assumed that the L0 MV is always associated with the first GPM partition, and the L1 MV is always associated with the second GPM partition. On the other hand, in Table 8, the inter prediction syntax, i.e., inter_pred_idc, is signaled before the GPM flag (i.e., gpm_flag), and thus the value of inter_pred_idc can be used to adjust the presence of gpm_flag. Specifically, the flag gpm_flag needs to be signaled only when inter_pred_idc is equal to PRED_BI (i.e., bi-prediction) and both inter_affine_flag and sym_mvd_flag are equal to 0 (i.e., the CU is not coded by either affine or SMVD mode). When the flag gpm_flag is not signaled, its value is always inferred to be 0 (i.e., GPM mode is disabled). When gpm_flag is 1, another syntax element gpm_partition_idx is further signaled to indicate the currently selected GPM mode (out of a total of 64 GPM partitions) for the CU.
[0124] [Table 17] TIFF2024524402000019.tif109170
[0125] Another method proposes to place the signaling of flag gpm_flag before other inter signaling syntax elements, so that the value of gpm_flag can be used to determine whether other inter syntax elements need to be present or not. Table 9 shows the corresponding syntax table when such a method is applied, with the newly added syntax elements shown in bold italics. As can be seen, in Table 9, gpm_flag is signaled first. When gpm_flag is equal to 1, the corresponding signaling of inter_pred_idc, inter_affine_flag, and sym_mvd_flag can be bypassed. Instead, the corresponding values of the three syntax elements can be inferred as PRED_BI, 0, and 0, respectively.
[0126] [Table 18] TIFF2024524402000021.tif110170
[0127] In both Table 8 and Table 9, the SMVD mode cannot be combined with the GPM mode. In another example, it is proposed to allow the SMVD mode when the current CU is coded by the GPM mode. When such a combination is allowed, by following the same design of the SMVD, the MVDs of the two GPM partitions are assumed to be symmetrical, and therefore only the MVD of the first GPM partition needs to be signaled, and the MVD of the second GPM partition is always symmetrical with the first MVD. When such a method is applied, the corresponding signaling condition of sym_mvd_flag for gpm_flag can be removed.
[0128] As shown above, in the first solution, we always assume that L0 MVs are used for the first GPM partition and L1 MVs are used for the second GPM partition. Such a design may be suboptimal in the sense that this method prohibits MVs of two GPM partitions from one and the same prediction list (L0 or L1). To solve such a problem, one alternative GPM-EMS scheme, solution 2, is proposed with a signaling design shown in Table 10. In Table 10, the newly added syntax elements are shown in italic bold. As shown in Table 10, first a flag gpm_flag is signaled. When this flag is equal to 1 (i.e., GPM is enabled), the syntax gpm_partition_idx is signaled to specify the selected GPM mode. Then, one additional flag gpm_pred_dir_flag0 is signaled to indicate the corresponding prediction list from which the MV of the first GPM partition comes. When the flag gpm_pred_dir_flag0 is equal to 1, it indicates that the MV of the first GPM partition comes from L1, otherwise (flag equal to 0), it indicates that the MV of the first GPM partition comes from L0. Then, the existing syntax elements ref_idx_l0, mvp_l0_flag, and mvd_coding() are utilized to signal the values of the reference picture index, mvp index, and MVD of the first GPM partition. On the other hand, similar to the first partition, another syntax element gpm_pred_dir_flag1 is introduced to select the corresponding prediction list of the second GPM partition, and subsequently the existing syntax elements ref_idx_l1, mvp_l1_flag, and mvd_coding() are used to derive the MV of the second GPM partition.
[0129] [Table 19] TIFF2024524402000023.tif152170
[0130] Finally, it should be mentioned that, provided that the GPM mode consists of two unidirectional prediction partitions (except for mixed samples on the partition edge), when the proposed GPM-EMS scheme is enabled for one inter-CU, some existing coding tools in VVC and AVS3 that are specifically designed for bidirectional prediction, e.g., bidirectional optical flow, decoder-side motion vector refinement (DMVR), and bidirectional prediction with CU weight (BCW), can be automatically bypassed. For example, provided that BCW cannot be applied to the GPM mode, in order to reduce the signaling overhead, when one of the proposed GPM-EMS is enabled for one CU, the corresponding BCW weight does not need to be further signaled for the CU.
[0131] Combination of GPM-MVR and GPM-EMS In this section, it is proposed to combine GPM-MVR and GPM-EMS for one CU with geometry partition. Specifically, unlike GPM-MVR or GPM-EMS in which only one of merge-based motion signaling or explicit signaling can be applied to signal unidirectional predicted MVs of two GPM partitions, the proposed scheme allows 1) one partition to use GPM-MVR-based motion signaling and the other to use GPM-EMS-based motion signaling, or 2) two partitions to use GPM-MVR-based motion signaling, or 3) two partitions to use GPM-EMS-based motion signaling. Using the GPM-MVR signaling in Table 4 and the GPM-EMS in Table 10, Table 11 shows the corresponding syntax table after the proposed GPM-MVR and GPM-EMS are combined. In Table 11, the newly added syntax elements are shown in bold italics. As shown in Table 11, two additional syntax elements gpm_merge_flag0 and gpm_merge_flag1 are introduced for partitions #1 and #2, respectively, to specify the corresponding partitions that use GPM-MVR-based merge signaling or GPM-EMS-based explicit signaling. When this flag is 1, it means that GPM-MVR-based signaling is enabled for the partitions where GPM unidirectional prediction motion is signaled by merge_gpm_idxX, gpm_mvr_partIdxX_enabled_flag, gpm_mvr_partIdxX_direction_idx, and gpm_mvr_partIdxX_distance_idx, where X=0,1. Otherwise, if this flag is 0, it means that the unidirectional predictive motion of this partition is explicitly signaled by the GPM-EMS scheme using the syntax elements gpm_pred_dir_flagX, ref_idx_lX, mvp_lX_flag, and mvd_lX, where X=0,1.
[0132] [Table 20]
[0133] Combining GPM-MVR and template matching In this chapter, different solutions are presented for combining GPM-MVR with template matching.
[0134] In Method 1, when one CU is coded in GPM mode, it is proposed to signal two separate flags for two GPM partitions, each flag indicating whether the unidirectional motion of the corresponding partition is further refined by template matching or not. When this flag is enabled, a template is generated using the top-left neighboring reconstructed sample of the current CU, and then the unidirectional motion of the partition is refined by minimizing the difference between the template and its reference sample, following the same procedure as introduced in the "Template Matching" chapter. Otherwise (when the flag is disabled), template matching is not applied to this partition, and GPM-MVR can be further applied. Using the GPM-MVR signaling method in Table 5 as an example, Table 12 shows the corresponding syntax table when GPM-MVR is combined with template matching. In Table 12, the newly added syntax elements are shown in bold italics.
[0135] [Table 21]
[0136] As shown in Table 12, in the proposed scheme, two additional flags, gpm_tm_enable_flag0 and gpm_tm_enable_flag1, are first signaled to indicate whether motion is refined for two GPM partitions, respectively. When this flag is 1, it indicates that TM is applied to refine unidirectional MV of one partition. When this flag is 0, one flag (gpm_mvr_partIdx0_enable_flag or gpm_mvr_partIdx0_enable_flag) is further signaled to indicate whether GPM-MVR is applied to the GPM partition, respectively. When the flag of one GPM partition is equal to 1, a distance index (indicated by syntax elements gpm_mvr_partIdx0_distance_idx and gpm_mvr_partIdx1_distance_idx) and a direction index (indicated by syntax elements gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx1_distance_idx) are signaled to specify the direction of the MVR. Then, the existing syntax merge_gpm_idx0 and merge_gpm_idx1 are signaled to identify the unidirectional MV for the two GPM partitions. Meanwhile, similar to the signaling conditions that apply in Table 5, the following conditions may be applied to ensure that the resulting MVs used for prediction of the two GPM partitions are not identical.
[0137] First, when the values of gpm_tm_enable_flag0 and gpm_tm_enable_flag1 are both equal to 1 (ie, TM is enabled for both two GPM partitions), the values of merge_gpm_idx0 and merge_gpm_idx1 cannot be the same.
[0138] Second, when one of gpm_tm_enable_flag0 and gpm_tm_enable_flag1 is 1 and the other is 0, the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be the same.
[0139] Otherwise, i.e., both gpm_tm_enable_flag0 and gpm_tm_enable_flag1 are equal to 1, firstly, when the values of gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are both equal to 0 (i.e., GPM-MVR is disabled for both of the two GPM partitions), the values of merge_gpm_idx0 and merge_gpm_idx1 cannot be the same, and secondly, when the values of gpm_mvr The values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical when gpm_mvr_partIdx0_enable_flag is equal to 1 (i.e., GPM-MVR is enabled for the first GPM partition) and gpm_mvr_partIdx1_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for the second GPM partition); third, when gpm_mvr_partIdx0_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for the first GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical. fourth, when the values of gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are both equal to 1 (i.e., GPM-MVR is enabled for the second GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical; and fourth, when the values of gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are both equal to 1 (i.e., GPM-MVR is enabled for the second GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical. When the MVR_IDX1 and MVR_PARTIDS are enabled for both, the determination of whether the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical depends on the values of MVR that apply to the two GPM partitions (indicated by gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx0_distance_idx, and gpm_mvr_partIdx1_direction_idx and gpm_mvr_partIdx1_distance_idx).If the two MVR values are equal, then merge_gpm_idx0 and merge_gpm_idx1 are not allowed to be the same. Otherwise (if the two MVR values are not equal), then merge_gpm_idx0 and merge_gpm_idx1 are allowed to be the same.
[0140] In the above method 1, TM and MVR are exclusively applied to GPM. In such a scheme, further application of MVR on top of the refined MV in TM mode is prohibited. Therefore, in order to further provide more MV candidates for GPM, method 2 is proposed to enable the application of MVR offset on top of TM refined MV. Table 13 shows the corresponding syntax table when GPM-MVR is combined with template matching. In Table 13, the newly added syntax elements are shown in italic bold.
[0141] [Table 22]
[0142] As shown in Table 13, unlike Table 12, the signaling conditions of gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag for gpm_tm_enable_flag0 and gpm_tm_enable_flag1 are removed. Therefore, regardless of whether TM is applied to refine the unidirectional movement of one GPM partition, MV refinement is always allowed to be applied to the MV of the GPM partition. Similar to above, the following conditions should be applied to ensure that the resulting MVs of the two GPM partitions are not identical.
[0143] First, when one of gpm_tm_enable_flag0 and gpm_tm_enable_flag1 is 1 and the other is 0, the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be the same.
[0144] Otherwise, i.e., when both gpm_tm_enable_flag0 and gpm_tm_enable_flag1 are equal to 1 or both of these flags are equal to 0, first, when the values of gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are both equal to 0 (i.e., GPM-MVR is disabled for both GPM partitions), the values of merge_gpm_idx0 and merge_gpm_idx1 cannot be the same. second, when gpm_mvr_partIdx0_enable_flag is equal to 1 (i.e., GPM-MVR is enabled for the first GPM partition) and gpm_mvr_partIdx1_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for the second GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical; and third, when gpm_mvr_partIdx0_enable_flag is equal to 0 (i.e., Fourth, when the values of gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are both equal to 1 (i.e., GPM-MVR is enabled for the second GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical; and fourth, when the values of gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are both equal to 1 (i.e., GPM-MVR is enabled for the second GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical. When a GPM partition is enabled for both GPM partitions, the determination of whether the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical depends on the value of MVR that applies to the two GPM partitions (indicated by gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx0_distance_idx, and gpm_mvr_partIdx1_direction_idx and gpm_mvr_partIdx1_distance_idx).If the two MVR values are equal, then merge_gpm_idx0 and merge_gpm_idx1 are not allowed to be the same. Otherwise (if the two MVR values are not equal), then merge_gpm_idx0 and merge_gpm_idx1 are allowed to be the same.
[0145] In the above two methods, two separate flags are required to be signaled to indicate whether TM is applied to each GPM partition or not. The added signaling may reduce the overall coding efficiency due to additional overhead, especially at low bit rates. To reduce the signaling overhead, instead of introducing additional signaling, method 3 is proposed to insert TM-based unidirectional MVs into the unidirectional MV candidate list of GPM mode. The TM-based unidirectional MVs are generated according to the same TM process as described in the “Template Matching” chapter, which uses the original unidirectional MV of GPM as the initial MV. With such a scheme, there is no need to further signal an extra control flag from the encoder to the decoder. Instead, the decoder can identify whether one MV is improved by TM or not by the corresponding merge index (i.e., merge_gpm_idx0 and merge_gpm_idx1) received from the bitstream. There may be different ways to arrange regular GPM MV candidates (i.e., non-TM) and TM-based MV candidates. In one method, it is proposed to place the TM-based MV candidates at the beginning of the MV candidate list, followed by the non-TM-based MV candidates. In another method, it is proposed to place the non-TM-based MV candidates first, followed by the TM-based candidates. In another method, it is proposed to place the TM-based MV candidates and the non-TM-based MV candidates alternately. For example, it can place the first N non-TM-based candidates, then all the TM-based candidates, and finally the remaining non-TM-based candidates. In another example, it can place the first N TM-based candidates, then all the non-TM-based candidates, and finally the remaining TM-based candidates. In another example, it is proposed to place the non-TM-based candidates and the TM-based candidates consecutively, i.e., one non-TM-based candidate, one TM-based candidate, etc.
[0146] In method 1, the two GPM template flags are signaled before the GPM-MVR flag. Specifically, in such a design, GPM-MVR may be enabled only for one given GPM partition by first signaling the GPM template flag of one partition equal to 0. The GPM template flag may be coded using an appropriate context model, but incurs a signaling penalty in the GPM-MVR mode. To solve such a problem, an embodiment of the present disclosure proposes to signal the GPM-MVR mode first and then the GPM-TM mode. Specifically, in this method, a GPM-MVR flag is first signaled for each GPM partition to indicate whether GPM-MVR applies to that partition. When the flag is equal to 1, the MVR syntax elements gpm_mvr_partIdx0_distance_idx / gpm_mvr_partIdx1_distance_idx and gpm_mvr_partIdx0_dierction_idx / gpm_mvr_partIdx1_direction_idx are further signaled to specify the corresponding values of the MVR magnitude and direction of the partition. Otherwise, when the GPM-MVR flag of the partition is equal to false, the GPM-TM flag is signaled to indicate whether the GPM-TM mode (which uses the left and top adjacent reconstructed samples to refine the MV of the partition) is applied. Table 22 shows the corresponding syntax table when the above signaling method is applied, with the newly added syntax elements in italic bold.
[0147] [Table 23]
[0148] Furthermore, in order to remove redundancy in signaling between GPM merge indexes, the following conditions should be applied:
[0149] First, when one of gpm_tm_enable_flag0 and gpm_tm_enable_flag1 is 1 and the other is 0, the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be the same.
[0150] Second, when both gpm_tm_enable_flag0 and gpm_tm_enable_flag1 are equal to 1, the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be the same.
[0151] In another example, when the values of gpm_tm_enable_flag0 and gpm_tm_enable_flag1 are both equal to 1, the values of merge_gpm_idx0 and merge_gpm_idx1 are not allowed to be the same.
[0152] Third, when both gpm_tm_enable_flag0 and gpm_tm_enable_flag1 are equal to 0, different conditions apply. When the values of both gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 0 (i.e., GPM-MVR is disabled for both two GPM partitions), the values of merge_gpm_idx0 and merge_gpm_idx1 cannot be the same. When gpm_mvr_partIdx0_enable_flag is equal to 1 (i.e., GPM-MVR is enabled for the first GPM partition) and gpm_mvr_partIdx1_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for the second GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical. When gpm_mvr_partIdx0_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for the first GPM partition) and gpm_mvr_partIdx1_enable_flag is equal to 1 (i.e., GPM-MVR is enabled for the second GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical. When the values of gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are both equal to 1 (i.e., GPM-MVR is enabled for both two GPM partitions), the determination of whether the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical depends on the value of the MVR that applies to the two GPM partitions (indicated by gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx0_distance_idx, and gpm_mvr_partIdx1_direction_idx and gpm_mvr_partIdx1_distance_idx). If the values of the two MVR are equal, merge_gpm_idx0 and merge_gpm_idx1 are not allowed to be identical.Otherwise (the two MVR values are not equal), the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical.
[0153] In another method, instead of using two separate GPM-TM flags, one single flag is proposed to jointly control the enable / disable of template matching for two GPM partitions. When the flag is true, it means that two unidirectional MVs of two GPM partitions need to be refined based on minimizing the difference between the template (i.e., the left and top adjacent reconstructed samples) and its corresponding reference sample by the template matching method. Specifically, similar to method 4, two GPM-MVR flags are first signaled to one GPM CU to indicate whether GPM-MVR is applied to one specific GPM partition. When the GPM-MVR flag of each partition is equal to true, the MVR magnitude and MVR direction are further signaled for that partition below. Furthermore, when both GPM-MVR flags of two GPM partitions are equal to false, the GPM-TM flag is further signaled to indicate whether GPM-TM is applied to both of the two GPM partitions. Table 23 shows the corresponding syntax table for GPM mode when such a design is applied, with the newly added syntax elements in italic bold.
[0154] [Table 24]
[0155] In another embodiment, it is proposed to signal the GPM-TM flag for two GPM partitions and then signal the two GPM-MVR flags. Correspondingly, the value of GPM-TM can be used to adjust the presence of the two GPM-MVR flags, such that the GPM-MVR flag is signaled only when the value of the GPM-TM flag is equal to 0 (i.e., GPM-TM is not applied to the two GPM partitions). Table 24 shows the corresponding syntax table of the GPM mode when such a signaling scheme is applied, with the newly added syntax elements in italic bold.
[0156] [Table 25]
[0157] In addition, for both of the two methods, the following conditions should be applied to remove the redundancy of signaling between GPM merge indexes:
[0158] First, the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be the same when gpm_tm_enable_flag is 1. Second, different conditions may apply when gpm_tm_enable_flag is equal to 0. For example, when the values of both gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 0 (i.e., GPM-MVR is disabled for both two GPM partitions), the values of merge_gpm_idx0 and merge_gpm_idx1 cannot be the same. Additionally, when gpm_mvr_partIdx0_enable_flag is equal to 1 (i.e., GPM-MVR is enabled for the first GPM partition) and gpm_mvr_partIdx1_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for the second GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical. Additionally, when gpm_mvr_partIdx0_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for the first GPM partition) and gpm_mvr_partIdx1_enable_flag is equal to 1 (i.e., GPM-MVR is enabled for the second GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical.Furthermore, when the values of both gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 1 (i.e., GPM-MVR is enabled for both two GPM partitions), the determination of whether the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical depends on the value of the MVR that applies to the two GPM partitions (indicated by gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx0_distance_idx, and gpm_mvr_partIdx1_direction_idx and gpm_mvr_partIdx1_distance_idx). If the values of the two MVR are equal, then merge_gpm_idx0 and merge_gpm_idx1 are not allowed to be identical. Otherwise (the two MVR values are not equal), the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be the same.
[0159] In another example, when gpm_tm_enable_flag is equal to 1, the values of merge_gpm_idx0 and merge_gpm_idx1 are not allowed to be the same.
[0160] When the template matching scheme is applied to the GPM mode, additional complexity is required for both the encoder and the decoder by performing computationally extensive motion estimation to identify the optimal unidirectional MV for each GPM partition. Such a non-negligible increase in complexity may make the GPM mode not viable for certain lower level encoders or for certain video applications such as live video streaming, video conferencing, and video gaming that require low video delay. Based on such consideration, it is proposed to add one control flag at certain higher coding levels such as sequence level, picture / slice level coding block group level, and additions for the CU to adaptively enable or disable the GPM-TM mode under its level. Assuming that the proposed adaptation is performed at the picture level, Table 25 shows the corresponding syntax elements signaled in the picture header, with the newly added syntax elements in italic bold.
[0161] [Table 26]
[0162] In the above syntax Table 25, the flag sps_dmvd_enable_flag is a sequence-level control flag that indicates whether template matching is enabled for coding of a video sequence, and ph_gpm_tm_enable_flag is a proposed GPM-TM control flag used to indicate whether GPM-TM may be applied to a CU within a picture.
[0163] Building GPM Candidate Lists by Pruning Motion Vectors As discussed in the introduction, to obtain the MVs of two geometric partitions, one unidirectional prediction candidate list is first derived directly from the regular merge candidate list generation process. Under the condition that the selection of the prediction direction of each GPM MV is based on the parity of the corresponding merge index, the MVs of the two geometric partitions can be identical, but it is obviously not reasonable since the geometric partition of the CU cannot provide any additional benefit compared to the partition-free case. In order to avoid such redundancy, it is proposed to apply a pruning of motion vectors when generating the unidirectional prediction MV candidate list of one GPM CU, such that only one MV can be added to the list only if it is not identical to any of the existing candidates in the list. In another scheme, it is further proposed that one MV threshold is applied when comparing two MVs. Specifically, by such a method, when the difference between the two MVs (horizontal and vertical directions, respectively) is smaller than one MV threshold, the two MVs are considered to be identical, otherwise (the MV difference in one direction is greater than or equal to the MV threshold), the two MVs are considered not to be identical. One method proposes to use one fixed MV threshold for all block sizes. Another method proposes to determine the value of the MV threshold based on the size of the coding block, such that a larger MV threshold is used for larger CUs and a smaller MV threshold is used for smaller CUs. In some examples, when the number of samples in the block N<64, the value of the MV threshold is set to 1 / 4 pel, when 64≦N<256, the value of the MV threshold is set to 1 / 2 pel, and when N≧256, the value of the MV threshold is set to 1 pel.
[0164] Improved Unidirectional MV Candidate List Construction An improved candidate list construction method is proposed to derive unidirectional MV candidates from the MV candidate list in the regular merge mode.
[0165] First, a parity-based unidirectional MV is obtained. Similar to existing GPM designs, multiple unidirectional MVs are first derived from a regular merge candidate list generation process. For example, as shown in Figure 5, n is denoted as the index of a unidirectional motion in the GPM MV candidate list. When X is equal to the parity of n, the LX motion vector of the nth merge candidate is used as the nth unidirectional MV in the candidate list. If the LX motion vector of the nth extended merge candidate does not exist, the L(1-X) MV of the same candidate is selected instead.
[0166] In some examples, anti-parity based unidirectional MVs are obtained. When the unidirectional MV candidate list is not complete, additional unidirectional MVs derived from the bi-predictive MVs in the regular merge candidate list are further added to the unidirectional MV candidate list. Specifically, for each bi-predictive MV with merge index n in the regular merge candidate list, L(1-X) MVs of the bi-predictive MV are further added to the unidirectional MV candidate list, where X is equal to the parity of n.
[0167] In some examples, a pair-averaged unidirectional MV is obtained. When the unidirectional MV candidate list is not full, one or more pair-averaged candidates are added to the list by averaging the first two unidirectional candidates in one reference picture list (L0 and L1) in the existing unidirectional MV candidate list. In some examples, the pair-averaged unidirectional MV is obtained after an anti-parity-based unidirectional MV is obtained and added to the unidirectional MV candidate list. In some other examples, after the pair-averaged unidirectional MV is obtained, an anti-parity-based unidirectional MV can be obtained and the anti-parity-based unidirectional MV can be added to the unidirectional MV candidate list.
[0168] For example, for one given reference picture list L0 / L1, the encoder / decoder may assume that the first MV candidate is defined as p0Cand and the second L0 MV candidate is defined as p1Cand. When two MV candidates point to the same reference picture, one pair average candidate is generated by averaging the two MV candidates that target the same reference; otherwise, when two MV candidates point to different reference pictures, the magnitude of the pair MV is calculated by averaging p0Cand and p1Cand, and the reference picture of the first MV candidate is selected as the reference picture of the resulting pair average MV. In another example, when two MV candidates point to different reference pictures, the reference picture of the second MV candidate is selected as the reference picture of the resulting pair average MV.
[0169] Furthermore, zero unidirectional MVs are obtained. If the unidirectional MV candidate list is not yet full, zero unidirectional MVs (targeting different reference pictures in the current picture's L0 / L1 reference picture list) are periodically added to the unidirectional MV candidate list until the maximum length of the list is reached.
[0170] In addition, an MV pruning process may be further applied in the above unidirectional MV candidate list generation scheme to remove redundant MV candidates from the list. In one or more embodiments, a default MV pruning method is applied such that only one MV is allowed to be added to the list if it is not identical to any of the existing candidates in that list. In some other embodiments, an alternative MV pruning method proposed in the section "Building a GPM Candidate List by Pruning Motion Vectors" is applied, where the MV threshold used to determine whether two MVs are identical depends on the block size of the current CU.
[0171] 9 illustrates a computing environment (or computing device) 910 coupled to a user interface 960. The computing environment 910 may be part of a data processing server. In some embodiments, the computing device 910 may perform any of the various methods or processes (e.g., encoding / decoding methods or processes) previously described herein in accordance with various examples of the present disclosure. The computing environment 910 may include a processor 920, a memory 940, and an I / O interface 950.
[0172] The processor 920 typically controls the overall operation of the computing environment 910, such as operations related to display, data acquisition, data communication, and image processing. The processor 920 may include one or more processors to execute instructions for performing all or some of the steps in the methods described above. Additionally, the processor 1020 may include one or more modules that facilitate interaction between the processor 920 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single chip machine, a GPU, or the like.
[0173] The memory 940 is configured to store various types of data to support the operation of the computing environment 910. The memory 940 may include predefined software 942. Examples of such data include instructions for any application or method operated on the computing environment 910, video data sets, image data, etc. The memory 940 may be implemented by using any type of volatile or non-volatile memory device, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic or optical disk, or a combination thereof.
[0174] The I / O interface 950 provides an interface between the processor 920 and peripheral interface modules such as a keyboard, a click wheel, buttons, etc. The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. The I / O interface 950 may be coupled to an encoder and a decoder.
[0175] In some embodiments, a non-transitory computer readable storage medium including a plurality of programs, such as those contained within memory 940, is also provided, the plurality of programs being executable by processor 920 within computing environment 910 to perform the methods described above. For example, the non-transitory computer readable storage medium can be a ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0176] A non-transitory computer-readable storage medium stores a plurality of programs for execution by a computing device having one or more processors, the plurality of programs, when executed by the one or more processors, causing the computing device to perform the aforementioned method for motion prediction.
[0177] In some embodiments, the computing environment 910 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), graphic processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0178] FIG. 8 is a flow diagram illustrating a method for decoding a video block with a GPM according to an example of this disclosure.
[0179] At step 801, the processor 920 may partition a video block into first and second geometric partitions.
[0180] In step 802, the processor 920 may build a unidirectional MV candidate list for the GPM by adding a number of regular merge candidates.
[0181] A number of canonical merge candidates may be derived from the canonical merge candidate list generation process. For example, as shown in FIG. 5, when X is equal to the parity of n, the LX motion vector of the nth merge candidate is used as the nth unidirectional MV in the candidate list.
[0182] In step 803, in response to determining that the unidirectional MV candidate list is not complete, the processor 920 may construct a first updated unidirectional MV candidate list by adding one or more additional unidirectional MVs derived from one or more bidirectional predicted MVs of the regular merge candidate list to the unidirectional MV candidate list.
[0183] In some examples, one or more additional unidirectional MVs may be derived by obtaining one or more MV candidates with odd merge indexes in the first reference picture list and one or more MVs with even merge indexes in the second reference picture list. For example, the first reference picture list may be L0, and the second reference picture list may be L1. The additional unidirectional MVs may include candidates with odd merge indexes in L0 and candidates with even merge indexes in L1. That is, for each bi-predicted MV with merge index n in the regular merge candidate list, L(1-X) MVs of the bi-predicted MVs are further added to the unidirectional MV candidate list, where X is equal to the parity of n. In some examples, the first reference picture list may be L1, and the second reference picture list may be L0.
[0184] In some examples, the processor 920 may construct a second updated unidirectional MV candidate list by adding one or more pair average candidates to the first updated unidirectional MV candidate list, as shown in step 804.
[0185] For example, the processor 920 may obtain a pair average candidate by obtaining the first two unidirectional MV candidates in the first reference picture list or the second reference picture list, and averaging the first two unidirectional MV candidates in response to determining that the first two unidirectional MV candidates indicate the same reference picture.
[0186] In some examples, the pair average candidate may be obtained by, in response to determining that the first two unidirectional MV candidates indicate different reference pictures, determining a magnitude of the pair average candidate by averaging the first two unidirectional MV candidates and determining the reference picture of the first unidirectional MV candidate as the reference picture of the pair average candidate.
[0187] In some examples, the pair average candidate may be obtained by, in response to determining that the first two unidirectional MV candidates indicate different reference pictures, determining a magnitude of the pair average candidate by averaging the first two unidirectional MV candidates and determining a reference picture of the second unidirectional MV candidate as a reference picture of the pair average candidate.
[0188] In some examples, in response to determining that the second updated unidirectional MV candidate list is not complete, as shown in step 805, the processor 920 may periodically add zero unidirectional MVs to the second updated unidirectional MV candidate list until a maximum length is reached.
[0189] In some examples, when constructing a unidirectional MV candidate list, redundant candidates may be removed from the unidirectional MV candidate list.
[0190] For example, the processor 920 may omit adding the additional unidirectional MV to the unidirectional MV candidate list in response to determining that the additional unidirectional MV is equal to a candidate in the unidirectional MV candidate list.
[0191] In some examples, an MV threshold is used to determine whether the additional unidirectional MV is equal to a candidate in the unidirectional MV candidate list. For example, in response to determining that a difference between the additional unidirectional MV and a candidate in the unidirectional MV candidate list is less than an MV threshold, the additional unidirectional MV is determined to be equal to the candidate in the unidirectional MV candidate list, where the MV threshold is a fixed threshold or a variable based on a block size of the video block.
[0192] Further, processor 920 may omit adding the pair average candidate to the first updated unidirectional MV candidate list in response to determining that the pair average candidate is equal to a candidate in the first updated unidirectional MV candidate list. In some examples, in response to determining that a difference between the pair average candidate and a candidate in the first updated unidirectional MV candidate list is less than an MV threshold, the pair average candidate is determined to be equal to the candidate, where the MV threshold is a fixed threshold or a variable based on a block size of the video block.
[0193] At step 806, the processor 920 may generate a unidirectional MV for the first geometric partition and a unidirectional MV for the second geometric partition.
[0194] In some examples, an apparatus for decoding a video block in a GPM is provided. The apparatus includes a processor 920 and a memory 940 configured to store instructions executable by the processor, the processor configured to, when executing the instructions, perform the method illustrated in FIG.
[0195] In some other examples, a non-transitory computer-readable storage medium having instructions stored thereon is provided that, when executed by the processor 920, cause the processor to perform the method illustrated in FIG.
[0196] Other examples of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the present disclosure disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, including departures from the present disclosure that are within known or customary practice in the art, in accordance with the general principles of the present disclosure. It is intended that the specification and examples be considered as illustrative only.
[0197] It will be understood that the present disclosure is not limited to the precise examples described above and illustrated in the accompanying drawings, and that various modifications and changes can be made thereto without departing from the scope of the present disclosure.
Claims
1. A method for decoding a video block in a geometric partitioning mode (GPM), comprising: partitioning the video block into first and second geometric partitions; constructing a unidirectional motion vector (MV) candidate list for the GPM by adding a plurality of normal merge candidates; in response to determining that the unidirectional MV candidate list is not complete, constructing a first updated unidirectional MV candidate list by adding one or more additional unidirectional MVs derived from one or more bidirectional prediction MVs in a normal merge candidate list to the unidirectional MV candidate list; in response to determining that the first updated unidirectional MV candidate list is not complete, constructing a second updated unidirectional MV candidate list by adding one or more pair average candidates to the first updated unidirectional MV candidate list; in response to determining that the second updated unidirectional MV candidate list is not complete, periodically adding zero unidirectional MVs to the second updated unidirectional MV candidate list until a maximum length is reached; generating a unidirectional MV for the first geometric partition and a unidirectional MV for the second geometric partition.
2. The method of claim 1, further comprising deriving the one or more additional unidirectional MVs by obtaining one or more MV candidates having odd merge indices in a first reference picture list and one or more MVs having even merge indices in a second reference picture list.
3. The method of claim 2, wherein the first reference picture list is reference picture L0 and the second reference picture list is L1.
4. The method of claim 2, wherein the first reference picture list is reference picture L1 and the second reference picture list is reference picture L0.
5. Adding the one or more pair average candidates to the first updated unidirectional MV candidate list comprises: obtaining first two unidirectional MV candidates in a first reference picture list or a second reference picture list; and in response to determining that the first two unidirectional MV candidates indicate the same reference picture, obtaining a pair average candidate by averaging the first two unidirectional MV candidates. The method of claim 1.
6. In response to determining that the first two unidirectional MV candidates indicate different reference pictures, determining the magnitude of the pair average candidate by averaging the first two unidirectional MV candidates, and determining the reference picture of the first unidirectional MV candidate as the reference picture of the pair average candidate, thereby obtaining the pair average candidate, further comprising the method according to claim 5.
7. In response to determining that the first two unidirectional MV candidates indicate different reference pictures, determining the magnitude of the pair average candidate by averaging the first two unidirectional MV candidates, and determining the reference picture of the second unidirectional MV candidate as the reference picture of the pair average candidate, thereby obtaining the pair average candidate, further comprising the method according to claim 5.
8. Further comprising removing redundant candidates from the unidirectional MV candidate list, the method according to claim 1.
9. In response to determining that an additional unidirectional MV is equal to a candidate in the unidirectional MV candidate list, further comprising omitting adding the additional unidirectional MV to the unidirectional MV candidate list, the method according to claim 8.
10. In response to determining that the difference between the additional unidirectional MV and the candidate in the unidirectional MV candidate list is less than an MV threshold, further comprising determining that the additional unidirectional MV is equal to the candidate in the unidirectional MV candidate list, wherein the MV threshold is a fixed threshold or a variable based on the block size of the video block, the method according to claim 9.
11. In response to determining that the pair average candidate is equal to a candidate in the first updated unidirectional MV candidate list, further comprising omitting adding the pair average candidate to the first updated unidirectional MV candidate list, the method according to claim 5.
12. In response to determining that the difference between the pair average candidate and the candidate in the first updated unidirectional MV candidate list is less than an MV threshold, further comprising determining that the pair average candidate is equal to the candidate, wherein the MV threshold is a fixed threshold or a variable based on the block size of the video block, the method according to claim 11.
13. An apparatus for video decoding, one or more processors, A non-transitory computer-readable storage medium configured to store instructions executable by the one or more processors, An apparatus, wherein when the one or more processors execute the instructions, the apparatus is configured to perform the method according to any one of claims 1 to 12.
14. A non-transitory computer-readable storage medium storing computer-executable instructions, wherein when the computer-executable instructions are executed by one or more computer processors, the one or more computer processors are caused to perform the method according to any one of claims 1 to 12.
15. A computer program product including a plurality of programs executed by an apparatus including one or more processors, wherein when the plurality of programs are executed by the one or more processors, the apparatus is caused to execute the method according to any one of claims 1 to 12.
16. A bitstream decoded by the method according to any one of claims 1 to 12.