Geometric Partitioning Mode by Motion Vector Improvement
The GPM-MVR method addresses the inefficiencies in the Geometric Partitioning Mode of existing video encoding standards by applying motion vector refinement and optimized signaling, resulting in improved coding efficiency and reduced overhead.
Patent Information
- Application Number
- JP2023579398
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-23
- Filing Date
- 2022-06-22
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-06-22
AI Technical Summary
Existing video encoding standards, such as VVC and AVS3, face challenges in achieving optimal coding efficiency for the Geometric Partitioning Mode (GPM) due to limitations in motion vector refinement and signaling overhead.
The proposed method, Geometric Partition Mode with Motion Vector Refinement (GPM-MVR), enhances the coding efficiency of GPM by applying motion refinement to each GPM partition and signaling this refinement using a modified syntax, allowing for more accurate motion vectors while minimizing signaling overhead.
GPM-MVR improves the coding efficiency of GPM by providing more accurate motion vectors for each partition, leading to reduced bitrates and preserved video quality, while also reducing signaling overhead.
Smart Images

Figure 0007699677000030 
Figure 0007699677000031 
Figure 0007699677000032
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 214,230, filed on June 23, 2021, the disclosure of which is hereby incorporated by reference in its entirety for all purposes.
[0002] This disclosure relates to video encoding and compression. More particularly, this disclosure relates to methods and apparatus for improving the coding efficiency of a Geometric Partitioning (GPM) mode, also known as an Angular Weight Prediction (AWP) mode.
Background Art
[0003] To compress video data, various video encoding techniques can be used. Video encoding is performed according to one or more video encoding standards. For example, today, some well-known video encoding standards include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC, also known as H.265 or MPEG-H Part 2), and Advanced Video Coding (AVC, also known as H.264 or MPEG-4 Part 10), which were jointly developed by ISO / IEC MPEG and ITU-T VECG. AOMedia Video 1 (AV1) was developed by the Alliance for Open Media (AOM) as a successor to the previous standard VP9. Audio Video Coding (AVS) refers to digital audio and digital video compression standards and is another series of video compression standards developed by the Audio and Video Coding Standard Workgroup of China. Most of the existing video encoding standards are built based on well-known hybrid video encoding frameworks, that is, using block-based prediction methods (e.g., inter prediction, intra prediction) to reduce the redundancy present in video images or sequences, and using transform coding to compress the energy of the prediction error. An important goal of video encoding techniques is to compress video data into a form that uses a lower bitrate while avoiding or minimizing the degradation of video quality.
Summary of the Invention
Problems to be Solved by the Invention
[0004] The present disclosure provides a method and apparatus for video encoding, as well as a non-transitory computer-readable storage medium.
Means for Solving the Problems
[0005] According to a first aspect of the present disclosure, a method for decoding a video block with GPM is provided. The method can include partitioning the video block into first and second geometric partitions. The method can include receiving a GPM (GPM-MVR) activation flag with a first motion vector refinement for the first geometric partition and receiving a second GPM-MVR activation flag for the second geometric partition. The method can include receiving a joint template matching (TM) activation flag for the first and second geometric partitions, where the joint TM activation flag can jointly indicate whether the unidirectional motion of the first partition is improved by TM and whether the unidirectional motion of the second partition is improved by TM. The method can include receiving a first merge GPM index for the first geometric partition and a second merge GPM index for the second geometric partition. The method can include constructing a unidirectional motion vector (MV) candidate list for GPM. The method can include generating a unidirectional MV for the first geometric partition and a unidirectional MV for the second geometric partition.
[0006] According to a second aspect of the present disclosure, an apparatus for video decoding is provided. The apparatus can include one or more processors and a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium is configured to store instructions executable by the one or more processors. The one or more processors are configured to implement the method of the first aspect when executing the instructions.
[0007] According to a third aspect of the present disclosure, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium can store computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to implement the method of the first aspect.
[0008] The accompanying drawings are incorporated herein and form a part of this specification, showing examples consistent with the present disclosure and serving to explain the principles of the present disclosure together with the description.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 3C
Figure 3D
Figure 3E
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 7
Figure 8
Figure 9
Figure 10
[0010] Embodiments illustrated by the accompanying drawings are next referred to in detail. The following description refers to the accompanying drawings, and the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementation examples described in the following description of the embodiments do not represent all implementation examples consistent with the present disclosure. Instead, these implementation examples are merely examples of apparatuses and methods consistent with aspects of the present disclosure described in the appended claims.
[0011] The terms used in the present disclosure are for the purpose of describing particular embodiments only and are not intended to limit the present disclosure. When used in the present disclosure and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein means any and all possible combinations of one or more of the associated listed items and is intended to encompass them.
[0012] For the purpose of explaining various information, terms such as "first," "second," "third," etc. may be used herein, but it should be understood that the information should not be limited by these terms. These terms are used only to distinguish one category of information from another. For example, without departing from the scope of the present disclosure, the first information may be referred to as the second information, and similarly, the second information may be referred to as the first information. As used herein, the term "if" is to be understood to mean "when" or "upon" or "in response to a judgment," depending on the context.
[0013] The first-generation AVS standard includes the Chinese national standard "Information Technology, Advanced Audio Video Coding, Part 2: Video" (also known as AVS1) and "Information Technology, Advanced Audio Video Coding Part 16: Radio Television Video" (also known as AVS+). It can provide about 50% bitrate savings at the same perceptual quality compared to the MPEG-2 standard. The video part of the AVS1 standard was published as a Chinese national standard in February 2006. The second-generation AVS standard includes a series of Chinese national standards "Information Technology, Efficient Multimedia Coding" (also known as AVS2), mainly targeted at the transmission of additional HD TV programs. The coding efficiency of AVS2 is twice that of AVS+. In May 2016, AVS2 was issued as a Chinese national standard. On the other hand, the video part of the AVS2 standard was submitted by the Institute of Electrical and Electronics Engineers (IEEE) as an international standard for application. The AVS3 standard is a new-generation video coding standard for the application of UHD video, aiming to exceed the coding efficiency of the latest international standard HEVC. In March 2019, at the 68th AVS meeting, the AVS3-P2 baseline was completed, which provides about 30% bitrate savings compared to the HEVC standard. Currently, to demonstrate the reference implementation of the AVS3 standard, a reference software called High Performance Model (HPM) is maintained by the AVS group.
[0014] Similar to HEVC, the AVS3 standard is constructed based on a block-based hybrid video coding framework.
[0015] FIG. 10 is a block diagram showing an exemplary system 10 for encoding and decoding video blocks in parallel according to some implementation examples of the present disclosure. As shown in FIG. 1, system 10 includes a source device 12 that generates video data and encodes it to be decoded later by a destination device 14. The source device 12 and the destination device 14 can each comprise any of a variety of electronic devices including desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, and the like. In some implementation examples, the source device 12 and the destination device 14 have wireless communication capabilities.
[0016] In some implementation examples, the destination device 14 can receive the encoded video data to be decoded via link 16. Link 16 can comprise any type of communication medium or device capable of moving the encoded video data from the source device 12 to the destination device 14. In one example, link 16 can comprise a communication medium for enabling the source device 12 to transmit the encoded video data directly to the destination device 14 in real time. The encoded video data can be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device 14. The communication medium can include any wireless or wired communication medium such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or any other device that can be useful in facilitating communication from the source device 12 to the destination device 14.
[0017] In some other implementation examples, the encoded video data can be transmitted from the output interface 22 to the storage device 32. Next, the encoded video data in the storage device 32 can be accessed by the destination device 14 via the input interface 28. The storage device 32 can include any of a variety of distributed or local access type data storage media, such as a hard drive, Blu-ray disk, digital versatile disk (DVD), compact disk read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data. In a further example, the storage device 32 can correspond to a file server or another intermediate storage device that can hold the encoded video data generated by the source device 12. The destination device 14 can access the video data stored in the storage device 32 via streaming or downloading. The file server can be any type of computer capable of storing the encoded video data and transmitting the encoded video data to the destination device 14. Exemplary file servers can include a web server (e.g., for a website), a file transfer protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive. The destination device 14 can access the encoded video data via any standard data connection, including a wireless channel (e.g., a wireless fidelity (Wi-Fi) connection) suitable for accessing the encoded video data stored on the file server, a wired connection (e.g., a digital subscriber line (DSL), cable modem, etc.), or a combination of both. The transmission of the encoded video data from the storage device 32 can be a streaming transmission, a download transmission, or a combination of both.
[0018] As shown in FIG. 10, source device 12 includes video source 18, video encoder 20, and output interface 22. Video source 18 can include sources such as a video capture device, e.g., a video camera, a video archive containing previously captured video, a video supply interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or a combination of such sources. As an example, when video source 18 is a video camera of a security monitoring system, source device 12 and destination device 14 can form a camera phone or a video phone. However, the implementation examples described in this application can generally be applicable to video encoding and can be applied to wireless and / or wired applications.
[0019] Captured, pre-captured, or computer-generated video can be encoded by video encoder 20. The encoded video data can be transmitted directly to destination device 14 via output interface 22 of source device 12. The encoded video data can also (or alternatively) be stored on storage device 32 for later access by destination device 14 or other devices for decoding and / or playback. Output interface 22 can further include a modem and / or a transmitter.
[0020] The destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. The input interface 28 can include a receiver and / or a modem and can receive encoded video data via link 16. The encoded video data communicated via link 16 or provided on the storage device 32 can include various syntax elements generated by the video encoder 20 for use by the video decoder 30 in decoding the video data. Such syntax elements can be included within the encoded video data transmitted over a communication medium, stored on a storage medium, or stored on a file server.
[0021] In some implementations, the destination device 14 can include a display device 34, which can be an integrated display device and an external display device configured to communicate with the destination device 14. The display device 34 is for displaying the decoded video data to the user and can comprise any of various display devices such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0022] The video encoder 20 and the video decoder 30 can operate according to proprietary or industry standards such as VVC, HEVC, MPEG-4, Part 10, AVC, or extensions of such standards. It should be understood that the present application is not limited to a particular video encoding / decoding standard and may be applicable to other video encoding / decoding standards. Generally, it is contemplated that the video encoder 20 of the source device 12 can be configured to encode video data according to any of these current or future standards. Similarly, it is generally contemplated that the video decoder 30 of the destination device 14 can be configured to decode video data according to any of these current or future standards.
[0023] Video encoder 20 and video decoder 30 can each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When implemented partially in software, the electronic device can store instructions for the software in a suitable non-transitory computer-readable medium and use one or more processors to execute those instructions in hardware to perform the video encoding / decoding operations disclosed in this disclosure. Each of video encoder 20 and video decoder 30 may be included within one or more encoders or decoders, all of which may be integrated as part of a combined encoder / decoder (CODEC) in their respective devices.
[0024] FIG. 1 shows a schematic diagram of a block-based video encoder for VVC. Specifically, FIG. 1 shows a typical encoder 100. Encoder 100 can be the video encoder 20 shown in FIG. 10. Encoder 100 has a video input 110, motion compensation 112, motion estimation 114, intra / inter mode decision 116, block predictor 140, adder 128, transform 130, quantization 132, prediction related information 142, intra prediction 118, picture buffer 120, inverse quantization 134, inverse transform 136, adder 126, memory 124, in-loop filter 122, entropy encoding 138, and bitstream 144.
[0025] Within encoder 100, a video frame is partitioned into a plurality of video blocks for processing. For each given video block, a prediction is formed based on an inter prediction technique or an intra prediction technique.
[0026] The prediction residual representing the difference between the current video block, a part of the video input 110 and its predictor, and a part of the block predictor 140 is transmitted from the adder 128 to the transform 130. Then the transform coefficients are transmitted from the transform 130 to the quantization 132 for entropy reduction. Then the quantized coefficients are sent to the entropy coding 138 to generate a compressed video bit stream. As shown in FIG. 1, prediction related information 142 from the intra / inter mode decision 116, such as video block partition information, motion vector (MV), reference picture index, and intra prediction mode, is also sent by the entropy coding 138 and stored in the compressed bit stream 144. The compressed bit stream 144 includes the video bit stream.
[0027] Within the encoder 100, decoder related circuitry is also required to reconstruct pixels for prediction purposes. First, the prediction residual is reconstructed by the inverse quantization 134 and the inverse transform 136. This reconstructed prediction residual is combined with the block predictor 140 to generate the unfiltered reconstructed pixels for the current video block.
[0028] Spatial prediction (or "intra prediction") uses pixels from samples of already coded adjacent blocks (referred to as reference samples) within the same video frame as the current video block to predict the current video block.
[0029] Temporal prediction (also referred to as "inter prediction") predicts the current video block using the reconstructed pixels from already-coded video pictures. Temporal prediction reduces the temporal redundancy inherent in the video signal. Typically, the temporal prediction signal for a given coding unit (CU) or coding block is signaled by one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and its temporal reference. Further, when multiple reference pictures are involved, one reference picture index is additionally transmitted and used to identify from which reference picture in the reference picture store the temporal prediction signal originated.
[0030] Motion estimation 114 takes in signals from video input 110 and picture buffer 120 and outputs a motion estimation signal to motion compensation 112. Motion compensation 112 takes in signals from video input 110, picture buffer 120, and the motion estimation signal from motion estimation 114 and outputs a motion compensation signal to intra / inter mode decision 116.
[0031] After spatial and / or temporal prediction is performed, the intra / inter mode decision 116 within the encoder 100 selects the best prediction mode, for example, based on a rate distortion optimization method. Next, the block predictor 140 is derived from the current video block, and the resulting prediction residual is decorrelated using the transform 130 and quantization 132. The resulting quantized residual coefficients are inverse quantized by the inverse quantization 134 and inverse transformed by the inverse transform 136 to form a reconstructed residual, and then the reconstructed residual is added back to the predicted block to form the reconstructed signal of the CU. Further, in-loop filtering 122 such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive loop filter (ALF) may be applied to the reconstructed CU, and then placed in the reference picture store of the picture buffer 120 and used to code future video blocks. To form the output video bitstream 144, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit 138, further compressed, and packed to form a bitstream.
[0032] Figure 1 gives a block diagram of a general-purpose block-based hybrid video coding system. The input video signal is processed block by block (referred to as a coding unit (CU)). Different from HEVC which partitions blocks based only on quad-trees, in AVS3, one coding tree unit (CTU) is divided into CUs to adapt to the varying local characteristics based on quad-trees / binary trees / extended quad-trees. In addition, the concept of multiple partition unit types in HEVC is removed, that is, in AVS3, there is no separation between CUs, prediction units (PUs), and transform units (TUs). Instead, each CU is always used as the basic unit for both prediction and transform without further partitioning. In the tree partitioning structure of AVS3, one CTU is first partitioned based on a quad-tree structure. Then each quad-tree leaf node may be further partitioned based on binary tree and extended quad-tree structures.
[0033] As shown in FIGS. 3A, 3B, 3C, 3D, and 3E, there are five partitioning types, namely, four-division, horizontal two-division, vertical two-division, horizontally extended quadtree division, and vertically extended quadtree division.
[0034] FIG. 3A shows a diagram showing a four-division of a block in a multi-type tree structure according to the present disclosure.
[0035] FIG. 3B shows a diagram showing a vertical two-division of a block in a multi-type tree structure according to the present disclosure.
[0036] FIG. 3C shows a diagram showing a horizontal two-division of a block in a multi-type tree structure according to the present disclosure.
[0037] FIG. 3D shows a diagram showing a vertical three-division of a block in a multi-type tree structure according to the present disclosure.
[0038] FIG. 3E shows a diagram showing a horizontal three-division of a block in a multi-type tree structure according to the present disclosure.
[0039] In FIG. 1, spatial prediction and / or temporal prediction can be performed. Spatial prediction (or "intra prediction") uses pixels from samples of already-coded adjacent blocks (referred to as reference samples) within the same video picture / slice to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also called "inter prediction" or "motion-compensated prediction") uses reconstructed pixels from already-coded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. Usually, the temporal prediction signal for a given CU is signaled by one or more motion vectors (MVs) that indicate the amount and direction of motion between the current CU and its temporal reference. Also, when multiple reference pictures are involved, one reference picture index is further transmitted and used to identify from which reference picture in the reference picture store the temporal prediction signal originated. After spatial and / or temporal prediction, a mode decision block in the encoder selects the best prediction mode, for example, based on a rate-distortion optimization method. Then the prediction block is subtracted from the current video block, and the prediction residual is decorrelated using a transform and then quantized. The quantized residual coefficients are inverse quantized and inverse transformed to form a reconstructed residual, and then the reconstructed residual is added back to the prediction block to form the reconstructed signal of the CU. Further, in-loop filtering such as deblocking filter, sample adaptive offset (SAO), and adaptive loop filter (ALF) may be applied to the reconstructed CU, and then it is placed in the reference picture store and used as a reference for coding future video blocks. To form the output video bitstream, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to an entropy coding unit and further compressed and packed.
[0040] Figure 2 shows a schematic block diagram of a video decoder for VVC. Specifically, Figure 2 shows a block diagram of a typical decoder 200. The block-based video decoder 200 can be the video decoder 30 shown in Figure 10. The decoder 200 has a bitstream 210, entropy decoding 212, inverse quantization 214, inverse transform 216, an adder 218, intra / inter mode selection 220, intra prediction 222, a memory 230, in-loop filter 228, motion compensation 224, a picture buffer 226, prediction-related information 234, and a video output 232.
[0041] The decoder 200 is similar to the reconstruction-related part existing in the encoder 100 of Figure 1. In the decoder 200, the incoming video bitstream 210 is first decoded by the entropy decoding 212 to derive the quantized coefficient levels and prediction-related information. Then, the quantized coefficient levels are processed by the inverse quantization 214 and the inverse transform 216 to obtain the reconstructed prediction residuals. A block prediction mechanism is implemented within the intra / inter mode selection unit 220 and is configured to perform intra prediction 222 or motion compensation 224 based on the decoded prediction information. A set of unfiltered reconstructed pixels is obtained by using the adder 218 to sum the reconstructed prediction residuals from the inverse transform 216 and the prediction output generated by the block prediction mechanism.
[0042] The reconstructed blocks can further pass through the in-loop filter 228 and are then stored in the picture buffer 226 that functions as a reference picture store. The reconstructed video in the picture buffer 226 may be transmitted to drive a display device and may also be used to predict future video blocks. With the in-loop filter 228 turned on, a filtering operation is performed on these reconstructed pixels to derive the final reconstructed video output 232.
[0043] Figure 2 gives a schematic block diagram of a block-based video decoder. First, a video bitstream is entropy decoded by an entropy decoding unit. Coding mode and prediction information are sent to a spatial prediction unit (when intra-coded) or a temporal prediction unit (when inter-coded) to form a prediction block. Residual transform coefficients are sent to an inverse quantization unit and an inverse transform unit to reconstruct a residual block. Then, the prediction block and the residual block are added together. The reconstructed block can further pass through an in-loop filter and is then stored in a reference picture buffer. Next, the reconstructed video in the reference picture buffer is sent for display and is also used to predict future video blocks.
[0044] The focus of this disclosure is to improve the coding performance of the geometric partitioning mode (GPM) used in both the VVC standard and the AVS3 standard. In AVS3, the tool is also known as angular weighted prediction (AWP), which follows the same design concept as GPM but has some minor differences in specific design details. To facilitate the description of this disclosure, hereinafter, the existing GPM design in the VVC standard is used as an example to explain the main aspects of the GPM / AWP tool. On the other hand, another existing inter-prediction technique called merge mode with motion vector difference (MMVD) applied in both the VVC standard and the AVS3 standard is also briefly considered, provided that it is closely related to the technique proposed in this disclosure. Then, some drawbacks of the current GPM / AWP design are identified. Finally, the proposed method is provided in detail. Throughout this disclosure, the existing GPM design in the VVC standard is used as an example, but those skilled in the art of modern video coding technology should note that the proposed technique can also be applied to other GPM / AWP designs or other coding tools with the same or similar design concept.
[0045] Geometric Partitioning Mode (GPM) In VVC, a geometric partitioning mode is supported for inter prediction. The geometric partitioning mode is signaled as one special merge mode by one CU level flag. In the current GPM design, for each possible CU size where both width and height are 8 or more and 64 or less, except for 8×64 and 64×8, a total of 64 partitions are supported by the GPM mode.
[0046] When this mode is used, the CU is geometrically divided into two parts by a line located as shown in Fig. 4 (the explanation will be given later). The location of the dividing line is mathematically derived from the angles and offset parameters of the specific partition. Each part of the geometric partition within the CU is inter predicted using its own motion, and only uni - directional prediction is allowed for each partition, that is, each part has one motion vector and one reference index. Similar to the conventional bi - directional prediction, motion constraints by uni - directional prediction are applied to ensure that only two motion - compensated predictions are required for each CU. When the geometric partitioning mode is used for the current CU, a geometric partition index indicating the partition mode (angle and offset) of the geometric partition, and two merge indexes (one for each partition) are further signaled. The number of maximum GPM candidate sizes is explicitly signaled at the sequence level.
[0047] Fig. 4 shows the allowed GPM partitions, and the division within each picture has one identical division direction.
[0048] Uni - directional prediction candidate list structure To derive the uni - directional prediction motion vector for one geometric partition, first, one uni - directional prediction candidate list is directly derived from the normal merge candidate list generation process. Let n denote the index of the uni - directional prediction motion in the geometric uni - directional prediction candidate list. The LX motion vector of the nth merge candidate is used as the uni - directional prediction motion vector for the geometric partitioning mode, where X is equal to the parity of n.
[0049] These motion vectors are indicated by "x" in FIG. 5 (described later). If there is no corresponding LX motion vector for the n-th extended merge candidate, the L(1-X) motion vector of the same candidate is used instead as the unidirectional prediction motion vector for the geometric partitioning mode.
[0050] FIG. 5 shows the selection of the unidirectional prediction motion vector from the motion vectors of the merge candidate list for GPM.
[0051] Mixing along the geometric partitioning edge After each geometric partition is obtained using its own motion, mixing is applied to two unidirectional prediction signals to derive samples around the geometric partition edge. The mixing weights for each position of the CU are derived based on the distance from each individual sample position to the corresponding partition edge.
[0052] GPM signaling design According to the current GPM design, the use of GPM is indicated by signaling one flag at the CU level. This flag is signaled only when the current CU is coded in merge mode or skip mode. Specifically, when this flag is equal to 1, it indicates that the current CU is predicted by GPM. Otherwise (the flag is equal to 0), the CU is coded by another merge mode such as normal merge mode, merge mode with motion vector difference, combination of inter and intra prediction, etc. When GPM is enabled for the current CU, one syntax element, namely merge_gpm_partition_idx, is further signaled to indicate the applied geometric partition mode (specifying the direction and offset of the straight line from the CU center that divides the CU into two partitions as shown in Figure 4). Then, two syntax elements merge_gpm_idx0 and merge_gpm_idx1 are signaled to indicate the indexes of the uni-directional prediction merge candidates used for the first and second GPM partitions. More specifically, these two syntax elements are used to determine the uni-directional MVs of the two GPM partitions from the uni-directional prediction merge list described in the chapter of "Uni-directional Prediction Merge List Structure". According to the current GPM design, in order to make the two uni-directional MVs more different, the two indexes cannot be made the same. Based on such prior knowledge, first, the uni-directional prediction merge index of the first GPM partition is signaled and used as a prediction unit to reduce the signaling overhead of the uni-directional prediction merge index of the second GPM partition. Specifically, when the second uni-directional prediction merge index is smaller than the first uni-directional prediction merge index, its original value is directly signaled. Otherwise (the second uni-directional prediction merge index is larger than the first uni-directional prediction merge index), after 1 is subtracted from its value, it is signaled to the bitstream. On the decoder side, first, the first uni-directional prediction merge index is the decoder.Next, for decoding the second unidirectional prediction merge index, if the parsed value is smaller than the first unidirectional prediction merge index, the second unidirectional prediction merge index is set equal to the parsed value; otherwise (if the parsed value is equal to or greater than the first unidirectional prediction merge index), the second unidirectional prediction merge index is set equal to the value obtained by adding 1 to the parsed value. Table 1 shows the existing syntax elements used in the GPM mode in the current VVC specification.
[0053]
Table 1
[0054] On the other hand, in the current GPM design, for binarization of two unidirectional prediction merge indexes, namely merge_gpm_idx0 and merge_gpm_idx1, a truncated unary code is used. In addition, since the two unidirectional prediction merge indexes cannot be made the same, different maximum values are used to shorten the codewords of the two unidirectional prediction merge indexes, and the two unidirectional prediction merge indexes are set equal to MaxGPMMergeCand - 1 and MaxGPMMergeCand - 2 for merge_gpm_idx0 and merge_gpm_idx1 respectively. MaxGPMMergeCand is the number of candidates in the unidirectional prediction merge list.
[0055] When the GPM / AWP mode is applied, two different binarization methods are applied to translate the syntax merge_gpm_partition_idx into a 2-bit string. Specifically, syntax elements are binarized by fixed-length codes and shortened binary codes in the VVC standard and the AVS3 standard, respectively. On the other hand, in the case of the AWP mode in AVS3, different maximum values are used for the binarization of the values of syntax elements. Specifically, in AVS3, the number of permitted GPM / AWP partition modes is 56 (i.e., the maximum value of merge_gpm_partition_idx is 55), and in VVC, the number is increased to 64 (i.e., the maximum value of merge_gpm_partition_idx is 63).
[0056] Merge mode with motion vector difference (MMVD) In addition to the conventional merge mode that derives the motion information of one current block from its spatial / temporal neighbors, the MMVD / UMVE mode is introduced as one special merge mode in both the VVC standard and the AVS standard. Specifically, in both VVC and AVS3, the mode is signaled by one MMVD flag at the coding block level. In the MMVD mode, the first two candidates in the merge list for the normal merge mode are selected as the two basic merge candidates for MMVD. After one basic merge candidate is selected and signaled, additional syntax elements are signaled to indicate the motion vector difference (MVD) that is added to the motion of the selected merge candidate. The MMVD syntax elements include a merge candidate flag for selecting the basic merge candidate, a distance index for specifying the magnitude of the MVD, and a direction index for indicating the direction of the MVD.
[0057] In the existing MMVD design, the distance index specifies the magnitude of the MVD defined based on a set of predefined offsets from the starting point. As shown in FIGS. 6A and 6B, the offsets are added to the horizontal or vertical component of the starting MV (i.e., the MV of the selected basic merge candidate).
[0058] Figure 6A shows the MMVD mode with respect to the L0 reference. Figure 6B shows the MMVD mode with respect to the L1 reference.
[0059] Table 2 shows the MVD offsets applied in AVS3 respectively.
[0060]
Table 2
[0061] As shown in Table 3, the direction index is used to specify the sign of the signaled MVD. Note that the meaning of the MVD sign may vary according to the starting MV. When the starting MV is a unidirectional prediction MV or a bidirectional prediction MV, the MV points to two reference pictures, and the POCs of both are greater than the POC of the current picture or both are less than the POC of the current picture, the signaled sign is the sign of the MVD added to the starting MV. When the starting MV is a bidirectional prediction MV that points to two reference pictures, the POC of one picture is greater than the current picture, and the POC of the other picture is less than the current picture, the signaled sign is applied to the L0 MVD, and the inverse value of the signaled sign is applied to the L1 MVD.
[0062]
Table 3
[0063] Motion Signaling for Normal Inter-Mode Similar to the HEVC standard, in addition to the merge / skip mode, in both VVC and AVS3, it is permitted that one inter-CU explicitly specifies its motion information in the bitstream. Overall, the signaling of motion information in both VVC and AVS3 remains the same as that of the HEVC standard. Specifically, first, one inter-prediction syntax, i.e., inter_pred_idc, is signaled to indicate whether the prediction signal is from list L0, L1, or both. For each reference list used, the corresponding reference picture is identified by signaling one reference picture index ref_idx_lx (x = 0, 1) for the corresponding reference list, and the corresponding MV is represented by one MVP index mvp_lx_flag (x = 0, 1) used to select the MV prediction unit (MVP) and then the motion vector difference (MVD) between the target MV and the selected MVP. In addition, in the VVC standard, one control flag mvd_l1_zero_flag is signaled at the slice level. When mvd_l1_zero_flag is equal to 0, the L1 MVD is signaled in the bitstream; otherwise (when the mvd_l1_zero_flag flag is equal to 1), the L1 MVD is not signaled and its value is always inferred as 0 at the encoder and decoder.
[0064] Bi-directional prediction by CU-level weighting In standards prior to VVC and AVS3, when weighted prediction (WP) is not applied, the bi-directional prediction signal is generated by averaging the unidirectional prediction signals obtained from two reference pictures. In VVC, to improve the efficiency of bi-directional prediction, one tool coding, i.e., bi-directional prediction by CU-level weighting (BCW), is introduced. Specifically, instead of simple averaging, the bi-directional prediction of BCW is extended as shown below by allowing the weighted average of two prediction signals. P’(i,j)=((8 - w)·P0(i,j)+w·P1(i,j)+4)≫3
[0065] In VVC, when the current picture is a low-latency picture, the weight of one BCW coding block is permitted to be selected from a set of predefined weight values w ∈ {-2, 3, 4, 5, 10}, and weight 4 represents the conventional bi-directional prediction case where two uni-directional prediction signals are equally weighted. In the case of low latency, only three weights w ∈ {3, 4, 5} are permitted. Generally, there are some design similarities between WP and BCW, but the two coding tools are targeted at solving the problem of illumination change at different granularities. However, the interaction between WP and BCW may complicate the VVC design in some cases, so it is not allowed for the two tools to be enabled simultaneously. Specifically, when WP is enabled for a slice, the BCW weight for all bi-directional prediction CUs within the slice is not signaled and is inferred to be 4 (i.e., equal weights are applied).
[0066] Template matching Template matching (TM) is a decoder-side MV derivation method for improving the motion information of the current CU by finding the best match between a template consisting of the reconstructed samples above and to the left of the current CU and a reference block (i.e., the same size as the template) in the reference picture. As shown in Figure 7, within the search range of [-8, +8] pels, one MV is searched around the initial motion vector of the current CU. The best match can be defined as the MV that achieves the lowest matching cost, such as sum of absolute differences (SAD), sum of absolute transformed differences (SATD), etc., between the current template and the reference template. There are two different methods for applying the TM mode to inter-coding.
[0067] In the AMVP mode, based on the template matching difference, the MVP candidate is determined to select the one that reaches the minimum difference between the template of the current block and the template of the reference block. Then, the TM is only performed on this specific MVP candidate for MV improvement. By using an iterative diamond search, the TM improves this MVP candidate from 1-pel MVD accuracy (or 4 pels in the case of 4-pel AMVR mode) within the [-8, +8] search range. The AMVP candidate can be further improved by using a cross search with 1-pel MVD accuracy (or 4 pels in the case of 4-pel AMVR mode) according to the AMVR mode specified in Table 13 below, followed by sequentially using 1 / 2-pel and 1 / 4-pel ones. This search process ensures that the MVP candidate still maintains the same MV accuracy as indicated by the AMVR mode after the TM process.
[0068]
Table 4
[0069] In the merge mode, a similar search method is applied to the merge candidate indicated by the merge index. As shown in the above table, the TM can perform all up to 1 / 8-pel MVD accuracy or skip the part behind 1 / 2-pel MVD accuracy depending on whether an alternative interpolation filter (used when AMVR is in 1 / 2-pel mode) is used according to the merged motion information.
[0070] As described above, the unidirectional motion used to generate the predicted samples of two GPM partitions is directly obtained from the normal merge candidates. If there is no strong correlation between the MVs of the spatial / temporal adjacent blocks, the derived unidirectional MVs from the merge candidates may not be accurate enough to capture the true motion of each GPM partition. Motion estimation can provide more accurate motion, but at the expense of signaling overhead that cannot be ignored by any motion refinement that can be applied on top of the existing unidirectional MVs. On the other hand, the MVMD mode has been proven to be an efficient signaling mechanism for reducing the MVD signaling overhead and is used in both the VVC standard and the AVS3 standard. Therefore, it may also be beneficial to combine GPM with the MMVD mode. Such a combination can, in some cases, improve the overall coding efficiency of the GPM tool by providing more accurate MVs to capture the individual motion of each GPM partition.
[0071] As discussed earlier, in both the VVC standard and the AVS3 standard, the GPM mode is only applied to the merge / skip modes. Such a design may not be optimal in terms of coding efficiency considering that not all non-merge inter CUs can benefit from the flexible non-square partitions of GPM. On the other hand, for the same reasons as described above, the unidirectional prediction motion candidates derived from the normal merge / skip modes are not always accurate enough to capture the true motion of the two geometric partitions. Based on such an analysis, a proper extension of the GPM mode to the non-merge inter mode (i.e., the CU that explicitly signals the motion information in the bitstream) may expect additional coding gains. However, the improvement in MV accuracy comes at the expense of an increased signaling overhead. Therefore, it should be important to identify an effective signaling method that can minimize the signaling cost while providing more accurate MVs for the two geometric partitions in order to efficiently apply the GPM mode to the explicit inter mode.
[0072] Proposed method In the present disclosure, a method is proposed to further improve the coding efficiency of GPM by applying further motion improvements on top of the existing unidirectional MVs applied to each GPM partition. The proposed method is referred to as Geometric Partition Mode with Motion Vector Refinement (GPM-MVR). In addition, in the proposed scheme, motion refinement is signaled in a similar way to one of the existing MMVD designs, i.e., based on a set of predefined MVD sizes and directions of motion refinement.
[0073] In another aspect of the present disclosure, a solution is provided to extend the GPM mode to an explicit inter-mode. For ease of explanation, these schemes are referred to as Geometric Partition Mode with Explicit Motion Signaling (GPM-EMS). Specifically, in the proposed GPM-EMS scheme, the existing motion signaling mechanisms, i.e., MVP and MVD, are utilized to specify the corresponding unidirectional MVs of two geometric partitions in order to achieve better harmony with the normal inter-mode.
[0074] Geometric partition mode with separate motion vector refinement To improve the coding efficiency of GPM, in this chapter, one improved geometric partitioning mode with separate motion vector refinement is proposed. Specifically, considering the GPM partition, the proposed method first uses the existing syntax merge_gpm_idx0 and merge_gpm_idx1 to identify the uni-directional MVs for two GPM partitions from the existing uni-directional prediction merge candidate list and uses these as the base MVs. After two base MVs are determined, two new sets of syntax elements are introduced to separately specify the values of the motion refinement applied on the base MVs of the two GPM partitions. Specifically, first, two flags, namely gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag, are signaled to indicate whether GPM-MVR is applied to the first and second GPM partitions, respectively. When the flag of one GPM partition is equal to 1, the corresponding value of the MVR applied to the base MV of that partition is signaled in MMVD format, i.e., one distance index (indicated by the syntax elements gpm_mvr_partIdx0_distance_idx and gpm_mvr_partIdx1_distance_idx) specifies the magnitude of the MVR, and one direction index (indicated by the syntax elements gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx1_distance_idx) specifies the direction of the MVR. Table 4 shows the syntax elements introduced by the proposed GPM-MVR method.
[0075]
Table 5
[0076] Based on the proposed syntax elements shown in Table 4, in the decoder, the last MV used to generate the unidirectional prediction samples for each GPM segment is equal to the sum of the signaled motion vector refinement and the corresponding base MV. In practice, different sets of MVR sizes and directions may be pre-defined and applied to the proposed GPM-MVR scheme, which can provide various trade-offs between motion vector accuracy and signaling overhead. In one particular example, it is proposed to reuse the eight MVD offsets (i.e., 1 / 4, 1 / 2, 1, 2, 4, 8, 16, and 32 pels) and four MVD directions (i.e., ±x and y axes) used in the VVC standard for the proposed GPM-MVR scheme. In another example, the existing five MVD offsets {1 / 4, 1 / 2, 1, 2, and 4 pels} and four MVD directions (i.e., ±x and y axes) used in the AVS3 standard are applied in the proposed GPM-MVR scheme.
[0077] As discussed in the "GPM Signaling Design" chapter, since the unidirectional MVs used for two GPM segments cannot be the same, in the existing GPM design, one constraint is applied to make the two unidirectional prediction merge indexes different. However, in the proposed GPM-MVR scheme, further motion refinement is applied on top of the existing GPM unidirectional MVs. Therefore, even when the base MVs of two GPM segments are the same, the last unidirectional MVs used to predict the two segments will still be different unless the values of the two motion vector refinements are the same. Based on the above considerations, this constraint (limiting the two unidirectional prediction merge indexes to be different) is removed when the proposed GPM-MVR scheme is applied. In addition, since the two unidirectional prediction merge indexes are allowed to be the same, the same maximum value MaxGPMMergeCand - 1 is used for the binarization of both merg_gpm_idx0 and merge_gpm_idx1, where MaxGPMMergeCand is the number of candidates in the unidirectional prediction merge list.
[0078] As analyzed above, when the unidirectional prediction merge indexes of two GPM partitions (i.e., merge_gpm_idx0 and merge_gpm_idx1) are the same, the values of the two motion vector refinements cannot be made the same to ensure that the last MVs used for the two partitions are different. Based on such conditions, in one embodiment of the present disclosure, when the unidirectional prediction merge indexes of two GPM partitions are the same (i.e., merge_gpm_idx0 is equal to merge_gpm_idx1), one signaling redundancy removal method is proposed to reduce the signaling overhead of the MVR of the second GPM partition using the MVR of the first GPM partition. In one example, the following signaling conditions are applied:
[0079] First, when the flag gpm_mvr_partIdx0_enable_flag is equal to 0 (i.e., GPM-MVR is not applied to the first GPM partition), the flag of gpm_mvr_partIdx1_enable_flag is not signaled but is inferred to be 1 (i.e., GPM-MVR is applied to the second GPM partition).
[0080] Second, when both flags gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 1 (i.e., GPM-MVR is applied to two GPM partitions), and gpm_mvr_partIdx0_direction_idx is equal to gpm_mvr_partIdx1_direction_idx (i.e., the MVRs of the two GPM partitions have the same direction), the magnitude of the MVR of the first GPM partition (i.e., gpm_mvr_partIdx0_distance_idx) is used to predict the magnitude of the MVR of the second GPM partition (i.e., gpm_mvr_partIdx1_distance_idx). Specifically, if gpm_mvr_partIdx1_distance_idx is smaller than gpm_mvr_partIdx0_distance_idx, its original value is directly signaled. Otherwise (gpm_mvr_partIdx1_distance_idx is larger than gpm_mvr_partIdx0_distance_idx), 1 is subtracted from the value and signaled to the bitstream. On the decoder side, to decode the value of gpm_mvr_partIdx1_distance_idx, if the parsed value is smaller than gpm_mvr_partIdx0_distance_idx, gpm_mvr_partIdx1_distance_idx is set equal to the parsed value, and otherwise (the parsed value is equal to or larger than gpm_mvr_partIdx0_distance_idx), gpm_mvr_partIdx1_distance_idx is set equal to the parsed value plus 1. In such a case, to further reduce the overhead, different maximum values MaxGPMMVRDistance - 1 and MaxGPMMVRDistance - 2 may be used for the binarization of gpm_mvr_partIdx0_distance_idx and gpm_mvr_partIdx1_distance_idx, where MaxGPMMVRDistance is the number of allowed magnitudes for motion vector refinement.
[0081] In another embodiment, it is proposed to switch the signaling order to gpm_mvr_partIdx0_direction_idx / gpm_mvr_partIdx1_direction_idx and gpm_mvr_partIdx0_distance_idx / gpm_mvr_partIdx1_distance_idx so that the size of the MVR is signaled before the size of the MVR. Accordingly, following the same logic as above, the encoder / decoder can adjust the signaling of the MVR direction of the second GPM section using the MVR direction of the first GPM section. In another embodiment, it is proposed to first signal the size and direction of the MVR of the second GPM section and use these to adjust the signaling of the size and direction of the MVR of the second GPM section.
[0082] In another embodiment, it is proposed to signal the syntax elements related to GPM-MVR before signaling the existing GPM syntax elements. Specifically, in such a design, first, two flags, gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag, are signaled to indicate whether GPM-MVR is applied to the first and second GPM sections, respectively. When the flag of one GPM section is equal to 1, the distance index (indicated by the syntax elements gpm_mvr_partIdx0_distance_idx and gpm_mvr_partIdx1_distance_idx) and the direction index (indicated by the syntax elements gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx1_distance_idx) specify the direction of the MVR. Then, the existing syntax elements merge_gpm_idx0 and merge_gpm_idx1 are signaled to identify the unidirectional MV, i.e., the base MV, for the two GPM sections. Table 5 shows the proposed GPM-MVR signaling method.
[0083] [Table 6]
[0084] Similar to the signaling method in Table 4, when the GPM-MVR signaling method in Table 5 is applied, specific conditions may be applied to ensure that the resulting MVs used for prediction of two GPM sections are not the same. Specifically, depending on the values of MVR applied to the first and second GPM sections, the following conditions for suppressing the signaling of the unidirectional prediction merge indices merge_gpm_idx0 and merge_gpm_idx1 are proposed.
[0085] First, when the values of both gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 0 (i.e., GPM-MVR is disabled for both of the two GPM sections), the values of merge_gpm_idx0 and merge_gpm_idx1 cannot be made the same.
[0086] Second, when gpm_mvr_partIdx0_enable_flag is equal to 1 (i.e., GPM-MVR is enabled for the first GPM section) and gpm_mvr_partIdx1_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for the second GPM section), the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same.
[0087] Third, when gpm_mvr_partIdx0_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for the first GPM section) and gpm_mvr_partIdx1_enable_flag is equal to 1 (i.e., GPM-MVR is enabled for the second GPM section), the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same.
[0088] Fourth, when the values of both gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 1 (i.e., GPM-MVR is enabled for both of the two GPM partitions), the determination of whether the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same depends on the values of MVR applied to the two GPM partitions (indicated by gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx0_distance_idx, as well as gpm_mvr_partIdx1_direction_idx and gpm_mvr_partIdx1_distance_idx). If the values of the two MVRs are equal, it is not permitted for merge_gpm_idx0 and merge_gpm_idx1 to be the same. Otherwise (if the values of the two MVRs are not equal), the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same.
[0089] In the above four cases, when the values of merge_gpm_idx0 and merge_gpm_idx1 are not allowed to be the same, the index value of one section can be used as a predictor of the index value of another section. In one method, it is proposed to first signal merge_gpm_idx0 and use its value to predict merge_gpm_idx1. Specifically, in the encoder, when merge_gpm_idx1 is greater than merge_gpm_idx0, the value of merge_gpm_idx1 transmitted to the decoder is reduced by 1. In the decoder, when the received value of merge_gpm_idx1 is equal to or greater than the received value of merge_gpm_idx0, the value of merge_gpm_idx1 is increased by 1. In another method, it is proposed to first signal merge_gpm_idx1 and use its value to predict merge_gpm_idx0. Thus, in such a case, in the encoder, when merge_gpm_idx0 is greater than merge_gpm_idx1, the value of merge_gpm_idx0 transmitted to the decoder is reduced by 1. In the decoder, when the received value of merge_gpm_idx0 is equal to or greater than the received value of merge_gpm_idx1, the value of merge_gpm_idx0 is increased by 1. In addition, similar to the existing GPM signaling design, different maximum values MaxGPMMergeCand-1 and MaxGPMMergeCand-2 can be used for the binarization of the first and second index values, respectively, according to the signaling order. On the other hand, since there is no correlation between the two index values, when the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be the same, the same maximum value MaxGPMMergeCand-1 is used for the binarization of both of the two index values.
[0090] In the above method, in order to reduce signaling cost, different maximum values can be applied to the binarization of merge_gpm_idx0 and merge_gpm_idx1. The selection of the corresponding maximum value depends on the decoded values of the MVR (indicated by gpm_mvr_partIdx0_enable, gpm_mvr_partIdx1_enable, gpm_mvr_partIdx0_direction_idx, gpm_mvr_partIdx1_direction_idx, gpm_mvr_partIdx0_distance_idx, and gpm_mvr_partIdx1_distance_idx). Such a design brings unwanted parsing dependencies between different GPM syntax elements and may affect the overall parsing. To solve such problems, in one embodiment, always the same maximum value (e.g., MaxGPMMergeCand - 1) is proposed for the parsed values of merge_gpm_idx0 and merge_gpm_idx1. When such a method is used, one bitstream conformity constraint can be used to prevent the two decoded MVs of the two GPM sections from being the same. In another method, such non - conformity constraints can also be removed so that the decoded MVs of the two GPM sections are allowed to be the same. On the other hand, when such a method is applied (i.e., using the same maximum value for merge_gpm_idx0 and merge_gpm_idx1), there is no parsing dependency between merge_gpm_idx0 / merge_gpm_idx1 and other GPM - MVR syntax elements. Therefore, the signaling order of these syntax elements is no longer a problem. In one example, it is proposed to move the signaling of merge_gpm_idx0 / merge_gpm_idx1 before the signaling of gpm_mvr_partIdx0_enable, gpm_mvr_partIdx1_enable, gpm_mvr_partIdx0_direction_idx, gpm_mvr_partIdx1_direction_idx, gpm_mvr_partIdx0_distance_idx, and gpm_mvr_partIdx1_distance_idx.
[0091] Geometric Partitioning Mode with Symmetric Motion Vector Refinement In the case of the GPM-MVR method discussed above, two separate MVR values are signaled, and one is applied to improve the base MV of only one GPM partition. Such a method can be efficient in terms of improving prediction accuracy by allowing independent motion refinement for each GPM partition. However, such flexible motion refinement comes at the expense of increasing signaling overhead, under the condition that two different sets of GMP-MVR syntax elements need to be sent from the encoder to the decoder. To reduce the signaling overhead, in this chapter, one geometric partitioning mode with symmetric motion vector refinement is proposed. Specifically, in this method, according to the symmetric relationship between the picture order count (POC) values of the current picture and the reference picture associated with two GPM partitions, one single MVR value is signaled for one GPM CU and used for both of the two GPM partitions. Table 6 shows the syntax elements when the proposed method is applied.
[0092]
Table 7
[0093] As shown in Table 6, after the base MVs of the two GPM partitions are selected (based on merge_gpm_idx0 and merge_gpm_idx1), one flag gpm_mvr_enable_flag is signaled to indicate whether the GPM-MVR mode is applied to the current GPM CU. When this flag is equal to 1, this indicates that motion refinement is applied to enhance the base MVs of the two GPM partitions. Otherwise (when the flag is equal to 0), this indicates that motion refinement is not applied to either of the two partitions. When the GPM-MVR mode is enabled, additional syntax elements are further signaled by the direction index gpm_mvr_direction_idx and the size index gpm_mvr_distance_idx to specify the value of the applied MVR. In addition, similar to the MMVD mode, the meaning of the MVR sign can vary according to the relationship between the current picture of the GPM partition and the POCs of the two reference pictures. Specifically, when both POCs of the two reference pictures are greater than or less than the POC of the current picture, the signaled sign is the sign of the MVR added to both of the two base MVs. Otherwise (when the POC of one reference picture is greater than the POC of the current picture and the POC of the other reference picture is less than the POC of the current picture), the signaled sign is applied to the MVR of the first GPM partition, and the opposite sign is applied to the second GPM partition. In Table 6, it is permitted that the values of merge_gpm_idx0 and merge_gpm_idx1 are the same.
[0094] In another embodiment, it is proposed to signal two different flags to separately control the enabling / disabling of the GPM-MVR mode for two separate GPM partitions. However, when the GPM-MVR mode is enabled, only one MVR is signaled based on the syntax elements gpm_mvr_direction_idx and gpm_mvr_distance_idx. The corresponding syntax table for such a signaling method is shown in Table 7.
[0095] [Table 8]
[0096] When the signaling method of Table 7 is applied, the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be identical. However, to ensure that the resulting MVs applied to the two GPM partitions are not redundant, when the flag gpm_mvr_partIdx0_enable_flag is equal to 0 (i.e., GPM-MVR is not applied to the first GPM partition), the flag gpm_mvr_partIdx1_enable_flag is not signaled but is inferred to be 1 (i.e., GPM-MVR is applied to the second GPM partition).
[0097] Adaptation of the Permitted MVR to the GPM-MVR In the GPM-MVR method discussed above, a set of fixed MVR values is used for GPM CUs at both the encoder and the decoder in one video sequence. Such a design is suboptimal for video content with high resolution or high motion. In those cases, the MV tends to be much larger, so that the fixed MVR values may not be optimal for capturing the true motion of those blocks. To further improve the coding performance of the GPM-MVR mode, this disclosure proposes to support the adaptation of the MVR values allowed to be selected by the GPM-MVR mode at various coding levels, such as sequence level, picture / slice picture, coding block group level, etc. For example, multiple MVR sets as well as corresponding codewords may be derived offline according to the specific motion characteristics of different video sequences. The encoder can select the best MVR set and signal the corresponding index of the selected set to the decoder.
[0098] In a particular embodiment of the present disclosure, in addition to the default MVR offsets including eight offset magnitudes (i.e., 1 / 4, 1 / 2, 1, 2, 4, 8, 16, and 32 pels) and four MVR directions (i.e., the ±x and y axes), another MVR offset defined in the following table is proposed for the GPM-MVR mode.
[0099]
Table 9
[0100]
Table 10
[0101] In Tables 15 and 16 above, the values +1 / 2 and -1 / 2 on the x-axis and y-axis indicate the diagonal directions in the horizontal and vertical directions (+45° and -45°). As shown in Tables 15 and 16, compared with the existing MVR offset set, the second MVR offset set introduces two new offset magnitudes (i.e., 3 pels and 6 pels) and four offset directions (45°, 135°, 225°, and 315°). The newly added MVR offsets make the second MVR offset set more suitable for coding video blocks with advanced motion. In addition, to enable an adaptive switch between the two MVR offset sets, one control flag is proposed to be signaled at one specific coding level (e.g., sequence, picture, slice, CTU, and coding block, etc.) to indicate which set of MVR offsets is selected for the GPM-MVR mode applied under the coding level. Assuming that the proposed adaptation is implemented at the picture level, Table 17 below shows the corresponding syntax elements signaled in the picture header.
[0102]
Table 11
[0103] In Table 17 above, the new flag ph_gpm_mvr_offset_set_flag is used to indicate the selection of the corresponding GPM MVR offset used for that picture. When this flag is equal to 0, this means that the default MVR offsets (i.e., sizes of 1 / 4, 1 / 2, 1, 2, 4, 8, 16, and 32 pels, and the 4 MVR directions of ±x and y axes) are applied in the GPM - MVR mode within this picture. Otherwise, when this flag is equal to 1, this means that the second MVR offset (i.e., sizes of 1 / 4, 1 / 2, 1, 2, 3, 4, 6, 8, 16 pels, and the 8 MVR directions of ±x, y axes, and 45°, 135°, 225°, and 315°) are applied in the GPM - MVR mode within this picture.
[0104] Different methods may be applied to signal the MVR offset. First, it is proposed to binarize the MVR direction using a fixed - length codeword, provided that the MVR directions are usually statistically uniformly distributed. Taking the default MVR offset as an example, there are a total of 4 directions, and codewords of 00, 01, 10, and 11 can be used to represent these 4 directions. On the other hand, since the MVR offset size may have a varying distribution adapted to the specific motion characteristics of the video content, it is proposed to binarize the MVR size using a variable - length codeword. Table 18 below shows one specific codeword table that can be used for the binarization of the MVR sizes of the default MVR offset set and the second MVR offset set.
[0105]
Table 12
[0106] In other embodiments, variable codewords of different fixed lengths may be applied to binarize the MVR offset sizes of the default and second MVR offset sets. For example, the bins "0" and "1" in the above codeword table may be exchanged to adapt to various 0 / 1 statistical information of the context adaptive binary arithmetic coding (CABAC) engine.
[0107] In one particular example, two different codeword tables are provided to binarize the value of the MVR size. The following table shows the corresponding codewords of the default and secondary MVR offset sets applied in the first and second codeword tables. Table 19 shows the codewords of the MVR offset size in the first codeword table. Table 20 shows the codewords of the MVR offset size in the second codeword table.
[0108]
Table 13
[0109]
Table 14
[0110] To enable an adaptive switch between the two codeword tables, one specific coding level (e.g., sequence, picture, slice, CTU, and coding block, etc.) is signaled, and one indication flag is proposed to specify which codeword table is used to binarize the MVR size under that coding level. Assuming that the proposed adaptation is performed at the picture level, Table 21 below shows the corresponding syntax elements signaled in the picture header, and the newly added syntax elements are in italic bold.
[0111]
Table 15
[0112] In the above syntax table, a new flag ph_gpm_mvr_step_codeword_flag is used to indicate the selection of the corresponding codeword table used for binarization of the MVR size of a picture. When this flag is equal to 0, this indicates that the first codeword table is applied to the picture, and otherwise (i.e., when the flag is equal to 1), this indicates that the second codeword table is applied to the picture.
[0113] In another embodiment, it is proposed to always use one codeword table for binarizing the MVR offset size during encoding / decoding of an entire video sequence. In one example, it is proposed to always use the first codeword table for binarization of the MVR size. In another example, it is proposed to always use the second codeword table for binarization of the MVR size. In another way, it is proposed to use one fixed codeword table (e.g., the second codeword table) for binarization of all MVR sizes.
[0114] In other ways, one statistical-based binarization method can be applied to adaptively design the optimal codeword for the MVR offset size on-the-fly rather than signaling it. The statistical information used to determine the optimal codeword can be, but is not limited to, the probability distribution of the MVR offset sizes collected over a plurality of previously coded pictures, slices, and / or coding blocks. The codeword can be re-determined / updated at various frequency levels. For example, the update can be performed each time a CU is coded in the GPM-MVR mode. In another example, the update can be re-determined and / or updated each time a plurality, e.g., 8 or 16, of CUs are coded in the GPM-MVR mode.
[0115] In other methods, instead of redesigning a new set of codewords, the proposed statistics-based method can also be used to reorder the MVR size values again based on the same set of codewords, such that shorter codewords are assigned to the more used sizes and longer codewords are assigned to the less used sizes. Taking the following table as an example, assuming that statistical information is collected at the picture level, the "Usage" row shows the corresponding ratios of the different MVR offset sizes used by the GPM-MVR coding blocks within the previously coded picture. According to the values (i.e., shortened unary codewords) in the "Usage" row using the same binarization method, the encoder / decoder can order the MVR size values based on their usage, and then the encoder / decoder can assign the shortest codeword (i.e., "1") to the most frequently used MVR size (i.e., 1 pel), the next shortest codeword (i.e., "01") to the next most frequently used MVR size (i.e., 1 / 2 pel), and the longest codewords (i.e., "0000001" and "0000000") to the two least used MVR sizes (i.e., 16 pel and 32 pel). Thus, with such a reordering scheme, the same set of codewords can be freely reordered again to adapt to the dynamic changes in the statistical distribution of the MVR sizes.
[0116]
Table 16
[0117] Encoder acceleration logic for GPM-MVR rate distortion optimization In the case of the proposed GPM-MVR method, in order to determine the optimal MVR for each GPM section, the encoder may need to test the rate distortion cost of each GPM section multiple times, varying the MVR value applied each time. This can significantly increase the complexity of the coding in the GPM mode. To address the issue of coding complexity, the following high-speed coding logic is proposed in this chapter.
[0118] First, due to the quadtree / binary tree / ternary tree block partitioning structures applied in VVC and AVS3, one and the same coding block can be identified during the rate distortion optimization (RDO) process and can be partitioned by each of one different partitioning path. In current VTM / HPM encoder implementation examples, the GPM and GPM-MVR modes are always tested whenever one and the same CU is obtained by a different combination of block partitioning together with other inter and intra coding modes. Generally, for different partitioning paths, only the adjacent blocks of one CU can be different, but should have a relatively minor impact on the optimal coding mode selected by one CU. Based on such considerations, it is proposed to remember the decision as to whether the GPM mode is selected when the RD cost of one CU is first identified in order to reduce the total number of GPM RDOs applied. Thereafter, when the same CU is re-identified by the RDO process (by a different partitioning path), the RD cost of GPM (including GPM-MVR) is only identified if GPM was selected for that CU for the first time. If GPM is not selected for the initial RD identification of one CU, only GPM (without GPM-MVR) is tested when the same CU is realized by a different partitioning path. In another way, when GPM is not selected for the initial RD identification of one CU and when the same CU is realized by a different partitioning path, neither GPM nor GPM-MVR is tested.
[0119] Second, in order to reduce the number of GPM partitions for the GPM-MVR mode, it is proposed to maintain the first M GPM partition modes without the minimum RD cost when the RD cost of one CU is first identified. Thereafter, when the same CU is re-identified by the RDO process (by a different partitioning path), only those M GPM partition modes are tested for the GPM-MVR mode.
[0120] Thirdly, for each GPM section, in order to reduce the number of GPM sections to be tested for one initial RDO process, it is proposed to first calculate the sum of absolute differences (SAD) values when using different unidirectional prediction merge candidates for two GPM sections. Then, for each GPM section under one specific section mode, select the best unidirectional prediction merge candidate having the minimum SAD value, and calculate the corresponding SAD value of the section mode equal to the sum of the SAD values of the best unidirectional prediction merge candidates for the two GPM sections. Then, for the subsequent RD process, only the first N section modes having the best SAD value for the previous step are tested for the GPM-MVR mode.
[0121] Geometric section with explicit motion signaling In this chapter, a plurality of methods are proposed for extending the GPM mode to the normal inter-mode bidirectional prediction in which two unidirectional MVs of the GPM mode are explicitly signaled from the encoder to the decoder.
[0122] In the first solution (Solution 1), it is proposed to completely reuse the existing motion signaling for bidirectional prediction to signal the two unidirectional MVs in the GPM mode. Table 8 shows the modified syntax table of the proposed method, and the newly added syntax elements are shown in italic bold. As shown in Table 8, in this solution, all existing syntax elements that signal L0 and L1 motion information are completely reused again to indicate the unidirectional MVs of the two GPM partitions, respectively. In addition, it is assumed that the L0 MV is always associated with the first GPM partition and the L1 MV is always associated with the second GPM partition. On the other hand, in Table 8, the inter-prediction syntax, i.e., inter_pred_idc, is signaled before the GPM flag (i.e., gpm_flag), and thus the value of inter_pred_idc can be used to condition the presence of the gpm_flag. Specifically, the flag gpm_flag needs to be signaled only when inter_pred_idc is equal to PRED_BI (i.e., bidirectional prediction) and both inter_affine_flag and sym_mvd_flag are equal to 0 (i.e., the CU is not coded by either the affine mode or the SMVD mode). When the flag gpm_flag is not signaled, its value is always inferred to be 0 (i.e., the GPM mode is disabled). When gpm_flag is 1, another syntax element gpm_partition_idx is further signaled to indicate the GPM mode selected for the current CU (from a total of 64 GPM partitions).
[0123]
Table 17
[0124] In another way, it is proposed to place the signaling of the flag gpm_flag before other inter-signaling syntax elements so that the value of gpm_flag can be used to determine whether other inter-syntax elements need to exist. Table 9 shows the corresponding syntax table when such a method is applied, and the newly added syntax elements are shown in italic bold. As can be seen, in Table 9, first gpm_flag is signaled. When gpm_flag is equal to 1, the corresponding signaling of inter_pred_idc, inter_affine_flag, and sym_mvd_flag can be bypassed. Instead, the corresponding values of the three syntax elements can be inferred as PRED_BI, 0, and 0 respectively.
[0125]
Table 18
[0126] In both Table 8 and Table 9, the SMVD mode cannot be combined with the GPM mode. In another example, it is proposed to permit the SMVD mode when the current CU is coded by the GPM mode. When such a combination is permitted, by following the same design of SMVD, it is assumed that the MVDs of the two GPM partitions are symmetric, and therefore only the MVD of the first GPM partition needs to be signaled, and the MVD of the second GPM partition is always symmetric to the first MVD. When such a method is applied, the corresponding signaling condition of sym_mvd_flag for gpm_flag can be removed.
[0127] As described above, in the first solution, it is always assumed that L0 MV is used for the first GPM partition and L1 MV is used for the second GPM partition. Such a design may not be optimal in the sense that this method prohibits the MVs of the two GPM partitions from coming from the same prediction list (L0 or L1). To solve such a problem, a signaling design shown in Table 10 proposes an alternative GPM-EMS method, Solution 2. In Table 10, the newly added syntax elements are shown in italic bold. As shown in Table 10, first, the flag gpm_flag is signaled. When this flag is equal to 1 (i.e., GPM is enabled), the syntax gpm_partition_idx is signaled to specify the selected GPM mode. Then, an additional flag gpm_pred_dir_flag0 is signaled to indicate the corresponding prediction list from which the MV of the first GPM partition comes. When the flag gpm_pred_dir_flag0 is equal to 1, this indicates that the MV of the first GPM partition comes from L1, and otherwise (when the flag is equal to 0), this indicates that the MV of the first GPM partition comes from L0. Subsequently, the existing syntax elements ref_idx_l0, mvp_l0_flag, and mvd_coding() are utilized to signal the reference picture index, mvp index, and the value of the MVD of the first GPM partition. On the other hand, similar to the first partition, another syntax element gpm_pred_dir_flag1 is introduced to select the corresponding prediction list of the second GPM partition, followed by the existing syntax elements ref_idx_l1, mvp_l1_flag, and mvd_coding() being used to derive the MV of the second GPM partition.
[0128]
Table 19
[0129] Finally, it should be noted that when the proposed GPM-EMS scheme is enabled for one inter-CU, several existing coding tools in VVC and AVS3, which are specifically designed for bidirectional prediction such as bidirectional optical flow, decoder-side motion vector refinement (DMVR), and bidirectional prediction by CU weight (BCW), can be automatically bypassed, provided that the GPM mode consists of two unidirectional prediction segments (excluding the mixed samples on the split edge). For example, when one of the proposed GPM-EMSs is enabled for one CU to reduce the signaling overhead, on condition that BCW cannot be applied to the GPM mode, the corresponding BCW weight does not need to be further signaled for the CU.
[0130] Combination of GPM-MVR and GPM-EMS In this chapter, it is proposed to combine GPM-MVR and GPM-EMS for one CU having geometric shape partitions. Specifically, unlike GPM-MVR or GPM-EMS where only one of merge-based motion signaling or explicit signaling is applicable to signal the unidirectional prediction MVs of two GPM partitions, in the proposed method, 1) one partition uses GPM-MVR-based motion signaling and the other uses GPM-EMS-based motion signaling, or 2) both partitions use GPM-MVR-based motion signaling, or 3) both partitions use GPM-EMS-based motion signaling are permitted. Using the GPM-MVR signaling in Table 4 and GPM-EMS in Table 10, Table 11 shows the corresponding syntax table after the proposed GPM-MVR and GPM-EMS are combined. In Table 11, the newly added syntax elements are shown in italic bold. As shown in Table 11, two additional syntax elements gpm_merge_flag0 and gpm_merge_flag1 are introduced into Partition #1 and #2 that specify the corresponding partitions using GPM-MVR-based merge signaling or GPM-EMS-based explicit signaling, respectively. When this flag is 1, it means that GPM-MVR-based signaling is enabled for the partition where the GPM unidirectional prediction motion is signaled by merge_gpm_idxX, gpm_mvr_partIdxX_enabled_flag, gpm_mvr_partIdxX_direction_idx, and gpm_mvr_partIdxX_distance_idx, where X = 0, 1. Otherwise, when this flag is 0, it means that the unidirectional prediction motion of this partition is explicitly signaled by the GPM-EMS method using the syntax elements gpm_pred_dir_flagX, ref_idx_lX, mvp_lX_flag, and mvd_lX, where X = 0, 1.
[0131]
Table 20
[0132] Combination of GPM-MVR and template matching In this chapter, different solutions for combining GPM-MVR with template matching are provided.
[0133] In Method 1, when one CU is coded in GPM mode, it is proposed to signal two separate flags for two GPM segments, where each flag indicates whether the unidirectional motion of the corresponding segment is further improved by template matching. When this flag is enabled, a template is generated using the reconstructed samples adjacent to the top-left of the current CU, and then the unidirectional motion of the segment is improved by minimizing the difference between the template and its reference samples according to the same procedure introduced in the "Template Matching" chapter. Otherwise (when the flag is disabled), template matching is not applied to this segment and GPM-MVR can be further applied. Using the GPM-MVR signaling method in Table 5 as an example, Table 12 shows the corresponding syntax table when GPM-MVR is combined with template matching. In Table 12, the newly added syntax elements are shown in italic bold.
[0134]
Table 21
[0135] As shown in Table 12, in the proposed scheme, two additional flags, gpm_tm_enable_flag0 and gpm_tm_enable_flag1, are first signaled for each of the two GPM partitions to indicate whether the movement is improved. When this flag is 1, it indicates that TM is applied to improve the unidirectional MV of one partition. When this flag is 0, one flag (gpm_mvr_partIdx0_enable_flag or gpm_mvr_partIdx0_enable_flag) is further signaled for each to indicate whether GPM-MVR is applied to the GPM partition. When the flag of one GPM partition is equal to 1, the distance index (indicated by the syntax elements gpm_mvr_partIdx0_distance_idx and gpm_mvr_partIdx1_distance_idx) and the direction index (indicated by the syntax elements gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx1_distance_idx) are signaled to specify the direction of MVR. Then, the existing syntax merge_gpm_idx0 and merge_gpm_idx1 are signaled to identify the unidirectional MVs for the two GPM partitions. On the other hand, similar to the signaling conditions applied in Table 5, the following conditions may be applied to ensure that the resulting MVs used for prediction of the two GPM partitions are not the same.
[0136] First, when both values of gpm_tm_enable_flag0 and gpm_tm_enable_flag1 are equal to 1 (i.e., TM is enabled for both of the two GPM partitions), the values of merge_gpm_idx0 and merge_gpm_idx1 cannot be the same.
[0137] Second, when one of gpm_tm_enable_flag0 and gpm_tm_enable_flag1 is 1 and the other is 0, the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same.
[0138] Otherwise, that is, when both gpm_tm_enable_flag0 and gpm_tm_enable_flag1 are equal to 1, and first, when both the values of gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 0 (i.e., GPM-MVR is disabled for both of the two GPM sections), the values of merge_gpm_idx0 and merge_gpm_idx1 cannot be made the same; second, when gpm_mvr_partIdx0_enable_flag is equal to 1 (i.e., GPM-MVR is enabled for the first GPM section) and gpm_mvr_partIdx1_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for the second GPM section), the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same; third, when gpm_mvr_partIdx0_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for the first GPM section) and gpm_mvr_partIdx1_enable_flag is equal to 1 (i.e., GPM-MVR is enabled for the second GPM section), the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same; fourth, when both the values of gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 1 (i.e., GPM-MVR is enabled for both of the two GPM sections), the determination of whether the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same depends on the values of MVR applied to the two GPM sections (indicated by gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx0_distance_idx, as well as gpm_mvr_partIdx1_direction_idx and gpm_mvr_partIdx1_distance_idx).When the values of the two MVRs are equal, merge_gpm_idx0 and merge_gpm_idx1 are not permitted to be the same. Otherwise (when the values of the two MVRs are not equal), the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same.
[0139] In the above Method 1, TM and MVR are exclusively applied to GPM. In such a manner, it is prohibited to further apply MVR on the improved MV in TM mode. Therefore, in order to further provide more MV candidates for GPM, Method 2 for enabling the application of the MVR offset on the TM improved MV is proposed. Table 13 shows the corresponding syntax table when GPM-MVR is combined with template matching. In Table 13, the newly added syntax elements are shown in italic bold.
[0140]
Table 22
[0141] As shown in Table 13, different from Table 12, the signaling conditions of gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag for gpm_tm_enable_flag0 and gpm_tm_enable_flag1 are removed. Therefore, regardless of whether TM is applied to improve the unidirectional movement of one GPM section, MV improvement is always permitted to be applied to the MV of the GPM section. Similar to the above, in order to ensure that the resulting MVs of the two GPM sections are not the same, the following conditions should be applied.
[0142] First, when one of gpm_tm_enable_flag0 and gpm_tm_enable_flag1 is 1 and the other is 0, the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same.
[0143] Otherwise, that is, when both gpm_tm_enable_flag0 and gpm_tm_enable_flag1 are equal to 1, or both of these flags are equal to 0, first, when both values of gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 0 (that is, GPM-MVR is disabled for both of the two GPM sections), the values of merge_gpm_idx0 and merge_gpm_idx1 cannot be made the same. Second, when gpm_mvr_partIdx0_enable_flag is equal to 1 (that is, GPM-MVR is enabled for the first GPM section), and gpm_mvr_partIdx1_enable_flag is equal to 0 (that is, GPM-MVR is disabled for the second GPM section), the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same. Third, when gpm_mvr_partIdx0_enable_flag is equal to 0 (that is, GPM-MVR is disabled for the first GPM section), and gpm_mvr_partIdx1_enable_flag is equal to 1 (that is, GPM-MVR is enabled for the second GPM section), the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same. Fourth, when both values of gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 1 (that is, GPM-MVR is enabled for both of the two GPM sections), the determination of whether the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same depends on the values of MVR applied to the two GPM sections (indicated by gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx0_distance_idx, as well as gpm_mvr_partIdx1_direction_idx and gpm_mvr_partIdx1_distance_idx).When the values of two MVRs are equal, merge_gpm_idx0 and merge_gpm_idx1 are not permitted to be the same. Otherwise (when the values of the two MVRs are not equal), the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same.
[0144] In the above two methods, two separate flags need to be signaled to indicate whether the TM is applied to each GPM section. The additional signaling may reduce the overall coding efficiency due to the additional overhead, especially at a particularly low bit rate. To reduce the signaling overhead, instead of introducing additional signaling, Method 3 is proposed for inserting TM-based unidirectional MVs into the unidirectional MV candidate list of the GPM mode. The TM-based unidirectional MV is generated according to the same TM process described in the "Template Matching" chapter, which uses the original unidirectional MV of the GPM as the initial MV. In such a way, there is no need to further signal an extra control flag from the encoder to the decoder. Instead, the decoder can identify whether one MV is improved by the TM according to the corresponding merge index received from the bitstream (i.e., merge_gpm_idx0 and merge_gpm_idx1). Different methods may exist for arranging the regular GPM MV candidates (i.e., non-TM) and the TM-based MV candidates. In one method, it is proposed to arrange the TM-based MV candidates at the beginning of the MV candidate list, followed by the non-TM-based MV candidates. In another method, it is proposed to first arrange the non-TM-based MV candidates at the beginning, followed by the TM-based candidates. In another method, it is proposed to arrange the TM-based MV candidates and the non-TM-based MV candidates alternately. For example, this can be done by arranging the first N non-TM-based candidates, then all the TM-based candidates, and finally the remaining non-TM-based candidates. In another example, this can be done by arranging the first N TM-based candidates, then all the non-TM-based candidates, and finally the remaining TM-based candidates. In another example, it is proposed to arrange the non-TM-based candidates and the TM-based candidates continuously, i.e., one non-TM-based candidate, one TM-based candidate, etc.
[0145] In Method 1, two GPM template flags are signaled before the GPM-MVR flag. Specifically, in such a design, GPM-MVR can be enabled only for a given GPM section by first signaling a GPM template flag for a section equal to 0. The GPM template flag can be coded using an appropriate context model, but it incurs a signaling penalty in the GPM-MVR mode. To solve such a problem, in one embodiment of the present disclosure, it is proposed to signal the GPM-MVR mode first and then the GPM-TM mode. Specifically, in this method, the GPM-MVR flag is first signaled for each GPM section to indicate whether GPM-MVR is applied to that section. When the flag is equal to 1, the MVR syntax elements gpm_mvr_partIdx0_distance_idx / gpm_mvr_partIdx1_distance_idx and gpm_mvr_partIdx0_dierction_idx / gpm_mvr_partIdx1_direction_idx are further signaled to specify the corresponding values of the MVR size and direction of that section. Otherwise, when the GPM-MVR flag of that section is equal to false, the GPM-TM flag is signaled to indicate whether the GPM-TM mode (improving the MV of the section using the left and upper adjacent reconstructed samples) is applied. Table 22 shows the corresponding syntax table when the above signaling method is applied, and the newly added syntax elements are in italic boldface.
[0146]
Table 23
[0147] Furthermore, the following conditions should be applied to remove the redundancy of signaling between GPM merge indexes.
[0148] First, when one of gpm_tm_enable_flag0 and gpm_tm_enable_flag1 is 1 and the other is 0, the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same.
[0149] Second, when both gpm_tm_enable_flag0 and gpm_tm_enable_flag1 are equal to 1, the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same.
[0150] Third, when both gpm_tm_enable_flag0 and gpm_tm_enable_flag1 are equal to 0, different conditions apply. When the values of both gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 0 (i.e., GPM-MVR is disabled for both of the two GPM partitions), the values of merge_gpm_idx0 and merge_gpm_idx1 cannot be made the same. When gpm_mvr_partIdx0_enable_flag is equal to 1 (i.e., GPM-MVR is enabled for the first GPM partition) and gpm_mvr_partIdx1_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for the second GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same. When gpm_mvr_partIdx0_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for the first GPM partition) and gpm_mvr_partIdx1_enable_flag is equal to 1 (i.e., GPM-MVR is enabled for the second GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same. When the values of both gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 1 (i.e., GPM-MVR is enabled for both of the two GPM partitions), the determination of whether the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same depends on the values of MVR applied to the two GPM partitions (indicated by gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx0_distance_idx, and gpm_mvr_partIdx1_direction_idx and gpm_mvr_partIdx1_distance_idx). If the values of the two MVRs are equal, it is not permitted for merge_gpm_idx0 and merge_gpm_idx1 to be the same.Otherwise (the values of the two MVRs are not equal), the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be the same.
[0151] Alternatively, instead of using two separate GPM-TM flags, a single flag is proposed to jointly control the enabling / disabling of template matching for two GPM partitions. When the flag is true, this means that the two unidirectional MVs of the two GPM partitions need to be improved based on minimizing the difference between the template (i.e., the left and upper adjacent reconstructed samples) and its corresponding reference sample by the template matching method. Specifically, similar to Method 4, two GPM-MVR flags are first signaled for one GPM CU to indicate whether the GPM-MVR is applied to one specific GPM partition. When the GPM-MVR flags of each partition are equally true, hereinafter, the MVR magnitude and the MVR direction are further signaled for that partition. Further, when both GPM-MVR flags of the two GPM partitions are equally false, the GPM-TM flag is further signaled to indicate whether the GPM-TM is applied to both of the two GPM partitions. Table 23 shows the corresponding syntax table of the GPM mode when such a design is applied, and the newly added syntax elements are in italic bold.
[0152]
Table 24
[0153] In another embodiment, it is proposed to signal the GPM-TM flag for two GPM sections and then signal the two GPM-MVR flags. Correspondingly, the value of GPM-TM can be used to adjust the presence of the two GPM-MVR flags such that the GPM-MVR flags are signaled only when the value of the GPM-TM flag is equal to 0 (i.e., GPM-TM is not applied to the two GPM sections). Table 24 shows the corresponding syntax table of the GPM mode when such a signaling method is applied, and the newly added syntax elements are in italic boldface.
[0154]
Table 25
[0155] In addition, for both of the two methods, the following conditions should be applied to remove the redundancy of signaling between GPM merge indices.
[0156] First, when gpm_tm_enable_flag is 1, the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same. Second, when gpm_tm_enable_flag is equal to 0, different conditions may apply. For example, when the values of both gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 0 (i.e., GPM-MVR is disabled for both of the two GPM partitions), the values of merge_gpm_idx0 and merge_gpm_idx1 cannot be made the same. Further, when gpm_mvr_partIdx0_enable_flag is equal to 1 (i.e., GPM-MVR is enabled for the first GPM partition) and gpm_mvr_partIdx1_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for the second GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same. Further, when gpm_mvr_partIdx0_enable_flag is equal to 0 (i.e., GPM-MVR is disabled for the first GPM partition) and gpm_mvr_partIdx1_enable_flag is equal to 1 (i.e., GPM-MVR is enabled for the second GPM partition), the values of merge_gpm_idx0 and merge_gpm_idx1 are permitted to be the same.Furthermore, when the values of both gpm_mvr_partIdx0_enable_flag and gpm_mvr_partIdx1_enable_flag are equal to 1 (i.e., GPM-MVR is enabled for both of the two GPM partitions), the determination of whether the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be the same depends on the values of MVR applied to the two GPM partitions (indicated by gpm_mvr_partIdx0_direction_idx and gpm_mvr_partIdx0_distance_idx, as well as gpm_mvr_partIdx1_direction_idx and gpm_mvr_partIdx1_distance_idx). When the values of the two MVRs are equal, it is not allowed for merge_gpm_idx0 and merge_gpm_idx1 to be the same. Otherwise (when the values of the two MVRs are not equal), the values of merge_gpm_idx0 and merge_gpm_idx1 are allowed to be the same.
[0157] When the template matching method is applied to the GPM mode, additional complexity is required for both the encoder and the decoder by performing computationally extensive motion estimation to identify the optimal unidirectional MV for each GPM partition. Such a non-negligible increase in complexity may make the GPM mode infeasible for certain lower encoders or certain video applications that require low video latency, such as live video streaming, video conferencing, and video gaming. Based on such considerations, it is proposed to add one control flag at a specific high coding level such as sequence level, picture / slice level, coding block group level, and additional levels for the CU to adaptively enable or disable the GPM-TM mode below that level. Assuming that the proposed adaptation is implemented at the picture level, Table 25 shows the corresponding syntax elements signaled in the picture header, and the newly added syntax elements are in italic bold.
[0158]
Table 26
[0159] In the above syntax Table 25, the flag sps_dmvd_enable_flag is a sequence-level control flag indicating whether template matching is enabled for video sequence coding, and the ph_gpm_tm_enable_flag is a proposed GPM-TM control flag used to indicate whether GPM-TM can be applied to CUs within a picture.
[0160] Construction of the GPM candidate list by motion vector pruning As discussed in the introduction part, to obtain the MVs of two geometric partitions, one unidirectional prediction candidate list is first directly derived from the normal merge candidate list generation process. Although the MVs of two geometric partitions can be the same under the condition that the selection of the prediction direction of each GPM MV is based on the parity of the corresponding merge index, it is obvious that this does not make sense because no additional benefit can be provided compared to the case where the geometric partition of the CU has no partition. To avoid such redundancy, when generating the unidirectional prediction MV candidate list for one GPM CU, motion vector pruning is proposed to be applied such that only one MV can be added to the list only when it is not the same as any of the existing candidates in the list. In another approach, it is further proposed that one MV threshold is applied when comparing two MVs. Specifically, in such a method, when the difference between two MVs (in the horizontal and vertical directions respectively) is smaller than one MV threshold, the two MVs are considered the same, and otherwise (the MV difference in one direction is greater than or equal to the MV threshold), the two MVs are considered not the same. In one method, it is proposed to use one fixed MV threshold for all block sizes. In another method, it is proposed to determine the value of the MV threshold based on the size of the coding block such that a larger MV threshold is used for larger CUs and a smaller MV threshold is used for smaller CUs. In some examples, when the number of samples in the block is N < 64, the value of the MV threshold is set to 1 / 4 pel, when 64 ≤ N < 256, the value of the MV threshold is set to 1 / 2 pel, and when N ≥ 256, the value of the MV threshold is set to 1 pel.
[0161] FIG. 9 shows a computing environment (or computing device) 910 coupled to a user interface 960. The computing environment 910 can be part of a data processing server. In some embodiments, the computing device 910 can execute any of the various methods or processes (such as encoding / decoding methods or processes) described herein according to the various examples of the present disclosure. The computing environment 910 can include a processor 920, a memory 940, and an I / O interface 950.
[0162] The processor 920 typically controls the overall operation of the computing environment 910, such as operations related to display, data acquisition, data communication, and image processing. The processor 920 can include one or more processors to execute instructions for performing all or some of the steps in the aforementioned methods. Further, the processor 1020 can include one or more modules that facilitate the interaction between the processor 920 and other components. The processor can be a central processing unit (CPU), a microprocessor, a single-chip machine, a GPU, etc.
[0163] The memory 940 is configured to store various types of data to support the operation of the computing environment 910. The memory 940 can include a predetermined software 942. Examples of such data include instructions for any application or method operating on the computing environment 910, video data sets, image data, etc. The memory 940 can be implemented by using any type of volatile or non-volatile memory device, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic or optical disks, or a combination thereof.
[0164] The I / O interface 950 provides an interface between the processor 920 and peripheral interface modules such as a keyboard, click wheel, buttons, etc. The buttons can include, but are not limited to, a home button, a scan start button, and a scan stop button. The I / O interface 950 can be coupled to an encoder and a decoder.
[0165] In some embodiments, a non-transitory computer-readable storage medium including a plurality of programs, such as those included in the memory 940, is also provided, and the plurality of programs are executable by the processor 920 in the computing environment 910 to implement the aforementioned methods. For example, the non-transitory computer-readable storage medium can be a ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0166] The non-transitory computer-readable storage medium stores a plurality of programs for execution by a computing device having one or more processors, and when the plurality of programs are executed by one or more processors, the computing device is caused to implement the aforementioned methods for motion prediction.
[0167] In some embodiments, the computing environment 910 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components to implement the above methods.
[0168] FIG. 8 is a flowchart showing a method for decoding a video block in a GPM according to an example of the present disclosure.
[0169] In step 801, the processor 920 can partition the video block into first and second geometric partitions.
[0170] In step 802, the processor 920 can receive a first GPM-MVR enable flag for the first geometric partition and a second GPM-MVR enable flag for the second geometric partition. As shown in Tables 23 and 24, the first GPM-MVR enable flag can be gpm_mvr_partIdx0_enable_flag, and the second GPM-MVR enable flag can be gpm_mvr_partIdx1_enable_flag.
[0171] In step 803, the processor 920 can receive a joint TM enable flag for the first and second geometric partitions, and the joint TM enable flag can jointly indicate whether the unidirectional motion of the first partition is improved by TM and whether the unidirectional motion of the second partition is improved by TM. As shown in Tables 23 and 24, the joint TM enable flag can be gpm_tm_enable_flag. As shown in Table 23, the first and second GPM-MVR enable flags can be signaled before the joint TM enable flag. As shown in Table 24, the joint TM enable flag can be signaled before the first and second GPM-MVR enable flags.
[0172] In step 804, the processor 920 can receive a first merge GPM index for the first geometric partition and a second merge GPM index for the second geometric partition.
[0173] In some examples, the first merge GPM index identifies the unidirectional MV for the first geometric partition, and the second merge GPM index identifies the unidirectional MV for the second geometric partition.
[0174] In some examples, the first merge GPM index can be the syntax element merge_gpm_idx0 shown in Table 11 or Table 12, and the second merge GPM index can be the syntax element merge_gpm_idx1 shown in Table 23 or Table 24.
[0175] In step 805, the processor 920 can construct a unidirectional MV candidate list for GPM.
[0176] In step 806, the processor 920 can generate a unidirectional MV for the first geometric partition and a unidirectional MV for the second geometric partition.
[0177] In some examples, in response to determining that both the first and second GPM-MVR enable flags are equal to 0, i.e., MVR is not applied to the first or second geometric partition, the processor 920 can receive a joint TM enable flag for the first and second geometric partitions.
[0178] In some examples, in response to determining that the joint TM enable flag is equal to 0, the processor 920 can receive a first GPM-MVR enable flag for the first geometric partition and a second GPM-MVR enable flag for the second geometric partition.
[0179] In some examples, the processor 920 can suppress the first merge GPM index and the second merge GPM index based on the joint TM enable flag.
[0180] In some examples, in response to determining that the joint TM enable flag is equal to 1, the processor 920 can determine that it is permissible for the first merge GPM index and the second merge GPM index to be the same.
[0181] In some examples, in response to determining that the joint TM enable flag is equal to 0 and both the first and second GPM-MVR enable flags are equal to 0, the processor 920 can determine that the first merge GPM index and the second merge GPM index are different.
[0182] In some examples, in response to determining that one of the first and second GPM-MVR enable flags is 0 and the other of the first and second GPM-MVR enable flags is 1, the processor 920 can determine that it is permitted for the first merge GPM index and the second merge GPM index to be the same.
[0183] In some examples, in response to determining that both the first and second GPM-MVR enable flags are equal to 1, based on the first and second MVRs applied to the first and second geometric partitions respectively, the processor 920 can determine the first merge GPM index and the second merge GPM index.
[0184] In some examples, in response to determining that the first MVR is equal to the second MVR, the processor 920 can determine that the first merge GPM index and the second merge GPM index are different.
[0185] In some examples, in response to determining that the first MVR is not equal to the second MVR, the processor 920 can determine that it is permitted for the first merge GPM index and the second merge GPM index to be the same.
[0186] In some examples, an apparatus for decoding video blocks with GPM is provided. The apparatus includes a processor 920 and a memory 940 configured to store instructions executable by the processor. When executing the instructions, the processor is configured to implement the method shown in FIG. 8.
[0187] In some other examples, a non-transitory computer-readable storage medium storing instructions is provided. When the instructions are executed by a processor 920, the instructions cause the processor to perform the method shown in FIG. 8.
[0188] Other examples of the present disclosure will be apparent to those skilled in the art in view of the present specification and by practicing the present disclosure disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, including departures from the present disclosure within the scope of well-known or customary practice in the art, in accordance with the general principles of the present disclosure. The specification and examples are intended to be considered as illustrative only.
[0189] It will be understood that the present disclosure is not limited to the exact examples described above and illustrated in the accompanying drawings, and that various modifications and changes can be made without departing from the scope of the present disclosure.
Claims
Claim 1 A method for decoding a video block in a geometric partitioning mode (GPM), comprising: partitioning the video block into first and second geometric partitions; receiving a GPM with first motion vector refinement (GPM-MVR) enable flag for the first geometric partition and a second GPM-MVR enable flag for the second geometric partition; receiving a joint template matching (TM) enable flag for the first and second geometric partitions, wherein the joint TM enable flag jointly indicates whether the unidirectional motion of the first partition is improved by the TM and whether the unidirectional motion of the second partition is improved by the TM; receiving a first merge GPM index for the first geometric partition and a second merge GPM index for the second geometric partition; constructing a unidirectional motion vector (MV) candidate list for the GPM; generating a unidirectional MV for the first geometric partition and a unidirectional MV for the second geometric partition. Claim 2 The method of claim 1, further comprising receiving the joint TM enable flag for the first and second geometric partitions in response to determining that both the first and second GPM-MVR enable flags are equal to 0. Claim 3 The method of claim 1, further comprising receiving the first GPM-MVR enable flag for the first geometric partition and the second GPM-MVR enable flag for the second geometric partition in response to determining that the joint TM enable flag is equal to 0. Claim 4 The method of claim 1, further comprising suppressing the first merge GPM index and the second merge GPM index based on the joint TM enable flag. Claim 5 Suppressing the first merge GPM index and the second merge GPM index based on the joint TM enable flag is The method of claim 4, further comprising determining that it is permitted for the first merge GPM index and the second merge GPM index to be the same in response to determining that the joint TM enable flag is equal to 1. Claim 6 Suppressing the first merge GPM index and the second merge GPM index based on the joint TM activation flag is, The method according to claim 4, comprising determining that the first merge GPM index and the second merge GPM index are different in response to determining that the joint TM activation flag is equal to 0 and both the first and second GPM-MVR activation flags are equal to 0. **Claim 7** Suppressing the first merge GPM index and the second merge GPM index based on the joint TM activation flag is, The method according to claim 4, comprising determining that it is permitted for the first merge GPM index and the second merge GPM index to be the same in response to determining that one of the first and second GPM-MVR activation flags is 0 and the other of the first and second GPM-MVR activation flags is 1. **Claim 8** Suppressing the first merge GPM index and the second merge GPM index based on the joint TM activation flag is, The method according to claim 4, comprising determining the first merge GPM index and the second merge GPM index based on first and second MVRs respectively applied to the first and second geometric partitions in response to determining that both the first and second GPM-MVR activation flags are equal to 1. **Claim 9** Determining the first merge GPM index and the second merge GPM index based on the first and second MVRs applied to the first and second geometric partitions is, Determining that the first merge GPM index and the second merge GPM index are different in response to determining that the first MVR is equal to the second MVR, The method according to claim 8, comprising determining that it is permitted for the first merge GPM index and the second merge GPM index to be the same in response to determining that the first MVR is not equal to the second MVR. **Claim 10** An apparatus for video encoding, comprising one or more processors, A non-transitory computer-readable storage medium configured to store instructions executable by the one or more processors, An apparatus, wherein when the one or more processors execute the instructions, the apparatus is configured to perform the method according to any one of claims 1 to 9. **Claim 11** A non-transitory computer-readable storage medium storing computer-executable instructions, wherein when the computer-executable instructions are executed by one or more computer processors, the one or more computer processors are caused to perform the method according to any one of claims 1 to 9. **Claim 12** A method for receiving a bitstream decoded by the video encoding device according to claim 10.