Method, apparatus, and non-temporary computer-readable storage medium for refining motion vectors related to GPM (geometric partition mode).
By refining motion vectors in GPM using template matching, the method addresses the challenge of high compression efficiency in video encoding, enhancing performance in line with VVC standards.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- ALIBABA (CHINA) CO LTD
- Filing Date
- 2022-04-11
- Publication Date
- 2026-05-25
AI Technical Summary
Existing video encoding standards face challenges in achieving high compression efficiency, particularly with the introduction of newer standards like VVC, where refining motion vectors for geometric partition mode (GPM) can enhance encoding performance.
The method involves receiving a bitstream encoded in GPM, decoding a parameter indicating template matching application, and refining motion information using template matching if applicable, implemented through an apparatus with processors and memory, or via executable instructions on a computer-readable storage medium.
This approach enhances video encoding efficiency by improving motion vector refinement, aligning with the goals of newer standards like VVC, reducing bandwidth requirements while maintaining quality.
Smart Images

Figure 0007864734000008 
Figure 0007864734000009 
Figure 0007864734000010
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications
[0001] This disclosure claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 173,540, filed Apr. 12, 2021; U.S. Provisional Patent Application No. 63 / 194,260, filed May 29, 2021; and U.S. Provisional Patent Application No. 63 / 215,519, filed Jun. 27, 2021, which are hereby incorporated by reference in their entireties.
[0002] Technical Field
[0002] This disclosure generally relates to video processing, and more particularly to methods and systems for motion vector refinement for GPM (geometric partition mode).
Background Art
[0003] Background
[0003] Video is a set of static pictures (or “frames”) that capture visual information. To reduce memory and transmission bandwidth, video can be compressed before storage or transmission and restored before display. The compression process is usually referred to as encoding, and the restoration process is usually referred to as decoding. Most commonly, there are various video encoding formats that use standardized video encoding techniques based on prediction, transformation, quantization, entropy encoding, and in - loop filtering. Video encoding standards such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266), and the standard AVS standard, which specify a particular video encoding format, have been developed by standardization organizations. As evolving video encoding techniques are increasingly adopted in video standards, the encoding efficiency of new video encoding standards becomes higher and higher.
Summary of the Invention
Means for Solving the Problems
[0004] Summary of this disclosure
[0004] Embodiments of the present disclosure provide a method for processing video data. The method includes receiving a bitstream containing coding units encoded by GPM (geometric partition mode), decoding a first parameter associated with the coding unit, the first parameter indicating whether template matching is applied to the coding unit, and determining motion information relating to the coding unit according to the first parameter, wherein if the first parameter indicates that template matching is applied to the coding unit, the motion information is refined using template matching.
[0005]
[0005] Embodiments of the present disclosure provide an apparatus for processing video data. The apparatus includes a memory configured to store instructions and one or more processors, the one or more processors are configured to execute instructions to receive a bitstream containing encoding units encoded by GPM (geometric partition mode), decode a first parameter associated with the encoding unit, wherein the first parameter indicates whether template matching is applied to the encoding unit, and determine motion information relating to the encoding unit according to the first parameter, wherein if the first parameter indicates that template matching is applied to the encoding unit, the motion information is refined using template matching.
[0006]
[0006] Embodiments of the present disclosure provide a non-temporary computer-readable storage medium that stores a set of instructions executable by one or more processors of an apparatus for causing the apparatus to initiate a method for processing video data, the method comprising: receiving a bitstream including coding units encoded by GPM (geometric partition mode); decoding a first parameter associated with the coding unit, the first parameter indicating whether template matching is applied to the coding unit; and determining motion information relating to the coding unit according to the first parameter, wherein if the first parameter indicates that template matching is applied to the coding unit, the motion information is refined using template matching.
[0007] Brief explanation of the drawing
[0007] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and accompanying drawings. Various features shown in the drawings are not depicted to scale. [Brief explanation of the drawing]
[0008] [Figure 1]
[0008] Figure 1 is a schematic diagram showing the structure of an exemplary video sequence according to some embodiments of the present disclosure. [Figure 2A]
[0009] Figure 2A is a schematic diagram illustrating an exemplary coding process of a hybrid video coding system according to an embodiment of the present disclosure. [Figure 2B]
[0010] Figure 2B is a schematic diagram showing another exemplary coding process of a hybrid video coding system according to an embodiment of the present disclosure. [Figure 3A]
[0011] Figure 3A is a schematic diagram illustrating an exemplary decoding process of a hybrid video coding system according to an embodiment of the present disclosure. [Figure 3B]
[0012] Figure 3B is a schematic diagram showing another exemplary decoding process of a hybrid video coding system according to an embodiment of the present disclosure. [Figure 4]
[0013] Figure 4 is a block diagram of an exemplary apparatus for encoding or decoding video, according to some embodiments of the present disclosure. [Figure 5]
[0014] This document shows exemplary geometric partition mode (GPM) divisions grouped by the same angle, according to some embodiments of this disclosure. [Figure 6]
[0015] The present disclosure describes an exemplary unidirectional predictive motion vector (MV) selection process for a GPM according to some embodiments of this disclosure. [Figure 7]
[0016] This disclosure shows an exemplary generation of weights w0 using GPM according to some embodiments of this disclosure. [Figure 8]
[0017] This disclosure describes an exemplary process for weight prediction of angle-weighted prediction (AWP) according to some embodiments of this disclosure. [Figure 9]
[0018] Eight exemplary intra-prediction angles supported in AWP mode, according to some embodiments of this disclosure, are shown. [Figure 10]
[0019] Seven different exemplary weight array configurations in AWP mode, according to some embodiments of this disclosure, are shown. [Figure 11]
[0020] The following are some embodiments of the present disclosure that demonstrate template matching performed on a search region around an initial MV. [Figure 12]
[0021] An exemplary flowchart of a method for applying template matching to a GPM to refine motion, according to some embodiments of this disclosure, is shown. [Figure 13A]
[0022] Exemplary templates for a GPM according to some embodiments of this disclosure are shown. [Figure 13B]
[0022] Shows an exemplary template for GPM according to some embodiments of the present disclosure. [Figure 13C]
[0022] Shows an exemplary template for GPM according to some embodiments of the present disclosure. [Figure 14]
[0023] Shows another exemplary flowchart of a method for applying template matching to GPM to refine movement according to some embodiments of the present disclosure. [Figure 15A]
[0024] Shows another modified form of an exemplary template for GPM according to some embodiments of the present disclosure. [Figure 15B]
[0024] Shows another modified form of an exemplary template for GPM according to some embodiments of the present disclosure. [Figure 16]
[0025] Shows an exemplary relationship between the GPM classification mode and the GPM classification angle according to some embodiments of the present disclosure. [Figure 17]
[0026] Shows exemplary angles regarding a 16×16 block according to some embodiments of the present disclosure. [Figure 18A]
[0027] Shows exemplary weights of each sample of the GPM classification mode for each classification angle shown in FIG. 17 regarding a 16×16 block according to some embodiments of the present disclosure. [Figure 18B]
[0027] Shows exemplary weights of each sample of the GPM classification mode for each classification angle shown in FIG. 17 regarding a 16×16 block according to some embodiments of the present disclosure. [Figure 18C]
[0027] Shows exemplary weights of each sample of the GPM classification mode for each classification angle shown in FIG. 17 regarding a 16×16 block according to some embodiments of the present disclosure. [Figure 18D]
[0027] Shows exemplary weights of each sample of the GPM classification mode for each classification angle shown in FIG. 17 regarding a 16×16 block according to some embodiments of the present disclosure. [Figure 18E]
[0027] The following are exemplary weights for each sample of the GPM segmentation mode for each segmentation angle shown in Figure 17, relating to a 16 × 16 block according to some embodiments of the present disclosure. [Figure 18F]
[0027] The following are exemplary weights for each sample of the GPM segmentation mode for each segmentation angle shown in Figure 17, relating to a 16 × 16 block according to some embodiments of the present disclosure. [Figure 18G]
[0027] The following are exemplary weights for each sample of the GPM segmentation mode for each segmentation angle shown in Figure 17, relating to a 16 × 16 block according to some embodiments of the present disclosure. [Figure 18H]
[0027] The following are exemplary weights for each sample of the GPM segmentation mode for each segmentation angle shown in Figure 17, relating to a 16 × 16 block according to some embodiments of the present disclosure. [Figure 18I]
[0027] The following are exemplary weights for each sample of the GPM segmentation mode for each segmentation angle shown in Figure 17, relating to a 16 × 16 block according to some embodiments of the present disclosure. [Figure 18J]
[0027] The following are exemplary weights for each sample of the GPM segmentation mode for each segmentation angle shown in Figure 17, relating to a 16 × 16 block according to some embodiments of the present disclosure. [Figure 18K]
[0027] The following are exemplary weights for each sample of the GPM segmentation mode for each segmentation angle shown in Figure 17, relating to a 16 × 16 block according to some embodiments of the present disclosure. [Figure 18L]
[0027] The following are exemplary weights for each sample of the GPM segmentation mode for each segmentation angle shown in Figure 17, relating to a 16 × 16 block according to some embodiments of the present disclosure. [Figure 18M]
[0027] The following are exemplary weights for each sample of the GPM segmentation mode for each segmentation angle shown in Figure 17, relating to a 16 × 16 block according to some embodiments of the present disclosure. [Figure 18N]
[0027] The following are exemplary weights for each sample of the GPM segmentation mode for each segmentation angle shown in Figure 17, relating to a 16 × 16 block according to some embodiments of the present disclosure. [Figure 18O]
[0027] The following are exemplary weights for each sample of the GPM segmentation mode for each segmentation angle shown in Figure 17, relating to a 16 × 16 block according to some embodiments of the present disclosure. [Figure 18P]
[0027] The following are exemplary weights for each sample of the GPM segmentation mode for each segmentation angle shown in Figure 17, relating to a 16 × 16 block according to some embodiments of the present disclosure. [Figure 18Q]
[0027] The following are exemplary weights for each sample of the GPM segmentation mode for each segmentation angle shown in Figure 17, relating to a 16 × 16 block according to some embodiments of the present disclosure. [Figure 18R]
[0027] The following are exemplary weights for each sample of the GPM segmentation mode for each segmentation angle shown in Figure 17, relating to a 16 × 16 block according to some embodiments of the present disclosure. [Figure 18S]
[0027] The following are exemplary weights for each sample of the GPM segmentation mode for each segmentation angle shown in Figure 17, relating to a 16 × 16 block according to some embodiments of the present disclosure. [Figure 18T]
[0027] The following are exemplary weights for each sample of the GPM segmentation mode for each segmentation angle shown in Figure 17, relating to a 16 × 16 block according to some embodiments of the present disclosure. [Figure 19]
[0028] Another exemplary flowchart of a method for applying template matching to a GPM to refine motion, according to some embodiments of this disclosure, is shown. [Figure 20]
[0029] Another exemplary flowchart of a method for applying template matching to a GPM to refine motion, according to some embodiments of this disclosure, is shown. [Figure 21]
[0030] The following are illustrative flowcharts illustrating methods for applying GPM, TM, and MMVD according to some embodiments of this disclosure. [Figure 22]
[0031] The following are illustrative flowcharts illustrating methods for applying GPM, TM, and MMVD according to some embodiments of this disclosure. [Modes for carrying out the invention]
[0009] Detailed explanation
[0032] Next, exemplary embodiments, illustrated in the accompanying drawings, will be described in detail. The following description refers to the accompanying drawings, where the same reference numerals in different drawings represent the same or similar elements unless otherwise indicated. The implementations shown in the following description of exemplary embodiments do not represent all implementations according to the present invention. Rather, they are merely examples of apparatus and methods according to aspects relating to the present invention as enumerated in the accompanying claims. Specific aspects of this disclosure will be described in more detail below. In the event of any conflict between terms and / or definitions incorporated by reference and those provided herein, the terms and definitions provided herein shall prevail.
[0010]
[0033] In July 2020, the Multipurpose Video Coding (VVC / H.266) standard, developed by the ITU-T Video Coding Expert Group (ITU-T VCEG) and the ISO / IEC Moving Picture Expert Group (ISO / IEC MPEG) Joint Video Experts Team (JVET), was completed and published as an international standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.
[0011]
[0034] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET is developing techniques that surpass HEVC using joint exploration model (JEM) reference software. Because the coding techniques are incorporated into JEM, JEM achieves substantially higher coding performance than HEVC.
[0012]
[0035] The VVC standard is a relatively recent development and continues to incorporate more encoding techniques to deliver better compression performance. VVC is based on the same hybrid video encoding system used in modern video compression standards such as HEVC, H.264 / AVC, MPEGT2, and H.263.
[0013]
[0036] Following the completion of the VVC standard, JVET began exploring new encoding tools to further improve the encoding performance of the VVC standard. In January 2021, the Enhanced Compression Model (ECM) was proposed and is being used as a new software foundation for development tools that surpass the VVC standard.
[0014]
[0037] Video is a set of static pictures (or "frames") arranged in chronological order to store visual information. A video capture device (e.g., a camera) can be used to capture and store these pictures in chronological order, and a video playback device (e.g., a television, computer, smartphone, tablet computer, video player, or any end-user terminal with display capabilities) can be used to display such pictures in chronological order. Depending on the application, the video capture device can also transmit the captured video in real time to a video playback device (e.g., a computer with a monitor) for purposes such as surveillance, conference hosting, or live broadcasting.
[0015]
[0038] To reduce the memory space and transmission bandwidth required for such applications, video can be compressed before storage and transmission, and decompressed before display. Compression and decompression can be performed by software executed by a processor (e.g., a general-purpose computer processor) or by specialized hardware. The module for compression is generally called an "encoder," and the module for decompression is generally called a "decoder." Encoders and decoders can be collectively called a "codec." Encoders and decoders can be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders can include circuit mechanisms such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic circuits, or any combination thereof. Software implementations of encoders and decoders can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression can be performed using various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, or similar. Depending on the application, a codec can decompress video from a first encoding standard and then recompress the decompressed video using a second encoding standard. In this case, the codec may be referred to as a "transcoder."
[0016]
[0039] A video encoding process can identify and maintain useful information that can be used to reconstruct a picture, while ignoring information that is not important for reconstruction. If the ignored, unimportant information cannot be fully reconstructed, such an encoding process may be called "lossy." Otherwise, it may be called "lossy." Most encoding processes are lossy, which is a trade-off to reduce the required memory space and transmission bandwidth.
[0017]
[0040] Useful information about the picture being encoded (referred to as the "current picture") includes changes relative to the reference picture (e.g., a previously encoded and reconstructed picture). Such changes can include changes in pixel position, brightness, or color, of which changes in position are of greatest interest. Changes in the position of a group of pixels representing an object can reflect the movement of the object between the reference picture and the current picture.
[0018]
[0041] A picture encoded without referencing another picture (i.e., the picture is its own reference picture) is called an "I-picture". A picture is called a "P-picture" when some or all of the blocks within it (e.g., blocks that generally refer to a portion of a video picture) are predicted using intra-prediction or inter-prediction with one reference picture (e.g., unidirectional prediction). A picture is called a "B-picture" when at least one block within it is predicted using two reference pictures (e.g., bidirectional prediction).
[0019]
[0042] Figure 1 shows the structure of an exemplary video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 may be live video or captured and archived video. The video 100 may be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). The video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., video files stored in a storage device), or a video supply interface for receiving video from a video content provider (e.g., a video broadcast transceiver).
[0020]
[0043] As shown in Figure 1, the video sequence 100 may include a series of pictures arranged in time along a timeline, including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, with further pictures between pictures 106 and 108. In Figure 1, picture 102 is an I picture, and its reference picture is picture 102 itself. Picture 104 is a P picture, and its reference picture is picture 102, as indicated by the arrows. Picture 106 is a B picture, and its reference pictures are pictures 104 and 108, as indicated by the arrows. Depending on the embodiment, the reference picture of a picture (e.g., picture 104) may not be immediately before or after that picture. For example, the reference picture of picture 104 may be the picture before picture 102. Please note that the reference pictures 102-106 are merely examples, and this disclosure does not limit the embodiments of the reference pictures to the examples shown in Figure 1.
[0021]
[0044] Typically, video codecs do not encode or decode an entire picture at once due to the computational complexity of such tasks. Rather, they can divide the picture into basic segments and encode or decode the picture segment by segment. Such basic segments are referred to in this disclosure as basic processing units ("BPUs"). For example, structure 110 in Figure 1 shows an exemplary structure of a picture (e.g., any of pictures 102-108) in video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, their boundaries shown as dashed lines. Depending on the embodiment, basic processing units may be referred to as "macroblocks" in some video encoding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC) or as "coding tree units" ("CTUs") in some other video encoding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing unit can have a variable size in the picture, or any shape and size of pixels, such as 128×128, 64×64, 32×32, 16×16, 4×8, or 16×32. The size and shape of the basic processing unit can be selected for the picture based on a balance between encoding efficiency and the level of detail that should be maintained in the basic processing unit.
[0022]
[0045] A basic processing unit can be a logical unit that can contain groups of different types of video data stored in computer memory (for example, in a video frame buffer). For example, a basic processing unit for a color picture may include a luminance component (Y) representing colorless luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luminance and chroma components may have the same size as the basic processing unit. The luminance and chroma components may be referred to as a "coding tree block" (CTB) in some video encoding standards (e.g., H.265 / HEVC or H.266 / VVC). Any operations performed on the basic processing unit can be repeated on each of its luminance and chroma components.
[0023]
[0046] Video encoding has multiple computational stages, examples of which are shown in Figures 2A-2B and 3A-3B. For each stage, the size of the basic processing unit may still be too large for processing and therefore may be further divided into segments referred to in this disclosure as “basic processing subunits.” Depending on the embodiment, a basic processing subunit may be referred to as a “block” in some video encoding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC) or as an “encoding unit” (“CU”) in some other video encoding standards (e.g., H.265 / HEVC or H.266 / VVC). A basic processing subunit may be the same size as or smaller than a basic processing unit. Like a basic processing unit, a basic processing subunit is also a logical unit that may contain groups of different types of video data (e.g., Y, Cb, Cr, and associated syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on the basic processing subunit can be repeated on each of its luma and chroma components. Note that such divisions can be carried out to further levels as needed for processing. Also note that different stages can divide the basic processing unit using different methods.
[0024]
[0047] For example, during the mode determination phase (an example of which is shown in Figure 2B), the encoder can determine which prediction mode (e.g., intrapicture prediction or interpicture prediction) to use for the basic processing unit, but the basic processing unit may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing subunits (e.g., CUs, as in the case of H.265 / HEVC or H.266 / VVC) and determine the type of prediction for each basic processing subunit.
[0025]
[0048] As another example, in the prediction phase (an example of which is shown in Figures 2A and 2B), the encoder can perform prediction calculations at the level of the basic processing subunit (e.g., CU). However, in some cases, the basic processing subunit may still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (e.g., referred to as "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), and prediction calculations can be performed at that level.
[0026]
[0049] As another example, in the transformation stage (an example of which is shown in Figures 2A and 2B), the encoder can perform transformation operations for residual basic processing subunits (e.g., CUs). However, in some cases, the basic processing subunit may still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (e.g., referred to as "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), and transformation operations can be performed at that level. Note that the division method for the same basic processing subunit may differ between the prediction and transformation stages. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.
[0027]
[0050] In the structure 110 of Figure 1, the basic processing unit 112 is further divided into 3x3 basic processing subunits, and their boundaries are shown as dotted lines. Different basic processing units of the same picture may be divided into basic processing subunits in different ways.
[0028]
[0051] Depending on the implementation, a single picture can be divided into processing regions to bring parallel processing and error tolerance to video encoding and decoding. This means that the encoding or decoding process does not have to rely on information from any other regions of the picture for any given region of the picture. In other words, each region of a picture can be processed independently. By doing so, the codec can process different regions of a picture in parallel, thus increasing encoding efficiency. Furthermore, if the data in a region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thus providing error tolerance. Some video encoding standards allow a picture to be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC offer two types of regions: "slices" and "tiles". It should also be noted that different pictures in video sequence 100 may have different partitioning schemes for dividing the picture into regions.
[0029]
[0052] For example, in Figure 1, structure 110 is divided into three regions 114, 116, and 118, their boundaries shown as solid lines within structure 110. Region 114 contains four basic processing units. Regions 116 and 118 each contain six basic processing units. Note that the basic processing units, basic processing subunits, and regions of structure 110 in Figure 1 are merely examples, and this disclosure does not limit its embodiments.
[0030]
[0053] Figure 2A shows a schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure. For example, the encoding process 200A may be performed by an encoder. As shown in Figure 2A, the encoder can encode a video sequence 202 into a video bitstream 228 according to process 200A. Similar to video sequence 100 in Figure 1, video sequence 202 may include a set of pictures arranged in chronological order (referred to as “original pictures”). Similar to structure 110 in Figure 1, each original picture in video sequence 202 may be divided by the encoder into a basic processing unit, basic processing subunit, or region for processing. In some embodiments, the encoder can perform process 200A at the level of a basic processing unit for each original picture in video sequence 202. For example, the encoder can perform process 200A in an iterative manner, in which case the encoder can encode a basic processing unit in a single iteration of process 200A. Depending on the embodiment, the encoder can perform process 200A in parallel for each region of the original picture in the video sequence 202 (e.g., regions 114-118).
[0031]
[0054] In Figure 2A, the encoder can supply the basic processing unit (referred to as the "original BPU") of the original picture of the video sequence 202 to the prediction stage 204, generating prediction data 206 and prediction BPU 208. The encoder can subtract the prediction BPU 208 from the original BPU to generate residual BPU 210. The encoder can supply the residual BPU 210 to the conversion stage 212 and the quantization stage 214, generating quantization conversion coefficients 216. The encoder can supply the prediction data 206 and quantization conversion coefficients 216 to the binary coding stage 226, generating video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as the "forward path". During process 200A, after the quantization stage 214, the encoder can supply the quantization conversion coefficients 216 to the inverse quantization stage 218 and the inverse conversion stage 220 to generate the reconstructed residual BPU 222. The encoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate the prediction criterion 224, which will be used in the prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as the “reconstruction path”. The reconstruction path may be used to ensure that both the encoder and the decoder use the same reference data for prediction.
[0032]
[0055] The encoder can iteratively perform process 200A to encode each original BPU of the original picture (in the forward path) and to generate a prediction criterion 224 (in the reconstruction path) for encoding the next original BPU of the original picture. After encoding all the original BPUs of the original picture, the encoder can proceed to encode the next picture in the video sequence 202.
[0033]
[0056] Referring to process 200A, the encoder can receive a video sequence 202 generated by a video acquisition device (e.g., a camera). As used herein, the term “receive” can mean receiving, inputting, acquiring, obtaining, getting, reading, accessing, or any act in any way for inputting data.
[0034]
[0057] In prediction stage 204, in the current iteration, the encoder receives the original BPU and prediction criterion 224, performs prediction calculations, and can generate prediction data 206 and prediction BPU 208. The prediction criterion 224 may be generated from the reconstruction path of a previous iteration of process 200A. The objective of prediction stage 204 is to reduce information redundancy by extracting prediction data 206, which can be used to reconstruct the original BPU as prediction BPU 208 from the prediction data 206 and prediction criterion 224.
[0035]
[0058] Ideally, the predicted BPU 208 can be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 is generally slightly different from the original BPU. To record such differences, after generating the predicted BPU 208, the encoder can subtract it from the original BPU to generate the residual BPU 210. For example, the encoder can subtract the pixel values (e.g., grayscale or RGB values) of the predicted BPU 208 from the corresponding pixel values of the original BPU. Each pixel of the residual BPU 210 may have a residual value resulting from such a subtraction between the original BPU and the corresponding pixel of the predicted BPU 208. Compared to the original BPU, the predicted data 206 and residual BPU 210 may have fewer bits, but they can be used to reconstruct the original BPU without significant quality degradation. Therefore, the original BPU is compressed.
[0036]
[0059] To further compress the residual BPU210, in transformation step 212, the encoder can reduce the spatial redundancy of the residual BPU210 by decomposing it into a set of two-dimensional "basis patterns," each basis pattern associated with a "transformation coefficient." The basis patterns can have the same size (e.g., the size of the residual BPU210). Each basis pattern can represent the changing frequency components (e.g., the frequency of brightness changes) of the residual BPU210. No basis pattern can be reconstructed from any combination (e.g., a linear combination) of any other basis pattern. In other words, the decomposition can decompose the changes in the residual BPU210 into the frequency domain. Such a decomposition is analogous to the discrete Fourier transform of a function, in which case the basis patterns are analogous to the basis functions of the discrete Fourier transform (e.g., trigonometric functions), and the transformation coefficients are analogous to the coefficients associated with the basis functions.
[0037]
[0060] Different transformation algorithms can use different basis patterns. For example, various transformation algorithms such as discrete cosine transform, discrete sine transform, or similar can be used in transformation stage 212. The transformation in transformation stage 212 is inversely operable. That is, the encoder can recover the residual BPU 210 by the inverse operation of the transformation (referred to as the "inverse transform"). For example, to recover the pixels of the residual BPU 210, the inverse transform can be generated by multiplying the values of the corresponding pixels in the basis pattern by their respective associated coefficients, adding the products, and generating a weighted sum. For video coding standards, both the encoder and decoder can use the same transformation algorithm (and therefore the same basis pattern). Therefore, the encoder can record only the transformation coefficients, and the decoder can reconstruct the residual BPU 210 from the transformation coefficients without receiving the basis pattern from the encoder. Compared to the residual BPU 210, the transformation coefficients may have fewer bits, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. Therefore, the residual BPU210 is further compressed.
[0038]
[0061] The encoder can further compress the conversion coefficients in the quantization stage 214. In the conversion process, different basis patterns can represent different change frequencies (e.g., brightness change frequencies). Since the human eye is generally better at recognizing low-frequency changes, the encoder can ignore information about high-frequency changes without causing significant quality degradation in decoding. For example, in the quantization stage 214, the encoder can generate quantization conversion coefficients 216 by dividing each conversion coefficient by an integer value (referred to as the "quantization scale factor") and rounding the quotient to its nearest integer. After such an operation, some conversion coefficients of high-frequency basis patterns may be converted to 0, and conversion coefficients of low-frequency basis patterns may be converted to smaller integers. The encoder can ignore quantization conversion coefficients 216 that are 0, thereby further compressing the conversion coefficients. The quantization process is also inversely operable, in which case the quantization conversion coefficients 216 can be reconstructed into conversion coefficients in the inverse operation of quantization (referred to as "inverse quantization").
[0039]
[0062] Because the encoder rounds off the remainder of such divisions, the quantization stage 214 can be irreversible. Typically, the quantization stage 214 can contribute the greatest information loss in process 200A. The greater the information loss, the fewer bits the quantization conversion coefficient 216 requires. To obtain different levels of information loss, the encoder can use different values for the quantization parameter or any other parameter of the quantization process.
[0040]
[0063] In the binary coding stage 226, the encoder can encode the predicted data 206 and the quantization conversion coefficients 216 using a binary coding technique such as entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. Depending on the embodiment, in addition to the predicted data 206 and the quantization conversion coefficients 216, the encoder can encode other information in the binary coding stage 226, such as the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the type of conversion in the conversion stage 212, the parameters of the quantization process (e.g., quantization parameters), the encoder control parameters (e.g., bitrate control parameters), or the like. The encoder can generate a video bitstream 228 using the output data from the binary coding stage 226. Depending on the embodiment, the video bitstream 228 can be further packetized for network transmission.
[0041]
[0064] Referring to the reconstruction path of process 200A, in the inverse quantization step 218, the encoder can perform inverse quantization on the quantization transformation coefficients 216 to generate reconstruction transformation coefficients. In the inverse transformation step 220, the encoder can generate reconstruction residual BPU 222 based on the reconstruction transformation coefficients. The encoder can add the reconstruction residual BPU 222 to the prediction BPU 208 to generate a prediction criterion 224, which will be used in the next iteration of process 200A.
[0042]
[0065] It should be noted that other variations of process 200A may also be used to encode the video sequence 202. In some embodiments, the steps of process 200A may be performed in a different order by the encoder. In some embodiments, one or more steps of process 200A may be combined into a single step. In some embodiments, a single step of process 200A may be divided into multiple steps. For example, the conversion step 212 and the quantization step 214 may be combined into a single step. In some embodiments, process 200A may include additional steps. In some embodiments, process 200A may omit one or more steps in Figure 2A.
[0043]
[0066] Figure 2B shows a schematic diagram of another exemplary encoding process 200B according to an embodiment of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used with an encoder compliant with a hybrid video encoding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B additionally includes a mode determination stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.
[0044]
[0067] Generally, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") can use pixels from one or more already encoded neighboring BPUs within the same picture to predict the current BPU. That is, the prediction criterion 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") can use regions from one or more already encoded pictures to predict the current BPU. That is, the prediction criterion 224 in temporal prediction can include encoded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.
[0045]
[0068] Referring to process 200B, within the forward path, the encoder performs prediction calculations in spatial prediction stage 2042 and temporal prediction stage 2044. For example, in spatial prediction stage 2042, the encoder may perform intra-prediction. For the original BPU of the picture being encoded, the prediction criterion 224 may include one or more neighboring BPUs encoded (within the forward path) and reconstructed (within the reconstruction path) within the same picture. The encoder may generate a prediction BPU 208 by extrapolating neighboring BPUs. The extrapolation technique may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, or similar. In some embodiments, the encoder may perform extrapolation at the pixel level, for example, by extrapolating the values of the corresponding pixels for each pixel of the prediction BPU 208. The adjacent BPUs used for extrapolation can be positioned relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., to the lower left, lower right, upper left, or upper right of the original BPU), or any direction defined in the video encoding standard used. For intra-prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the adjacent BPUs used, the size of the adjacent BPUs used, the extrapolation parameters, the orientation of the adjacent BPUs used relative to the original BPU, or similar.
[0046]
[0069] As another example, in the temporal prediction stage 2044, the encoder can perform interpretation. For the original BPU of the current picture, the prediction criterion 224 may include one or more pictures (referred to as "reference pictures") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures may be encoded and reconstructed for each BPU. For example, the encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs for the same picture have been generated, the encoder can generate a reconstructed picture as a reference picture. The encoder may perform a "motion estimation" operation to search for matching regions within a range of the reference picture (referred to as a "search window"). The location of the search window in the reference picture may be determined based on the location of the original BPU of the current picture. For example, the search window may have its center in the reference picture at a location with the same coordinates as the original BPU in the current picture and may extend outward over a predetermined distance. When the encoder identifies a region similar to the original BPU within the search window (for example, by using a pixel-recursive (pel-recursive) algorithm, a block-matching algorithm, or similar), the encoder can determine such a region as a matching region. The matching region may have different dimensions from the original BPU (for example, smaller than, equal to, larger than, or of a different shape than the original BPU). Because the reference picture and the current picture are temporally separated in the timeline (for example, as shown in Figure 1), the matching region can be considered to "move" to the location of the original BPU over time. The encoder can record the direction and distance of such movement as a "motion vector". When multiple reference pictures are used (for example, picture 106 in Figure 1), the encoder can search for a matching region for each reference picture and determine its associated motion vector. In some embodiments, the encoder can assign weights to the pixel values of the matching region of each matching reference picture.
[0047]
[0070] Motion estimation can be used to identify various types of motion, such as translation, rotation, zooming, or similar. For interpretation, the prediction data 206 may include, for example, the location of the matching region (e.g., coordinates), the motion vector associated with the matching region, the number of reference pictures, the weights associated with the reference pictures, or similar.
[0048]
[0071] To generate a predicted BPU 208, the encoder can perform a “motion compensation” operation. Motion compensation can be used to reconstruct the predicted BPU 208 based on prediction data 206 (e.g., motion vectors) and prediction criteria 224. For example, the encoder can move the matching region of a reference picture according to the motion vector, in which case the encoder can predict the original BPU of the current picture. When multiple reference pictures are used (e.g., picture 106 in Figure 1), the encoder can move the matching region of each reference picture according to its respective motion vector and average the pixel values of the matching region. In some embodiments, if the encoder has weighted the pixel values of the matching region of each matching reference picture, the encoder can add the weighted sum of the pixel values of the moved matching region.
[0049]
[0072] Depending on the embodiment, interpretation can be unidirectional or bidirectional. Unidirectional interpretation can use one or more reference pictures in the same time direction relative to the current picture. For example, picture 104 in Figure 1 is a unidirectional interpretation picture in which the reference picture (i.e., picture 102) precedes picture 104. Bidirectional interpretation can use one or more reference pictures in both time directions relative to the current picture. For example, picture 106 in Figure 1 is a bidirectional interpretation picture in which the reference pictures (i.e., pictures 104 and 108) are in both time directions relative to picture 104.
[0050]
[0073] Still referring to the forward path of process 200B, after the spatial prediction 2042 and the temporal prediction stage 2044, in the mode determination stage 230, the encoder may select a prediction mode (e.g., one of intra-prediction or inter-prediction) for the current iteration of process 200B. For example, the encoder may perform a rate-distortion optimization technique. In this technique, the encoder may select a prediction mode to minimize the value of the cost function, depending on the bit rate of the candidate prediction mode and the distortion of the reconstructed reference picture under the candidate prediction mode. Depending on the selected prediction mode, the encoder may generate the corresponding prediction BPU 208 and prediction data 206.
[0051]
[0074] If the intra-prediction mode is selected in the forward path within the reconstruction path of process 200B, after generating the prediction criterion 224 (e.g., the current BPU encoded and reconstructed in the current picture), the encoder can directly feed the prediction criterion 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). The encoder can also feed the prediction criterion 224 to the loop filtering stage 232, where the encoder can apply a loop filter to the prediction criterion 224 to reduce or eliminate distortions (e.g., blocking artifacts) introduced during the encoding of the prediction criterion 224. The encoder can apply various loop filtering techniques in the loop filtering stage 232, such as deblocking, sample-adaptive offset, adaptive loop filtering, or similar. The loop-filtered reference picture can be stored in the buffer 234 (or "decoded picture buffer") for later use (e.g., to be used as an inter-prediction reference picture for future pictures in the video sequence 202). The encoder can store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. Depending on the embodiment, the encoder can encode the parameters of the loop filter (e.g., loop filter strength) together with the quantization transformation coefficients 216, the prediction data 206, and other information in the binary coding stage 226.
[0052]
[0075] Figure 3A shows a schematic diagram of an exemplary decoding process 300A according to an embodiment of the present disclosure. Process 300A may be a decompression process corresponding to the compression process 200A in Figure 2A. Depending on the embodiment, process 300A may be similar to the reconstruction path of process 200A. The decoder can decode the video bitstream 228 into a video stream 304 according to process 300A. The video stream 304 may be very similar to the video sequence 202. However, due to information loss in the compression and decompression processes (e.g., the quantization stage 214 in Figures 2A-2B), the video stream 304 is generally not identical to the video sequence 202. Similar to processes 200A and 200B in Figures 2A-2B, the decoder may perform process 300A at the level of a basic processing unit (BPU) for each picture encoded within the video bitstream 228. For example, the decoder can perform process 300A in an iterative manner, in which case the decoder can decode the basic processing unit in a single iteration of process 300A. Depending on the embodiment, the decoder can perform process 300A in parallel for each region of picture encoded in the video bitstream 228 (e.g., regions 114-118).
[0053]
[0076] In Figure 3A, the decoder can supply a portion of the video bitstream 228 associated with the basic processing unit of the encoded picture (referred to as the "encoded BPU") to the binary decoding stage 302. In the binary decoding stage 302, the decoder can decode this portion into prediction data 206 and quantization conversion coefficients 216. The decoder can supply the quantization conversion coefficients 216 to the inverse quantization stage 218 and the inverse conversion stage 220 to generate the reconstructed residual BPU 222. The decoder can supply the prediction data 206 to the prediction stage 204 to generate the prediction BPU 208. The decoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate the prediction criterion 224. Depending on the embodiment, the prediction criterion 224 can be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder can supply the prediction criterion 224 to the prediction stage 204 for the prediction calculation of process 300A.
[0054]
[0077] The decoder can iteratively perform process 300A to decode each encoding BPU of the encoded picture and generate a prediction criterion 224 for encoding the next encoding BPU of the encoded picture. After decoding all encoding BPUs of the encoded picture, the decoder can output the picture to the video stream 304 for display and proceed to decode the next encoded picture in the video bitstream 228.
[0055]
[0078] In the binary decoding stage 302, the decoder can perform the inverse operation of the binary coding technique used by the encoder (e.g., entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). Depending on the embodiment, in addition to the prediction data 206 and quantization conversion coefficients 216, the decoder can decode other information in the binary decoding stage 302, such as, for example, the prediction mode, the parameters of the prediction operation, the type of conversion, the parameters of the quantization process (e.g., quantization parameters), the encoder control parameters (e.g., bitrate control parameters), or the like. Depending on the embodiment, if the video bitstream 228 is transmitted in the form of packets over the network, the decoder can depacket the video bitstream 228 before supplying it to the binary decoding stage 302.
[0056]
[0079] Figure 3B shows a schematic diagram of another exemplary decoding process 300B according to an embodiment of the present disclosure. Process 300B may be modified from process 300A. For example, process 300B may be used with a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 300A, process 300B further divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.
[0057]
[0080] In process 300B, with respect to the encoding base processing unit ("current BPU") of the encoded picture being decoded ("current picture"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 may include various types of data, depending on which prediction mode was used by the encoder to encode the current BPU. For example, if intra-prediction was used by the encoder to encode the current BPU, the prediction data 206 may include intra-prediction, parameters of the intra-prediction operation, or a prediction mode indicator (e.g., a flag value) indicating the same. Parameters of the intra-prediction operation may include, for example, the location (e.g., coordinates) of one or more neighboring BPUs used as a reference, the size of the neighboring BPUs, extrapolation parameters, the orientation of the neighboring BPUs relative to the original BPU, or the same. As another example, if inter-prediction was used by the encoder to encode the current BPU, the prediction data 206 may include inter-prediction, parameters of the inter-prediction operation, or a prediction mode indicator (e.g., a flag value) indicating the same. The parameters for the interpretation calculation may include, for example, the number of reference pictures associated with the current BPU, the weights associated with each reference picture, the location (e.g., coordinates) of one or more matching regions within each reference picture, one or more motion vectors associated with each matching region, or similar.
[0058]
[0081] Based on the prediction mode indicator, the decoder can determine whether to perform a spatial prediction (e.g., intra-prediction) in the spatial prediction stage 2042 or a temporal prediction (e.g., inter-prediction) in the temporal prediction stage 2044. Details of performing such spatial or temporal predictions are shown in Figure 2B and will not be repeated below. After performing such spatial or temporal predictions, the decoder can generate a prediction BPU 208. The decoder can then add the prediction BPU 208 and the reconstructed residual BPU 222 to generate a prediction criterion 224, as described in Figure 3A.
[0059]
[0082] In process 300B, the decoder may supply the prediction criterion 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 in order to perform prediction calculations in the next iteration of process 300B. For example, if the current BPU is decoded using intra-prediction in the spatial prediction stage 2042, after generating the prediction criterion 224 (e.g., the decoded current BPU), the decoder may directly supply the prediction criterion 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If the current BPU is decoded using inter-prediction in the temporal prediction stage 2044, after generating the prediction criterion 224 (e.g., the reference picture with all BPUs decoded), the encoder may supply the prediction criterion 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may apply the loop filter to the prediction criterion 224 in the manner described in Figure 2B. Loop-filtered reference pictures may be stored in buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., to be used as inter-prediction reference pictures for future encoded pictures of the video bitstream 228). The decoder may store one or more reference pictures in buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the prediction data may further include loop filter parameters (e.g., loop filter strength). In some embodiments, the prediction data includes loop filter parameters if the prediction mode indicator in prediction data 206 indicates that inter-prediction was used to encode the current BPU.
[0060]
[0083] Figure 4 is a block diagram of an exemplary apparatus 400 for encoding or decoding video according to an embodiment of the present disclosure. As shown in Figure 4, the apparatus 400 may include a processor 402. When the processor 402 executes instructions as described herein, the apparatus 400 can become a specialized machine for video encoding or decoding. The processor 402 can be any type of circuit mechanism having the ability to manipulate or process information. For example, the processor 402 may include any number or any combination of a central processing unit (or "CPU"), graphics processing unit (or "GPU"), neural processing unit ("NPU"), microcontroller unit ("MCU"), optical processor, programmable logic controller, microcontroller, microprocessor, digital signal processor, intellectual property (IP) core, programmable logic array (PLA), programmable array logic (PAL), generic array logic (GAL), composite programmable logic unit (CPLD), field programmable gate array (FPGA), system-on-a-chip (SoC), application-specific integrated circuit (ASIC), or similar. Depending on the embodiment, the processor 402 may also be a set of processors grouped as a single logical component. For example, as shown in Figure 4, the processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0061]
[0084] The device 400 may also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, or the like). For example, as shown in Figure 4, the stored data may include program instructions (e.g., program instructions for carrying out steps in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 may access the program instructions and data for processing (e.g., via bus 410), execute the program instructions, and perform arithmetic or operations on the data for processing. The memory 404 may include a high-speed random-access storage device or a non-volatile storage device. Depending on the embodiment, the memory 404 may include any number and any combination of random-access memory (RAM), read-only memory (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, CompactFlash® (CF) cards, or the like. Memory 404 can also be a group of memories grouped together as a single logical component (not shown in Figure 4).
[0062]
[0085] Bus 410 can be a communication device that transfers data between internal components of the device 400, such as an internal bus (e.g., a CPU-memory bus), an external bus (e.g., a universal serial bus port, a peripheral component interconnect express port), or similar.
[0063]
[0086] To facilitate explanation without creating ambiguity, the processor 402 and other data processing circuits are collectively referred to as “data processing circuits” in this disclosure. The data processing circuits may be implemented entirely as hardware, or as a combination of software, hardware, or firmware. In addition, the data processing circuits may be a single, standalone module, or may be fully or partially combined with any other component of the device 400.
[0064]
[0087] The device 400 may further include a network interface 406 for providing wired or wireless communication to a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, or the same). Depending on the embodiment, the network interface 406 may include any number or any combination of a network interface controller (NIC), radio frequency (RF) module, transponder, transceiver, modem, router, gateway, wired network adapter, wireless network adapter, Bluetooth® adapter, infrared adapter, near-field communication ("NFC") adapter, cellular network chip, or the same.
[0065]
[0088] Depending on the embodiment, the apparatus 400 may optionally further include a peripheral interface 408 for providing connectivity to one or more peripheral devices. As shown in Figure 4, the peripheral devices may include, but are not limited to, cursor control devices (e.g., mouse, touchpad, or touchscreen), keyboards, displays (e.g., cathode ray tube displays, liquid crystal displays, or light-emitting diode displays), video input devices (e.g., cameras, or input interfaces coupled to video archives), or similar.
[0066]
[0089] It should be noted that the video codec (for example, the codec that performs processes 200A, 200B, 300A, or 300B) may be implemented as any combination of any software or hardware modules within the device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more software modules of the device 400, such as program instructions that can be loaded into memory 404. As another example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of the device 400, such as special data processing circuits (e.g., FPGA, ASIC, NPU, or similar).
[0067]
[0090] This disclosure provides a method for performing motion vector refinement with respect to the GPM.
[0068]
[0091] VVC supports GPM (geometric partition mode) for inter-prediction. GPM is signaled using a CU level flag as a type of merge mode, along with other merge modes such as normal merge mode, motion vector difference merge mode (MMVD) mode, combined inter-intra prediction (CIIP) mode, and sub-block merge mode. Possible CU size w×h=2 m ×2 n A total of 64 divisions are supported by the GPM, where m,n ∈ {3···6}, and 8×64 and 64×8 are excluded.
[0069]
[0092] When GPM is used, the CU is divided into two parts by a geometrically positioned straight line. Figure 5 shows an exemplary GPM (geometric partition mode) division grouped by the same angle, according to some embodiments of this disclosure. As shown in Figure 5, the position of the dividing line is mathematically derived from the angle and offset parameters of a particular division. Each part of the geometric division within the CU is interpreted using its own motion, and only unidirectional predictions are allowed for each division; that is, each part has one motion vector and one reference index. This unidirectional prediction motion constraint is applied to ensure that only two motion-compensated predictions are needed for each CU (which is the same as conventional bidirectional predictions). The unidirectional prediction motion for each division is derived using a unidirectional prediction candidate list construction process, which is described in more detail below.
[0070]
[0093] When GPM is used for the current CU, the prediction signal for the entire CU is expressed as follows: Geometric segment indices (angle and offset) indicating the segmentation mode of the geometric segment are signaled. Then two merge indices (one for each segment) are further signaled. The size of the maximum number of GPM candidates is explicitly signaled within the SPS, specifying the binarization of the syntax for the GPM merge indices. After predicting each of the segments of the geometric segment, sample values along the edges of the geometric segment are adjusted using a mixed process with adaptive weights, which is described in more detail below. The transformation and quantization processes applied to the entire CU are the same as those applied in other prediction modes. Finally, the motion field of the CU predicted using GPM is stored. The detailed process for storing the motion field with respect to GPM is described in more detail below.
[0071]
[0094] The process for constructing the unidirectional prediction candidate list is described below in detail. The unidirectional prediction candidate list is derived directly from the merge candidate list, which is typically configured for merge mode. The index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list is denoted by n. The motion vector of the nth extended merge candidate, denoted by LX (where X is equal to the parity of n), is used as the nth unidirectional prediction motion vector for the GPM. Figure 6 shows an exemplary unidirectional prediction motion vector (MV) selection process for the GPM according to some embodiments of this disclosure. As shown in Figure 6, the motion vectors are marked with "x". If a corresponding LX motion vector for the nth extended merge candidate does not exist, the L(1-X) motion vector of the same candidate is used instead as the unidirectional prediction motion vector for the GPM.
[0072]
[0095] Regarding the process for mixing along the edges of a geometric division, after predicting each part of the geometric division using its own motion, mixing is applied to the two predicted signals of each part to derive samples around the edges of the geometric division. The mixing weights for each CU location are derived based on the distance between the individual location and the edge of the division.
[0073]
[0096] Figure 7 shows an exemplary generation of weight w0 using GPM according to some embodiments of this disclosure. As shown in Figure 7, the distance (x,y) to the edge of the section is derived as follows:
number
[0074]
[0097] The weights of each part of the geometric division are derived as follows:
number
[0075]
[0098] Regarding the memory of fields related to GPM, MV1 from the first part of the geometric division, MV2 from the second part of the geometric division, and the combined MV of MV1 and MV2 are stored in the movement field of the GPM encoded CU.
[0076]
[0099] The type of motion vector stored for each individual position within the motion field is determined as follows: sType=abs(motionIdx)<32?2:(motionIdx≦0?(1-partIdx):partIdx) (8) However, motionIdx is equal to d(4x+2,4y+2) recalculated from equation (1). partIdx depends on the angle index i.
[0077]
[0100] If sType is equal to 0 or 1, MV1 or MV2 is stored in the corresponding motion field; otherwise, if sType is equal to 2, a combined MV from MV1 and MV2 is stored. The combined MV is generated using the following process: (1) if MV1 and MV2 are from different reference picture lists (one from L0 and the other from L1), MV1 and MV2 are simply combined to form a bidirectional predicted motion vector; and (2) otherwise, if MV1 and MV2 are from the same list, only the unidirectional predicted motion MV2 is stored.
[0078]
[0101] Similar to GPM in VVC, the Audio Video Coding Standard 3 (AVS3) employs a tool called Angular Weighted Prediction (AWP). The AVS3 video standard was developed by the AVS Working Group, established in China in 2002. Its predecessors, AVS1 and AVS2, were published as Chinese national standards in 2006 and 2016, respectively. AVS3 supports AWP mode for both skip and direct modes. AWP mode is signaled using a CU level flag as a type of skip or direct mode. In AWP mode, a motion vector candidate list containing five different unidirectional predictive motion vectors is constructed by deriving motion vectors from spatially adjacent blocks and temporal motion vector predictors. Two unidirectional predictive motion vectors are then selected from the motion vector candidate list to predict the current block. Unlike bidirectional predictive intermode, where all samples have equal weights, each sample encoded in AWP mode can have different weights. The weight of each sample is predicted from a weight array with values from 0 to 8.
[0079]
[0102] Figure 8 shows an exemplary process for angle-weighted prediction (AWP) weight prediction according to some embodiments of the present disclosure. Figure 9 shows eight exemplary intra-prediction angles supported in AWP mode according to some embodiments of the present disclosure. Figure 10 shows seven different exemplary weight array settings in AWP mode according to some embodiments of the present disclosure. As shown in Figure 8, angle-weighted prediction is similar to the process in intra-prediction mode. Possible CU size w×h=2 where m,n∈{3···6} m ×2 nEach AWP mode supports a total of 56 different types of weights, including 8 intra-prediction angles (shown in Figure 9) and 7 different weight array configurations (shown in Figure 10). It should be noted that AWP mode signals directly to the decoder without prediction. AWP mode indices are binarized using truncated binary; that is, indices from 0 to 7 are encoded using 5 bits, and indices from 8 to 55 are encoded using 6 bits.
[0080]
[0103] Assume that the two selected unidirectional predicted motion vectors are MV1 and MV2. Two predicted blocks P0 and P1 are obtained by performing motion compensation using MV1 and MV2, respectively. The final predicted block P is calculated as follows: P=(P0×w0+P1×(8-w0))>>3 (9) However, w0 is the weight matrix derived by the weight prediction method described above.
[0081]
[0104] After prediction, the unidirectional predicted motion vectors are stored in a 4x4 subdivision. For each 4x4 unit, one of the two unidirectional predicted motion vectors is stored.
[0082]
[0105] Template matching (TM) is a decoder-side MV derivation method for refining motion information of the current CU by finding the closest match between a template in the current picture (e.g., adjacent blocks above and / or to the left of the current CU) and a block in a reference picture (e.g., the same size as the template). Figure 11 shows template matching performed on a search area around an initial MV according to some embodiments of the present disclosure. As shown in Figure 11, a better MV is searched around the initial motion of the current CU within a search area of [-8, +8] pixels. The TM mode can be applied to merge mode and AMVP mode.
[0083]
[0106] When applied to merge mode, the merge candidate indicated by the signaled merge index is used as the initial movement. The search methods shown in Table 1 are performed to refine the movement. Depending on whether an alternative interpolation filter (used when AMVR is in half-pixel mode) is used according to the movement information being merged, TM can be performed up to 1 / 8 pixel MVD precision, or anything exceeding half-pixel MVD precision can be skipped.
[0084] [Table 1]
[0085]
[0107] In AMVP mode, the MVP candidate is determined based on the template matching error to obtain the one that reaches the minimum difference between the current block's template and the reference block's template, and then the TM performs MV refinement only on this particular MVP candidate. The TM refines this MVP candidate by using an iterative diamond search, starting with 1-pixel MVD precision (or 4-pixel for 4-pixel AMVR mode) within a search range of [-8, +8] pixels. Depending on the AMVR mode specified in Table 1, the AMVP candidate can be further refined by using a cross search with 1-pixel MVD precision (or 4-pixel for 4-pixel AMVR mode), followed by half-pixel and quarter-pixel searches. This search process ensures that the MVP candidate continues to maintain the same MV precision indicated by the AMVR mode after the TM process.
[0086]
[0108] A Merge Mode by Motion Vector Difference (MMVD) is introduced to the VVC, which signals MVD for merge candidates. The MMVD flag is signaled immediately after the normal merge flag is sent to specify whether MMVD mode is used in the CU. In MMVD, after a merge candidate is selected, the candidate is further refined by the signaled MVD information. The further information includes a merge candidate flag, an index to indicate the magnitude of the motion, and an index to indicate the direction of the motion. In MMVD mode, one of the first two candidates in the merge list is selected to be used as the basis for the MV. The MMVD candidate flag is signaled to indicate which of the first and second merge candidates to use.
[0087]
[0109] The distance index specifies information about the magnitude of the movement and indicates a default offset from the starting point. In MMVD mode, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the default offset is specified in Table 2 below.
[0088] [Table 2]
[0089]
[0110] The direction index represents the direction of the MVD relative to the starting point. The direction index can represent four directions, as shown in the table below. It should be noted that the meaning of the sign of the MVD can change depending on the information of the starting MV. When both lists point to the same side of the current picture (for example, the picture order counts (POCs) of both references are either greater than or less than the current picture's POC), and the starting MV is a unidirectional or bidirectional predicted MV, the sign in the table below specifies the sign of the MV offset added to the starting MV. When the two MVs point to different sides of the current picture (for example, the POC of one reference is greater than the current picture's POC, and the POC of the other reference is less than the current picture's POC), and the starting MV is a bidirectional predicted MV, and the difference in POCs in List 0 is greater than the difference in POCs in List 1, the sign in the table below specifies the sign of the MV offset added to the MV component of List 0 of the starting MV, and the sign of the MV in List 1 has the opposite value. If the difference in POCs within List 1 is greater than that in List 0, the signs in the table below specify the sign of the MV offset added to the MV component of List 1 for the starting MV, and the signs of the MVs in List 0 have the opposite value.
[0090] [Table 3]
[0091]
[0111] Recently, it has been proposed that MMVD be applied to GPM. When CU is encoded using GPM mode, each geometric segment can freely choose whether or not its motion is refined by the signaled MVD information. Two additional flags are signaled to indicate whether MMVD is applied to each of two geometric segments. To allow for more flexible combinations of refining the MV of two GPM segments, it should be noted that the following conditions can be applied to the two selected MVs of the two GPM segments: a. If MV refinement is not applied to either the first GPM category or the second GPM category, the two selected MVs in the two GPM categories cannot be considered identical. b. If MV refinement is applied to one of the two GPM categories and not to the other, it is recognized that the two selected MVs in the two GPM categories are the same. c. When both GPM categories apply MV refinement, the two selected MVs are considered the same if the MV refinements of the two categories are different, but they are not considered the same if the MV refinements of the two categories are identical.
[0092]
[0112] TM mode refines motion on the decoder side without signaling motion vector differences. However, this is only applicable to the normal merge mode and not to GPM. Therefore, GPM cannot benefit from TM mode, which can provide more accurate motion predictions for one or both GPM segments.
[0093]
[0113] This disclosure proposes a method for applying template matching to GPM in order to refine the motion.
[0094]
[0114] Figure 12 shows an exemplary flowchart of Method 1200 for applying template matching to a GPM to refine motion, according to some embodiments of the present disclosure. Method 1200 may be performed as part of a video encoding process (e.g., process 200A in Figure 2A or process 200B in Figure 2B) or by one or more software or hardware components of an apparatus (e.g., apparatus 400 in Figure 4). For example, a processor (e.g., processor 402 in Figure 4) may perform Method 1200. Depending on the embodiment, Method 1200 may be performed by a computer program product embodied in a computer-readable medium containing computer-executable instructions, such as program code, which is executed by a computer (e.g., apparatus 400 in Figure 4). Referring to Figure 12, Method 1200 may include the following steps 1202 to 1206.
[0095]
[0115] In step 1202, it is determined whether CU is encoded by GPM.
[0096]
[0116] In step 1204, in response to the CU being encoded by the GPM which divides the CU into a first and second section, a parameter is signaled to indicate whether the TM applies. For example, the parameter could be a flag used to indicate whether the TM applies to the entire CU (e.g., this flag is signaled at the CU level), or multiple parameters (e.g., flags) could be used to indicate whether the TM applies to different sections, each separately. Further details regarding the parameters are described below.
[0097]
[0117] In step 1206, in response to the application of TM, the motion of the GPM segment is refined using TM. If TM is not applied to the CU, the motion is not refined. In some embodiments, if TM is not applied to the CU, the motion can be refined using other methods.
[0098]
[0118] Depending on the embodiment, to obtain further flexibility, it is determined whether to apply TM to each segment individually. For example, if the coding unit is coded by GPM, a first parameter (e.g., a first flag) is signaled to indicate whether the first movement of the first segment (indicated by a first merge index) is refined using TM. Then, a second parameter (e.g., a second flag) is signaled to indicate whether the second movement of the second segment (indicated by a second merge index) is refined using TM. An example is shown in Table 4.
[0099] [Table 4]
[0100]
[0119] If both the first and second parameters are equal to 0, TM does not apply to the two divisions. If the first parameter is equal to 0 and the second parameter is equal to 1, TM applies only to the second division. If the first parameter is equal to 1 and the second parameter is equal to 0, TM applies only to the first division. If both the first and second parameters are equal to 1, TM applies to both divisions.
[0101]
[0120] As shown in Table 5, it should be noted that the first and second parameters can be joined to a third parameter (e.g., an index).
[0102] [Table 5]
[0103]
[0121] If the third parameter is equal to 0, TM is not applied to the two divisions. Neither the first movement of the first division nor the second movement of the second division is refined using TM. If the third parameter is equal to 1, TM is applied only to the second division. The second movement of the second division is refined using TM, while the first movement of the first division is not. If the third parameter is equal to 2, TM is applied only to the first division. The first movement of the first division is refined using TM, while the second movement of the second division is not. If the third parameter is equal to 3, TM is applied to both divisions. Both the first and second movements are refined using TM. In some embodiments, if the third parameter is equal to 1, TM is applied only to the first division. If the third parameter is equal to 2, TM is applied only to the second division.
[0104]
[0122] Depending on the embodiment, when refining the movement of the GPM, a template is constructed from adjacent samples to the left and / or above. Figures 13A to 13C show three exemplary templates for a GPM according to several embodiments of the present disclosure.
[0105]
[0123] As shown in Figure 13A, the template consists of both left neighbor samples and upper neighbor samples. For further flexibility, depending on the embodiment, instead of always using both upper and left neighbor samples, only left neighbor samples or only upper neighbor samples can be used. As shown in Figure 13B, the template consists of only upper neighbor samples. As shown in Figure 13C, the template consists of only left neighbor samples. Furthermore, when a TM is applied to a GPM coded block, several parameters are signaled to indicate which of the three templates is used. For example, an index is signaled to indicate the template. If the index is equal to 0, both upper and left neighbor samples are used to construct the template (as shown in Figure 13A). If the index is equal to 1, only upper neighbor samples are used (as shown in Figure 13B). If the index is equal to 2, only left neighbor samples are used (as shown in Figure 13C). Depending on the embodiment, three parameters are signaled to indicate each of the three templates.
[0106]
[0124] It is assumed that each category can individually select a template. For example, Category 1 can select the adjacent sample above as its template, and Category 2 can select the adjacent sample to the left as its template.
[0107]
[0125] The embodiments described above can be combined in any suitable manner. Figure 14 shows another exemplary flowchart of Method 1400 for applying template matching to a GPM to refine motion, according to some embodiments of the present disclosure. As shown in Figure 14, Method 1400 may include the following steps 1402-1410.
[0108]
[0126] In step 1402, it is determined whether CU is encoded by GPM.
[0109]
[0127] In step 1404, in response to the fact that the CU is encoded by a GPM that divides the CU into a first segment and a second segment, a first parameter (e.g., a first flag) is signaled to indicate whether the TM applies to the first segment.
[0110]
[0128] In step 1406, a second parameter (e.g., a second flag) is signaled to indicate whether TM applies to the second category. Thus, it is possible to individually determine whether TM applies to each category.
[0111]
[0129] In step 1408, in response to the application of TM to the first segment, a first index is signaled to indicate which of the three templates is used to refine the movement of the first segment. If TM is not applied to the first segment, the first index is not signaled.
[0112]
[0130] In step 1410, in response to the application of TM to the second segment, a second index is signaled to indicate which of the three templates is used to refine the movement of the second segment. If TM is not applied to the second segment, the second index is not signaled.
[0113]
[0131] This allows us to refine the movement of the first segment using TM along with the template determined by the first index, and to refine the movement of the second segment using TM along with the template determined by the second index.
[0114]
[0132] Depending on the embodiment, instead of signaling which template to use, we propose deriving the template based on the GPM.
[0115]
[0133] Figures 15A and 15B show another modification of an exemplary template for a GPM according to some embodiments of the present disclosure. Generally, instead of signaling an index of templates for a given GPM segment, the selection of a template (left neighbor sample, upper neighbor sample, or both left and upper neighbor samples) may depend on the segmentation mode of the GPM. Taking any of the GPM segments as an example, if the GPM segment has only upper neighbor samples, only the upper template is used; if the GPM segment has only left neighbor samples, only the left template is used; and if the GPM segment has both upper and left neighbor samples, both the upper and left templates are used.
[0116]
[0134] As shown in Figure 15A, in the first section 1510A, only the upper template 1511A is used to refine the movement. In the second section 1520A, only the left template 1521A is used to refine the movement.
[0117]
[0135] As shown in Figure 15B, in the first section 1510B, both the left template and the upper template 1511B are used to refine the movement. In the second section 1520B, only the upper template 1521B is used to refine the movement.
[0118]
[0136] In some embodiments, templates are derived based on GPM partition angles. Figure 16 shows an exemplary relationship between GPM partition modes and GPM partition angles in some embodiments of the present disclosure. Within the syntax, a GPM partition mode can be denoted as merge_gpm_partition_idx, a GPM partition angle as angleIdx, and a distance as distanceIdx. As shown in Figure 16, there are a total of 64 partition modes, including 20 angles and 4 distances.
[0119]
[0137] During the refinement of the GPM movement, the template is initially selected from only the left adjacent sample, only the upper adjacent sample, or both the left and upper adjacent samples, according to the segment angle. The basic principle for template selection is as follows: For a segment, the upper and left adjacent samples are determined. If only the upper adjacent sample is available for that segment and the left adjacent sample is unavailable, the template is selected from only the upper adjacent sample. If only the left adjacent sample is available for that segment and the upper adjacent sample is unavailable, the template is selected from only the left adjacent sample. If both the left and upper adjacent samples are available, the template is selected from both the left and upper adjacent samples.
[0120]
[0138] Figure 17 shows exemplary angles for a 16x16 block according to some embodiments of the present disclosure. As shown in Figure 17, the block can be divided into various angles having different numerical values, for example, shown as division angles. Figures 18A to 18T show exemplary weights for each sample of various GPM division modes for each division angle shown in Figure 17, for a 16x16 block according to some embodiments of the present disclosure. As shown in Figures 18A to 18T, the CU is divided into a first division 1810 and a second division 1820 by a division line. Adaptive weights are used to adjust the sample values along the edges of the geometric divisions (division lines). The first division 1810 is predicted using a first movement indicated by a first merge index, and the second division 1820 is predicted using a second movement indicated by a second merge index. The first and second merge indices are signaled when the block is encoded using GPM. In particular, as shown in Figures 18A, 18B, 18C, 18I, 18J, 18K, 18L, 18M, 18S, and 18T, for each of the segment angles 0, 2, 3, 13, 14, 16, 18, 19, 29, and 30, the upper adjacent sample is selected as the template for the first segment predicted using the first motion, and the left adjacent sample and the upper adjacent sample are selected as the templates for the second segment predicted using the second motion. As shown in Figures 18D and 18N, for each of the segment angles 4 and 20, the upper adjacent sample is selected as the template for the first segment, and the left adjacent sample is selected as the template for the second segment. As shown in Figures 18E, 18F, 18G, 18O, 18P, and 18Q, for division angles 5, 8, 11, 21, 24, and 27, both the left adjacent sample and the upper adjacent sample are selected as templates for the first division, and the left adjacent sample is selected as the template for the second division. For division angles 12 and 28, both the left adjacent sample and the upper adjacent sample are selected as templates for the first and second divisions.
[0121]
[0139] When refining the GPM movement, the search pattern can be any one of the patterns shown in Table 1. Depending on the embodiment, the search method may be the same as the method used in merge mode with the alternative interpolation filter turned off.
[0122]
[0140] Figure 19 shows another exemplary flowchart of Method 1900 for applying template matching to a GPM to refine motion, according to some embodiments of the present disclosure. Method 1900 may be performed as a video encoding process (e.g., process 200A in Figure 2A or process 200B in Figure 2B) or by one or more software or hardware components of an apparatus (e.g., apparatus 400 in Figure 4). For example, a processor (e.g., processor 402 in Figure 4) may perform Method 1900. Depending on the embodiment, Method 1900 may be performed by a computer program product embodied in a computer-readable medium containing computer-executable instructions such as program code, which is executed by a computer (e.g., apparatus 400 in Figure 4). Referring to Figure 19, Method 1900 may include the following steps 1902-1912.
[0123]
[0141] In step 1902, it is determined whether the CU is encoded in TM mode.
[0124]
[0142] In step 1904, in response to the fact that the CU is encoded in TM mode, a flag is signaled to indicate whether the CU is divided into two parts and predicted using the GPM. For example, if the flag is equal to 1, the CU is divided into a first and a second part, and the GPM is used for prediction. If the flag is equal to 0, the GPM is not applied, and the CU is not divided.
[0125]
[0143] In step 1906, in response to the application of GPM to the CU, one partitioning mode and two merge indices are further signaled. Thus, when GPM is applied, the CU is partitioned based on the partitioning mode. Two merge indices are signaled, each representing two movements of the first partition and the second partition.
[0126]
[0144] In step 1908, TM is used to refine the two movements indicated by the two merge indices.
[0127]
[0145] Depending on the embodiment, method 1900 may further include steps 1910 and 1912.
[0128]
[0146] In step 1910, motion compensation is performed using the refined motion.
[0129]
[0147] In step 1912, a mixing process is applied along the edges of the geometric divisions according to the division mode.
[0130]
[0148] Depending on the embodiment, steps 1910 and 1912 may be carried out in any other way, for example, in methods 1200 and 1400 in which TM is used to refine the movement of CU encoded by GPM.
[0131]
[0149] It should be noted that the method of applying TM to GPM can also be applied to AWP mode in the AVS3 standard.
[0132]
[0150] Figure 20 shows another exemplary flowchart of Method 2000 for applying template matching to a GPM to refine motion, according to some embodiments of the present disclosure. Method 2000 may be performed as part of a video decoding process (e.g., process 300A in Figure 3A or process 300B in Figure 3B) or by one or more software or hardware components of an apparatus (e.g., apparatus 400 in Figure 4). For example, a processor (e.g., processor 402 in Figure 4) may perform Method 2000. Depending on the embodiment, Method 2000 may be performed by a computer program product embodied in a computer-readable medium containing computer-executable instructions such as program code, which is executed by a computer (e.g., apparatus 400 in Figure 4). Referring to Figure 20, Method 2000 may include the following steps 2002-2008.
[0133]
[0151] In step 2002, the bitstream containing the encoding unit (for example, the video bitstream 228 in Figure 3B) is received by the decoder.
[0134]
[0152] In step 2004, it is determined whether the CU is encoded by GPM.
[0135]
[0153] In step 2006, in response to the block being encoded by the GPM which divides the CU into a first and second segment, parameters (e.g., flags) indicating whether a TM is applied are decoded. In some embodiments, the parameters may include multiple flags indicating whether a TM is applied to each segment (see Table 4 again). In some embodiments, the parameters may include an index indicating the combination of segments to which the TM is applied (see Table 5 again).
[0136]
[0154] In step 2008, in response to the application of TM, the motion of the GPM segment is refined using TM. If TM is not applied to the CU, the motion is not refined. Depending on the embodiment, if TM is not applied to the CU, the motion may be refined using other methods.
[0137]
[0155] This disclosure further provides a method for combining GPM with MMVD and TM.
[0138]
[0156] Depending on the embodiment, MMVD and TM cannot be applied to the same CU. When a CU is encoded using GPM, a first parameter (e.g., a first flag) is signaled to indicate whether TM is applied to the CU. If TM is applied, the GPM and two merge indices are further signaled. The movements of both GPM segments are then refined using TM. If TM is not applied to the CU, a second parameter is signaled to indicate whether MMVD is applied to the GPM segment. If MMVD is applied to the GPM segment, MVD information is further signaled, and the movements are refined using the signaled MVD information. In one example, the second parameter includes a second flag and a third flag. The second flag indicates whether MMVD is applied to the first GPM segment, and the third flag indicates whether MMVD is applied to the second GPM segment. In another example, the second parameter contains only the fourth flag, which indicates whether MMVD applies to both of the two GPM segments.
[0139]
[0157] Depending on the embodiment, MMVD and TM cannot be applied to the same GPM segment. When CU is encoded by GPM, CU is divided into two GPM segments. For each GPM segment, a first parameter is signaled to indicate whether TM is applied to the GPM segment. If TM is applied, the motion of the GPM segment is refined using TM. If TM is not applied, a second parameter is signaled to indicate whether MMVD is applied to the GPM segment. Note that if the first parameter indicates that TM is applied to the GPM segment, the second parameter is not signaled. Note that the two GPM segments can also individually choose whether to use TM or MMVD to refine their motion. That is, the motion of one GPM segment can be refined using TM, and the motion of the other GPM segment can be refined using MMVD.
[0140]
[0158] Depending on the embodiment, the signaling order of two parameters (a parameter indicating whether TM is applied and another parameter indicating whether MMVD is applied) can be rearranged. The parameter indicating whether MMVD is applied can be signaled before the parameter indicating whether TM is applied. If the parameter signals MMVD first and indicates whether MMVD is applied to a GPM segment, then the parameter indicating whether TM is applied is no longer signaled and is inferred to be off. Therefore, TM is not applied to the GPM segment. Figure 21 shows an exemplary flowchart of method 2100 for applying GPM, TM, and MMVD according to some embodiments of the present disclosure. Method 2100 may be performed as part of a video coding process (e.g., process 200A in Figure 2A or process 200B in Figure 2B) or by one or more software or hardware components of an apparatus (e.g., apparatus 400 in Figure 4). For example, a processor (e.g., processor 402 in Figure 4) may perform method 2100. Depending on the embodiment, Method 2100 may be carried out by a computer program product that is implemented in a computer-readable medium containing computer-executable instructions, such as program code, and is executed by a computer (e.g., Apparatus 400 in Figure 4). Referring to Figure 21, Method 2100 may include the following steps 2102 to 2110.
[0141]
[0159] In step 2102, it is determined whether the CU is encoded by GPM.
[0142]
[0160] In step 2104, in response to the fact that the CU is encoded by the GPM which divides the CU into a first and a second segment, a first parameter is signaled to indicate whether MMVD is applied to the first segment. For example, the first parameter may be a first flag. If the first flag is equal to 1, MMVD is applied to the first segment. If the first flag is equal to 0, MMVD is not applied to the first segment.
[0143]
[0161] In step 2106, a second parameter is signaled to indicate whether MMVD applies to the second category. For example, the second parameter could be a second flag. If the second flag is equal to 1, MMVD applies to the second category. If the second flag is equal to 0, MMVD does not apply to the second category.
[0144]
[0162] In step 2108, in response to the determination that MMVD does not apply to either the first or second category, a third parameter is signaled to indicate whether TM applies to both the first and second categories. For example, the third parameter may be a third flag. If the third flag is equal to 1, TM applies to both categories. If the third flag is equal to 0, TM does not apply to either category.
[0145]
[0163] In step 2110, in response to the application of TM to both the first and second segments, the motion of the first and second segments is refined using TM. Furthermore, in some embodiments, the template used to refine the motion is determined according to the segment mode or segment angle.
[0146]
[0164] Figure 22 shows an exemplary flowchart of Method 2200 for applying GPM, TM, and MMVD according to some embodiments of the present disclosure. Method 2100 may be performed as part of a video decoding process (e.g., process 300A in Figure 3A or process 300B in Figure 3B) or by one or more software or hardware components of an apparatus (e.g., apparatus 400 in Figure 4). For example, a processor (e.g., processor 402 in Figure 4) may perform Method 2200. Depending on the embodiment, Method 2200 may be performed by a computer program product embodied in a computer-readable medium containing computer-executable instructions such as program code, which is executed by a computer (e.g., apparatus 400 in Figure 4). Referring to Figure 22, Method 2200 may include the following steps 2202 to 2212.
[0147]
[0165] In step 2202, the bitstream containing the encoding unit (for example, the video bitstream 228 in Figure 3B) is received by the decoder.
[0148]
[0166] In step 2204, it is determined whether the CU is encoded by GPM.
[0149]
[0167] In step 2206, in response to the block being encoded by the GPM which divides the CU into a first and a second segment, a first parameter is decoded that indicates whether MMVD is applied to the first segment. For example, the first parameter can be a first flag, where if the first flag is equal to 1, MMVD is applied to the first segment, and if the first flag is equal to 0, MMVD is not applied to the first segment.
[0150]
[0168] In step 2208, a second parameter is decoded that indicates whether MMVD is applied to the second segment. For example, the second parameter can be a second flag, where if the second flag is equal to 1, MMVD is applied to the second segment, and if the second flag is equal to 0, MMVD is not applied to the second segment.
[0151]
[0169] In step 2210, in response to the determination that MMVD is not applied to the first or second category, a third parameter is decoded to indicate whether TM is applied to both the first and second categories. For example, the third parameter can be a third flag, where if the third flag is equal to 1, TM is applied to both categories, and if the third flag is equal to 0, TM is not applied to either category.
[0152]
[0170] In step 2210, in response to the application of TM to both the first and second segments, the motion of the first and second segments is refined using TM. Furthermore, in some embodiments, the template used to refine the motion is determined according to the segment mode or segment angle. In some embodiments, the template used to refine the motion is determined by decoding multiple parameters.
[0153]
[0171] In some embodiments, MMVD and TM can be applied to the same GPM segment. For each GPM segment, a first parameter and a second parameter are signaled to indicate whether TM and MMVD are applicable. In one example, when both TM and MMVD are applied to a GPM segment, the motion is first refined using TM. The refined motion is then further modified by adding the signaled MVD information. In another example, when both TM and MMVD are applied to a GPM segment, the signaled MVD information is first added to the motion. The modified motion is then used as the starting point for TM and refined by TM.
[0154]
[0172] In some embodiments, MMVD and TM can be applied to the same GPM segment only if the signaled MVD information satisfies several conditions. In one example, the conditions include the distance index of the signaled MVD information being less than a default value (e.g., 1, i.e., the MVD offset being less than 1 / 2 pixel). For each GPM segment, a second parameter is signaled to indicate whether MMVD is applicable. If MMVD is applicable, MVD information including the distance index and direction index is further signaled. If the distance index is less than a default value, a first parameter is signaled to indicate whether TM is applicable. If MMVD is not applicable, the first parameter is always signaled to indicate whether TM is applicable. In another example, the conditions include the distance index of the signaled MVD information being greater than a default value.
[0155]
[0173] In some embodiments, MMVD and TM can be applied to the same GPM segment only if the size / encoding mode of the GPM-encoded CU meets certain conditions. In one example, if the width and / or height of the CU exceeds a predetermined threshold (e.g., 16 or 32), both MMVD and TM can be applied to the same GPM segment. In another example, if the aspect ratio of the CU is less than a predetermined threshold, both MMVD and TM can be applied to the same GPM segment. The aspect ratio can be defined as CU_width / CU_height if CU_width > CU_height, or CU_height / CU_width if CU_height >= CU_width. In yet another example, if the CU is encoded using merge mode instead of skip mode, both MMVD and TM can be applied to the same GPM segment. In yet another example, if the CU is encoded using skip mode instead of merge mode, both MMVD and TM can be applied to the same GPM segment.
[0156]
[0174] Embodiments can be further described using the following clauses. 1. Receiving a bitstream containing encoded units encoded by GPM (geometric partition mode), Decoding a first parameter related to a coding unit, wherein the first parameter indicates whether template matching is applied to the coding unit, and Determining motion information relating to a coding unit according to a first parameter, wherein if the first parameter indicates that template matching is applied to the coding unit, the motion information is refined using template matching. A video decoding method, including the above. 2. The coding unit is divided into a first section and a second section, the motion information of the coding unit includes the first motion of the first section and the second motion of the second section, the first parameter includes the first flag, and the motion information is refined using template matching. Using template matching according to the first flag, refine both the first movement of the first segment and the second movement of the second segment. The method described in Clause 1, further including the method described in Clause 1. 3. The coding unit is divided into a first section and a second section, and the first parameter includes a first flag indicating whether template matching is applied to the first section and a second flag indicating whether template matching is applied to the second section, and the refinement of motion information is According to the first flag, template matching is used to refine the first movement of the first segment, and Using template matching according to the second flag, refine the second movement of the second segment. The method described in Clause 1, further including the method described in Clause 1. 4. The coding unit is divided into a first section and a second section, and the first parameter includes an index indicating whether template matching is applied to the first and second sections, and the refinement of motion information is performed. Refining the first movement of the first division, the second movement of the second division, or both the first movement of the first division and the second movement of the second division, according to the index value. The method described in Clause 1, further including the method described in Clause 1. 5. The encoding unit is divided into a first section and a second section, and the motion information includes the first motion in the first section and the second motion in the second section, and the refinement of the motion information is performed. To constitute a first template for a first category, wherein the first template consists of a first set of adjacent samples. To construct a second template for a second category, wherein the second template consists of a second set of adjacent samples. Refine the first and second movements, respectively, using the first and second templates. It further includes, Each of the first set of adjacent samples and the second set of adjacent samples is: Only the adjacent sample on the left, Only the adjacent samples above, or Both the adjacent sample on the left and the adjacent sample above Includes one or more adjacent samples selected from, The method described in any one of the clauses 1 to 4. 6. The first set of adjacent samples is selected based on the availability of the adjacent samples to the left and above in the first section. A second set of neighboring samples is selected based on the availability of the left neighboring sample and the above neighboring sample in the second section. The method described in Article 5. 7. Configure the first template and the second template based on the GPM classification mode. The method described in Clause 5 or 6, further including the method described in Clause 5 or 6. 8. Configure the first template and the second template based on the GPM division angle. The method described in Clause 7, further including the method described in Clause 7. 9. Decode the second parameter associated with the first template. Decode the third parameter related to the second template. The second and third parameters further include the first set of neighboring samples and the second set of neighboring samples, Only the adjacent sample on the left, Only the adjacent samples above, or Both the adjacent sample on the left and the adjacent sample above The method described in any one of clauses 5 to 8, indicating whether each is selected from the options. 10. The method described in any one of clauses 5 to 9, wherein the first set of adjacent samples is different from the second set of adjacent samples. 11. The method described in any one of the clauses 5 to 9, wherein the first set of adjacent samples is the same as the second set of adjacent samples. 12. Perform motion compensation using refined motion, and Applying a mixing process along the edges of geometric divisions according to the GPM division mode. The method described in any one of the clauses 1 to 11, further including the method described in any one of the clauses 1 to 11. 13. In response to the fact that template matching has not been applied to the coding unit, decode a second parameter indicating whether the merge mode by motion vector difference (MMVD) has been applied. In response to the application of MMVD, decode the motion vector difference (MVD) information, and Refining motion using MVD information The method described in any one of the clauses 1 to 12, further including the method described in any one of the clauses 1 to 12. 14. The method according to Clause 13, wherein the second parameter includes a first flag indicating whether MMVD is applied to the first category, and a second flag indicating whether MMVD is applied to the second category. 15. The encoding unit is divided into a first section and a second section, and the method is, Decode the second parameter indicating whether MMVD is applied to the first category. Decode the third parameter indicating whether MMVD is applied to the second category. In response to the fact that MMVD is not applied to either the first or second category, determine whether template matching is applied to both the first and second categories according to the first parameter. In response to the application of template matching to both the first and second sections, refine the motion information of the first and second sections using template matching. The method described in any one of the clauses 1 to 12, further including the method described in any one of the clauses 1 to 12. 16. The method according to clause 15, wherein the template used to refine the motion is determined based on the segmentation mode. 17. A device for processing video data, Memory configured to store instructions, It includes one or more processors, and one or more processors execute instructions. Receiving a bitstream containing encoded units encoded by GPM (geometric partition mode), Decoding a first parameter related to a coding unit, wherein the first parameter indicates whether template matching is applied to the coding unit, and Determining motion information relating to a coding unit according to a first parameter, wherein if the first parameter indicates that template matching is applied to the coding unit, the motion information is refined using template matching. A device configured to perform a certain action. 18. The coding unit is divided into a first section and a second section, the motion information of the coding unit includes the first motion of the first section and the second motion of the second section, and the first parameter includes the first flag. The processor executes the instruction, Using template matching according to the first flag, refine both the first movement of the first segment and the second movement of the second segment. The device is configured to perform the following action. The apparatus described in Article 17. 19. The coding unit is divided into a first section and a second section, and the first parameter includes a first flag indicating whether template matching is applied to the first section and a second flag indicating whether template matching is applied to the second section. The processor executes the instruction, According to the first flag, template matching is used to refine the first movement of the first segment, and Using template matching according to the second flag, refine the second movement of the second segment. The device is further configured to perform the following actions: The apparatus described in Article 17. 20. The coding unit is divided into a first section and a second section, and the first parameter includes an index indicating whether template matching is applied to the first section and the second section. The processor executes the instruction, Refining the first movement of the first division, the second movement of the second division, or both the first movement of the first division and the second movement of the second division, according to the index value. The device is further configured to perform the following actions: The apparatus described in Article 17. 21. The encoding unit is divided into a first section and a second section, and the motion information includes the first motion of the first section and the second motion of the second section. The processor executes the instruction, To constitute a first template for a first category, wherein the first template consists of a first set of adjacent samples. To construct a second template for a second category, wherein the second template consists of a second set of adjacent samples. Refine the first and second movements, respectively, using the first and second templates. It is further configured to have the device perform the following: Each of the first set of adjacent samples and the second set of adjacent samples is: Only the adjacent sample on the left, Only the adjacent samples above, or Both the adjacent sample on the left and the adjacent sample above Includes one or more adjacent samples selected from, The apparatus described in any one of clauses 17 to 20. 22. The first set of adjacent samples is selected based on the availability of the adjacent samples to the left and above in the first section. A second set of neighboring samples is selected based on the availability of the left neighboring sample and the above neighboring sample in the second section. The apparatus described in Clause 21. 23. The processor executes the instruction. Configure the first template and the second template based on the GPM classification mode. The device is further configured to perform the following actions: The apparatus described in Clause 21 or 22. 24. The processor executes the instruction. The first template and the second template are constructed based on the GPM division angle. The device is further configured to perform the following actions: The apparatus described in Clause 23. 25. The processor executes the instruction. Decoding the second parameter associated with the first template, Decode the third parameter related to the second template. It is further configured to have the device perform the following: The second and third parameters are determined by the first set of neighboring samples and the second set of neighboring samples. Only the adjacent sample on the left, Only the adjacent samples above, or Both the adjacent sample on the left and the adjacent sample above The apparatus described in any one of clauses 21 to 24, indicating whether it is selected from each of the above. 26. The apparatus according to any one of clauses 21 to 25, wherein the first set of adjacent samples is different from the second set of adjacent samples. 27. The apparatus according to any one of clauses 21 to 25, wherein the first set of adjacent samples is the same as the second set of adjacent samples. 28. The processor executes the instruction, Using refined motion to perform motion compensation, and Applying a mixing process along the edges of geometric divisions according to the GPM division mode. The device is further configured to perform the following actions: The apparatus described in any one of clauses 17 to 27. 29. The processor executes the instruction. In response to the fact that template matching is not applied to the coding unit, a second parameter is decoded to indicate whether the merge mode by motion vector difference (MMVD) is applied. In response to the application of MMVD, decode the motion vector difference (MVD) information, and Refining motion using MVD information The device is further configured to perform the following actions: The apparatus described in any one of clauses 17 to 28. 30. The apparatus as described in Clause 29, wherein the second parameter includes a first flag indicating whether MMVD is applied to the first category, and a second flag indicating whether MMVD is applied to the second category. 31. The encoding unit is divided into a first section and a second section, The processor executes the instruction, Decode the second parameter indicating whether MMVD is applied to the first category. Decode the third parameter indicating whether MMVD is applied to the second category. In response to the fact that MMVD is not applied to either the first or second category, determine whether template matching is applied to both the first and second categories according to the first parameter. In response to the application of template matching to both the first and second sections, refine the motion information of the first and second sections using template matching. The device is further configured to perform the following actions: The apparatus described in any one of clauses 17 to 28. 32. The apparatus described in Clause 31, wherein the template used to refine the motion is determined based on the segmentation mode. 33. A non-temporary computer-readable medium storing a set of instructions executable by one or more processors of a device for initiating a method for processing video data, wherein the method Receiving a bitstream containing encoded units encoded by GPM (geometric partition mode), Decoding a first parameter related to a coding unit, wherein the first parameter indicates whether template matching is applied to the coding unit, and Determining motion information relating to a coding unit according to a first parameter, wherein if the first parameter indicates that template matching is applied to the coding unit, the motion information is refined using template matching. Non-temporary computer-readable media, including [specific examples of such media]. 34. The coding unit is divided into a first section and a second section, the motion information of the coding unit includes the first motion of the first section and the second motion of the second section, and the first parameter includes the first flag. The set of instructions is, Use template matching according to the first flag to refine both the first movement of the first segment and the second movement of the second segment. A non-temporary computer-readable medium as described in Clause 33, which is executable by one or more processors of the device in order to cause the device to perform further actions. 35. The coding unit is divided into a first section and a second section, and the first parameter includes a first flag indicating whether template matching is applied to the first section and a second flag indicating whether template matching is applied to the second section. The set of instructions is, According to the first flag, template matching is used to refine the first movement of the first segment, and Using template matching according to the second flag, refine the second movement of the second segment. A non-temporary computer-readable medium as described in Clause 33, which is executable by one or more processors of the device in order to cause the device to perform further actions. 36. The coding unit is divided into a first section and a second section, and the first parameter includes an index indicating whether template matching is applied to the first section and the second section. The set of instructions is, Refining the first movement of the first division, the second movement of the second division, or both the first movement of the first division and the second movement of the second division, according to the index value. A non-temporary computer-readable medium as described in Clause 33, which is executable by one or more processors of the device in order to cause the device to perform further actions. 37. The encoding unit is divided into a first section and a second section, and the motion information includes the first motion in the first section and the second motion in the second section. The set of instructions is, To constitute a first template for a first category, wherein the first template consists of a first set of adjacent samples. To construct a second template for a second category, wherein the second template consists of a second set of adjacent samples. Refine the first and second movements, respectively, using the first and second templates. This can be performed by one or more processors of the device in order to have the device perform further tasks. Each of the first set of adjacent samples and the second set of adjacent samples is: Only the adjacent sample on the left, Only the adjacent samples above, or Both the adjacent sample on the left and the adjacent sample above Includes one or more adjacent samples selected from, Non-temporary computer-readable media as described in any one of clauses 33 to 36. 38. The first set of adjacent samples is selected based on the availability of the adjacent samples to the left and above in the first section. A second set of neighboring samples is selected based on the availability of the left neighboring sample and the above neighboring sample in the second section. Non-temporary computer-readable media as defined in Article 37. 39. The set of instructions is, Configure the first template and the second template based on the GPM classification mode. A non-temporary computer-readable medium as described in Clause 37 or 38, which can be executed by one or more processors of the device to cause the device to perform further actions. 40. The set of instructions is, The first template and the second template are constructed based on the GPM division angle. A non-temporary computer-readable medium as described in Clause 39, which is executable by one or more processors of the device in order to cause the device to perform further actions. 41. The set of instructions is, Decoding the second parameter associated with the first template, Decode the third parameter related to the second template. This can be performed by one or more processors of the device in order to have the device perform further tasks. The second and third parameters are determined by the first set of neighboring samples and the second set of neighboring samples. Only the adjacent sample on the left, Only the adjacent samples above, or Both the adjacent sample on the left and the adjacent sample above A non-temporary computer-readable medium as defined in any one of clauses 37-40, indicating whether it is selected from the above. 42. A non-temporary computer-readable medium as described in any one of clauses 37 to 41, wherein the first set of adjacent samples differs from the second set of adjacent samples. 43. A non-temporary computer-readable medium as described in any one of clauses 37 to 41, wherein the first set of adjacent samples is the same as the second set of adjacent samples. 44. The set of instructions is, Using refined motion to perform motion compensation, and Applying a mixing process along the edges of geometric divisions according to the GPM division mode. A non-temporary computer-readable medium as described in any one of clauses 33 to 43, which is executable by one or more processors of the device to cause the device to perform further actions. 45. The set of instructions is, In response to the fact that template matching is not applied to the coding unit, a second parameter is decoded to indicate whether the merge mode by motion vector difference (MMVD) is applied. In response to the application of MMVD, decode the motion vector difference (MVD) information, and Refining motion using MVD information A non-temporary computer-readable medium as described in any one of clauses 33 to 44, which is executable by one or more processors of the device to cause the device to perform further actions. 46. The non-temporary computer-readable media described in Clause 45, wherein the second parameter includes a first flag indicating whether MMVD is applied to the first section, and a second flag indicating whether MMVD is applied to the second section. 47. The encoding unit is divided into a first section and a second section, The set of instructions is, Decode the second parameter indicating whether MMVD is applied to the first category. Decode the third parameter indicating whether MMVD is applied to the second category. In response to the fact that MMVD is not applied to either the first or second category, determine whether template matching is applied to both the first and second categories according to the first parameter. In response to the application of template matching to both the first and second sections, refine the motion information of the first and second sections using template matching. A non-temporary computer-readable medium as described in any one of clauses 33 to 44, which is executable by one or more processors of the device to cause the device to perform further actions. 48. A non-temporary computer-readable medium as described in Clause 47, in which the template used to refine the motion is determined based on the segmentation mode. 49. A non-temporary computer-readable medium for storing a bitstream, wherein the bitstream includes a first parameter related to an encoding unit, The first parameter indicates whether template matching is applied, and the encoding unit is encoded by GPM (geometric partition mode) in a non-temporal, computer-readable medium. 50. The coding unit is divided into a first section and a second section, and the first parameter is: A first flag indicating whether template matching is applied to the first category, and A second flag indicating whether template matching is applied to the second category. Non-temporary computer-readable media as described in Clause 49, including, further. 51. The bitstream includes a second parameter related to the encoding unit, The second parameter indicates the template for template matching, and the template is, Only the adjacent sample on the left, Only the adjacent samples above, or Both the adjacent sample on the left and the adjacent sample above Consists of one or more adjacent samples selected from, Non-temporary computer-readable media as defined in Article 49. 52. The bitstream is A second parameter related to the encoding unit, which indicates whether the motion vector difference merge mode (MMVD) is applied. Non-temporary computer-readable media as described in Clause 49, including, further. 53. The encoding unit is divided into a first section and a second section, and the bitstream is, A second parameter related to the coding unit, which indicates whether the motion vector difference merge mode (MMVD) is applied to the first segment, and A third parameter related to video data, which indicates whether MMVD is applied to the second category. Non-temporary computer-readable media as described in Clause 49, including, further.
[0157]
[0175] Depending on the embodiment, non-temporary computer-readable storage media containing instructions are also provided. Depending on the embodiment, the medium may store all or part of a video bitstream having one or more flags indicating GPM, TM, or MMVD applied to a CU or section. Depending on the embodiment, the medium may store instructions that can be executed by a device (such as the disclosed encoder and decoder) to perform the above method. Common forms of non-temporary media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tapes, or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media having a pattern of holes, RAM, PROMs, and EPROMs, FLASH®-EPROMs, or any other flash memory, NVRAMs, caches, registers, any other memory chips or cartridges, and networked versions thereof. A device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.
[0158]
[0176] It should be noted that relational terms in this specification, such as "first" and "second," are used merely to distinguish one entity or action from another, and do not imply or require any actual relationship or order between these entities or actions. Furthermore, the words "comprising," "having," "containing," and "including," as well as other similar forms, are intended to be open-ended in that they are equivalent in meaning, and the elements or groups of elements that follow any of these words do not mean to be an exhaustive enumeration of such elements or groups of elements, or to be limited only to the enumerated elements or groups of elements.
[0159]
[0177] As used herein, unless otherwise specified, the term "or" encompasses all possible combinations, except in cases where it is not feasible. For example, if it is stated that a database may contain A or B, then unless otherwise specified or it is not feasible, the database may contain A, B, A and B. As a second example, if it is stated that a database may contain A, B, or C, then unless otherwise specified or it is not feasible, the database may contain A, B, C, A and B, A and C, B and C, A and B and C.
[0160]
[0178] It will be understood that the embodiments described above may be implemented by hardware, software (program code), or a combination of hardware and software. When implemented by software, it may be stored in the computer-readable medium described above. The software can perform the methods of this disclosure when executed by a processor. The computing units and other functional units described in this disclosure may be implemented by hardware, software, or a combination of hardware and software. Those skilled in the art will also understand that some of the above modules / units may be combined into a single module / unit, and each of the above modules / units may be further divided into some submodules / subunits.
[0161]
[0179] In the above-described specification, embodiments have been described with reference to numerous specific details that may differ depending on the implementation. Specific adaptations and modifications of the above-described embodiments can be made. Other embodiments may become apparent to those skilled in the art from the considerations herein and the implementations of the invention disclosed herein. The specification and examples are intended to be considered as examples only, and the true scope and spirit of the invention are indicated by the appended claims. Furthermore, the arrangement of steps shown in the figures is for illustrative purposes only and is not intended to limit the invention to any particular arrangement of steps. Therefore, those skilled in the art will understand that these steps may be performed in different orders while carrying out the same method.
[0162]
[0180] Exemplary embodiments are disclosed in the drawings and specification. However, many variations and modifications can be made to these embodiments. Therefore, where certain terms are used, they are merely for general descriptive purposes and not for limiting purposes.
Claims
1. Receiving a bitstream containing encoded units encoded by GPM (geometric partition mode), Decoding a first parameter related to the coding unit, wherein the first parameter indicates whether template matching is applied to the coding unit, and Determining motion information relating to the coding unit according to the first parameter, wherein if the first parameter indicates that the template matching is applied to the coding unit, the motion information is refined using the template matching. A video decoding method, including the above.
2. The coding unit is divided into a first section and a second section, the motion information of the coding unit includes a first motion of the first section and a second motion of the second section, the first parameter includes a first flag, and the motion information is refined using template matching. Using the template matching according to the first flag, refine both the first movement of the first segment and the second movement of the second segment. The method according to claim 1, further comprising:
3. The encoding unit is divided into a first section and a second section, and the first parameter includes a first flag indicating whether the template matching is applied to the first section and a second flag indicating whether the template matching is applied to the second section, and the refinement of the motion information is, Using the template matching according to the first flag, the first movement of the first segment is refined, and Using the template matching according to the second flag, the second movement of the second segment is refined. The method according to claim 1, further comprising:
4. The encoding unit is divided into a first section and a second section, the first parameter includes an index indicating whether the template matching is applied to the first section and the second section, and the refinement of the motion information is The first movement of the first division, the second movement of the second division, or both the first movement of the first division and the second movement of the second division are refined according to the index value. The method according to claim 1, further comprising:
5. The encoding unit is divided into a first section and a second section, the motion information includes a first motion in the first section and a second motion in the second section, and the refinement of the motion information is To constitute a first template for the first division, wherein the first template consists of a first set of adjacent samples. To constitute a second template for the second division, wherein the second template consists of a second set of adjacent samples. The first and second movements are refined using the first and second templates, respectively. It further includes, Each of the first set of adjacent samples and the second set of adjacent samples is: Only the adjacent sample on the left, Only the adjacent samples above, or Both the left adjacent sample and the upper adjacent sample Includes one or more adjacent samples selected from, The method according to claim 1.
6. In response to the fact that the template matching is not applied to the encoding unit, a second parameter is decoded indicating whether the merge mode by motion vector difference (MMVD) is applied. In response to the application of the aforementioned MMVD, the motion vector difference (MVD) information is decoded, and The motion information is refined using the aforementioned MVD information. The method according to claim 1, further comprising:
7. The encoding unit is divided into a first section and a second section, and the method is Decoding a second parameter indicating whether MMVD is applied to the first category, Decoding a third parameter indicating whether the MMVD is applied to the second category, In response to the fact that the MMVD is not applied to the first or second category, determine whether the template matching is applied to both the first and second categories according to the first parameter. In response to the application of the template matching to both the first and second sections, the motion information of the first and second sections is refined using the template matching. The method according to claim 1, further comprising:
8. An encoding method, Receiving a video sequence, The encoding unit of the video sequence is encoded using GPM (geometric partition mode). Encoding a first parameter related to the encoding unit, wherein the first parameter indicates whether template matching is applied to the encoding unit, and Determining motion information relating to the coding unit according to the first parameter, wherein if the first parameter indicates that the template matching is applied to the coding unit, the motion information is refined using the template matching. Methods that include...
9. The coding unit is divided into a first section and a second section, the motion information of the coding unit includes a first motion of the first section and a second motion of the second section, the first parameter includes a first flag, and the motion information is refined using template matching. Using the template matching according to the first flag, refine both the first movement of the first segment and the second movement of the second segment. Further including, The method according to claim 8.
10. The coding unit is divided into a first section and a second section, and the first parameter includes a first flag indicating whether the template matching is applied to the first section and a second flag indicating whether the template matching is applied to the second section. The refinement of the motion information is Using the template matching according to the first flag, the first movement of the first segment is refined, and Using the template matching according to the second flag, the second movement of the second segment is refined. Further including, The method according to claim 8.
11. The coding unit is divided into a first section and a second section, and the first parameter includes an index indicating whether the template matching is applied to the first section and the second section. The refinement of the motion information is The first movement of the first division, the second movement of the second division, or both the first movement of the first division and the second movement of the second division are refined according to the index value. Further including, The method according to claim 8.
12. The encoding unit is divided into a first section and a second section, and the motion information includes a first motion in the first section and a second motion in the second section. The refinement of the motion information is To constitute a first template for the first division, wherein the first template consists of a first set of adjacent samples. To constitute a second template for the second division, wherein the second template consists of a second set of adjacent samples. The first and second movements are refined using the first and second templates, respectively. It further includes, Each of the first set of adjacent samples and the second set of adjacent samples is: Only the adjacent sample on the left, Only the adjacent samples above, or Both the left adjacent sample and the upper adjacent sample Includes one or more adjacent samples selected from, The method according to claim 8.
13. In response to the fact that the template matching is not applied to the encoding unit, encode a second parameter indicating whether the merge mode by motion vector difference (MMVD) is applied. In response to the application of the aforementioned MMVD, the motion vector difference (MVD) information is encoded, and The motion information is refined using the aforementioned MVD information. The method according to claim 8, further comprising:
14. The encoding unit is divided into a first section and a second section, The method described above is Encoding a second parameter indicating whether MMVD is applied to the first category, Encoding a third parameter indicating whether the MMVD is applied to the second category, In response to the fact that the MMVD is not applied to the first or second category, determine whether the template matching is applied to both the first and second categories according to the first parameter, and In response to the application of the template matching to both the first and second sections, the motion information of the first and second sections is refined using the template matching. Further including, The method according to claim 8.
15. A method for storing a bitstream of a video sequence, wherein the method is: Receiving a video sequence, Encoding one or more pictures of the aforementioned video sequence, To generate a bitstream, and The bitstream is stored in a non-temporary computer-readable storage medium. Including the above encoding, The encoding unit of the video sequence is encoded using GPM (geometric partition mode). Encoding a first parameter related to the encoding unit, wherein the first parameter indicates whether template matching is applied to the encoding unit, and Determining motion information relating to the coding unit according to the first parameter, wherein if the first parameter indicates that the template matching is applied to the coding unit, the motion information is refined using the template matching. Methods that include...
16. The coding unit is divided into a first section and a second section, the motion information of the coding unit includes a first motion of the first section and a second motion of the second section, the first parameter includes a first flag, and the motion information is refined using template matching. Using the template matching according to the first flag, refine both the first movement of the first segment and the second movement of the second segment. Further including, The method according to claim 15.
17. The coding unit is divided into a first section and a second section, and the first parameter includes a first flag indicating whether the template matching is applied to the first section and a second flag indicating whether the template matching is applied to the second section. The refinement of the motion information is Using the template matching according to the first flag, the first movement of the first segment is refined, and Using the template matching according to the second flag, the second movement of the second segment is refined. Further including, The method according to claim 15.
18. The coding unit is divided into a first section and a second section, and the first parameter includes an index indicating whether the template matching is applied to the first section and the second section. The refinement of the motion information is The first movement of the first division, the second movement of the second division, or both the first movement of the first division and the second movement of the second division are refined according to the index value. Further including, The method according to claim 15.
19. The encoding unit is divided into a first section and a second section, and the motion information includes a first motion in the first section and a second motion in the second section. The refinement of the motion information is To constitute a first template for the first division, wherein the first template consists of a first set of adjacent samples. To constitute a second template for the second division, wherein the second template consists of a second set of adjacent samples. The first and second movements are refined using the first and second templates, respectively. It further includes, Each of the first set of adjacent samples and the second set of adjacent samples is: Only the adjacent sample on the left, Only the adjacent samples above, or Both the left adjacent sample and the upper adjacent sample Includes one or more adjacent samples selected from, The method according to claim 15.
20. The encoding unit is divided into a first section and a second section, The refinement of the motion information is Encoding a second parameter indicating whether MMVD is applied to the first category, Encoding a third parameter indicating whether the MMVD is applied to the second category, In response to the fact that the MMVD is not applied to the first or second category, determine whether the template matching is applied to both the first and second categories according to the first parameter. In response to the application of the template matching to both the first and second sections, the motion information of the first and second sections is refined using the template matching. Further including, The method according to claim 15.