MMVD candidate refinement method

The method for refining motion vector differences in video coding techniques addresses the challenge of improving compression efficiency by using processing circuitry to derive refined motion vectors, resulting in enhanced coding efficiency and reduced redundancy in video data.

JP2025514580APending Publication Date: 2025-05-09TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024515689
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-31
Filing Date
2022-11-04
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Current video coding techniques face challenges in efficiently refining motion vector differences (MVDs) for improved compression efficiency and reduced redundancy in video data.

Method used

The proposed solution involves a method for video encoding/decoding that includes processing circuitry to extract candidate information for motion vector differences (MVDs). This circuitry generates refinement offsets for MVD candidates based on a refined step size and multiple refinement locations, allowing for the derivation of refined motion vectors. The processing circuit reconstructs blocks using reference blocks indicated by these refined motion vectors.

Benefits of technology

This approach enhances compression efficiency by refining MVDs, leading to improved coding efficiency and reduced redundancy in video data, thereby optimizing storage and bandwidth requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025514580000001_ABST
    Figure 2025514580000001_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide a method and apparatus for video encoding / decoding. In some examples, the apparatus for video decoding includes a processing circuit. The processing circuit extracts, from a bitstream, merge by motion vector difference (MMVD) candidate information for a current block in a current picture. The processing circuit generates a first MV refinement offset associated with a first motion vector of the MMVD candidate based on a refined step size and a plurality of refinement positions. The processing circuit derives a first refined motion vector (MV) value associated with the MMVD candidate according to the MMVD candidate information and the generated first MV refinement offset. The processing circuit reconstructs the current block according to a first reference block in a first reference picture, the first reference block being indicated by the derived first refined MV value.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims benefit of priority to U.S. Provisional Application No. 63 / 331,936, entitled "Method and Apparatus for Merge with Motion Vector Difference Candidate Refinement," filed April 18, 2022, which claims benefit of priority to U.S. Provisional Application No. 17 / 978,107, entitled "MMVD CANDIDATE REFINEMENT METHODS," filed October 31, 2022. The disclosures of the prior applications are incorporated herein by reference in their entireties.

[0002] This disclosure describes embodiments generally related to video coding. [Background technology]

[0003] The background discussion provided herein is intended to generally present the context of the present disclosure. The inventors' work, to the extent that it is described in this background section, and aspects of the description that may not be admitted as prior art at the time of filing, are not admitted expressly or impliedly as prior art to the present disclosure.

[0004] Uncompressed digital images and / or videos may include a sequence of pictures, each having spatial dimensions of, for example, 1920x1080 luma samples and associated chroma samples. The sequence of pictures may have a fixed or variable picture rate (also informally known as frame rate), for example, 60 pictures per second or 60Hz. Uncompressed images and / or videos have specific bitrate requirements. For example, 1080p 60 4:2:0 video (1920x1080 luma sample resolution at a frame rate of 60Hz) with 8 bits per sample requires a bandwidth approaching 1.5Gbit / s. One hour of such video requires more than 600GByte of storage space.

[0005] One objective of image and / or video coding and decoding may be the reduction of redundancy in the input image and / or video signal through compression. Compression may help reduce the aforementioned bandwidth and / or storage space requirements, possibly by more than one order of magnitude. The description herein uses video encoding / decoding as an illustrative example, but the same techniques may be applied to image encoding / decoding in a similar manner without departing from the spirit of this disclosure. Both lossless and lossy compression, and combinations thereof, may be employed. Lossless compression refers to techniques where an exact copy of the original signal may be reconstructed from a compressed original signal. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough to make the reconstructed signal useful for the intended application. For video, lossy compression is widely adopted. The amount of acceptable distortion depends on the application, for example, a user of a particular consumer streaming application may tolerate higher distortion than a user of a television distribution application. The achievable compression ratio may represent that the higher the acceptable / tolerable distortion, the higher the compression ratio that can be obtained.

[0006] Video encoders and decoders can utilize techniques from a number of broad categories, including, for example, motion compensation, transform processing, quantization, and entropy coding.

[0007] Video codec techniques can include a technique known as intra-coding. In intra-coding, sample values ​​are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into blocks of samples. If all blocks of samples are coded in intra mode, the picture may be an intra picture. Intra pictures and their derivatives, such as independent decoder refresh pictures, are used to reset the decoder state and therefore may be used as the first picture in a coded video bitstream and video session or as a still image. Samples of an intra block are transformed and the transform coefficients may be quantized before entropy coding. Intra prediction may be a technique that minimizes sample values ​​in the pre-transform domain. In some cases, the smaller the DC value after the transform and the smaller the AC coefficients, the fewer bits are needed for a given quantization step size to represent the block after entropy coding.

[0008] Conventional intra-coding, for example as used in MPEG-2 generation coding techniques, does not use intra-prediction. However, some newer video compression techniques include techniques that attempt to make predictions based on surrounding sample data and / or metadata obtained during encoding / decoding of a block of data. Such techniques are hereafter referred to as "intra-prediction" techniques. It should be noted that at least in some cases, intra-prediction uses only reference data from the current picture being reconstructed, and not from a reference picture.

[0009] Intra prediction may take many different forms. When more than one of such techniques may be used in a given video coding technique, the particular technique in use may be coded as a particular intra prediction mode using the particular technique. In certain cases, an intra prediction mode may have sub-modes and / or parameters that may be coded separately or included in a mode codeword that defines the prediction mode used. Which codeword is used for a given mode, sub-mode, and / or parameter combination may affect the coding efficiency gains via intra prediction, and therefore may also affect the entropy coding technique used to convert the codeword into a bitstream.

[0010] A specific mode of intra prediction was introduced in H.264, refined in H.265, and further refined in novel coding techniques such as the Joint Search Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). A predictor block may be formed using neighboring sample values ​​of already available samples. The sample values ​​of the neighboring samples are copied to the predictor block according to the direction. The reference to the direction in use may be coded in the bitstream or may itself be predicted.

[0011] Referring to FIG. 1A, depicted at the bottom right is a subset of 9 predictor directions known from the 33 possible predictor directions defined in H.265 (corresponding to the 33 angle modes out of the 35 intra modes). The point where the arrows converge (101) represents the sample being predicted. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from the sample to the top right, at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from the sample to the bottom left of sample (101), at an angle of 22.5 degrees from the horizontal.

[0012] 1A, at the top left is shown a square block (104) of 4×4 samples (indicated by a thick dashed line). The square block (104) contains 16 samples, each labeled with "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample (from the top) in the Y dimension and the first sample (from the left) in the X dimension. Similarly, sample S44 is the fourth sample in the block (104) in both the Y and X dimensions. Since the block is 4×4 samples in size, S44 is at the bottom right. Also shown are reference samples that follow a similar numbering scheme. The reference samples are labeled with R, their Y position (e.g., row index), and X position (column index) relative to the block (104). In both H.264 and H.265, the predicted samples are in the neighborhood of the block being reconstructed, so there is no need to use negative values.

[0013] Intra-picture prediction can work by copying reference sample values ​​from neighboring samples indicated by the signaled prediction direction. For example, assume that the coded video bitstream includes signaling for this block indicating a prediction direction that coincides with the arrow (102), i.e., the sample is predicted from the upper right sample at an angle of 45 degrees from the horizontal. In that case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.

[0014] In certain cases, to calculate a reference sample, especially when the orientation is not evenly divisible by 45 degrees, the values ​​of multiple reference samples may be combined, for example by interpolation.

[0015] The number of possible directions has increased as video coding techniques have developed. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013). Currently, JEM / VVC / BMS can support up to 65 directions. Experiments have been carried out to identify the most likely directions, and certain techniques of entropy coding are used to represent those likely directions with a small number of bits, accepting a certain penalty for less likely directions. Furthermore, the direction itself may be predicted from nearby directions used in nearby already decoded blocks.

[0016] FIG. 1B shows a schematic diagram (110) depicting 65 intra prediction directions with JEM to show the increasing number of prediction directions over time.

[0017] The mapping of intra-prediction direction bits representing directions in the coded video bitstream may vary from one video coding technique to another. Such mappings may range, for example, from simple direct mappings to complex adaptive schemes including codeword most probable modes and similar techniques. However, in most cases, there may be certain directions that are statistically less likely to occur in the video content than certain other directions. Since the goal of video compression is to reduce redundancy, in a well-performing video coding technique, these less likely directions are represented by more bits than the more likely directions.

[0018] Image and / or video coding and decoding may be performed using inter-picture prediction with motion compensation. Motion compensation may be a lossy compression technique and may refer to a technique in which blocks of sample data from a previously reconstructed picture or part thereof (reference picture) are spatially shifted in a direction indicated by a motion vector (hereinafter MV) and then used to predict a newly reconstructed picture or part of a picture. In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions X and Y, or three dimensions, with the third dimension being an indication of the reference picture in use (the third dimension may indirectly be a temporal dimension).

[0019] In some video compression techniques, the MV applicable to a particular region of sample data may be predicted from other MVs, e.g., from MVs associated with other regions of sample data that are spatially adjacent to the region being reconstructed and that precede that MV in decoding order. Doing so may significantly reduce the amount of data required to code the MV, thereby eliminating redundancy and improving compression ratios. For example, when coding an input video signal derived from a camera (known as natural video), MV prediction may work effectively because there is a statistical likelihood that regions larger than the region to which a single MV is applicable move in similar directions and therefore, in some cases, may be predicted using similar motion vectors derived from MVs of nearby regions. As a result, the MV found for a given region will be similar or the same as the MV predicted from the surrounding MVs, and thus, after entropy coding, may be represented with fewer bits than would be used if the MV were coded directly. In some cases, MV prediction may be an example of lossless compression of a signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, the MV prediction itself may be lossy, e.g., due to rounding errors when computing a predictor from several surrounding MVs.

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (iTU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms offered by H.265, the one hereafter referred to as "spatial merging" is described with reference to Fig. 2.

[0021] Referring to Figure 2, a current block (201) contains samples that are found by the encoder during the motion search process to be predictable from a previous block of the same size but spatially shifted. Instead of coding its MV directly, the MV may be derived from metadata associated with one or more reference pictures, e.g., the most recent reference picture (in decoding order), using MVs associated with any one of five surrounding samples, denoted A0, A1, and B0, B1, B2 (202-206, respectively). In H.265, MV prediction may use predictors from the same reference picture that neighboring blocks are using. Summary of the Invention [Means for solving the problem]

[0022] Aspects of the present disclosure provide a method and apparatus for video encoding / decoding. In some examples, the apparatus for video decoding includes a processing circuit. The processing circuit extracts (e.g., parses) merge by motion vector difference (MMVD) candidate information for a current block in a current picture from a bitstream. The processing circuit generates a first MV refinement offset associated with a first motion vector of the MMVD candidate based on a refined step size and a plurality of refinement positions. The processing circuit derives a first refined motion vector (MV) value associated with the MMVD candidate according to the MMVD candidate information and the generated first MV refinement offset. The processing circuit reconstructs the current block according to a first reference block in a first reference picture, the first reference block being indicated by the derived first refined MV value.

[0023] In some examples, the first MV refinement offset is a portion of the motion vector differential applied to the base candidate to form the MMVD candidate. In one example, the refinement step of the first MV refinement offset is ¼ of the MMVD step of the motion vector differential.

[0024] In one example, the first MV refinement offset corresponds to a refinement position among four potential refinement positions for a first motion vector associated with the MMVD candidate. In another example, the first MV refinement offset corresponds to a refinement position among eight potential refinement positions for a first motion vector associated with the MMVD candidate.

[0025] In some examples, the MMVD candidate is a uni-predictive candidate.

[0026] In some examples, the MMVD candidate is a bi-predictive candidate. The processing circuit derives a second refined MV value associated with the MMVD candidate, the second refined MV value being generated by applying a second MV refinement offset to a second motion vector associated with the MMVD candidate. The processing circuit can reconstruct the current block according to a first reference block in the first reference picture and a second reference block in the second reference picture, the second reference block being indicated by the second refined MV value.

[0027] In one example, the second MV refinement offset is equal to the first MV refinement offset.

[0028] In another example, the second MV refinement offset is a mirrored offset to the first MV refinement offset.

[0029] In another example, the processing circuit determines that the second reference picture and the first reference picture are on the same time side of the current picture, and then the processing circuit derives a second refined MV value associated with the MMVD candidate. The second refined MV value is generated by applying a second MV refinement offset to a second motion vector associated with the MMVD candidate, where the second MV refinement offset is equal to the first MV refinement offset.

[0030] In another example, the processing circuit determines that the second reference picture is on a different temporal side of the current picture than the first reference picture, and then derives a second refined MV value associated with the MMVD candidate. The second refined MV value is generated by applying a second MV refinement offset to a second motion vector associated with the MMVD candidate, and the second MV refinement offset is a mirrored offset of the first MV refinement offset.

[0031] In another example, the processing circuit determines a scaling factor based on a first temporal distance from the current picture to the first reference picture and a second temporal distance from the current picture to the second reference picture. Then, the processing circuit derives a second refined MV value associated with the MMVD candidate, the second refined MV value being generated by applying a second MV refinement offset to a second motion vector associated with the MMVD candidate, the second MV refinement offset being a scaled offset from the first MV refinement offset according to the scaling factor.

[0032] In some embodiments, to derive the first refined MV value, the processing circuit determines an MMVD candidate from the MMVD candidate information, determines a first MV refinement offset, and applies the first MV refinement offset to the MMVD candidate.

[0033] In some examples, the MMVD candidate information includes a first index indicating a base candidate from the merge candidate list, the base candidate providing a starting motion vector. The MMVD candidate information further includes a second index indicating a distance of the motion vector difference from the starting motion vector and a third index indicating a direction of the motion vector difference. In one example, to determine the MMVD candidate, the motion vector difference is applied to the starting motion vector of the base candidate.

[0034] In some examples, the MMVD candidate information includes a first index indicating a base candidate from the merge candidate list, where the base candidate provides a starting motion vector. The MMVD candidate information also includes a second index indicating an MMVD candidate from a sorted list of the multiple MMVD candidates. In one example, the processing circuit applies the potential motion vector differentials to the starting motion vector of the base candidate to generate the multiple MMVD candidates. Furthermore, the processing circuit calculates respective template matching costs for the multiple MMVD candidates. The processing circuit sorts the multiple MMVD candidates into the sorted list according to the template matching costs. The processing circuit selects the MMVD candidate from the sorted list according to the second index.

[0035] In some examples, to determine the first MV refinement offset, the processing circuitry decodes a signal indicating an MV refinement position that corresponds to the first MV refinement offset from the bitstream.

[0036] In some examples, to determine the first MV refinement offset, the processing circuitry applies the potential MV refinement offsets to the MMVD candidates, respectively, to generate refined candidates corresponding to the potential MV refinement offsets. The processing circuitry calculates template matching costs for the refined candidates, respectively. The processing circuitry determines a best template matching cost (e.g., the lowest template matching cost) from the template matching costs. The processing circuitry selects the first MV refinement offset from the potential MV refinement offsets, and the refined candidate corresponding to the first MV refinement offset has the best template matching cost.

[0037] In some embodiments, the MMVD candidate information includes a first index indicating a base candidate from the merge candidate list, where the base candidate provides a starting motion vector, and the MMVD candidate information also includes a second index indicating a refined candidate from the sorted list of refined candidates. To derive the first refined MV value, the processing circuit applies a potential motion vector differential to the base candidate to generate potential MMVD candidates, and the processing circuit also applies a potential MV refinement offset to each of the potential MMVD candidates to generate potential refined candidates for each of the potential MMVD candidates. The processing circuit determines the refined candidates according to a template matching cost for each of the potential MMVD candidates. For example, the first refined candidate for the first potential MMVD candidate is selected from the first potential refined candidates for the first potential MMVD candidate in response to the first refined candidate having the best template matching cost among the first potential refined candidates. The processing circuit sorts the refined candidates according to the template matching costs of the refined candidates to form a sorted list. The processing circuit selects a particular refined candidate from the sorted list according to a second index. A first refined MV value is derived according to the particular refined candidate.

[0038] In some examples, the MMVD candidate information includes a first index indicating a base candidate from the merge candidate list, where the base candidate provides a starting motion vector. The MMVD candidate also includes a second index indicating a refined candidate from the sorted list of refined candidates. To derive the first refined MV value, in some examples, the processing circuit applies a potential motion vector differential to the base candidate to generate a potential MMVD candidate. The processing circuit applies a potential MV refinement offset to each of the potential MMVD candidates, respectively, to generate a potential refined candidate for the potential MMVD candidate. The potential refined candidates are sorted into a sorted potential list according to the template matching cost of the potential refined candidate. The processing circuit selects a portion of the sorted potential list to form a sorted list of refined candidates. The processing circuit can select a particular refined candidate from the sorted list according to the second index, where the first refined MV value is determined according to the particular refined candidate.

[0039] To select a portion of the sorted potential list to form the sorted list of refined candidates, in one example, the top potential refined candidates in the sorted potential list are selected, hi another example, the top potential refined candidates for each of the potential MMVD candidates are selected.

[0040] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method for video decoding.

[0041] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Brief description of the drawings]

[0042] [Figure 1A]FIG. 2 is a schematic diagram of an example subset of intra-prediction modes. [Figure 1B] FIG. 2 is a diagram of an example intra-prediction direction. [Diagram 2] FIG. 2 is a schematic diagram of a current block and its surrounding spatial merge candidates in one example. [Diagram 3] FIG. 3 is a schematic diagram of a simplified block diagram of a communication system (300) according to one embodiment. [Figure 4] FIG. 4 is a schematic diagram of a simplified block diagram of a communication system (400) according to one embodiment. [Diagram 5] FIG. 2 is a schematic diagram of a simplified block diagram of a decoder according to one embodiment. [Figure 6] FIG. 2 is a schematic diagram of a simplified block diagram of an encoder according to one embodiment. [Figure 7] FIG. 4 is a block diagram of an encoder according to another embodiment. [Figure 8] FIG. 4 is a block diagram of a decoder according to another embodiment. [Figure 9A] FIG. 13 is a schematic diagram of a four-parameter affine model according to another embodiment. [Figure 9B] FIG. 13 is a schematic diagram of a six-parameter affine model according to another embodiment. [Figure 10] FIG. 11 is a schematic diagram of an affine motion vector field associated with sub-blocks within a block according to another embodiment; [Figure 11] FIG. 13 is a schematic diagram of exemplary locations of spatial merging candidates according to another embodiment. [Figure 12] FIG. 11 is a schematic diagram of control point motion vector inheritance according to another embodiment. [Figure 13] FIG. 13 is a schematic diagram of candidate positions for constructing an affine merge mode according to another embodiment. [Figure 14] FIG. 1 is a schematic diagram of prediction refinement by optical flow (PROF) according to another embodiment. [Figure 15] FIG. 4 is a schematic diagram of an affine motion estimation process according to another embodiment. [Figure 16]FIG. 13 illustrates a flowchart of an affine motion estimation search according to another embodiment. [Figure 17] FIG. 1 is a schematic diagram of an extended coding unit (CU) region for bidirectional optical flow (BDOF) according to another embodiment. [Figure 18] FIG. 13 is an exemplary schematic diagram of decoder-side motion vector refinement; [Figure 19] 1 is an example of a search process in one embodiment. [Figure 20] 13 is an example of a search point in one embodiment. [Figure 21] 1 is an example of template matching in some examples. [Figure 22] 1 is an example of template matching in affine merge mode in one example. [Figure 23] FIG. 11 illustrates the direction of adding motion vector differentials in some examples. [Figure 24] FIG. 1 shows four refinement positions in one example. [Diagram 25] FIG. 1 shows eight refinement positions in one example. [Figure 26] FIG. 1 illustrates template matching calculations in some examples. [Figure 27] 1 is a flowchart outlining a process according to some embodiments of the present disclosure. [Figure 28] 10 is a flowchart outlining another process in accordance with some embodiments of the present disclosure. [Figure 29] FIG. 1 is a schematic diagram of a computer system according to one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0043] FIG. 3 illustrates an example block diagram of a communication system (300). The communication system (300) includes a plurality of terminal devices capable of communicating with each other, for example, via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) perform unidirectional transmission of data. For example, the terminal device (310) can code video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The encoded video data can be transmitted in the form of one or more coded video bitstreams. The terminal device (320) can receive the coded video data from the network (350), decode the coded video data to reconstruct the video pictures, and display the video pictures according to the reconstructed video data. Unidirectional data transmission may be common, such as in media serving applications.

[0044] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) for bidirectional transmission of coded video data, for example during a video conference. In the case of bidirectional transmission of data, in one example, each of the terminal devices (330) and (340) can code video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) over the network (350). Each of the terminal devices (330) and (340) can also receive coded video data transmitted by the other of the terminal devices (330) and (340), can decode the coded video data to recover the video pictures, and can display the video pictures on an accessible display device according to the recovered video data.

[0045] In the example of FIG. 3, terminal devices (310), (320), (330), and (340) are shown as a server, a personal computer, and a smartphone, respectively, although the principles of the present disclosure may not be so limited. Embodiments of the present disclosure apply to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network (350) represents any number of networks that convey coded video data between terminal devices (310), (320), (330), and (340), including, for example, wired (cable) and / or wireless communication networks. Communication network (350) may exchange data over circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this description, the architecture and topology of network (350) may not be important to the operation of the present disclosure unless otherwise described herein below.

[0046] 4 shows a video encoder and a video decoder in a streaming environment as an example of an application for the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, streaming services, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0047] The streaming system may include a capture subsystem (413) that may include a video source (401), such as a digital camera, that generates a stream of uncompressed video pictures (402). In one example, the stream of video pictures (402) includes samples taken by a digital camera. The stream of video pictures (402), depicted as a thick line to emphasize the amount of data compared to the encoded video data (404) (or coded video bitstream), may be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (404) (or encoded video bitstream), depicted as a thin line to emphasize the amount of data compared to the stream of video pictures (402), may be stored in a streaming server (405) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to obtain copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include a video decoder (410), for example, in an electronic device (430). The video decoder (410) decodes the incoming copy of the encoded video data (407) and creates an outgoing stream of video pictures (411) that can be rendered on a display (412) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., a video bitstream) can be encoded according to a particular video coding / compression standard. Examples of these standards include ITU-T Recommendation H.265.In one example, a developing video coding standard is informally known as Versatile Video Coding (VVC), and the disclosed subject matter may be used in the context of VVC.

[0048] It should be noted that the electronic devices (420) and (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).

[0049] 5 shows an example block diagram of a video decoder (510). The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used in place of the video decoder (410) of the example of FIG.

[0050] The receiver (531) may receive one or more coded video sequences to be decoded by the video decoder (510). In one embodiment, one coded video sequence is received at a time, and the decoding of each coded video sequence is independent of the decoding of the other coded video sequences. The coded video sequences may be received from a channel (501), which may be a hardware / software link to a storage device that stores the encoded video data. The receiver (531) may receive coded video data with other data, e.g., coded audio data and / or auxiliary data streams, which may be transferred to each other using an entity (not shown). The receiver (531) may separate the coded video sequences from the other data. To combat network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter, "parser (520)"). In certain applications, the buffer memory (515) is part of the video decoder (510). In other cases, the buffer memory (515) may be external to the video decoder (510) (not shown). In still other cases, there may be a buffer memory (not shown) external to the video decoder (510), for example to combat network jitter, and another buffer memory (515) internal to the video decoder (510), for example to handle playback timing. When the receiver (531) receives data from a storage / forwarding device with sufficient bandwidth and controllability, or from an asynchronous network, the buffer memory (515) may be unnecessary or may be small. For use with best-effort packet networks such as the Internet, the buffer memory (515) may be required and may be relatively large, may be advantageously adaptively sized, and may be implemented at least partially within an operating system or similar element (not shown) external to the video decoder (510).

[0051] The video decoder (510) may include a parser (520) that reconstructs symbols (521) from the coded video sequence. These categories of symbols include information used to manage the operation of the video decoder (510) and potential information for controlling a rendering device, such as a render device (512) (e.g., a display screen) that is not an integral part of the electronic device (530) but may be coupled to the electronic device (530) as shown in FIG. 5. The rendering device control information may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (520) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context dependency, etc. The parser (520) may extract a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroups may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser (520) may also extract information from the coded video sequence information, such as transform coefficients, quantization parameter values, motion vectors, etc.

[0052] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to produce symbols (521).

[0053] The reconstruction of the symbols (521) can involve several different units, depending on the type of coded video picture or part thereof (inter-picture and intra-picture, inter-block and intra-block, etc.), as well as other factors. Which units are involved and how may be controlled by subgroup control information parsed by the parser (520) from the coded video sequence. The flow of such subgroup control information between the parser (520) and the following units is not shown for clarity.

[0054] Beyond the functional blocks already mentioned, the video decoder (510) may be conceptually subdivided into a number of functional units, as described below. In an actual implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, a conceptual subdivision into the following functional units is appropriate:

[0055] The first unit is a scalar / inverse transform unit (551), which receives quantized transform coefficients as well as control information from the parser (520) including which transform to use, block size, quantization coefficients, quantization scaling matrices, etc. as symbols (521). The scalar / inverse transform unit (551) can output blocks containing sample values ​​that can be input to an aggregator (555).

[0056] In some cases, the output samples of the scaler / inverse transform unit (551) may relate to intra-coded blocks. Intra-coded blocks are blocks that do not use prediction information from a previously reconstructed picture, but may use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates a block of the same size and shape as the block being reconstructed using surrounding already reconstructed information fetched from a current picture buffer (558). The current picture buffer (558) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. The aggregator (555) may append the prediction information generated by the intra-prediction unit (552) to the output sample information from the scaler / inverse transform unit (551) on a sample-by-sample basis.

[0057] In other cases, the output samples of the scalar / inverse transform unit (551) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion compensated prediction unit (553) may access the reference picture memory (557) to fetch samples used for prediction. After motion compensating the fetched samples according to the symbols (521) related to the block, these samples may be added by the aggregator (555) to the output of the scalar / inverse transform unit (551) (in this case referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (557) from which the motion compensated prediction unit (553) fetches the prediction samples may be controlled by motion vectors available to the motion compensated prediction unit (553), for example, in the form of symbols (521) that may have X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​fetched from the reference picture memory (557) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, and the like.

[0058] The output samples of the aggregator (555) may be subjected to various loop filtering techniques in a loop filter unit (556). Video compression techniques may include in-loop filter techniques controlled by parameters contained in the coded video sequence (also called a coded video bitstream) and available to the loop filter unit (556) as symbols (521) from the parser (520). Video compression may also be responsive to meta-information obtained during decoding of previous portions (in decoding order) of the coded picture or coded video sequence, and to previously reconstructed loop filtered sample values.

[0059] The output of the loop filter unit (556) may be a sample stream that may be output to a render device (512) and may also be stored in a reference picture memory (557) for use in future inter-picture prediction.

[0060] Once fully reconstructed, a particular coded picture may be used as a reference picture for future prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) may become part of the reference picture memory (557), and a new current picture buffer may be reallocated before beginning reconstruction of the next coded picture.

[0061] The video decoder (510) may perform decoding operations according to a given video compression technique or standard, such as ITU-T Recommendation H.265. The coded video sequence may comply with the syntax specified by the video compression technique or standard being used, in the sense that the coded video sequence adheres to both the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. Specifically, the profile may select certain tools from among all tools available in the video compression technique or standard as tools that are only available to them under the profile. Also, compliance may require that the complexity of the coded video sequence be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may be further limited in some cases by the specification of a hypothetical reference decoder (HRD) and metadata for HRD buffer management signaled within the coded video sequence.

[0062] In one embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0063] 6 shows an example block diagram of a video encoder (603). The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmitting circuit). The video encoder (603) may be used in place of the video encoder (403) of the example of FIG.

[0064] The video encoder (603) can receive video samples from a video source (601) (which is not part of the electronic device (620) in the example of FIG. 6) that can capture video images to be coded by the video encoder (603). In other examples, the video source (601) is part of the electronic device (620).

[0065] The video source (601) may provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream that may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (601) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures that give motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0066] According to one embodiment, the video encoder (603) may code and compress pictures of a source video sequence into a coded video sequence (643) in real time or under any other time constraint required. Enforcing an appropriate coding rate is one function of the controller (650). In some embodiments, the controller (650) controls and is operatively coupled to other functional units described below. This coupling is not depicted for clarity. Parameters set by the controller (650) may include rate control related parameters (picture skip, quantizer, lambda value for rate distortion optimization techniques, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) may be configured with other appropriate functions associated with the video encoder (603) optimized for a particular system design.

[0067] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As an oversimplified explanation, in one example, the coding loop may include a source coder (630) (responsible for creating symbols, such as a symbol stream, based on an input picture to be coded and a reference picture, for example) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to that which a (remote) decoder would also create. The reconstructed sample stream (sample data) is input to a reference picture memory (634). Since decoding of the symbol stream results in bit-exact results regardless of the location of the decoder (local or remote), the contents of the reference picture memory (634) are also bit-exact between the local and remote encoders. In other words, the predictive part of the encoder "sees" exactly the same sample values ​​as the decoder would "see" when using prediction during decoding as reference picture samples. This basic principle of reference picture synchrony (and the resulting drift when synchrony cannot be maintained, for example due to channel errors) is also used in some related techniques.

[0068] The operation of the "local" decoder (633) may be the same as that of a "remote" decoder, such as the video decoder (510) already described in detail above in connection with Figure 5. However, with brief reference also to Figure 5, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder (645) and parser (520) may be lossless, the entropy decoding portion of the video decoder (510), including the buffer memory (515) and parser (520), may not be fully implemented in the local decoder (633).

[0069] In one embodiment, the decoder techniques, except for parsing / entropy decoding, present in the decoder are present in the corresponding encoder in the same or substantially the same functional form. Thus, the disclosed subject matter focuses on the operation of the decoder. The description of the encoder techniques may be omitted since they are the inverse of the decoder techniques described generically. Only in certain areas is a more detailed description provided below.

[0070] In operation, in some examples, the source coder (630) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as “reference pictures.” In this manner, the coding engine (632) codes differences between pixel blocks of the input picture and pixel blocks of reference pictures that may be selected as predictive references for the input picture.

[0071] The local video decoder (633) may decode the coded video data of pictures that may be designated as reference pictures based on the symbols created by the source coder (630). The operation of the coding engine (632) may advantageously be a lossy process. When the coded video data may be decoded in a video decoder (not shown in FIG. 6), the reconstructed video sequence may usually be a copy of the source video sequence with some errors. The local video decoder (633) may replicate the decoding process that may be performed by the video decoder on the reference pictures and store the reconstructed reference pictures in the reference picture memory (634). In this way, the video encoder (603) may locally store copies of reconstructed reference pictures that have common content with the reconstructed reference pictures obtained by the far-end video decoder (without transmission errors).

[0072] The predictor (635) may perform a predictive search for the coding engine (632). That is, for a new picture to be coded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors, block shapes, etc., that may serve as suitable predictive references for the new picture. The predictor (635) may operate on sample blocks, pixel block by pixel block, to find suitable predictive references. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have predictive references drawn from multiple reference pictures stored in the reference picture memory (634).

[0073] The controller (650) may manage the coding operations of the source coder (630), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0074] The output of all the aforementioned functional units may undergo entropy coding in an entropy coder (645), which converts the symbols produced by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0075] The transmitter (640) may buffer the coded video sequence created by the entropy coder (645) in preparation for transmission over a communication channel (660), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (640) may merge the coded video data from the video encoder (603) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0076] The controller (650) may manage the operation of the video encoder (603). During coding, the controller (650) may assign a particular coded picture type to each coded picture, which may affect the coding technique that may be applied to the respective picture. For example, pictures may often be assigned as one of the following picture types:

[0077] It should be noted that intra-pictures (I-pictures) can be coded and decoded without using other pictures in the sequence as a source of prediction. Some video codecs allow different types of intra-pictures, including, for example, independent decoder refresh ("IDR") pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.

[0078] A predictive picture (P picture) may be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict the sample values ​​of each block.

[0079] Bidirectionally predicted pictures (B-pictures) may be coded and decoded using intra- or inter-prediction, which uses up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predicted pictures may use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0080] A source picture may generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and coded block by block. A block may be predictively coded with reference to other (already coded) blocks determined by the coding assignment applied to the block's respective picture. For example, a block of an I picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial or intra prediction). A pixel block of a P picture may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. A block of a B picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.

[0081] The video encoder (603) may perform coding operations in accordance with a given video coding technique or standard, such as ITU-T Rec. H.265. In its operations, the video encoder (603) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.

[0082] In one embodiment, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include such data as part of the coded video sequence. The additional data may comprise temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0083] A video may be captured in time sequence as multiple source pictures (video pictures). Intra-picture prediction (often abbreviated as intra-prediction) uses spatial correlation within a given picture, while inter-picture prediction uses correlation (temporal or other) between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. When a block in the current picture is similar to a reference block in a reference picture that has been previously coded and is still buffered in the video, the block in the current picture may be coded by a vector, called a motion vector. The motion vector points to a reference block in the reference picture, and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0084] In some embodiments, bi-prediction techniques may be used in inter-picture prediction. According to bi-prediction techniques, two reference pictures, such as a first reference picture and a second reference picture, both of which are before the current picture in the video in decoding order (but may be past and future, respectively, in display order), are used. A block in the current picture may be coded by a first motion vector that points to a first reference block in the first reference picture and by a second motion vector that points to a second reference block in the second reference picture. A block may be predicted by a combination of the first reference block and the second reference block.

[0085] Furthermore, merge mode techniques may be used in inter-picture prediction to improve coding efficiency.

[0086] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed on a block-by-block basis. For example, according to the HEVC standard, a picture in a sequence of video pictures is divided into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, and 16×16 pixels. In general, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU can be recursively quad-tree partitioned into one or more coding units (CUs). For example, a CTU of 64×64 pixels can be partitioned into one CU of 64×64 pixels, or into four CUs of 32×32 pixels, or into 16 CUs of 16×16 pixels. In one example, each CU is analyzed to determine a prediction type of the CU, such as an inter prediction type or an intra prediction type. A CU is divided into one or more prediction units (PUs) according to temporal predictability and / or spatial predictability. Generally, each PU includes one luma prediction block (PB) and two chroma PBs. In one embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of values ​​(e.g., luma values) for pixels of 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0087] 7 shows an example diagram of a video encoder (703). The video encoder (703) is configured to receive a processed block of sample values ​​(e.g., a predictive block) in a current video picture in a sequence of video pictures and to encode the processed block into a coded picture that is part of a coded video sequence. In one example, the video encoder (703) is used in place of the video encoder (403) of the example of FIG. 4.

[0088] In an HEVC example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as a predictive block of 8x8 samples. The video encoder (703) determines whether the processing block is best coded using intra-mode, inter-mode, or bi-predictive mode, for example, using rate-distortion optimization. When the processing block is to be coded in intra-mode, the video encoder (703) may encode the processing block into a coded picture using intra-prediction techniques, and when the processing block is to be coded in inter-mode or bi-predictive mode, the video encoder (703) may encode the processing block into a coded picture using inter-prediction techniques or bi-prediction techniques, respectively. In certain video coding techniques, the merge mode may be an inter-picture prediction sub-mode in which a motion vector is derived from one or more motion vector predictors without the aid of coded motion vector components outside the predictors. In certain other video coding techniques, there may be motion vector components applicable to the current block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown), to determine the mode of the processing block.

[0089] In the example of FIG. 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725), coupled together as shown in FIG.

[0090] The inter-encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block to one or more reference blocks in a reference picture (e.g., blocks in previous and subsequent pictures), generate inter-prediction information (e.g., a description of redundant information due to inter-encoding techniques, motion vectors, merge mode information), and calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture that is decoded based on the encoded video information.

[0091] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), possibly compare the block to already coded blocks in the same picture, generate transformed quantized coefficients, and possibly also generate intra prediction information (e.g., intra prediction direction information according to one or more intra encoding techniques). In one example, the intra encoder (722) also calculates intra prediction results (e.g., predicted blocks) based on the intra prediction information and reference blocks in the same picture.

[0092] The generic controller (721) is configured to determine generic control data and control other components of the video encoder (703) based on the generic control data. In one example, the generic controller (721) determines the mode of the block and provides a control signal to the switch (726) based on the mode. For example, if the mode is an intra mode, the generic controller (721) controls the switch (726) to select the intra mode result used by the residual calculator (723) and controls the entropy encoder (725) to select the intra prediction information and include the intra prediction information in the bitstream; if the mode is an inter mode, the generic controller (721) controls the switch (726) to select the inter prediction result used by the residual calculator (723) and controls the entropy encoder (725) to select the inter prediction information and include the inter prediction information in the bitstream.

[0093] The residual calculator (723) is configured to calculate a difference (residual data) between a received block and a prediction result selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) is configured to operate based on the residual data to encode the residual data to generate transform coefficients. In one example, the residual encoder (724) is configured to transform the residual data from the spatial domain to the frequency domain to generate transform coefficients. The transform coefficients then undergo a quantization process to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data may be used by the intra-encoder (722) and the inter-encoder (730) as appropriate. For example, the inter-encoder (730) may generate decoded blocks based on the decoded residual data and the inter-prediction information, and the intra-encoder (722) may generate decoded blocks based on the decoded residual data and the intra-prediction information. The decoded blocks are appropriately processed to generate decoded pictures, which may be buffered in a memory circuit (not shown) and may be used as reference pictures in some examples.

[0094] The entropy encoder (725) is configured to format a bitstream to include the encoded block. The entropy encoder (725) is configured to include various information in the bitstream according to an appropriate standard, such as the HEVC standard. In one example, the entropy encoder (725) is configured to include in the bitstream general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information. It should be noted that, in accordance with the disclosed subject matter, when coding a block in a merged sub-mode of either an inter mode or a bi-prediction mode, the residual information is not present.

[0095] 8 shows an example diagram of a video decoder (810). The video decoder (810) is configured to receive coded pictures that are part of a coded video sequence and decode the coded pictures to generate reconstructed pictures. In one example, the video decoder (810) is used in place of the video decoder (410) of the example of FIG. 4.

[0096] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872), coupled to each other as shown in FIG. 8.

[0097] The entropy decoder (871) may be configured to reconstruct from the coded picture certain symbols representing syntax elements that make up the coded picture. Such symbols may include, for example, prediction information (e.g., intra-mode, inter-mode, bi-predictive mode, etc., where inter-mode and bi-predictive mode are in merged submode or other submode) that may identify the mode in which the block is coded, as well as certain samples or metadata used for prediction by the intra-decoder (872) or inter-decoder (880), respectively (e.g., intra-predictive information or inter-predictive information, etc.). The symbols may also include, for example, residual information in the form of quantized transform coefficients, etc. In one example, when the prediction mode is an inter-mode or bi-predictive mode, the inter-predictive information is provided to the inter-decoder (880); when the prediction type is an intra-predictive type, the intra-predictive information is provided to the intra-decoder (872). The residual information may undergo inverse quantization and is provided to the residual decoder (873).

[0098] The inter decoder (880) is configured to receive the inter prediction information and to generate inter prediction results based on the inter prediction information.

[0099] The intra decoder (872) is configured to receive the intra prediction information and to generate a prediction result based on the intra prediction information.

[0100] The residual decoder (873) is configured to perform inverse quantization to extract inverse quantized transform coefficients, and to process the inverse quantized transform coefficients to transform the residual information from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (to include quantizer parameters (QP)), which may be provided by the entropy decoder (871) (data path not shown as this may be only low volume control information).

[0101] The reconstruction module (874) is configured to combine, in the spatial domain, the residual information output by the residual decoder (873) and the prediction result (possibly with the prediction result output by the inter- or intra-prediction module) to form a reconstructed block that may be part of a reconstructed picture, which in turn may be part of a reconstructed video. It should be noted that other suitable operations, such as a deblocking operation, may be performed to improve visual quality.

[0102] It should be noted that the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using any suitable technology. In one embodiment, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using one or more integrated circuits. In other embodiments, the video encoders (403), (603), and (603) and the video decoders (410), (510), and (810) may be implemented using one or more processors executing software instructions.

[0103] Aspects of this disclosure provide techniques for refining merging by motion vector difference (MMVD). In some examples, refinement is added on top of MMVD candidates.

[0104] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). In 2015, the two standardization bodies jointly formed the Joint Video Research Team (JVET) to explore the possibility of developing the next video coding standard beyond HEVC. In October 2017, the two standardization bodies announced a Joint Call for Proposals on Video Compression with Capability beyond HEVC (CfP). By February 15, 2018, 22 CfP responses had been submitted for standard dynamic range (SDR), 12 for high dynamic range (HDR), and 12 for the 360 ​​video category. In April 2018, all received CfP responses were evaluated at the 122 MPEG / 10th JVET meeting. As a result of the meeting, JVET formally launched the standardization process for next-generation video coding beyond HEVC. This new standard was named Versatile Video Coding (VVC) and JVET was renamed the Joint Video Experts Team. In 2020, ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the VVC video coding standard (version 1).

[0105] In inter prediction, motion parameters are needed for each inter predicted coding unit (CU), e.g., to code the VVC features used for inter predicted sample generation. The motion parameters may include motion vectors, reference picture indexes, reference picture list usage indexes, and / or additional information. The motion parameters may be signaled in an explicit or implicit manner. If a CU is coded in skip mode, the CU may be associated with one PU, and significant residual coefficients, coded motion vector deltas, and / or reference picture indexes may not be needed. If a CU is coded in merge mode, the motion parameters of the CU may be obtained from neighboring CUs. The neighboring CUs may include spatial and temporal candidates, as well as additional schedules (or additional candidates) as introduced in VVC. The merge mode may be applied to any inter predicted CU, not just to skip mode. An alternative to the merge mode is explicit transmission of motion parameters, where the motion vectors, the corresponding reference picture indexes of each reference picture list, the reference picture list usage flag, and / or other necessary information may be explicitly signaled for each CU.

[0106] In VVC, the VVC Test Model (VTM) reference software can include several new refined inter-predictive coding tools, which can include one or more of the following: (1) Enhanced Merge Prediction (2) Merge by Motion Vector Difference (MMVD) (3) AMVP mode with symmetric MVD signaling (4) Affine motion compensation prediction (5) Sub-block based temporal motion vector prediction (SbTMVP) (6) Adaptive Motion Vector Resolution (AMVR) (7) Motion field memory: 1 / 16 luminance sample MV memory and 8x8 motion field compression (8) Bi-prediction with CU-level weights (BCW) (9) Bidirectional Optical Flow (BDOF) (10) Decoder-side Motion Vector Refinement (DMVR) (11) Combined Inter- and Intra-Prediction (CIIP) (12) Geometric Partition Mode (GPM)

[0107] In HEVC, a translational motion model is applied to motion compensated prediction (MCP). In the real world, many types of motion can exist, such as zoom in / out, rotation, perspective motion, and other irregular motion. Block-based affine transform motion compensated prediction can be applied to VTM, etc. Figure 9A shows an affine motion field of a block (902) described by motion information of two control points (4 parameters). Figure 9B shows an affine motion field of a block (904) described by three control point motion vectors (6 parameters).

[0108] As shown in FIG. 9A, in a four-parameter affine motion model, the motion vector at a sample location (x,y) in a block (902) may be derived in accordance with equation (1) as follows:

number

number

[0109] As shown in FIG. 9B, in a six-parameter affine motion model, the motion vector at a sample location (x,y) within a block (904) may be derived in accordance with equation (3) as follows:

number

number

[0110] As shown in FIG. 10, to simplify the motion compensation prediction, a block-based affine transformation prediction may be applied. To derive the motion vector of each 4×4 luma sub-block, the motion vector of the center sample (e.g., (1002)) of each sub-block (e.g., (1004)) in the current block (1000) may be calculated according to Equations (1)-(4) and rounded to 1 / 16 fractional precision. A motion compensation interpolation filter may be applied to generate a prediction of each sub-block using the derived motion vector. The sub-block size of the chroma components may also be set as 4×4. The MV of a 4×4 chroma sub-block may be calculated as the average of the MVs of the four corresponding 4×4 luma sub-blocks.

[0111] In affine merge prediction, the affine merge (AF_MERGE) mode may be applied to CUs with both width and height equal to or greater than 8. The CPMV of the current CU may be generated based on the motion information of spatially neighboring CUs. Up to five CPMVP candidates may be applied to affine merge prediction, and an index may be signaled to indicate which of the five CPMVP candidates may be used for the current CU. In affine merge prediction, three types of CPMV candidates may be used to form the affine merge candidate list: (1) inherited affine merge candidates extrapolated from the CPMV of neighboring CUs, (2) affine merge candidates constructed with CPMVP derived using the translation MV of neighboring CUs, and (3) zero MV.

[0112] In VTM3, up to two inherited affine candidates may be applied. The two inherited affine candidates may be derived from the affine motion models of the neighboring blocks. For example, one inherited affine candidate may be derived from the left neighboring CU, and the other inherited affine candidate may be derived from the top neighboring CU. An exemplary candidate block may be shown in FIG. 11. As shown in FIG. 11, for the left predictor (or the left inherited affine candidate), the scanning order may be A0→A1, and for the top predictor (or the top inherited affine candidate), the scanning order may be B0→B1→B2. Therefore, only the first available inherited candidate from each side may be selected. No pruning check may be performed between the two inherited candidates. Once a neighboring affine CU is identified, the control point motion vector of the neighboring affine CU may be used to derive a CPMVP candidate in the affine merge list of the current CU. As shown in FIG. 12, when the neighboring lower-left block A of the current block (1204) is coded in affine mode, motion vectors v2, v3, and v4 of the upper-left corner, upper-right corner, and lower-left corner of the CU (1202) containing block A may be achieved. If block A is coded with a four-parameter affine model, two CPMVs of the current CU (1204) may be calculated according to v2 and v3 of the CU (1202). If block A is coded with a six-parameter affine model, three CPMVs of the current CU (1204) may be calculated according to v2, v3, and v4 of the CU (1202).

[0113] The constructed affine candidate of the current block may be a candidate constructed by combining the neighborhood translational motion information of each control point of the current block. The motion information of the control points may be derived from specific spatial and temporal neighborhoods, which may be shown in FIG. 13. As shown in FIG. 13, the CPMV k(k=1,2,3,4) represents the kth control point of the current block (1302). For CPMV1, the B2→B3→A2 block may be checked and the MV of the first available block may be used. For CPMV2, the B1→B0 block may be checked. For CPMV3, the A1→A0 block may be checked. If CPM4 is not available, TMVP may be used as CPMV4.

[0114] After the MVs of the four control points are achieved, affine merge candidates for the current block (1302) may be constructed based on the motion information of the four control points. For example, affine merge candidates may be constructed based on combinations of the MVs of the four control points in the following order: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}.

[0115] A combination of three CPMVs can constitute a six-parameter affine merge candidate, and a combination of two CPMVs can constitute a four-parameter affine merge candidate. To avoid the motion scaling process, related combinations of control point MVs can be discarded if the reference indices of the control points are different.

[0116] After the inherited and constructed affine merge candidates have been checked, if the list is not already full, zero MV may be inserted at the end of the list.

[0117] In affine AMVP prediction, affine AMVP mode may be applied to CUs with both width and height of 16 or more. A CU-level affine flag may be signaled in the bitstream to indicate whether affine AMVP mode is used, and then another flag may be signaled to indicate whether 4-parameter affine or 6-parameter affine is applied. In affine AMVP prediction, the difference between the CPMV of the current CU and the predictor of the CPMVP of the current CU may be signaled in the bitstream. The size of the affine AMVP candidate list may be 2, and the affine AMVP candidate list may be generated by using the four types of CPMV candidates in the following order: (1) Inherited affine AMVP candidates extrapolated from the CPMVs of nearby CUs; (2) A constructed affine AMVP candidate with CPMVP derived using the translational MVs of nearby CUs; (3) translational MVs from nearby CUs, and (40 Zero MV.

[0118] The check order of the inherited affine AMVP candidates may be the same as the check order of the inherited affine merge candidates. To determine the AVMP candidates, only affine CUs with the same reference picture as the current block may be considered. When an inherited affine motion predictor is inserted into the candidate list, the pruning process may not be applied.

[0119] The constructed AMVP candidates may be derived from the specified spatial neighborhood. As shown in FIG. 13, the same check order as in the affine merge candidate configuration may be applied. In addition, the reference picture indexes of the neighboring blocks may also be checked. The first block in the check order may be inter-coded and may have the same reference picture as the current CU (1302). If the current CU (1302) is coded in a four-parameter affine mode and both mv0 and mv1 are available, one constructed AMVP candidate may be determined. The constructed AMVP candidate may be further added to the affine AMVP list. If the current CU (1302) is coded in a six-parameter affine mode and all three CPMVs are available, the constructed AMVP candidate may be added as one candidate in the affine AMVP list. If not available, the constructed AMVP candidate may be set as unavailable.

[0120] After the inherited affine AMVP candidates and constructed AMVP candidates are checked, if there are still less than two candidates in the affine AMVP list, mv0, mv1, and mv2 may be added in order. mv0, mv1, and mv2 may serve as translation MVs to predict all control point MVs of the current CU (e.g., (1302)), if available. Finally, if the affine AMVP is not yet full, a zero MV may be used to fill the affine AMVP list.

[0121] Sub-block-based affine motion compensation can save memory access bandwidth and reduce computational complexity compared to pixel-based motion compensation, at the expense of a penalty in prediction accuracy. To achieve finer granularity of motion compensation, prediction refinement by optical flow (PROF) can be used to refine sub-block-based affine motion compensation prediction without increasing memory access bandwidth for motion compensation. In VVC, after sub-block-based affine motion compensation is performed, luma prediction samples can be refined by adding the difference derived by the optical flow equation. PROF can be described in the following four steps:

[0122] Step (1): Sub-block-based affine motion compensation may be performed to generate a sub-block prediction I(i,j).

[0123] Step (2): Spatial gradient of subblock prediction g x (i,j) and g y (i,j) can be calculated at each sample position using a 3-tap filter [-1,0,1]. The gradient calculation can be the same as the gradient calculation in BDOF. For example, the spatial gradient g x (i,j) and g y (i,j) can be calculated based on equations (5) and (6), respectively. g x (i,j)=(I(i+1,j)>>shift1)-(I(i-1,j)>>shift1) Equation (5) g y (i,j)=(I(i,j+1)>>shift1)-(I(i,j-1)>>shift1) Equation (6) As shown in equations (5) and (6), shift1 may be used to control the accuracy of the gradient. Sub-block (e.g., 4×4) prediction may be extended by one sample on each side for gradient calculation. To avoid additional memory bandwidth and additional interpolation calculations, the extended samples on the extension boundary may be copied from the nearest integer pixel location in the reference picture.

[0124] Step (3): The luminance prediction accuracy can be calculated by the optical flow equation as shown in Equation (7). ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) Equation (7) Here, Δv(i,j) is the sum of the sample MV calculated for sample position (i,j) represented by v(i,j) and the v SB FIG. 14 shows an example diagram of the difference between a sample MV and a subblock MV. As shown in FIG. 14, a subblock (1402) may be included in the current block (1400), and a sample (1404) may be included in the subblock (1402). The sample (1404) may include a sample motion vector v(i,j) corresponding to a reference pixel (1406). The subblock (1402) may be represented by a subblock motion vector v(i,j). SB The sub-block motion vector v SB Based on, the sample (1404) may correspond to the reference pixel (1408). The difference between the sample MV and the sub-block MV, denoted by Δv(i,j), may be denoted by the difference between the reference pixel (1406) and the reference pixel (1408). Δv(i,j) may be quantized in units of 1 / 32 luma sample precision.

[0125] Since the affine model parameters and sample positions relative to the sub-block center may not change from sub-block to other sub-blocks, Δv(i,j) may be calculated for a first sub-block (e.g., (1402)) and reused for other sub-blocks (e.g., (1410)) in the same CU (e.g., (1400)). Let dx(i,j) be the horizontal offset and dy(i,j) be the distance from the sample position (i,j) to the center of the sub-block (x SB ,y SB), then Δv(x,y) can be derived by the following equations (8) and (9):

number

number

[0126] To maintain accuracy, the center of the subblock (x SB ,y SB ) is ((W SB -1) / 2,(H SB -1) / 2), where W SB and H SB are the width and height of the sub-block, respectively.

[0127] Once Δv(x,y) is obtained, the parameters of the affine model can be obtained. For example, in the case of a four-parameter affine model, the parameters of the affine model can be shown in Equation (10).

number

number

[0128] Step (4): Finally, the luma prediction accuracy ΔI(i,j) may be added to the sub-block prediction I(i,j). The final prediction I′ may be generated as shown in equation (12). I'(i,j)=I(i,j)+ΔI(i,j) Equation (12)

[0129] PROF may not be applied to an affine coded CU in two cases: (1) when all control points MV are the same, indicating that the CU has only translational motion, and (2) when the affine motion parameters are larger than the specified limit since sub-block-based affine MC has been downgraded to CU-based MC to avoid large memory access bandwidth requirements.

[0130] Affine motion estimation (ME), such as the VVC reference software VTM, can be operated for both uni-prediction and bi-prediction. Uni-prediction can be performed for either reference list L0 or reference list L1, and bi-prediction can be performed for both reference list L0 and reference list L1.

[0131] FIG. 15 shows a schematic diagram of affine ME (1500). As shown in FIG. 15, in affine ME (1500), affine uni-prediction (S1502) may be performed on reference list L0 to obtain a prediction P0 of the current block based on an initial reference block in reference list L0. Affine uni-prediction (S1504) may also be performed on reference list L1 to obtain a prediction P1 of the current block based on an initial reference block in reference list L1. In (S1506), affine bi-prediction may be performed. Affine bi-prediction (S1506) may start with an initial prediction residual (2I-P0)-P1, where I may be an initial value of the current block. Affine bi-prediction (S1506) may search candidates in reference list L1 around the initial reference block in reference list L1 to find the best (or selected) reference block with the smallest prediction residual (2I-P0)-Px, where Px is the prediction of the current block based on the selected reference block.

[0132] Using the reference picture, for the current coding block, the affine ME process may first choose a set of control point motion vectors (CPMVs) as a base. An iterative method may be used to generate a prediction output of the current affine model corresponding to the set of CPMVs, calculate the gradient of the prediction sample, and then solve a linear equation to determine the delta CPMVs and optimize the affine prediction. The iteration may be stopped when all delta CPMVs are 0 or the maximum number of iterations is reached. The CPMV obtained from the iteration may be the final CPMV of the reference picture.

[0133] After the best affine CPVM of both reference lists L0 and L1 is determined for affine uni-prediction, an affine bi-prediction search may be performed using the best uni-prediction CPMV and the reference list on one side to search for the best CPMV of the other reference list to optimize the affine bi-prediction output. The affine bi-prediction search may be performed iteratively on the two reference lists to obtain the optimal result.

[0134] 16 shows an example affine ME process (1600) in which a final CPMV associated with a reference picture may be calculated. The affine ME process (1600) may begin at (S1602). At (S1602), a base CPMV of a current block may be determined. The base CPMV may be determined based on one of a merge index, an advanced motion vector prediction (AMVP) predictor index, an affine merge index, etc.

[0135] In (S1604), an initial affine prediction of the current block may be obtained based on the base CPMV. For example, according to the base CPMV, a four-parameter affine motion model of a six-parameter affine motion model may be applied to generate an initial affine prediction.

[0136] In (S1606), a gradient of the initial affine prediction may be obtained. For example, the gradient of the initial affine prediction may be obtained based on Equation (5) and Equation (6).

[0137] At (S1608), a delta CPMV may be determined. In some embodiments, the delta CPMV may be associated with a displacement between an initial affine prediction and a subsequent affine prediction, such as the first affine prediction. Based on a gradient of the initial affine prediction and the delta CPMV, the first affine prediction may be obtained. The first affine prediction may correspond to the first CPMV.

[0138] At (S1610), a decision may be made to check whether the delta CPMV is zero or the number of iterations is equal to or greater than a threshold. If the delta CPMV is zero or the number of iterations is equal to or greater than a threshold, at (S1612), a final (or selected) CPMV may be determined. The final (or selected) CPMV may be a first CPMV determined based on the gradient of the initial affine prediction and the delta CPMV.

[0139] Further referring to (S1610), if the delta CPMV is not zero or the number of iterations is less than a threshold, a new iteration may be started. In the new iteration, an updated CPMV (e.g., the first CPMV) may be provided to (S1604) to generate an updated affine prediction. The affine ME process (1600) may then proceed to (S1606), where a gradient of the updated affine prediction may be calculated. The affine ME process (1600) may then proceed to (S1608) to continue the new iteration.

[0140] In the affine motion model, the four-parameter affine motion model may be further described by an equation including rotation and zoom motion. For example, the four-parameter affine motion model may be rewritten in equation (13) as follows:

number

number

number

[0141] Bidirectional Optical Flow (BDOF) in VVC was previously called BIO in JEM. Compared to the JEM version, BDOF in VVC can be a simpler version that requires less computation, especially in terms of the number of multiplications and the size of the multipliers.

[0142] BDOF may be used to refine the bi-predictive signal of a CU at the 4×4 sub-block level. BDOF may be applied to a CU if the CU satisfies the following conditions: (1) The CU is coded using a “true” bi-prediction mode, i.e., one of the two reference pictures is before the current picture in display order, and the other is after the current picture in display order; (2) The distances (e.g., POC differences) from the two reference pictures to the current picture are the same; (3) both reference pictures are short-term reference pictures; (4) the CU is not coded using affine mode or SbTMVP merge mode; (5) The CU has more than 64 luma samples, (6) Both the CU height and the CU width are equal to or greater than 8 luma samples; (7) The BCW weight index indicates equal weights, (8) Weight position (WP) is not enabled for the current CU; (9) CIIP mode is not used for the current CU.

[0143] BDOF may be applied only to the luminance component. As the name BDOF suggests, the BDOF mode may be based on the concept of optical flow, which assumes that the motion of objects is smooth. For each 4 × 4 sub-block, the motion accuracy (v x ,v y ) may be calculated. Motion refinement may then be used to adjust the bi-predictive sample values ​​in the 4×4 sub-block. BDOF may include the following steps:

[0144] First, the horizontal gradients of the two predicted signals from reference list L0 and reference list L1

number

number

number

number

[0145] Next, the autocorrelation and cross-correlation of the gradients S1, S2, S3, S5, and S6 may be calculated according to the following equations (18)-(22): S1=Σ (i,j)∈Ω Abs(ψ x (i,j)) Equation (18) S2=Σ (i,j)∈Ω ψ x (i,j) Sign(ψ y (i,j)) Equation (19) S3=Σ (i,j)∈Ω θ(i,j)·Sign(ψ x (i,j)) Equation (20) S5=Σ (i,j)∈Ω Abs(ψ x (i,j)) Equation (21) S6=Σ (i,j)∈Ω θ(i,j)·Sign(ψ y (i,j)) Equation (22) Here, ψ x (i,j), ψ y (i,j) and θ(i,j) can be given by equations (23) to (25), respectively.

number

number

[0146] Then, motion refinement (v x ,v y ) may be derived using cross-correlation and auto-correlation terms using equations (26) and (27) as follows:

number

number

number

number

number

number

number

[0147] Finally, the BDOF samples of the CU may be calculated by adjusting the bi-predictive samples in equation (29) as follows: pred BDOF (x,y)(I (0) (x,y)+I (1) (x,y)+b(x,y)+o offset )>>shift Eq.(29) The values ​​may be selected such that the multipliers in the BDOF process do not exceed 15 bits and the maximum bit width of the intermediate parameters in the BDOF process may be kept within 32 bits.

[0148] To derive the gradient value, we first select some prediction samples I in list k (k=0, 1) outside the current CU boundary. (k) (i,j) needs to be generated. As shown in FIG. 17, BDOF in VVC can use one extended row / column (1702) around the boundary (1706) of the CU (1704). To control the computational complexity of generating out-of-bounds predicted samples, the predicted samples in the extended region (e.g., the unshaded region in FIG. 17) can be generated by taking reference samples of nearby integer positions directly (e.g., using a floor() operation on the coordinates) without interpolation, and a normal 8-tap motion compensation interpolation filter can be used to generate the predicted samples in the CU (e.g., the shaded region in FIG. 17). The extended sample values ​​can be used only in gradient calculation. In the remaining steps of the BDOF process, if samples and gradient values ​​outside the CU boundary are required, the samples and gradient values ​​can be padded (e.g., repeated) from the nearest neighbors of the samples and gradient values.

[0149] If the width and / or height of a CU is greater than 16 luma samples, the CU may be divided into sub-blocks with width and / or height equal to 16 luma samples, and the sub-block boundaries may be treated as CU boundaries in the BDOF process. The maximum unit size of the BDOF process may be limited to 16×16. For each sub-block, the BDOF process may be skipped. If the sum of absolute differences (SAD) between the initial L0 and L1 predicted samples is less than a threshold, the BDOF process may not be applied to the sub-block. The threshold may be set equal to (8*W*(H>>1), where W may indicate the width of the sub-block and H may indicate the height of the sub-block. To avoid additional complexity of the SAD calculation, the SAD between the initial L0 predicted samples and the L1 predicted samples calculated in the DMVR process may be reused in the BDOF process.

[0150] If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, bidirectional optical flow may be disabled. Similarly, if WP is enabled for the current block, i.e., the luma weight flag (e.g., luma_weight_lx_flag) is 1 for either of the two reference pictures, BDOF may also be disabled. If the CU is coded in symmetric MVD mode or CIIP mode, BDOF may also be disabled.

[0151] To improve the accuracy of the MV in the merge mode, a bidirectional matching (BM)-based decoder-side motion vector refinement such as in VVC may be applied. In bi-predictive operation, a refined MV may be searched around the initial MV in the reference picture list L0 and the reference picture list L1. In the BM method, the distortion between two candidate blocks in the reference picture list L0 and the list L1 may be calculated.

[0152] FIG. 18 shows an example schematic diagram of BM-based decoder-side motion vector refinement. As shown in FIG. 18, a current picture (1802) may include a current block (1808). The current picture may include a reference picture list L0 (1804) and a reference picture list L1 (1806). The current block (1808) may include an initial reference block (1812) in reference picture list L0 (1804) with an initial motion vector MV0 and an initial reference block (1814) in reference picture list L1 (1806) with an initial motion vector MV1. A search process may be performed around an initial MV0 in reference picture list L0 (1804) and an initial MV1 in reference picture list L1 (1806). For example, a first candidate reference block (1810) may be identified in reference picture list L0 (1804), and a first candidate reference block (1816) may be identified in reference picture list L1 (1806). The SAD between the candidate reference blocks (e.g., MV0 and MV1) based on each MV candidate (e.g., MV0′ and MV1′) around the initial MV (e.g., (1810) and (1816)) may be calculated. The MV candidate with the lowest SAD becomes the refined MV and may be used to generate a bi-predictive signal for predicting the current block (1808).

[0153] The application of DMVR may be restricted and may only be applied to CUs coded based on modes and features such as VVC as follows: (1) CU level merge mode with bi-predictive MVs; (2) For a current picture, one reference picture is in the past and another reference picture is in the future; (3) The distances (e.g., POC difference) from the two reference pictures to the current picture are the same; (4) both reference pictures are short-term reference pictures; (5) The CU has more than 64 luma samples, (6) Both the CU height and the CU width are equal to or greater than 8 luma samples; (7) The BCW weight index indicates equal weights, (8)WP is not enabled for the current block, (9) CIIP mode is not used for the current block.

[0154] The refined MVs derived by the DMVR process may be used to generate inter-prediction samples and may be used in temporal motion vector prediction for future picture coding, while the original MVs may be used in the deblocking process and may be used in spatial motion vector prediction for future CU coding.

[0155] In DVMR, the search points can surround the initial MV, and the MV offset can follow the MV difference mirroring rule. In other words, any point checked by DMVR, indicated by the candidate MV pair (MV0, MV1), can follow the MV difference mirroring rule shown in equations (30) and (31). MV0'=MV0+MV_offset Equation (30) MV1'=MV1-MV_offset formula (31) Here, MV_offset may represent a refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range may be two integer luma samples from the initial MV. The search may include an integer sample offset search stage and a fractional sample refinement stage.

[0156] For example, a 25-point full search may be applied to the integer sample offset search. The SAD of the initial MV pair may be calculated first. If the SAD of the initial MV pair is less than a threshold, the integer sample stage of the DMVR may be terminated. If not, the SAD of the remaining 24 points may be calculated and checked in a scan order, such as a raster scan order. The point with the smallest SAD may be selected as the output of the integer sample offset search stage. In order to reduce the penalty of uncertainty in DMVR refinement, the original MVs in the DMVR process may have a priority to be selected. The SAD between the reference blocks referenced by the initial MV candidates may be reduced by ¼ of the SAD value.

[0157] Following the integer sample search, fractional sample refinement can be performed. To reduce computational complexity, instead of additional search by SAD comparison, fractional sample refinement can be derived by using a parametric error surface equation. The fractional sample refinement can be conditionally invoked based on the output of the integer sample search stage. If the integer sample search stage ends at the center of the minimum SAD in either the first iteration search or the second iteration search, the fractional sample refinement can be further applied.

[0158] In parametric error surface based sub-pixel offset estimation, a 2D parabolic error surface equation can be fitted based on equation (32) using the central location cost and the costs at the four neighboring locations from the center: E(x,y)=A(xx min ) 2 +B(yy min ) 2 +C formula (32) Here, (x min ,y min ) can correspond to the minimum cost fractional position, and C can correspond to the minimum cost value. By solving equation (32) using the cost values ​​of the five search points, (x min ,y min ) can be calculated: x min=(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) Equation (33) y min =(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0))) Equation (34) x min and y min The value of x can be automatically constrained to be between -8 and 8 since all cost values ​​are positive and the minimum value is E(0,0). min and y min The value constraint of can correspond to a half pel (or pixel) offset with MV precision of 1 / 16 pel in VVC. The calculated fraction (x min , y min ) can be added to the integer distance refined MV to obtain the sub-pixel accurate refined delta MV.

[0159] In VVC and the like, secondary interpolation and sample padding may be applied. The resolution of the MV may be, for example, 1 / 16 luma samples. The fractional position samples may be interpolated using an 8-tap interpolation filter. In DMVR, the search points may surround the initial fractional pel MV with integer sample offsets, so the fractional position samples need to be interpolated for the DMVR search process. To reduce the computational complexity, a bilinear complementation filter may be used to generate the fractional samples for the search process in DMVR. In another important effect, by using a bilinear filter with a 2-sample search range, DVMR does not access more reference samples compared to a normal motion compensation process. After the refined MV is achieved with the DMVR search process, a normal 8-tap interpolation filter may be applied to generate the final prediction. Due to the lack of access to more reference samples compared to a normal MC process, samples that may not be necessary for the interpolation process based on the original MV, but may be necessary for the interpolation process based on the refined MV, may be padded from the available samples.

[0160] If the width and / or height of a CU is greater than 16 luma samples, the CU may be further divided into sub-blocks having widths and / or heights equal to 16 luma samples. The maximum unit size of the DMVR search process may be limited to 16x16.

[0161] In one embodiment, a merge by motion vector difference (MMVD) mode is used in VVC, etc., and samples of a CU (e.g., a current CU) can be predicted using implicitly derived motion information. The MMVD mode is used in either skip mode or merge mode in the motion vector representation method. For example, after signaling a skip flag or a merge flag, an MMVD merge flag can be signaled to specify whether the MMVD mode is used for a CU.

[0162] In some examples, MMVD reuses merge candidates. Candidates can be selected from among the merge candidates, and is further extended by a motion vector representation method. MMVD provides motion vector representation with simplified signaling. In some examples, the motion vector representation method includes a starting point, a motion magnitude, and a motion direction.

[0163] In some instances (e.g., VVC), the MMVD technique may use a merge candidate list to select the starting candidate, however, in one instance, only candidates that are the default merge type (MRG_TYPE_DEFAULT_N) are considered for the extension of the MMVD.

[0164] In some examples, a base candidate index is used to define a starting point. The base candidate index indicates the best candidate among the candidates in the list, as shown in Table 1. For example, the list is a merge candidate list with a motion vector predictor (MVP). The base candidate index can indicate the best candidate in the merge candidate list.

[0165] [Table 1]

[0166] Note that in one example, if the number of base candidates is equal to 1, the base candidate IDX is not signaled.

[0167] In MMVD mode, after a merge candidate (also called an MV basis or MV origin) is selected, the merge candidate may be refined by additional information such as signaled MVD information. The additional information may include an index used to specify the motion magnitude (e.g., a distance index, e.g., mmvd_distance_idx[x0][y0]) and an index used to indicate the motion direction (e.g., a direction index, e.g., mmvd_direction_idx[x0][y0]). In MMVD mode, one of the first two candidates in the merge list may be selected as the MV basis. For example, a merge candidate flag (e.g., mmvd_cand_flag[x0][y0]) indicates one of the first two candidates in the merge list. The merge candidate flag may be signaled to indicate (e.g., specify) which of the first two candidates was selected. The additional information may indicate the MVD (or motion offset) relative to the MV basis. For example, the motion magnitude indicates the magnitude of the MVD, and the motion direction indicates the direction of the MVD.

[0168] In one example, a merging candidate selected from the merging candidate list is used to provide a start point or MV start point in a reference picture. The motion vector of the current block may be represented by a start point and a motion offset (or MVD) that includes the magnitude and direction of motion relative to the start point. On the encoder side, the selection of the merging candidate and the determination of the motion offset may be based on a search process (evaluation process) as shown in Figure 19. On the decoder side, the selected merging candidate and the motion offset may be determined based on signaling from the encoder side.

[0169] FIG. 19 illustrates an example of a search process (1900) in MMVD mode. FIG. 20 illustrates an example of search points in MMVD mode. In some examples, a subset or the entire set of search points in FIG. 20 are used in the search process (1900) of FIG. 19. By performing the search process (1900), for example, at the encoder side, additional information may be determined for a current block (1901) in a current picture (or current frame), including a merge candidate flag (e.g., mmvd_cand_flag[x0][y0]), a distance index (e.g., mmvd_distance_idx[x0][y0]), and a direction index (e.g., mmvd_direction_idx[x0][y0]).

[0170] A first motion vector (1911) and a second motion vector (1921) belonging to a first merge candidate are shown. The first motion vector (1911) and the second motion vector (1921) are MV start points used in the search process (1900). The first merge candidate may be a merge candidate on a merge candidate list configured for the current block (1901). The first motion vector (1911) and the second motion vector (1921) may be associated with two reference pictures (1902) and (1903) in the reference picture lists L0 and L1, respectively. Referring to Figures 19 to 20, the first motion vector (1911) and the second motion vector (1921) may point to two start points (2011) and (2021) in the reference pictures (1902) and (1903), respectively, as shown in Figure 20.

[0171] Referring to Figure 20, two starting points (2011) and (2021) in Figure 20 may be determined in the reference pictures (1902) and (1903). In one example, based on the starting points (2011) and (2021), a number of predetermined points extending from the starting points (2011) and (2021) in the vertical direction (represented by +Y or -Y) or the horizontal direction (represented by +X and -X) in the reference pictures (1902) and (1903) may be evaluated. In one example, pairs of points that mirror each other with respect to the respective starting points (2011) or (2021), such as pair of points (2014) and (2024) (e.g., shown by a shift of 1S in FIG. 19), or pair of points (2015) and (2025) (e.g., shown by a shift of 2S in FIG. 19), can be used to determine pairs of motion vectors (e.g., MVs (1913) and (1923) in FIG. 19) that may form candidate predictor motion vectors for the current block (1901). Predictor motion vector candidates (e.g., MVs (1913) and (1923) in FIG. 19) determined based on predetermined points surrounding the starting points (2011) or (2021) can be evaluated.

[0172] The distance index (e.g., mmvd_distance_idx[x0][y0]) specifies motion magnitude information and may indicate a predefined offset (e.g., 1S or 2S in FIG. 19) from the starting point indicated by the merge candidate flag. Note that in one example, the predefined offset is also referred to as the MMVD step.

[0173] With reference to FIG. 19, an offset (e.g., MVD(1912) or MVD(1922)) may be applied (e.g., added) to a horizontal or vertical component of a starting MV (e.g., MV(1911) or (1921)). An example relationship between a distance index (IDX) and a pre-defined offset is specified in Table 2. When full-pel MMVD is off, e.g., when a full-pel MMVD flag (e.g., slice_fpel_mmvd_enabled_flag) is equal to 0, the pre-defined offset of the MMVD may range from ¼ luma sample to 32 luma samples. When full-pel MMVD is off, the pre-defined offset may have a non-integer value, such as a fraction of a luma sample (e.g., ¼ pixel or ½ pixel). When full-pel MMVD is on, e.g., when a full-pel MMVD flag (e.g., slice_fpel_mmvd_enabled_flag) is equal to 1, the predefined offset of the MMVD can range from 1 luma sample to 128 luma samples. In one example, when full-pel MMVD is on, the predefined offset has only integer values, such as one or more luma samples.

[0174] [Table 2]

[0175] The direction index can represent the direction (or motion direction) of the MVD relative to the starting point. In one example, the direction index represents one of the four directions shown in Table 3. The meaning of the MVD code in Table 3 may change depending on the information of the starting MV. In one example, if the starting MV is a uni-predictive MV or the starting MV is a bi-predictive MV with both reference lists pointing to the same side of the current picture (e.g., the POCs of the two reference pictures are both greater than the POC of the current picture, or the POCs of the two reference pictures are both less than the POC of the current picture), the MVD code in Table 3 specifies the code of the MV offset (or MVD) added to the starting MV.

[0176] If the start MV is a bi-predictive MV where the two MVs point to different sides of the current picture (e.g., the POC of one reference picture is larger than that of the current picture and the POC of the other reference picture is smaller than that of the current picture), the MVD code in Table 3 specifies the code of the MV offset (or MVD) added to the list0 MV component of the start MV, and the MVD code of the list1 MV has the opposite value. With reference to Figure 19, the start MVs (1911) and (1921) are bi-predictive MVs where the two MVs (1911) and (1921) point to different sides of the current picture. The POC of the L1 reference picture (1903) is larger than that of the current picture, and the POC of the L0 reference picture (1902) is smaller than that of the current picture. The MVD sign (e.g., the x-axis sign “+”) indicated by the direction index (e.g., 00) in Table 2 identifies the sign (e.g., the x-axis sign “+”) of the MVD (e.g., MVD(1912)) added to the list0 MV component of the starting MV (e.g., (1911)), and the MVD sign of MVD(1922) for the list1 MV component of the starting MV (e.g., (1921)) has the opposite value, such as the opposite sign “-” to the sign “+” of MVD(1912).

[0177] Referring to Table 3, directional index 00 indicates the positive direction of the x-axis, directional index 01 indicates the negative direction of the x-axis, directional index 10 indicates the positive direction of the y-axis, and directional index 11 indicates the negative direction of the y-axis.

[0178] [Table 3]

[0179] The syntax element mmvd_merge_flag[x0][y0] may be used to represent the MMVD merge flag of the current CU. In one example, an MMVD merge flag equal to 1 (e.g., mmvd_merge_flag[x0][y0]) specifies that the MMVD mode is used to generate inter prediction parameters of the current CU. An MMVD merge flag equal to 0 (e.g., mmvd_merge_flag[x0][y0]) specifies that the MMVD mode is not used to generate inter prediction parameters. The array indexes x0 and y0 may specify the position (x0, y0) of the top-left luma sample of the considered coding block (e.g., the current CB) relative to the top-left luma sample of the picture (e.g., the current picture).

[0180] If the MMVD merge flag (e.g., mmvd_merge_flag[x0][y0]) is not present for the current CU, the MMVD merge flag (e.g., mmvd_merge_flag[x0][y0]) may be inferred to be equal to 0 for the current CU.

[0181] In some examples, such as the VVC specification, a single context is used to signal the MMVD merge flag (e.g., mmvd_merge_flag). For example, a single context is used to code (e.g., encode and / or decode) the MMVD merge flag in context-adaptive binary arithmetic coding (CABAC).

[0182] The syntax element mmvd_cand_flag[x0][y0] may represent a merge candidate flag. In one example, the merge candidate flag (e.g., mmvd_cand_flag[x0][y0]) specifies whether to use the first (0) or second (1) candidate in the merge candidate list using the MVD derived from the distance index (e.g., mmvd_distance_idx[x0][y0]) and the direction index (e.g., mmvd_direction_idx[x0][y0]). The array indexes x0 and y0 may specify the position (x0, y0) of the top-left luma sample of the considered coding block (e.g., current CB) relative to the top-left luma sample of the picture (e.g., current picture).

[0183] If a merge candidate flag (e.g., mmvd_cand_flag[x0][y0]) is not present, then the merge candidate flag (e.g., mmvd_cand_flag[x0][y0]) may be inferred to be equal to 0.

[0184] The syntax element mmvd_distance_idx[x0][y0] may represent a distance index. In one example, the distance index (e.g., mmvd_distance_idx[x0][y0]) specifies an index used to derive MmvdDistance[x0][y0] as specified in Table 4. The array indexes x0 and y0 may specify the position (x0, y0) of the top-left luma sample of the considered coding block (e.g., the current CB) relative to the top-left luma sample of the picture (e.g., the current picture).

[0185] [Table 4]

[0186] The first column of Table 4 indicates the distance index (e.g., mmvd_distance_idx[x0][y0]). The second column of Table 4 indicates the motion magnitude (e.g., MmvdDistance[x0][y0]) when full-pel MMVD is off, e.g., when the full-pel MMVD flag (e.g., slice_fpel_mmvd_enabled_flag) is equal to 0. The third column of Table 4 indicates the motion magnitude (e.g., MmvdDistance[x0][y0]) when full-pel MMVD is on, e.g., when the full-pel MMVD flag (e.g., slice_fpel_mmvd_enabled_flag) is equal to 1.

[0187] In one example, the units in the second and third columns of Table 4 are 1 / 4 luma samples. Referring to the first row of Table 4, if the distance index (e.g., mmvd_distance_idx[x0][y0]) is 0, when full-pel MMVD is off (e.g., slice_fpel_mmvd_enabled_flag is 0), the motion magnitude (e.g., MmvdDistance[x0][y0]) is 1×1 / 4 luma samples or 1 / 4 luma samples. If the distance index (e.g., mmvd_distance_idx[x0][y0]) is 0, when full-pel MMVD is on (e.g., slice_fpel_mmvd_enabled_flag is 1), the motion magnitude (e.g., MmvdDistance[x0][y0]) is 4. The motion magnitude (eg, MmvdDistance[x0][y0]) is 4×1 / 4 luma samples or 1 luma sample.

[0188] In one example, the second column of Table 4 (in ¼ luma sample units) corresponds to the second row of Table 1 (in luma sample units), and the third column of Table 4 (in ¼ luma sample units) corresponds to the third row of Table 2 (in luma sample units).

[0189] The syntax element mmvd_direction_idx[x0][y0] can represent a direction index. In one example, the direction index (e.g., mmvd_direction_idx[x0][y0]) specifies an index used to derive a motion direction (e.g., MmvdSign[x0][y0]) as shown in Table 5. The array indexes x0 and y0 specify the position (x0, y0) of the top-left luma sample of the considered coding block (e.g., current CB) relative to the top-left luma sample of the picture (e.g., current picture). The first column of Table 4 indicates the direction index (e.g., mmvd_distance_idx[x0][y0]). The second column of Table 5 indicates the first component of the MVD (e.g., MVD x The third column of Table 5 indicates the first sign (e.g., MmvdSign[x0][y0][0]) of the MVD (e.g., MmvdOffset[x0][y0][0]). y Or MmvdOffset[x0][y0][1]), indicates the second sign (e.g., MmvdSign[x0][y0][1]).

[0190] [Table 5]

[0191] The first component (e.g., MmvdOffset[x0][y0][0]) and the second component (e.g., MmvdOffset[x0][y0][1]) of the MVD or offset MmvdOffset[x0][y0] may be derived as follows: MmvdOffset[x0][y0][0]=(MmvdDistance[x0][y0]<<2)×MmvdSign[x0][y0][0] (Equation 35) MmvdOffset[x0][y0][1]=(MmvdDistance[x0][y0]<<2)×MmvdSign[x0][y0][1] (Equation 36)

[0192] In one example, the distance index (e.g., mmvd_distance_idx[x0][y0]) is 3 and the direction index (e.g., mmvd_distance_idx[x0][y0]) is 2. Based on Table 5 and the direction index (e.g., mmvd_direction_idx[x0][y0]) being 2, the first component of the MVD (e.g., MVD x or the first sign of MmvdOffset[x0][y0][0] (e.g., MmvdSign[x0][y0][0]) is 0 and the second component of MVD (e.g., MVD y Or the second sign of MmvdOffset[x0][y0][1] (e.g., MmvdSign[x0][y0][1]) is "+1". In this example, the MVD is along the positive vertical direction (+y) and has no horizontal component.

[0193] When the full-pel MMVD flag (e.g., slice_fpel_mmvd_enabled_flag) is equal to 0 and full-pel MMVD is off, the motion magnitude indicated by MmvdDistance[x0][y0] is 8 based on Table 4 and the distance index (e.g., mmvd_distance_idx[x0][y0]) is 3. Based on Equations 10-11, the first component of MVD (e.g., MmvdOffset[x0][y0][0]) is (8<<2)×0=0 and the second component of MVD (e.g., MmvdOffset[x0][y0][1]) is (8<<2)×(+1)=2 (luminance samples).

[0194] When a full-pel MMVD flag (e.g., slice_fpel_mmvd_enabled_flag) is equal to 1 and full-pel MMVD is on, based on Table 4 and the distance index (e.g., mmvd_distance_idx[x0][y0]) is 3, the motion magnitude indicated by MmvdDistance[x0][y0] is 32. Based on (Equation 35) and (Equation 36), the first component of MVD (e.g., MmvdOffset[x0][y0][0]) is (32<<2)×0=0, and the second component of MVD (e.g., MmvdOffset[x0][y0][1]) is (32<<2)×(+1)=8 (luminance samples).

[0195] According to one aspect of the present disclosure, affine merging with motion vector difference (affine MMVD) may be used in video coding. Affine MMVD selects an available affine merge candidate from a subblock-based merge list as a base predictor. Affine MMVD applies a motion vector offset to the motion vector value of each control point from the base predictor. In one example, if no affine merge candidate is available, affine MMVD is not used. In some examples, a distance index and an offset direction index may be subsequently signaled.

[0196] In some examples, a distance index is signaled to indicate which distance offset to use from an offset table such as that shown in Table 6.

[0197] [Table 6]

[0198] In some examples, the direction index can represent four directions as shown in Table 7, and only the x or y direction can have MV differences, but not both directions.

[0199] [Table 7]

[0200] In some examples, the inter prediction is uni-predictive and the signaled distance offset is applied in the offset direction of each control point predictor to generate a result that includes an MV value for each control point.

[0201] In some examples, the inter prediction is bi-predictive, and the signaled distance offset may be applied in the signaled offset direction of the L0 motion vector of the control point predictor, and the offset applied to the L1 MV may be applied on a mirroring or scaling basis, as in the examples specified below.

[0202] In a specific example, the inter prediction is bi-predictive, and the signaled distance offset is applied in the signaled offset direction of the L0 motion vector of the control point predictor. For L1 CPMV, the offset is applied on a mirror basis, which means that the same amount of distance offset is applied in the opposite direction.

[0203] In another example, the POC distance-based offset mirroring method is used for bi-prediction. When the base candidate is bi-predicted, the offset applied to L0 is as signaled, and the offset for L1 depends on the temporal positions of the reference pictures in list L0 and list L1. For example, if both reference pictures are on the same temporal side of the current picture, the same distance offset and the same offset direction are applied to the CPMV of both L0 and L1. In another example, if the two reference pictures are on different sides of the current picture, the CPMV of L1 can apply a distance offset in the opposite offset direction.

[0204] In another example, a POC distance-based offset scaling method is used for bi-prediction. If the base candidate is bi-predicted, the offset applied to L0 is as signaled, and the offset for L1 may be scaled based on the temporal distance of the reference pictures of list 0 and list 1.

[0205] In some examples, the range of distance offset values ​​is extended. For example, three sets of distance offset values ​​may be provided, and a set of distance offset values ​​may be adaptively selected based on the picture resolution. In one example, the offset table is selected based on the picture resolution. Table 8 shows an example of an extended distance offset table including three sets of distance offset values ​​each associated with a different picture resolution. The set of distance offset values ​​may be selected based on the picture resolution.

[0206] [Table 8]

[0207] Template matching (TM) techniques may be used in video / image coding. To further improve the compression efficiency of the VVC standard, for example, the MV may be refined using TM. In one example, TM is used at the decoder side. In the TM mode, the MV may be refined by constructing a template (e.g., a current template) for a block (e.g., a current block) in a current picture, and determining the closest match between the template for the block in the current picture and multiple possible templates (e.g., multiple possible reference templates) in a reference picture. In one embodiment, the template for the block in the current picture may include reconstructed samples of the left neighborhood of the block and reconstructed samples of the upper neighborhood of the block. TM may be used for video / image coding other than VVC.

[0208] FIG. 21 illustrates an example of template matching (2100). The TM may be used to derive motion information for the current CU (2101) (e.g., derive final motion information from initial motion information such as initial MV 2102) by determining the closest match between a template (e.g., current template) (2121) of a current CU (e.g., current block) (2101) in a current picture (2110) and a template (e.g., reference template) among multiple possible templates (e.g., one of the multiple possible templates is template (2125)) in a reference picture (2111). The template (2121) of the current CU (2101) may have any suitable shape and any suitable size.

[0209] In one embodiment, the template (2121) of the current CU (2101) includes a top template (2122) and a left template (2123). Each of the top template (2122) and the left template (2123) may have any suitable shape and any suitable size.

[0210] The top template (2122) may include samples in one or more top-neighboring blocks of the current CU (2101). In one example, the top template (2122) includes four rows of samples in one or more top-neighboring blocks of the current CU (2101). The left template (2123) may include samples in one or more left-neighboring blocks of the current CU (2101). In one example, the left template (2123) includes four columns of samples in one or more left-neighboring blocks of the current CU (2101).

[0211] Each of the multiple possible templates (e.g., template (2125)) in the reference picture (2111) corresponds to a template (2121) in the current picture (2110). In one embodiment, the initial MV (2102) points from the current CU (2101) in the reference picture (2111) to the reference block (2103). Each of the multiple possible templates (e.g., template (2125)) in the reference picture (2111) and the template (2121) in the current picture (2110) may have the same shape and size. For example, the template (2125) of the reference block (2103) includes a top template (2126) in the reference picture (2111) and a left template (2127) in the reference picture (2111). The top template (2126) may include samples in one or more top-neighboring blocks of the reference block (2103). The left template (2127) may include samples in one or more left neighboring blocks of the reference block (2103).

[0212] The TM cost may be determined based on a pair of templates, such as a template (e.g., a current template) (2121) and a template (e.g., a reference template) (2125). The TM cost may indicate a match between the template (2121) and the template (2125). The optimized MV (or final MV) may be determined based on a search around the initial MV (2102) of the current CU (2101) in a search range (2115). The search range (2115) may have any suitable shape and any suitable number of reference samples. In one example, the search range (2115) in the reference picture (2111) includes a [-L,L]-pel range, where L is a positive integer such as 8 (e.g., 8 samples). For example, a difference (e.g., [0,1]) is determined based on the search range (2115), and an intermediate MV is determined by the sum of the initial MV (2102) and the difference (e.g., [0,1]). Based on the intermediate MV, an intermediate reference block in the reference picture (2111) and a corresponding template may be determined. A TM cost may be determined based on the template (2121) and the intermediate template in the reference picture (2111). The TM cost may correspond to a difference (e.g., [0,0], [0,1], etc., corresponding to the initial MV (2102)) determined based on a search range (2115). In one example, the difference corresponding to the minimum TM cost is selected, and the optimized MV is the sum of the difference corresponding to the minimum TM cost and the initial MV (2102). As described above, the TM can derive the final motion information (e.g., the optimized MV) from the initial motion information (e.g., the initial MV 2102).

[0213] In the example of FIG. 21, we may search for a better MV around the initial motion vector of the current CU within a search range such as [-8pel, +8pel].

[0214] TM may be applied in an affine mode, such as affine AMVP mode, affine merge mode, etc., and may be referred to as affine TM. FIG. 22 shows an example of a TM (2200) in an affine merge mode, etc. The template (2221) of a current block (e.g., current CU) (2201) may correspond to a template (e.g., template (2121) in FIG. 21) in a TM applied to a translational motion model. The reference template (2225) of a reference block in a reference picture may include multiple subblock templates (e.g., 4×4 subblocks) pointed to by MVs derived from control point MVs (CPMVs) of neighboring subblocks (e.g., A0-A3 and L0-L3 as shown in FIG. 22) at block boundaries.

[0215] The search process of the TM applied in affine mode (e.g., affine merge mode) may start from CPMV0 while keeping other CPMVs constant (e.g., (i) CPMV1 if a four-parameter model is used, or (ii) CPMV1 and CPMV2 if a six-parameter model is used). The search may be done towards the horizontal and vertical directions. In one example, the search continues in the diagonal direction only if the zero vector is not the best difference vector found from the horizontal and vertical search. The affine TM may repeat the same search process for CPMV1. The affine TM may repeat the same search process for CPMV2 if a six-parameter model is used. If the zero vector is not the best difference vector from the previous iteration and the search process has been iterated less than three times, the entire search process may be restarted from the refined CPMV0 based on the refined CPMV.

[0216] According to one aspect of the present disclosure, a candidate reordering technique based on template matching can be used to reduce signaling overhead. The candidate reordering technique based on template matching can be used in MMVD and affine MMVD.

[0217] In some examples, the MMVD offset is extended to more positions for the MMVD and affine MMVD modes.

[0218] FIG. 23 illustrates the directions in which refinement positions may be added for the MMVD. In FIG. 23, additional refinement positions along the k×π / 8 diagonal are added, where k is an integer. Position (2301) corresponds to the base candidate and may be the starting point, positions (2311)-(2314) are in the directions of 0, π / 2, π, 3π / 2, positions (2321)-(2324) are in the directions of π / 4, 3π / 4, 5π / 4, 7π / 4, and positions (2331)-(2338) are in the directions of π / 8, 3π / 8, 5π / 8, 7π / 8, 9π / 8, 11π / 8, 13π / 8, 15π / 8. Thus, the number of directions is increased from 4 to 16. Furthermore, in one example, each direction may have 6 MMVD refinement positions. The total number of possible MMVD refinement positions is 16×6.

[0219] According to one aspect of the present disclosure, the SAD cost between the current template (e.g., one row up and one column left of the current block) and the reference template may be calculated for each refinement position. Based on the refinement position's SAD cost, all possible MMVD refinement positions (16×6) for each base candidate are sorted. Then, the top of the refinement positions, such as the top 1 / 8 refinement positions (e.g., 12), having the smallest template SAD cost, are retained as available positions for the resulting MMVD index coding. The MMVD index is binarized by a Rice code with a parameter equal to 2.

[0220] In some examples, the refinement positions of the affine MMVD can be increased, and candidate reordering based on template matching can be applied to the affine MMVD reordering. For example, the affine MMVD refinement positions are in directions along the k×π / 4 diagonal, such as 8 directions of 0, π / 4, π / 2, 3π / 4, π, 5π / 4, 3π / 2, and 7π / 4, respectively. Each direction can have 6 affine MMVD refinement positions. The total number of possible affine MMVD refinement positions is 8×6. In one example, the SAD cost between the current template (e.g., one row above and one column to the left of the current block) and the reference template can be calculated for each refinement position. Based on the SAD cost of the refinement position, all possible affine MMVD refinement positions (8×6) for each base candidate are reordered. Then, the top of the refinement positions, such as the top 1 / 2 refinement positions (eg, 24), having the smallest template SAD cost, are retained as available positions for eventual affine MMVD index coding.

[0221] According to some aspects of the present disclosure, in MMVD, the MV offset is limited to a specific step and direction, even in the extended direction. Step- and direction-limited MMVD may not always capture the best motion candidates.

[0222] Some aspects of the present disclosure provide techniques for applying MV refinement to MMVD candidates.

[0223] In some embodiments, after the MMVD candidate is derived, an additional MV refinement offset is added on top of the MV of the MMVD candidate. In some examples, the additional MV refinement offset may be defined by a refinement step size denoted by M, and a refinement position may be defined relative to the MMVD candidate.

[0224] For example, if the current MV value of an MMVD candidate is (mv_hor, mv_ver), the MV refinement offset is (mv_offset_hor, mv_offset_ver), and the refined MV value, denoted as MV', may be written as (mv_hor+mv_offset_hor, mv_ver+mv_offset_ver).

[0225] In some examples, M (refinement step) is set to be a fraction of the current MMVD step size (e.g., offset in Table 2). For example, M is set equal to 1 / 4 of the current MMVD step size. In one example, the current MMVD step size is 1 / 4 pel and the MV refinement offset is 1 / 16 pel, and if the current MMVD step size is 32 pel, the MV refinement offset is 8 pel.

[0226] It should be noted that the refinement positions may be defined in a variety of configurations.

[0227] In one example, four refinement positions may be defined for an MMVD candidate. Figure 24 is a diagram (2400) showing four refinement positions (2411)-(2414) for an MMVD candidate (2401). The refinement positions (2411)-(2414) have MV refinement offsets (-M, 0), (M, 0), (0,-M), and (0,M), respectively.

[0228] In another example, eight refinement positions may be defined for an MMVD candidate. Figure 25 is a diagram (2500) showing eight refinement positions (2511) through (2518) for an MMVD candidate (2501). The refinement positions (2511) through (2518) have MV refinement offsets (-M,-M), (0,-M), (M,-M), (-M,0), (M,0), (-M,M), (0,M), and (M,M), respectively.

[0229] In one embodiment, if the current MMVD candidate is a uni-predictive candidate, the MV refinement offset may be added directly to the MV value of the current MMVD candidate.

[0230] In some embodiments, if the current MMVD candidate is a bi-predictive candidate, the MV refinement offset may be applied to the MV values ​​of the current MMVD candidate using various techniques. In one embodiment, the MV refinement offset is added to the MV values ​​of both reference lists L0 and L1 in the same way. In other embodiments, the MV refinement offset is added to the MV values ​​of reference list L0 (reference picture from reference list L0), and a mirrored MV refinement offset (meaning both horizontal and vertical components are multiplied by -1) is applied to the MV values ​​of reference list L1 (reference picture from reference list L1).

[0231] In other embodiments, the MV refinement offset is applied to the MV of the reference pictures from reference list L0 and reference list L1 based on the temporal position of the reference pictures relative to the current picture. In one example, if both reference pictures are on the same temporal side of the current picture, the MV refinement offset may be added to the MV of both reference pictures from the reference list in the same way. In another example, if the reference pictures are on different temporal sides of the current picture, the MV refinement offset may be added to the MV of reference list L0 (reference picture from reference list L0), and the mirrored MV refinement offset (meaning both horizontal and vertical components are multiplied by -1) is applied to the MV of reference list L1 (reference picture from reference list L1).

[0232] In another embodiment, the MV refinement offset is applied to the MV of the reference picture from reference list L0 and the scaled MV refinement offset is applied to the MV of the reference picture from reference list L1, the scaling factor being based on the temporal distance between the current picture and the two reference pictures in reference list L0 and reference list L1.

[0233] According to one aspect of the present disclosure, the final MV refinement offset for MMVD candidate refinement may be explicitly signaled or may be derived without signaling.

[0234] In some embodiments, the refinement position is signaled in the bitstream as a refinement index that indicates the final refinement position that corresponds to the final MV refinement offset to be applied.

[0235] In some embodiments, the refinement position is not signaled. In some examples, the original MMVD candidate signaling technique can be used to signal the combination of the MMVD candidate and the MV refinement offset. For example, for each MMVD candidate, the (potential) MV refinement offset (e.g., FIG. 24 or FIG. 25) can be combined with the MMVD candidate to form a new set of potential refined candidates, and then the candidate in the new set with the best template matching cost can be used as the refined candidate.

[0236] In one example, the template matching cost of a motion vector is calculated using a reference template derived from the current template and the motion vector.

[0237] FIG. 26 shows a diagram (2600) illustrating template matching cost calculation in some examples. In the example of FIG. 26, a current picture (2610) includes a current template (2621) for a current block (2601). The current template (2621) may include a top current template (2622) and a left current template (2623). The MV (2602) points to a reference block (2651) in a reference picture (2650). The reference picture (2650) includes a reference template (2671) for the reference block (2651). The reference template (2671) may include a top current template (2672) and a left current template (2673).

[0238] In some examples, the sum of absolute differences (SAD) between the current template (2621) and the reference template (2671) is used as the template matching cost associated with the MV (2602).

[0239] In some examples, a template matching cost-based candidate reordering may be applied. In one example, the reordering of MMVD candidates may be based on the template matching cost of unrefined MMVD candidates. The refinement of each MMVD candidate is independent of the reordering of the MMVD candidates. For example, both the encoder side and the decoder side perform MMVD candidate reordering based on the template matching cost of unrefined MMVD candidates to generate an MMVD candidate order. The encoder appropriately selects an MMVD candidate according to, for example, the best template matching cost among the refined candidates, and signals the selected MMVD candidate in the bitstream according to the order of the MMVD candidates. The decoder may determine the MMVD candidates based on the signal in the bitstream and the MMVD candidate order. The decoder may apply a (potential) MV refinement offset (e.g., FIG. 24 or FIG. 25) to the MMVD candidates to form a new set of potential refined candidates, and then the candidate in the new set with the best template matching cost may be used as the refined candidate. In one example, the new set includes the original MMVD candidates.

[0240] In another example, the reordering of MMVD candidates may be based on the template matching cost of the refined candidates. For example, both the encoder side and the decoder side can respectively determine refined candidates for the MMVD candidates. For example, both the encoder side and the decoder side can respectively apply (potential) MV refinement offsets (e.g., FIG. 24 or FIG. 25) to the MMVD candidates for each MMVD candidate to form a new set of potential refined candidates, and then the candidate in the new set with the best template matching cost can be determined as the refined candidate of the MMVD candidate. In one example, the new set includes the original MMVD candidates. Then, both the encoder side and the decoder side can perform MMVD candidate reordering based on the template matching cost of the refined candidates to generate a refined candidate order. The encoder appropriately selects the refined candidates and signals the selected refined candidates in the bitstream according to the refined candidate order. The decoder can determine the refined candidates based on the signal in the bitstream and the refined candidate order.

[0241] In other embodiments, the refinement position is not signaled. All MV refinement offsets are applied to each MMVD candidate to generate refined MV values, and the refined MV values ​​are combined with the original MMVD candidates to form a new candidate set. The new set of candidates is sorted according to template matching cost. In some examples, the top N candidates (e.g., with the lowest template matching cost) in the new set of candidates are used as candidates to be signaled (referred to as signaling candidates), where N is a positive integer. In some examples, only the N candidates of each MMVD base with the best template matching cost may be used as candidates to be signaled (referred to as signaling candidates), where N is a positive number. For example, if N is 2, then for each MMVD candidate, the top two candidates (with the best template matching cost) associated with the MMVD may be signaled. In one example, the signaling candidates may be further evaluated, for example, according to rate-distortion optimization, to select a signaling candidate for signaling.

[0242] FIG. 27 shows a flow chart outlining a process (2700) according to one embodiment of the present disclosure. The process (2700) may be used in a video encoder. In various embodiments, the process (2700) is performed by a processing circuit, such as the processing circuit of the terminal devices (310), (320), (330) and (340), a processing circuit performing the function of the video encoder (403), a processing circuit performing the function of the video encoder (603), a processing circuit performing the function of the video encoder (703), etc. In some embodiments, the process (2700) is implemented in software instructions, and thus the processing circuit performs the process (2700) when the processing circuit executes the software instructions. The process starts at (S2701) and proceeds to (S2710).

[0243] At (S2710), it is determined that MMVD candidate refinement is applied to code the current block in the current picture.

[0244] At (S2720), a first refined motion vector (MV) value associated with the MMVD candidate is derived for the current block. The first refined MV value is generated by applying a first MV refinement offset to a first motion vector associated with the MMVD candidate.

[0245] At (S2730), MMVD candidate information of the current block is determined according to the first refined MV value and signaled in the bitstream.

[0246] According to one aspect of the present disclosure, the first MV refinement offset is a fraction of the motion vector differential, and the motion vector differential is applied to the base candidate to form the MMVD candidate. In some examples, the refinement step (e.g., M) of the first MV refinement offset is set to ¼ of the MMVD step of the motion vector differential. In some examples, the first MV refinement offset corresponds to a refinement position in four potential refinement positions (e.g., (2411)-(2414) of FIG. 24) for the first motion vector associated with the MMVD candidate (e.g., (2401) of FIG. 24). In some examples, the first MV refinement offset corresponds to a refinement position in eight potential refinement positions (e.g., (2511)-(2518) of FIG. 25) for the first motion vector associated with the MMVD candidate (e.g., (2501) of FIG. 25).

[0247] In some examples, the MMVD candidate is a uni-predictive candidate.

[0248] In some examples, the MMVD candidate is a bi-predictive candidate. A second refined MV value associated with the MMVD candidate is derived, and the second refined MV value is generated by applying a second MV refinement offset to a second motion vector associated with the MMVD candidate. The second refined MV value indicates a second reference block in a second reference picture.

[0249] In one example, the second MV refinement offset is equal to the first MV refinement offset.

[0250] In another example, the second MV refinement offset is a mirrored offset to the first MV refinement offset.

[0251] In another example, the second reference picture and the first reference picture are determined to be on the same time side of the current picture, and then a second refined MV value associated with the MMVD candidate is derived. The second refined MV value is generated by applying a second MV refinement offset to a second motion vector associated with the MMVD candidate, and the second MV refinement offset is equal to the first MV refinement offset.

[0252] In another example, the second reference picture is determined to be on a different time side of the current picture than the first reference picture, and then a second refined MV value associated with the MMVD candidate is derived. The second refined MV value is generated by applying a second MV refinement offset to a second motion vector associated with the MMVD candidate, and the second MV refinement offset is a mirrored offset of the first MV refinement offset.

[0253] In another example, the scaling factor is determined based on a first temporal distance from the current picture to the first reference picture and a second temporal distance from the current picture to the second reference picture. Then, a second refined MV value associated with the MMVD candidate is derived, the second refined MV value being generated by applying a second MV refinement offset to a second motion vector associated with the MMVD candidate, the second MV refinement offset being a scaled offset from the first MV refinement offset according to the scaling factor.

[0254] In some examples, to derive a first refined MV value, an MMVD candidate is determined and a first MV refinement offset is determined. The first MV refinement offset is applied to the MMVD candidate to derive the first refined MV value.

[0255] In some examples, an MMVD candidate may be formed by applying a motion vector difference to a starting motion vector of a base candidate. The MMVD candidate information may include a first index indicating a base candidate from a merge candidate list, where the base candidate provides the starting motion vector. The MMVD candidate information also includes a second index indicating a distance of the motion vector difference from the starting motion vector (also referred to as an MMVD step) and a third index indicating a direction of the motion vector difference.

[0256] In some examples, the potential motion vector differentials are applied to the starting motion vector of a base candidate to generate a plurality of MMVD candidates, and respective template matching costs are calculated for the plurality of MMVD candidates. The plurality of MMVD candidates are sorted into a sorted list according to the template matching costs. The MMVD candidate is appropriately selected from the sorted list. The MMVD candidate information includes a first index indicating a base candidate from the merge candidate list, where the base candidate provides the starting motion vector. And, the MMVD candidate information includes a second index that is an index of the MMVD candidate in the sorted list of the plurality of MMVD candidates.

[0257] In some examples, a signal indicating the MV refinement position corresponding to the first MV refinement offset is directly encoded into the bitstream.

[0258] In some examples, the potential MV refinement offsets are respectively applied to the MMVD candidates to generate refined candidates corresponding to the potential MV refinement offsets. A template matching cost is respectively calculated for the refined candidates. A best template matching cost is determined from the template matching costs. A first MV refinement offset is selected from the potential MV refinement offsets, and the refined candidate corresponding to the first MV refinement offset has the best template matching cost.

[0259] In some examples, potential motion vector differentials are applied to the base candidates to generate potential MMVD candidates. Potential MV refinement offsets are applied to each of the potential MMVD candidates to generate potential refined candidates for each of the potential MMVD candidates. The refined candidates are determined for each of the potential MMVD candidates according to a template matching cost. For example, a first refined candidate for a first potential MMVD candidate is selected from the first potential refined candidates for the first potential MMVD candidate in response to the first refined candidate having the best template matching cost among the first potential refined candidates. The refined candidates are sorted to form a sorted list according to the template matching costs of the refined candidates. A particular refined candidate is selected from the sorted list. The MMVD candidate information includes a first index indicating a base candidate from the merge candidate list and a second index indicating a particular refined candidate in the sorted list.

[0260] In some examples, potential motion vector differentials are applied to the base candidate to generate potential MMVD candidates. Potential MV refinement offsets are applied to each of the potential MMVD candidates to generate potential refined candidates for the potential MMVD candidates. The potential refined candidates are sorted into a sorted potential list according to the template matching costs of the potential refined candidates. A portion of the sorted potential list is selected to form a sorted list of refined candidates. A particular refined candidate is selected from the sorted list. The MMVD candidate information includes a first index indicating the base candidate from the merge candidate list and a second index indicating the particular refined candidate from the sorted list of refined candidates. In one example, the top of all potential refined candidates in the sorted potential list is selected. In another example, the top of the potential refined candidate associated with each of the potential MMVD candidates is selected.

[0261] The process then proceeds to (S2799) and ends.

[0262] The process (2700) may be adapted as appropriate. Steps of the process (2700) may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0263] FIG. 28 shows a flow chart outlining a process (2800) according to one embodiment of the present disclosure. The process (2800) may be used in a video decoder. In various embodiments, the process (2800) is performed by a processing circuit, such as the processing circuit of the terminal devices (310), (320), (330) and (340), the processing circuit performing the function of the video decoder (410), the processing circuit performing the function of the video decoder (510), etc. In some embodiments, the process (2800) is implemented in software instructions, and thus the processing circuit performs the process (2800) when the processing circuit executes the software instructions. The process starts at (S2801) and proceeds to (S2810).

[0264] At (S2810), motion vector differential (MMVD) merge candidate information for the current block in the current picture is extracted (eg, parsed) from the bitstream.

[0265] At (S2820), a first refined motion vector (MV) value associated with the MMVD candidate is derived according to the MMVD candidate information. The first refined MV value is generated by applying a first MV refinement offset to the first motion vector associated with the MMVD candidate. In some examples, the first MV refinement offset associated with the first motion vector of the MMVD candidate is generated based on a refined step size and a plurality of refinement positions. The first refined motion vector (MV) value associated with the MMVD candidate is derived according to the MMVD candidate information and the generated first MV refinement offset.

[0266] At (S2830), the current block is reconstructed according to a first reference block in a first reference picture, the first reference block being indicated by the derived first refined MV value.

[0267] According to one aspect of the present disclosure, the first MV refinement offset is a fraction of the motion vector differential, and the motion vector differential is applied to the base candidate to form the MMVD candidate. In some examples, the refinement step (e.g., M) of the first MV refinement offset is set to ¼ of the MMVD step of the motion vector differential. In some examples, the first MV refinement offset corresponds to a refinement position in four potential refinement positions (e.g., (2411)-(2414) of FIG. 24) for the first motion vector associated with the MMVD candidate (e.g., (2401) of FIG. 24). In some examples, the first MV refinement offset corresponds to a refinement position in eight potential refinement positions (e.g., (2511)-(2518) of FIG. 25) for the first motion vector associated with the MMVD candidate (e.g., (2501) of FIG. 25).

[0268] In some examples, the MMVD candidate is a uni-predictive candidate.

[0269] In some examples, the MMVD candidate is a bi-predictive candidate. A second refined MV value associated with the MMVD candidate is derived, and the second refined MV value is generated by applying a second MV refinement offset to a second motion vector associated with the MMVD candidate. The current block is reconstructed according to a first reference block in the first reference picture and a second reference block in the second reference picture, and the second reference block is indicated by the second refined MV value.

[0270] In one example, the second MV refinement offset is equal to the first MV refinement offset.

[0271] In another example, the second MV refinement offset is a mirrored offset to the first MV refinement offset.

[0272] In another example, the second reference picture and the first reference picture are determined to be on the same time side of the current picture, and then a second refined MV value associated with the MMVD candidate is derived. The second refined MV value is generated by applying a second MV refinement offset to a second motion vector associated with the MMVD candidate, and the second MV refinement offset is equal to the first MV refinement offset.

[0273] In another example, the second reference picture is determined to be on a different time side of the current picture than the first reference picture, and then a second refined MV value associated with the MMVD candidate is derived. The second refined MV value is generated by applying a second MV refinement offset to a second motion vector associated with the MMVD candidate, and the second MV refinement offset is a mirrored offset of the first MV refinement offset.

[0274] In another example, the scaling factor is determined based on a first temporal distance from the current picture to the first reference picture and a second temporal distance from the current picture to the second reference picture. Then, a second refined MV value associated with the MMVD candidate is derived, the second refined MV value being generated by applying a second MV refinement offset to a second motion vector associated with the MMVD candidate, the second MV refinement offset being a scaled offset from the first MV refinement offset according to the scaling factor.

[0275] In some embodiments, to derive the first refined MV value, an MMVD candidate is determined from the MMVD candidate information, a first MV refinement offset is determined, and the first MV refinement offset is applied to the MMVD candidate.

[0276] In some examples, the MMVD candidate information includes a first index indicating a base candidate from the merge candidate list, the base candidate providing a starting motion vector. The MMVD candidate information further includes a second index indicating a distance of the motion vector difference from the starting motion vector and a third index indicating a direction of the motion vector difference. In one example, to determine the MMVD candidate, the motion vector difference is applied to the starting motion vector of the base candidate.

[0277] In some examples, the MMVD candidate information includes a first index indicating a base candidate from a merge candidate list, where the base candidate provides a starting motion vector. The MMVD candidate information also includes a second index indicating an MMVD candidate from a sorted list of the multiple MMVD candidates. In one example, the potential motion vector differential is applied to the starting motion vector of the base candidate to generate the multiple MMVD candidates. Respective template matching costs for the multiple MMVD candidates are calculated. The multiple MMVD candidates are sorted into the sorted list according to the template matching costs. The MMVD candidate is selected from the sorted list according to the second index.

[0278] In some examples, to determine the first MV refinement offset, a signal indicating an MV refinement position corresponding to the first MV refinement offset is decoded from the bitstream.

[0279] In some examples, to determine a first MV refinement offset, the potential MV refinement offsets are applied to the MMVD candidates respectively to generate refined candidates corresponding to the potential MV refinement offsets. Template matching costs are calculated respectively for the refined candidates. A best template matching cost (e.g., the lowest template matching cost) is determined from the template matching costs. The first MV refinement offset is selected from the potential MV refinement offsets, and the refined candidate corresponding to the first MV refinement offset has the best template matching cost.

[0280] In some embodiments, the MMVD candidate information includes a first index indicating a base candidate from the merge candidate list, where the base candidate provides a starting motion vector, and the MMVD candidate information also includes a second index indicating a refined candidate from the sorted list of refined candidates. To derive a first refined MV value, a potential motion vector differential is applied to the base candidate to generate a potential MMVD candidate, and a potential MV refinement offset is applied to each potential MMVD candidate to generate a potential refined candidate for each of the potential MMVD candidates. The refined candidates are determined for each potential MMVD candidate according to a template matching cost. For example, a first refined candidate for a first potential MMVD candidate is selected from the first potential refined candidates for the first potential MMVD candidate in response to the first refined candidate having the best template matching cost among the first potential refined candidates. The refined candidates are sorted to form a sorted list according to the template matching cost of the refined candidates. A particular refined candidate is selected from the sorted list according to the second index, and a first refined MV value is derived according to the particular refined candidate.

[0281] In some examples, the MMVD candidate information includes a first index indicating a base candidate from the merge candidate list, where the base candidate provides a starting motion vector. The MMVD candidate also includes a second index indicating a refined candidate from the sorted list of refined candidates. To derive a first refined MV value, a potential motion vector differential is applied to the base candidate to generate a potential MMVD candidate. A potential MV refinement offset is applied to each of the potential MMVD candidates to generate a potential refined candidate for the potential MMVD candidate. The potential refined candidates are sorted into a sorted potential list according to the template matching costs of the potential refined candidates. A portion of the sorted potential list is selected to form a sorted list of refined candidates. A particular refined candidate is selected from the sorted list according to a second index, and a first refined MV value is determined according to the particular refined candidate.

[0282] To select a portion of the sorted potential list to form the sorted list of refined candidates, in one example, the top potential refined candidates in the sorted potential list are selected, hi another example, the top potential refined candidates for each of the potential MMVD candidates are selected.

[0283] The process then proceeds to (S2899) and ends.

[0284] The process (2800) may be adapted as appropriate. Steps of the process (2800) may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0285] The techniques described above may be implemented as computer software using computer readable instructions and physically stored on one or more computer readable media. For example, Figure 29 illustrates a computer system (2900) suitable for implementing certain embodiments of the disclosed subject matter.

[0286] The computer software may be coded using any suitable machine code or computer language that can be subjected to mechanisms such as assembling, compiling, linking, etc. to produce code containing instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or that can be executed via interpretation, microcode execution, etc.

[0287] The instructions may be executed on various types of computers or components thereof including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, and the like.

[0288] The components illustrated in Figure 29 of the computer system (2900) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The arrangement of components should not be interpreted as having any dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of the computer system (2900).

[0289] The computer system (2900) may include certain human interface input devices. Such human interface input devices may be responsive to input by one or more human users via, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not shown). Human interface devices may also be used to capture certain media not necessarily directly associated with conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still image cameras), video (2D video, 3D video including stereoscopic video, etc.).

[0290] The input human interface devices may include one or more (only one of each is shown) of a keyboard (2901), a mouse (2902), a trackpad (2903), a touch screen (2910), a data glove (not shown), a joystick (2905), a microphone (2906), a scanner (2907), and a camera (2908).

[0291] The computer system (2900) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, by haptic output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touch screen (2910), data gloves (not shown), or joystick (2905), although there may also be haptic feedback devices that do not function as input devices), audio output devices (e.g., speakers (2909), headphones (not shown)), visual output devices (e.g., screens (2910), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capability, each with or without haptic feedback capability, some of which may be capable of outputting two-dimensional visual output or three-dimensional or higher output via means such as stereo output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0292] The computer system (2900) may also include human accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (2920) with media (2921) such as CDs / DVDs, thumb drives (2922), removable hard drives or solid state drives (2923), legacy magnetic media such as tapes and floppy disks (not shown), and dedicated ROM / ASIC / PLD based devices such as security dongles (not shown).

[0293] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass communication media, carrier waves, or other transitory signals.

[0294] The computer system (2900) may also include an interface (2954) to one or more communication networks (2955). The networks may be, for example, wireless, wired, optical. The networks may further be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial including CANBus, etc. Certain networks generally require an external network interface adapter attached to a particular general-purpose data port or peripheral bus (2949) (e.g., a USB port on the computer system (2900)), while other networks are generally integrated into the core of the computer system (2900) by attachment to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (2900) may communicate with other entities. Such communications may be, for example, one-way receive only (e.g., broadcast TV), one-way transmit only (e.g., CANbus to a particular CANbus device), or two-way, to other computer systems using local or wide area digital networks. Specific protocols and protocol stacks may be used in each of these networks and network interfaces, as described above.

[0295] The aforementioned human interface devices, human access storage devices, and network interfaces may be attached to the core (2940) of the computer system (2900).

[0296] The cores (2940) may include one or more central processing units (CPUs) (2941), graphics processing units (GPUs) (2942), dedicated programmable processing units in the form of field programmable gate areas (FPGAs) (2943), hardware accelerators for specific tasks (2944), graphics adapters (2950), and the like. These devices may be connected via a system bus (2948), along with read only memory (ROM) (2945), random access memory (2946), internal mass storage such as an internal non-user accessible hard drive, SSD, and the like (2947). In some computer systems, the system bus (2948) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, and the like. Peripherals may be attached directly to the core's system bus (2948) or via a peripheral bus (2949). In one example, a screen (2910) may be connected to the graphics adapter (2950). Peripheral bus architectures include PCI, USB, etc.

[0297] The CPU (2941), GPU (2942), FPGA (2943), and accelerator (2944) can execute certain instructions that may combine to constitute the above-mentioned computer code. The computer code may be stored in ROM (2945) or RAM (2946). Transient data may also be stored in RAM (2946), while persistent data may be stored, for example, in internal mass storage (2947). Rapid storage and retrieval from any of the memory devices may be enabled by the use of cache memory, which may be closely associated with one or more of the CPU (2941), GPU (2942), mass storage (2947), ROM (2945), RAM (2946), etc.

[0298] The computer-readable medium can bear computer code for performing various computer-implemented operations. The medium and computer code can be those specially designed and constructed for the purposes of the present disclosure, or they can be of the kind well known and available to those skilled in the computer software arts.

[0299] By way of example and not limitation, a computer system having the architecture (2900), and in particular the cores (2940), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage as described above, as well as media associated with specific storage of the cores (2940) of a non-transitory nature, such as the core internal mass storage (2947) or ROM (2945). Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the cores (2940). The computer-readable media can include one or more memory devices or chips, depending on the particular needs. The software can cause the cores (2940), and in particular the processors therein (including a CPU, GPU, FPGA, etc.) to perform certain processes or certain portions of certain processes described herein, including defining data structures stored in RAM (2946) and modifying such data structures according to the processes defined by the software. Additionally, or alternatively, the computer system may provide functionality as a result of logic embodied in hardwired or otherwise circuitry (e.g., accelerator (2944)), which may operate in place of or in conjunction with software to perform particular processes or particular portions of particular processes described herein. References to software may encompass logic, where appropriate, and vice versa. References to computer-readable media may encompass circuitry (such as integrated circuits (ICs)) that stores software for execution, circuitry that embodies logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware and software.

[0300] Appendix A: Acronyms JEM: Joint exploration model VVC: versatile video coding BMS: Benchmark set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplementary Enhancement Information VUI: Video Usability Information GOP: Group of Pictures TU: Transform Unit PU: Prediction Unit CTU: Coding Tree Unit CTB: Coding Tree Block PB: Prediction Block HRD: Hypothetical Reference Decoder SNR: Signal to Noise Ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: Cathode Ray Tube LCD: Liquid Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile communications LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Area SSD: Solid-state drive IC: Integrated Circuit CU: Coding Unit

[0301] While this disclosure describes several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure. [Explanation of symbols]

[0302] 101 sample, 102 arrow, 103 arrow, 104 square block, 201 current block, 300 communication system, 310 terminal device, 320 terminal device, 330 terminal device, 340 terminal device, 350 communication network, 400 communication system, 401 video source, 402 stream, 403 video encoder, 404 video data, 405 streaming server, 406 client subsystem, 407 incoming copy, 410 video decoder, 411 video picture, 412 display, 413 capture subsystem, 420 electronic device, 430 electronic device, 501 channel, 510 video decoder, 512 rendering device, 515 buffer memory, 520 parser, 521 symbol, 530 electronic device, 531 receiver, 551 inverse transform unit, 552 intra-picture prediction unit, 553 motion compensation prediction unit, 555 aggregator, 556 loop filter unit, 557 reference picture memory, 558 picture buffer, 601 video source, 603 video encoder, 620 electronic device, 630 source coder, 632 coding engine, 633 local decoder, 634 reference picture memory, 635 predictor, 640 transmitter, 643 video sequence, 645 entropy coder, 650 controller, 660 communication channel, 703 video encoder, 721 general controller, 722 intra encoder, 723 residual calculator, 724 residual encoder, 725 entropy encoder, 726 switch, 728 residual decoder, 730 inter encoder, 810 video decoder, 871 entropy decoder, 872 intra decoder, 873 residual decoder, 874 reconstruction module, 880 inter decoder, 902 Block, 904 Block, 1000 Block, 1204 Block, 1302 Block, 1400 Block, 1402 Sub-Block, 1404 Sample, 1406 Reference Pixel, 1408 Reference Pixel, 1500 Affine ME, 1600 Affine ME Process, 1702 Row / Column, 1704 CU, 1706 Boundary, 1802 Picture, 1808 Block, 1810candidate reference block, 1812 initial reference block, 1814 initial reference block, 1816 candidate reference block, 1900 search process, 1901 block, 1902 reference picture, 1903 reference picture, 1911 motion vector, 1921 motion vector, 2011 starting point, 2021 starting point, 2100 template matching, 2101 block, 2102 initial MV, 2103 reference block, 2110 picture, 2111 reference picture, 2115 search range, 2121 template, 2122 top template, 2123 left template, 2125 reference template, 2126 top template, 2127 left template, 2221 template, 2225 reference template, 2401 MMVD candidate, 2411 refinement position, 2412 refinement location, 2413 refinement location, 2414 refinement location, 2501 MMVD candidate, 2511 refinement location, 2512 refinement location, 2513 refinement location, 2514 refinement location, 2515 refinement location, 2516 refinement location, 2517 refinement location, 2518 refinement location, 2601 block, 2610 picture, 2621 template, 2622 template, 2623 template, 2650 reference picture, 2651 reference block, 2671 reference template, 2672 template, 2673 template, 2700 process, 2800 process, 2900 computer system, 2901 keyboard, 2902 mouse, 2903 trackpad, 2905 joystick, 2906 microphone, 2907 scanner, 2908 camera, 2909 speaker, 2910 touch screen, 2921 media, 2920 CD / DVD ROM / RW, 2922 thumb drive, 2923 solid state drive, 2940 core, 2941 CPU, 2942 GPU, 2943 FPGA, 2944 accelerator, 2945 ROM, 2946 memory, 2947 mass storage, 2948 system bus, 2949 peripheral bus, 2950 graphics adapter, 2954 interface, 2955 communication network, AF_MERGE affine merge, AMVP advanced motion vector prediction, AMVR adaptive motion vector resolution, BCWCU level weights, BDOF bidirectional optical flow, BM bidirectional matching, BMS benchmark set, CABAC context-adaptive binary arithmetic coding, CIIP combined inter and intra prediction, CPMV control point motion vector, CPU central processing unit, CTB coding tree block, CTU coding tree unit, CU coding unit, DMVR decoder side motion vector refinement, GPM geometric partition mode, GPU graphics processing unit, HDR high dynamic range, HRD virtual reference decoder, JEM joint search model, L0 reference picture list, L1 reference picture list, LAN local area network, MCP motion compensated prediction, ME motion estimation, MMVD merge by motion vector difference, MV motion vector, MVP motion vector predictor, PROF prediction refinement by optical flow, PU prediction unit, QP quantizer parameter, SAD sum of absolute difference, SDR standard dynamic range, SEI supplementary enhancement information, SNR signal to noise ratio, TM template matching, TU Transformation unit, VTM VVC test model, VUI video usability information, VVC versatile video coding, WP weighted position, S1502 affine uni-prediction, S1504 affine uni-prediction, S1506 affine bi-prediction

Claims

1. 1. A method of video processing in a decoder, comprising the steps of: extracting, from the bitstream, merge by motion vector difference (MMVD) candidate information for a current block in a current picture; generating a first MV refinement offset associated with a first motion vector of the MMVD candidate based on a refined step size and a plurality of refinement positions; deriving a first refined motion vector (MV) value associated with an MMVD candidate according to the MMVD candidate information and the generated first MV refinement offset; reconstructing the current block according to a first reference block in a first reference picture, the first reference block being indicated by the derived first refined MV value; A method comprising:

2. The method of claim 1 , wherein the first MV refinement offset is a fraction of a motion vector differential applied to a base candidate to form the MMVD candidate.

3. The method of claim 2 , wherein a refinement step of the first MV refinement offset is ¼ of an MMVD step of the motion vector differential.

4. The method of claim 1 , wherein the first MV refinement offset corresponds to a refinement position among four potential refinement positions for the first motion vector associated with the MMVD candidate.

5. The method of claim 1 , wherein the first MV refinement offset corresponds to a refinement position among eight potential refinement positions for the first motion vector associated with the MMVD candidate.

6. The method of claim 1 , wherein the MMVD candidates are uni-predictive candidates.

7. the MMVD candidate is a bi-predictive candidate; deriving a second refined MV value associated with the MMVD candidate, the second refined MV value being generated by applying a second MV refinement offset to a second motion vector associated with the MMVD candidate, the second MV refinement offset being equal to the first MV refinement offset; reconstructing the current block according to the first reference block in the first reference picture and a second reference block in a second reference picture, the second reference block being indicated by the second refined MV value; The method of claim 1, further comprising:

8. The MMVD candidates are bi-predictive candidates, and the method comprises: deriving a second refined MV value associated with the MMVD candidate, the second refined MV value being generated by applying a second MV refinement offset to a second motion vector associated with the MMVD candidate, the second MV refinement offset being a mirrored offset relative to the first MV refinement offset; reconstructing the current block according to the first reference block in the first reference picture and a second reference block in a second reference picture, the second reference block being indicated by the second refined MV value; The method of claim 1, further comprising:

9. The MMVD candidates are bi-predictive candidates, and the method comprises: determining that a second reference picture and the first reference picture are on the same time side of the current picture; deriving a second refined MV value associated with the MMVD candidate, the second refined MV value being generated by applying a second MV refinement offset to a second motion vector associated with the MMVD candidate, the second MV refinement offset being equal to the first MV refinement offset; reconstructing the current block according to the first reference block in the first reference picture and a second reference block in the second reference picture, the second reference block being indicated by the second refined MV value; The method of claim 1, further comprising:

10. The MMVD candidates are bi-predictive candidates, and the method comprises: determining that a second reference picture is on a different time side of the current picture than the first reference picture; deriving a second refined MV value associated with the MMVD candidate, the second refined MV value being generated by applying a second MV refinement offset to a second motion vector associated with the MMVD candidate, the second MV refinement offset being a mirrored offset of the first MV refinement offset; reconstructing the current block according to the first reference block in the first reference picture and a second reference block in the second reference picture, the second reference block being indicated by the second refined MV value; The method of claim 1, further comprising:

11. The MMVD candidates are bi-predictive candidates, and the method comprises: determining a scaling factor based on a first temporal distance from the current picture to the first reference picture and a second temporal distance from the current picture to a second reference picture; deriving a second refined MV value associated with the MMVD candidate, the second refined MV value being generated by applying a second MV refinement offset to a second motion vector associated with the MMVD candidate, the second MV refinement offset being a scaled offset from the first MV refinement offset according to the scaling factor; reconstructing the current block according to the first reference block in the first reference picture and a second reference block in the second reference picture, the second reference block being indicated by the second refined MV value; The method of claim 1, further comprising:

12. The step of deriving the first refined MV value comprises: determining the MMVD candidates from the MMVD candidate information; determining the first MV refinement offset; applying the first MV refinement offset to the MMVD candidate; The method of claim 1, further comprising:

13. The MMVD candidate information is a first index indicating a base candidate from a merge candidate list, the base candidate providing a starting motion vector; a second index indicating a distance of the motion vector differential from the starting motion vector; a third index indicating a direction of the motion vector differential; Including, The step of determining the MMVD candidates comprises: applying the motion vector differential to the starting motion vector of the base candidate to determine the MMVD candidate. The method of claim 12.

14. The MMVD candidate information is a first index indicating a base candidate from a merge candidate list, the base candidate providing a starting motion vector; a second index indicating the MMVD candidate from the sorted list of the plurality of MMVD candidates; said method comprising: applying potential motion vector differentials to the starting motion vector of the base candidate to generate the plurality of MMVD candidates; calculating a template matching cost for each of the multiple MMVD candidates; sorting the plurality of MMVD candidates into the sorted list according to the template matching cost; selecting the MMVD candidate from the sorted list according to the second index; The method of claim 12, comprising:

15. The step of determining the first MV refinement offset comprises: and decoding, from the bitstream, a signal indicative of an MV refinement position corresponding to the first MV refinement offset. The method of claim 12.

16. The step of determining the first MV refinement offset comprises: applying potential MV refinement offsets to the MMVD candidates, respectively, to generate refined candidates corresponding to the potential MV refinement offsets; calculating a template matching cost for each of the refined candidates; determining a best template matching cost from the template matching costs; selecting the first MV refinement offset from the potential MV refinement offsets, where the refined candidate corresponding to the first MV refinement offset has the best template matching cost; 13. The method of claim 12, further comprising:

17. The MMVD candidate information is a first index indicating a base candidate from a merge candidate list, the base candidate providing a starting motion vector; a second index indicating the refined candidate from the sorted list of refined candidates; Including, The step of deriving the first refined MV value comprises: applying potential motion vector differentials to the base candidates to generate potential MMVD candidates; applying a potential MV refinement offset to each of the potential MMVD candidates to generate a potential refined candidate for each of the potential MMVD candidates; determining the refined candidates according to a template matching cost for each of the potential MMVD candidates, wherein a first refined candidate of a first potential MMVD candidate is selected from the first potential refined candidates of the first potential MMVD candidate in response to the first refined candidate having a best template matching cost among the first potential refined candidates; sorting the refined candidates according to their template matching costs to form the sorted list; selecting a particular refined candidate from the sorted list according to the second index; deriving the first refined MV value according to the particular refined candidate; Further comprising: The method of claim 1.

18. The MMVD candidate information is a first index indicating a base candidate from a merge candidate list, the base candidate providing a starting motion vector; a second index indicating the refined candidate from the sorted list of refined candidates; Including, The step of deriving the first refined MV value comprises: applying potential motion vector differentials to the base candidates to generate potential MMVD candidates; applying a potential MV refinement offset to each of the potential MMVD candidates to generate potential refined candidates for the potential MMVD candidates; sorting the potential refined candidates into a sorted potential list according to their template matching costs; selecting a portion of the sorted potential list to form the sorted list of refined candidates; selecting a particular refined candidate from the sorted list according to the second index; and deriving the first refined MV value according to the particular refined candidate. The method of claim 1.

19. The step of selecting the portion of the sorted potential list to form the sorted list of refined candidates comprises: selecting the top of the potential refined candidates in the sorted potential list; selecting a top of the potential refined candidates for each of the potential MMVD candidates.

20. The method of claim 18, further comprising at least one of:

20. 1. An apparatus for video decoding, comprising: Extracting merge by motion vector difference (MMVD) candidate information for a current block in a current picture from the bitstream; generating a first MV refinement offset associated with a first motion vector of the MMVD candidate based on the refined step size and a plurality of refinement positions; Derive a first refined motion vector (MV) value associated with an MMVD candidate according to the MMVD candidate information and the generated first MV refinement offset; 16. An apparatus comprising: a processing circuit configured to reconstruct the current block according to a first reference block in a first reference picture, the first reference block being indicated by a first derived refined MV value.

Citation Information

Patent Citations

  • Image decoding device and image coding device

    JP2020145650A

  • Image processing device and method

    WO2019244669A1