Bilateral matching using affine motion

By employing affine bilateral matching to derive affine motion parameters for video decoding, the method addresses the challenges of redundancy reduction and compression efficiency in video coding, achieving improved performance in encoding and decoding complex video content.

JP7675206B2Active Publication Date: 2025-05-12TENCENT AMERICA LLC

Patent Information

Application Number
JP2023563068
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-09-23
Filing Date
2022-09-28
Publication Date
2025-05-12
Estimated Expiration
2042-09-28

Smart Images

  • Figure 0007675206000031
    Figure 0007675206000031
  • Figure 0007675206000032
    Figure 0007675206000032
  • Figure 0007675206000033
    Figure 0007675206000033
Patent Text Reader

Abstract

Prediction information of a current block in a current picture is decoded from a coded video bitstream, and the prediction information indicates that the current block should be predicted based on an affine model. A plurality of affine motion parameters of the affine model are derived by affine bilateral matching, in which the affine model is derived based on a reference block in a first reference picture and a second reference picture of the current picture. The plurality of affine motion parameters are not included in the coded video bitstream. A control point motion vector of the affine model is determined based on the derived plurality of affine motion parameters. The current block is reconstructed based on the derived affine model.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims benefit of priority to U.S. Patent Application No. 17 / 951,900, entitled "BILATERAL MATCHING WITH AFFINE MOTION," filed September 23, 2022, which claims benefit of priority to U.S. Provisional Application No. 63 / 329,835, entitled "Bilateral Matching with Affine Motion," filed April 11, 2022. The disclosures of the prior applications are incorporated herein by reference in their entireties.

[0002] This disclosure generally describes embodiments related to video coding. [Background technology]

[0003] The background discussion provided herein is intended to generally present the context of the present disclosure. The inventors' work, to the extent that it is described in this background section, and aspects of the description that may not be admitted as prior art at the time of filing, are not admitted expressly or impliedly as prior art to the present disclosure.

[0004] An uncompressed digital image and / or video may include a sequence of pictures, each having spatial dimensions of, for example, 1920x1080 luma samples and associated chroma samples. The sequence of pictures may have a fixed or variable picture rate (also informally known as frame rate) of, for example, 60 pictures per second, i.e., 60 Hz. Uncompressed images and / or videos have specific bitrate requirements. For example, 1080p60 4:2:0 video (1920x1080 luma sample resolution at a frame rate of 60 Hz) with 8 bits per sample requires a bandwidth approaching 1.5 Gbit / s. One hour of such video requires more than 600 GByte of storage space.

[0005] One objective of image and / or video coding and decoding can be the reduction of redundancy in the input image and / or video signal through compression. Compression can help reduce the aforementioned bandwidth and / or storage space requirements, in some cases by more than one order of magnitude. The description herein uses video encoding / decoding as an illustrative example, but the same techniques can be applied to image encoding / decoding in a similar manner without departing from the spirit of this disclosure. Both lossless and lossy compression, and combinations thereof, can be employed. Lossless compression refers to techniques where an exact copy of the original signal can be reconstructed from the compressed original signal. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough to make the reconstructed signal useful for the intended application. For video, lossy compression is widely adopted. The amount of acceptable distortion depends on the application, for example, a user of a particular consumer streaming application may tolerate higher distortion than a user of a television distribution application. The achievable compression ratio may reflect that the higher the tolerable / acceptable distortion, the higher the compression ratio can be.

[0006] Video encoders and decoders can utilize techniques from a number of broad categories, including, for example, motion compensation, transform processing, quantization, and entropy coding.

[0007] Video codec techniques can include a technique known as intra-coding. In intra-coding, sample values ​​are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into blocks of samples. When all blocks of samples are coded in intra mode, the picture can be an intra picture. Intra pictures and their derivatives, such as independent decoder refresh pictures, are used to reset the decoder state and can therefore be used as the first picture in a coded video bitstream and video session or as a still image. Samples of an intra block undergo a transform and the transform coefficients can be quantized before entropy coding. Intra prediction can be a technique that minimizes sample values ​​in the pre-transform domain. In some cases, the smaller the DC value after the transform and the smaller the AC coefficients, the fewer bits are needed for a given quantization step size to represent the block after entropy coding.

[0008] Conventional intra-coding, for example as used in MPEG-2 generation coding techniques, does not use intra-prediction. However, some newer video compression techniques include techniques that attempt to make predictions based on surrounding sample data and / or metadata obtained during encoding / decoding of a block of data. Such techniques are hereafter referred to as "intra-prediction" techniques. It should be noted that in at least some cases, intra-prediction uses only reference data from the current picture being reconstructed, and not from a reference picture.

[0009] Intra prediction can have many different forms. When more than one of such techniques can be used in a given video coding technique, the particular technique in use can be coded as a particular intra prediction mode that uses the particular technique. In certain cases, the intra prediction mode can have sub-modes and / or parameters that can be coded separately or included in a mode codeword that defines the prediction mode used. Which codeword is used for a given mode, sub-mode, and / or parameter combination can affect the coding efficiency gains via intra prediction, and thus the entropy coding technique used to convert the codeword into a bitstream.

[0010] A specific mode of intra prediction was introduced in H.264, improved in H.265, and further refined in new coding techniques such as the Joint Search Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). A predictor block can be formed using sample values ​​in the neighborhood of already available samples. The sample values ​​of the neighboring samples are copied to the predictor block according to a direction. A reference to the direction in use can be coded in the bitstream or can itself be predicted.

[0011] Referring to FIG. 1A, depicted at the bottom right is a subset of 9 predictor directions known from the 33 possible predictor directions defined in H.265 (corresponding to the 33 angle modes of the 35 intra modes). The point where the arrows converge (101) represents the sample being predicted. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples to the upper right at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples to the lower left of sample (101) at an angle of 22.5 degrees from the horizontal.

[0012] 1A, at the top left is shown a square block (104) of 4×4 samples (indicated by a thick dashed line). The square block (104) contains 16 samples, each labeled with "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample (from the top) in the Y dimension and the first sample (from the left) in the X dimension. Similarly, sample S44 is the fourth sample in both the Y and X dimensions in the block (104). Since the block is 4×4 samples in size, S44 is at the bottom right. Also shown are reference samples that follow a similar numbering scheme. The reference samples are labeled with R, their Y position (e.g., row index), and X position (column index) relative to the block (104). In both H.264 and H.265, the predicted samples are in the neighborhood of the block being reconstructed, so negative values ​​do not need to be used.

[0013] Intra-picture prediction can work by copying reference sample values ​​from nearby samples indicated by the signaled prediction direction. For example, assume that the coded video bitstream includes signaling for this block indicating a prediction direction that matches the arrow (102), i.e., the sample is predicted from the sample to the upper right at an angle of 45 degrees from the horizontal. In that case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.

[0014] In certain cases, to calculate a reference sample, especially when the orientation is not evenly divisible by 45 degrees, the values ​​of multiple reference samples may be combined, for example by interpolation.

[0015] The number of possible directions has increased as video coding techniques have developed. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013). Currently, JEM / VVC / BMS can support up to 65 directions. Experiments have been carried out to identify the most likely directions, and certain techniques of entropy coding are used to represent those likely directions with a small number of bits, accepting a certain penalty for less likely directions. Furthermore, in some cases, the direction itself can be predicted from nearby directions used in nearby, already decoded blocks.

[0016] FIG. 1B shows a schematic diagram (110) showing 65 intra prediction directions with JEM to illustrate the increasing number of prediction directions over time.

[0017] The mapping of intra-prediction direction bits, which represent directions in the coded video bitstream, can vary from video coding techniques. Such mappings can range from simple direct mappings to complex adaptive schemes including codewords, most probable modes, and similar techniques. However, in most cases, there may be certain directions that are statistically less likely to occur in the video content than certain other directions. Since the goal of video compression is to reduce redundancy, these less likely directions are represented with more bits than more likely directions in well-performing video coding techniques.

[0018] Image and / or video coding and decoding can be performed using inter-picture prediction with motion compensation. Motion compensation can be a lossy compression technique and can refer to a technique in which blocks of sample data from a previously reconstructed picture or part thereof (reference picture) are spatially shifted in a direction indicated by a motion vector (hereinafter MV) and then used to predict a newly reconstructed picture or part of a picture. In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV can have two dimensions X and Y, or three dimensions, with the third dimension being an indication of the reference picture in use (the third dimension can indirectly be a temporal dimension).

[0019] In some video compression techniques, the MV applicable to a particular area of ​​sample data can be predicted from other MVs, e.g., from MVs associated with other areas of sample data that are spatially adjacent to the area being reconstructed and that precede that MV in decoding order. Doing so can significantly reduce the amount of data required to code the MV, thereby eliminating redundancy and increasing the compression ratio. For example, when coding an input video signal derived from a camera (known as natural video), MV prediction can work effectively because there is a statistical likelihood that areas larger than the area to which a single MV is applicable move in similar directions and therefore, in some cases, can be predicted using similar motion vectors derived from MVs of nearby areas. As a result, the detected MV for a given area is similar or the same as the MV predicted from the surrounding MVs, which, after entropy coding, can be represented with fewer bits than would be used if the MVs were coded directly. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., MV) derived from an original signal (i.e., a sample stream). In other cases, the MV prediction itself can be lossy, e.g., due to rounding errors when computing a predictor from several surrounding MVs.

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms offered by H.265, the one hereafter referred to as "spatial merging" is described with reference to Fig. 2.

[0021] Referring to Figure 2, a current block (201) contains samples that have been discovered by the encoder during the motion search process to be predictable from a previous block of the same size but spatially shifted. Instead of coding its MV directly, the MV can be derived from metadata associated with one or more reference pictures, e.g., from the most recent reference picture (in decoding order), using MVs associated with any one of the five surrounding samples, denoted A0, A1, and B0, B1, B2 (202-206, respectively). In H.265, MV prediction can use predictors from the same reference picture that neighboring blocks are using. Summary of the Invention [Means for solving the problem]

[0022] Aspects of the present disclosure provide a method and apparatus for video encoding / decoding. In some examples, an apparatus for video decoding includes a processing circuit.

[0023] According to one aspect of the present disclosure, a method of video decoding performed in a video decoder is provided. In the method, prediction information of a current block in a current picture can be decoded from a coded video bitstream, and the prediction information can indicate that the current block should be predicted based on an affine model. A plurality of affine motion parameters of the affine model can be derived by affine bilateral matching, in which the affine model is derived based on a reference block in a first reference picture and a second reference picture of the current picture. The plurality of affine motion parameters may not be included in the coded video bitstream. A control point motion vector of the affine model can be determined based on the derived plurality of affine motion parameters. The current block can be reconstructed based on the derived affine model.

[0024] A plurality of affine motion parameters of the affine model can be derived based on a reference block pair from a plurality of candidate reference block pairs of reference blocks in a first reference picture and a second reference picture. The reference block pair can include a first reference block in the first reference picture and a second reference block in the second reference picture based on the affine model and a cost value constraint. The constraint can be associated with a temporal distance ratio based on (i) a first temporal distance between the current picture and the first reference picture, and (ii) a second temporal distance between the current picture and the second reference picture. The cost value can be based on a difference between the first reference block and the second reference block.

[0025] In some embodiments, the temporal distance ratio may be equal to the product of a weighting factor and a ratio of (i) a first temporal distance between the current picture and the first reference picture and (ii) a second temporal distance between the current picture and the second reference picture, where the weighting factor may be a positive integer.

[0026] In some embodiments, the affine model constraint may indicate that a first translation coefficient of a first affine motion vector from the current block to the first reference block is proportional to a temporal distance ratio. The affine model constraint may indicate that a second translation coefficient of a second affine motion vector from the current block to the second reference block is proportional to a temporal distance ratio.

[0027] In some embodiments, the constraints of the affine model may further indicate that a second zoom factor of a second affine motion vector from the current block to the second reference block is equal to a first zoom factor of a first affine motion vector from the current block to the first reference block relative to a power of the temporal distance ratio.

[0028] In some embodiments, the constraint of the affine model may indicate that a ratio of (i) a first delta zoom factor of a first affine motion vector from the current block to the first reference block and (ii) a second delta zoom factor of a second affine motion vector from the current block to the second reference block is equal to a temporal distance ratio. The first delta zoom factor may be equal to the first zoom factor -1, and the second delta zoom factor may be equal to the second zoom factor -1.

[0029] In some embodiments, the constraints of the affine model may further indicate that a ratio between (i) a first rotation angle of a first affine motion vector from the current block to the first reference block and (ii) a second rotation angle of a second affine motion vector from the current block to the second reference block is equal to a temporal distance ratio.

[0030] To determine the reference block pair, a plurality of candidate reference block pairs can be determined according to the constraints of the affine model. Each candidate reference block pair of the plurality of candidate reference block pairs can include a respective candidate reference block in a first reference picture and a respective candidate reference block in a second reference picture. A respective cost value can be determined for each candidate reference block pair of the plurality of candidate reference block pairs. The reference block pair associated with the minimum cost value can be determined as the candidate reference block pair of the plurality of candidate reference block pairs.

[0031] In some embodiments, the plurality of candidate reference block pairs may include a first candidate reference block pair, and the first candidate reference block pair may include a first candidate reference block in a first reference picture and a first candidate reference block in a second reference picture. To determine the reference block pair, an initial predictor associated with the first reference picture of the current block may be determined based on the initial reference block in the first reference picture. An initial predictor associated with the second reference picture of the current block may be determined based on the initial reference block in the second reference picture. The first predictor associated with the first reference picture of the current block may be determined based on the initial predictor associated with the first reference picture of the current block, and the first predictor associated with the first reference picture of the current block may be associated with the first candidate reference block in the first reference picture. The first predictor associated with the second reference picture of the current block can be determined based on an initial predictor associated with the second reference picture of the current block, and the first predictor associated with the second reference picture of the current block can be associated with a first candidate reference block in the second reference picture. The first cost value can be determined based on a difference between the first predictor associated with the first reference picture of the current block and the first predictor associated with the second reference picture of the current block.

[0032] In the method, the initial predictor associated with the first reference picture of the current block may be indicated by one of a merge index, an advanced motion vector prediction (AMVP) predictor index, and an affine merge index.

[0033] To determine a first predictor associated with the first reference picture of the current block, a first component in a first direction of a gradient value of an initial predictor associated with the first reference picture of the current block may be determined. A second component in a second direction of a gradient value of an initial predictor associated with the first reference picture of the current block may be determined. The second direction may be perpendicular to the first direction. A first component in a first direction of a displacement between an initial reference block in the first reference picture and a first candidate reference block in the first reference picture may be determined. A second component in a second direction of a displacement between an initial reference block in the first reference picture and a first candidate reference block in the first reference picture may be determined. A first predictor associated with the first reference picture of the current block may be determined to be equal to the sum of (i) an initial predictor associated with the first reference picture of the current block, (ii) a product of a first component of a gradient value of the initial predictor and a first component of the displacement, and (iii) a product of a second component of the gradient value of the initial predictor and a second component of the displacement.

[0034] The plurality of candidate reference block pairs may include an Nth candidate reference block pair, and the Nth candidate reference block pair may include an Nth candidate reference block in a first reference picture and an Nth candidate reference block in a second reference picture. To determine the reference block pair, an Nth predictor associated with the Nth candidate reference block in a first reference picture of the current block may be determined based on an (N-1)th predictor associated with the (N-1)th candidate reference block in a first reference picture of the current block. An Nth predictor associated with the Nth candidate reference block in a second reference picture of the current block may be determined based on an (N-1)th predictor associated with the (N-1)th candidate reference block in a second reference picture of the current block. An Nth cost value may be determined based on a difference between the Nth predictor associated with the first reference picture of the current block and the Nth predictor associated with the second reference picture of the current block.

[0035] To determine the N-th predictor associated with the first reference picture of the current block, a first component in a first direction of a gradient value of the (N-1)-th predictor associated with the first reference picture of the current block may be determined. A second component in a second direction of a gradient value of the (N-1)-th predictor associated with the first reference picture of the current block may be determined. The second direction may be perpendicular to the first direction. A first component of a displacement may be determined. The first component of the displacement may be a difference in a first direction between the N-th candidate reference block in the first reference picture and the (N-1)-th candidate reference block in the first reference picture. A second component of the displacement may be a difference in a second direction between the N-th candidate reference block in the first reference picture and the (N-1)-th candidate reference block in the first reference picture. The Nth predictor may be determined based on the first reference picture of the current block as being equal to the sum of (i) the (N-1)th predictor associated with the first reference picture of the current block, (ii) the product of a first component of the gradient value of the (N-1)th predictor and a first component of the displacement, and (iii) the product of a second component of the gradient value of the (N-1)th predictor and a second component of the displacement.

[0036] In some embodiments, the multiple candidate reference block pairs may include N candidate reference block pairs based on one of: (i) N is equal to an upper limit value; and (ii) a displacement between the Nth candidate reference block in the first reference picture and the (N+1)th candidate reference block in the first reference picture is 0.

[0037] In some embodiments, the delta motion vector of each sub-block in the current block may be less than or equal to a threshold.

[0038] The prediction information may include a flag indicating whether the current block is predicted based on affine bilateral matching using an affine model.

[0039] In the method, a list of candidate affine model motion vectors associated with the current block can be determined. The candidate motion vector list can include control point motion vectors associated with the reference block pair.

[0040] According to another aspect of the present disclosure, an apparatus is provided, the apparatus including a processing circuit, the processing circuit being configured to perform any of the methods for video encoding / decoding.

[0041] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to perform any of the methods for video encoding / decoding.

[0042] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Brief description of the drawings]

[0043] [Figure 1A] FIG. 2 is a schematic diagram of an example subset of intra-prediction modes. [Figure 1B] FIG. 2 is a diagram of an example intra-prediction direction. [Diagram 2] FIG. 2 is a schematic diagram of a current block and its surrounding spatial merge candidates in one example. [Diagram 3] FIG. 3 is a schematic diagram of a simplified block diagram of a communication system (300) according to one embodiment. [Figure 4] FIG. 4 is a schematic diagram of a simplified block diagram of a communication system (400) according to one embodiment. [Diagram 5] FIG. 2 is a schematic diagram of a simplified block diagram of a decoder according to one embodiment. [Figure 6] FIG. 2 is a schematic diagram of a simplified block diagram of an encoder according to one embodiment. [Figure 7] FIG. 4 is a block diagram of an encoder according to another embodiment. [Figure 8] FIG. 4 is a block diagram of a decoder according to another embodiment. [Figure 9A] FIG. 13 is a schematic diagram of a four-parameter affine model according to another embodiment. [Figure 9B] FIG. 13 is a schematic diagram of a six-parameter affine model according to another embodiment. [Figure 10] FIG. 11 is a schematic diagram of an affine motion vector field associated with sub-blocks within a block according to another embodiment; [Figure 11] FIG. 13 is a schematic diagram of exemplary locations of spatial merging candidates according to another embodiment. [Figure 12] FIG. 11 is a schematic diagram of control point motion vector inheritance according to another embodiment. [Figure 13] FIG. 13 is a schematic diagram of candidate positions for constructing an affine merge mode according to another embodiment. [Figure 14] FIG. 13 is a schematic diagram of prediction refinement using optical flow (PROF) according to another embodiment. [Figure 15] FIG. 4 is a schematic diagram of an affine motion estimation process according to another embodiment. [Figure 16] FIG. 13 illustrates a flowchart of an affine motion estimation search according to another embodiment. [Figure 17] FIG. 1 is a schematic diagram of an extended coding unit (CU) region for bidirectional optical flow (BDOF) according to another embodiment. [Figure 18] FIG. 11 is a schematic diagram of a decoding-side motion vector refinement according to another embodiment; [Figure 19] FIG. 2 shows a flowchart outlining an example decoding process according to some embodiments of the present disclosure. [Figure 20] FIG. 2 illustrates a flowchart outlining an exemplary encoding process according to some embodiments of the present disclosure. [Figure 21] 1 is a schematic diagram of a computer system, according to one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0044] FIG. 3 illustrates an example block diagram of a communication system (300). The communication system (300) includes a plurality of terminal devices that can communicate with each other, for example, via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) perform unidirectional transmission of data. For example, the terminal device (310) may code video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to the other terminal device (320) via the network (350). The encoded video data may be transmitted in the form of one or more coded video bitstreams. The terminal device (320) may receive the coded video data from the network (350), decode the coded video data to reconstruct the video pictures, and display the video pictures according to the reconstructed video data. Unidirectional data transmission may be common, such as in media serving applications.

[0045] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of coded video data, for example during a video conference. In the case of bidirectional transmission of data, in one example, each of the terminal devices (330) and (340) may code video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) over the network (350). Each of the terminal devices (330) and (340) may also receive coded video data transmitted by the other of the terminal devices (330) and (340), decode the coded video data to reconstruct the video pictures, and display the video pictures on an accessible display device according to the reconstructed video data.

[0046] In the example of FIG. 3, terminal devices (310), (320), (330), and (340) are shown as a server, a personal computer, and a smartphone, respectively, although the principles of the present disclosure may not be so limited. The embodiments of the present disclosure apply with laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network (350) represents any number of networks that convey decoded video data between terminal devices (310), (320), (330), and (340), including, for example, wired (wired) and / or wireless communication networks. Communication network (350) may exchange data over circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of network (350) may not be important to the operation of the present disclosure unless otherwise described herein below.

[0047] 4 shows a video encoder and a video decoder in a streaming environment as an example of an application for the disclosed subject matter. The disclosed subject matter can be equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, streaming services, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0048] The streaming system may include a capture subsystem (413) that may include a video source (401), such as a digital camera, that creates a stream of uncompressed video pictures (402). In one example, the stream of video pictures (402) includes samples taken by a digital camera. The stream of video pictures (402) is shown in bold to emphasize its large amount of data compared to the encoded video data (404) (or coded video bitstream) that may be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (404) (or encoded video bitstream) is shown in thin to emphasize its small amount of data compared to the stream of video pictures (402) that may be stored in a streaming server (405) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include a video decoder (410), for example within an electronic device (430). The video decoder (410) decodes an input copy (407) of the encoded video data and creates an output stream (411) of video pictures that can be rendered on a display (412) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., a video bitstream) can be encoded according to a particular video coding / compression standard. Examples of such standards include ITU-T Recommendation H.265.In one example, a developing video coding standard is informally known as Versatile Video Coding (VVC), and the disclosed subject matter may be used in conjunction with VVC.

[0049] It should be noted that the electronic devices (420) and (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).

[0050] 5 shows an example block diagram of a video decoder (510). The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). The video decoder (510) can be used in place of the video decoder (410) in the example of FIG. 4.

[0051] The receiver (531) may receive one or more coded video sequences that are decoded by the video decoder (510). In one embodiment, one coded video sequence is received at a time, and the decoding of each coded video sequence is independent of the decoding of the other coded video sequences. The coded video sequences are received from a channel (501), which may be a hardware / software link to a storage device that stores the encoded video data. The receiver (531) may receive the encoded video data along with other data, such as coded audio data and / or auxiliary data streams, which may be forwarded to their respective using entities (not shown). The receiver (531) may separate the coded video sequences from the other data. To combat network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter "parser (520)"). In certain applications, the buffer memory (515) is part of the video decoder (510). In other applications, the buffer memory (515) can be external to the video decoder (510) (not shown). In still other applications, there can be buffer memories (not shown) external to the video decoder (510), e.g., to combat network jitter, plus other buffer memories (515) internal to the video decoder (510), e.g., to handle playout timing. When the receiver (531) is receiving data from a storage / forwarding device of sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory (515) may not be needed or can be small. For use over a best-effort packet network such as the Internet, the buffer memory (515) may be needed and can be relatively large, advantageously adaptively sized, and implemented at least in part within an operating system or similar element (not shown) external to the video decoder (510).

[0052] The video decoder (510) may include a parser (520) that reconstructs symbols (521) from the coded video sequence. These categories of symbols include information used to manage the operation of the video decoder (510) and potentially information for controlling a rendering device, such as a rendering device (512) (e.g., a display screen) that is not an integral part of the electronic device (530) but may be coupled to the electronic device (530), as shown in FIG. 5. The control information for the rendering device(s) may be in the form of a supplemental enhancement information (SEI) message or a video usability information (VUI) parameter set fragment (not shown). The parser (520) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles including variable length coding with or without context dependency, Huffman coding, arithmetic coding, etc. The parser (520) may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroups may include Group of Pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (520) may also extract information from the coded video sequence, such as transform coefficients, quantization parameter values, motion vectors, etc.

[0053] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).

[0054] The reconstruction of the symbols (521) can involve several different units, depending on the type of coded video picture or part thereof (inter-picture and intra-picture, inter-block and intra-block, etc.), as well as other factors. Which units are involved and how can be controlled by subgroup control information parsed by the parser (520) from the coded video sequence. The flow of such subgroup control information between the parser (520) and the following units is not shown for clarity.

[0055] Beyond the functional blocks already mentioned, the video decoder (510) may be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate:

[0056] The first unit is a scalar / inverse transform unit (551), which receives quantized transform coefficients as well as control information from the parser (520) including which transform to use, block size, quantization coefficients, quantization scaling matrices, etc. as symbol(s) (521). The scalar / inverse transform unit (551) can output blocks containing sample values ​​that can be input to an aggregator (555).

[0057] In some cases, the output samples of the scaler / inverse transform unit (551) may relate to intra-coded blocks. Intra-coded blocks are blocks that do not use prediction information from a previously reconstructed picture, but may use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates a block of the same size and shape as the block being reconstructed using surrounding already reconstructed information fetched from a current picture buffer (558). The current picture buffer (558) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. The aggregator (555) adds, possibly on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).

[0058] In other cases, the output samples of the scalar / inverse transform unit (551) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion compensated prediction unit (553) may access the reference picture memory (557) to fetch samples used for prediction. After motion compensating the fetched samples according to the symbols (521) related to the block, these samples may be added by the aggregator (555) to the output of the scalar / inverse transform unit (551) (in this case referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (557) from which the motion compensated prediction unit (553) fetches the prediction samples may be controlled by motion vectors available to the motion compensated prediction unit (553), for example, in the form of symbols (521) that may have X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​fetched from the reference picture memory (557) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, and the like.

[0059] The output samples of the aggregator (555) may be subjected to various loop filtering techniques in the loop filter unit (556). Video compression techniques may include in-loop filter techniques controlled by parameters contained in the coded video sequence (also called the coded video bitstream) and available to the loop filter unit (556) as symbols (521) from the parser (520). Video compression may also be responsive to meta-information obtained during decoding of previous portions (in decoding order) of the coded picture or coded video sequence, and to previously reconstructed loop filtered sample values.

[0060] The output of the loop filter unit (556) can be a sample stream that can be output to a rendering device (512) and can also be stored in a reference picture memory (557) for use in future inter-picture prediction.

[0061] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a new current picture buffer can be reallocated before beginning reconstruction of the next coded picture.

[0062] The video decoder (510) may perform decoding operations according to a given video compression technique or standard, such as ITU-T Rec. H.265. The decoded video sequence may conform to the syntax specified by the video compression technique or standard being used, meaning that the coded video sequence adheres to both the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. Specifically, the profile may select certain tools from among all tools available in the video compression technique or standard as tools that are only available to them under that profile. Also, compliance may require that the complexity of the coded video sequence be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may be further limited in some cases by the specification of a hypothetical reference decoder (HRD) and metadata for HRD buffer management signaled in the coded video sequence.

[0063] In one embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence(s). The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0064] 6 shows an example block diagram of a video encoder (603). The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used in place of the video encoder (403) of the example of FIG.

[0065] The video encoder (603) can receive video samples from a video source (601) (which is not part of the electronic device (620) in the example of FIG. 6) that can capture the video image(s) to be coded by the video encoder (603). In other examples, the video source (601) is part of the electronic device (620).

[0066] The video source (601) may provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream that may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (601) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a number of separate pictures that give motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0067] According to one embodiment, the video encoder (603) can code and compress pictures of a source video sequence into a coded video sequence (643) in real-time or under any other time constraint as needed. Enforcing an appropriate coding rate is one function of the controller (650). In some embodiments, the controller (650) controls and is operatively coupled to other functional units described below. This coupling is not depicted for clarity. Parameters set by the controller (650) can include rate control related parameters (picture skip, quantizer, lambda value for rate distortion optimization techniques, ...), picture size, Group of Pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other appropriate functions associated with the video encoder (603) optimized for a particular system design.

[0068] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As an oversimplified explanation, in one example, the coding loop can include a source coder (630) (responsible for creating symbols, such as a symbol stream, based on an input picture to be coded and reference picture(s)) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to that which the (remote) decoder also creates. The reconstructed sample stream (sample data) is input to a reference picture memory (634). Since decoding of the symbol stream results in bit-exact results regardless of the location of the decoder (local or remote), the contents of the reference picture memory (634) are also bit-exact between the local and remote encoders. In other words, the predictive part of the encoder "sees" exactly the same sample values ​​as the decoder would "see" when using prediction during decoding as reference picture samples. This basic principle of reference picture synchrony (and the resulting drift when synchrony cannot be maintained, for example due to channel errors) is also used in some related techniques.

[0069] The operation of the "local" decoder (633) may be the same as that of a "remote" decoder, such as the video decoder (510) already described in detail above in connection with Figure 5. However, with brief reference also to Figure 5, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder (645) and parser (520) may be lossless, the entropy decoding portion of the video decoder (510), including the buffer memory (515) and parser (520), may not be fully implemented in the local decoder (633).

[0070] In one embodiment, the decoder techniques, except for parsing / entropy decoding, present in the decoder are present in the corresponding encoder in the same or substantially the same functional form. Thus, the disclosed subject matter focuses on the operation of the decoder. The description of the encoder techniques may be omitted, since they are the inverse of the decoder techniques described generically. In certain areas, more detailed descriptions are provided below.

[0071] In operation, in some examples, the source coder (630) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as “reference pictures.” In this manner, the coding engine (632) codes differences between pixel blocks of the input picture and pixel blocks of reference picture(s) that may be selected as predictive reference(s) for the input picture.

[0072] The local video decoder (633) may decode the coded video data of pictures that may be designated as reference pictures based on the symbols created by the source coder (630). The operation of the coding engine (632) may advantageously be a lossy process. When the coded video data may be decoded in a video decoder (not shown in FIG. 6), the reconstructed video sequence may usually be a copy of the source video sequence with some errors. The local video decoder (633) may replicate the decoding process that may be performed by the video decoder on the reference pictures and store the reconstructed reference pictures in the reference picture memory (634). In this way, the video encoder (603) may locally store copies of reconstructed reference pictures that have common content with the reconstructed reference pictures obtained by the far-end video decoder (without transmission errors).

[0073] The predictor (635) may perform a predictive search for the coding engine (632). That is, for a new picture to be coded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) that can serve as suitable predictive references for the new pixels, or for specific metadata such as reference picture motion vectors, block shapes, etc. The predictor (635) may operate on sample blocks on a pixel block by pixel block basis to find a suitable predictive reference. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have predictive references drawn from multiple reference pictures stored in the reference picture memory (634).

[0074] The controller (650) may manage the coding operations of the source coder (630), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0075] The output of all the aforementioned functional units may undergo entropy coding in an entropy coder (645), which converts the symbols produced by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, or arithmetic coding.

[0076] The transmitter (640) may buffer the coded video sequence(s) created by the entropy coder (645) in preparation for transmission over a communication channel (660), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (640) may merge the coded video data from the video encoder (603) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0077] The controller (650) may manage the operation of the video encoder (603). During coding, the controller (650) may assign a particular coded picture type to each coded picture, which may affect the coding technique that may be applied to the respective picture. For example, pictures may often be assigned as one of the following picture types:

[0078] An intra picture (I-picture) may be one that can be coded and decoded without using any other picture in a sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh ("IDR") pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.

[0079] A predictive picture (P picture) may be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict the sample values ​​of each block.

[0080] Bidirectionally predicted pictures (B-pictures) may be coded and decoded using intra- or inter-prediction, which uses up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predicted pictures may use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0081] A source picture may generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and coded block by block. A block may be predictively coded with reference to other (already coded) blocks determined by a coding assignment applied to the block's respective picture. For example, a block of an I picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial or intra prediction). A pixel block of a P picture may be predictively coded via spatial prediction with reference to one previously coded reference picture or via temporal prediction. A block of a B picture may be predictively coded via spatial prediction with reference to one or two previously coded reference pictures or via temporal prediction.

[0082] The video encoder (603) may perform coding operations in accordance with a given video coding technique or standard, such as ITU-T Rec. H.265. In its operations, the video encoder (603) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.

[0083] In one embodiment, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0084] A video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) uses spatial correlation within a given picture, while inter-picture prediction uses correlation (temporal or other) between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. When a block in the current picture is similar to a reference block in a reference picture that was previously coded and is still buffered in the video, the block in the current picture may be coded by a vector called a motion vector. A motion vector points to a reference block in a reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0085] In some embodiments, bi-prediction techniques may be used in inter-picture prediction. According to bi-prediction techniques, two reference pictures, such as a first reference picture and a second reference picture, are used, both of which are before the decoding order of the current picture in the video (but the display order may be past and future, respectively). A block in the current picture may be coded by a first motion vector pointing to a first reference block in the first reference picture and by a second motion vector pointing to a second reference block in the second reference picture. A block may be predicted by a combination of the first reference block and the second reference block.

[0086] Furthermore, to improve coding efficiency, merge mode techniques can be used in inter-picture prediction.

[0087] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed on a block-by-block basis. For example, according to the HEVC standard, a picture in a sequence of video pictures is divided into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, 16×16 pixels, etc. In general, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU can be recursively quad-tree partitioned into one or more coding units (CUs). For example, a CTU of 64×64 pixels can be partitioned into one CU of 64×64 pixels, or into four CUs of 32×32 pixels, or into 16 CUs of 16×16 pixels. In one example, each CU is analyzed to determine a prediction type of the CU, such as an inter prediction type or an intra prediction type. A CU is divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. In general, each PU includes one luma prediction block (PB) and two chroma PBs. In one embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of values ​​(e.g., luma values) for pixels of 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0088] 7 shows an example diagram of a video encoder (703). The video encoder (703) is configured to receive a processed block of sample values ​​(e.g., a predictive block) in a current video picture in a sequence of video pictures and encode the processed block into a coded picture that is part of a coded video sequence. In one example, the video encoder (703) is used in place of the video encoder (403) in the example of FIG. 4.

[0089] In an HEVC example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as a predictive block of 8×8 samples. The video encoder (703) determines whether the processing block is best coded using intra-mode, inter-mode, or bi-predictive mode, for example, using rate-distortion optimization. When the processing block is to be coded in intra-mode, the video encoder (703) may encode the processing block into a coded picture using intra-prediction techniques, and when the processing block is to be coded in inter-mode or bi-predictive mode, the video encoder (703) may encode the processing block into a coded picture using inter-prediction techniques or bi-prediction techniques, respectively. In certain video coding techniques, the merge mode may be an inter-picture prediction sub-mode in which a motion vector is derived from one or more motion vector predictors without the aid of coded motion vector components outside the predictors. In certain other video coding techniques, there may be motion vector components applicable to the current block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the mode of the processing blocks.

[0090] In the example of FIG. 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725), coupled together as shown in FIG.

[0091] The inter-encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block to one or more reference blocks in a reference picture (e.g., blocks in previous and subsequent pictures), generate inter-prediction information (e.g., a description of redundant information due to inter-encoding techniques, motion vectors, merge mode information), and calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture that is decoded based on the encoded video information.

[0092] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), possibly compare the block to already coded blocks in the same picture, generate transformed and quantized coefficients, and possibly also generate intra prediction information (e.g., intra prediction direction information according to one or more intra encoding techniques). In one example, the intra encoder (722) also calculates intra prediction results (e.g., a prediction block) based on the intra prediction information and reference blocks in the same picture.

[0093] The generic controller (721) is configured to determine generic control data and control other components of the video encoder (703) based on the generic control data. In one example, the generic controller (721) determines the mode of the block and provides a control signal to the switch (726) based on the mode. For example, if the mode is an intra mode, the generic controller (721) controls the switch (726) to select the intra mode result used by the residual calculator (723) and controls the entropy encoder (725) to select the intra prediction information and include the intra prediction information in the bitstream; if the mode is an inter mode, the generic controller (721) controls the switch (726) to select the inter prediction result used by the residual calculator (723) and controls the entropy encoder (725) to select the inter prediction information and include the inter prediction information in the bitstream.

[0094] The residual calculator (723) is configured to calculate a difference (residual data) between the received block and a prediction result selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) is configured to operate based on the residual data to encode the residual data to generate transform coefficients. In one example, the residual encoder (724) is configured to transform the residual data from the spatial domain to the frequency domain to generate transform coefficients. The transform coefficients then undergo a quantization process to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be used by the intra-encoder (722) and the inter-encoder (730) as appropriate. For example, the inter-encoder (730) can generate decoded blocks based on the decoded residual data and the inter-prediction information, and the intra-encoder (722) can generate decoded blocks based on the decoded residual data and the intra-prediction information. The decoded blocks are appropriately processed to generate decoded pictures, which can be buffered in a memory circuit (not shown) and used as reference pictures in some examples.

[0095] The entropy encoder (725) is configured to format a bitstream to include the encoded block. The entropy encoder (725) is configured to include various information in the bitstream according to an appropriate standard, such as the HEVC standard. In one example, the entropy encoder (725) is configured to include in the bitstream general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information. It should be noted that, according to the disclosed subject matter, there is no residual information when coding a block in a merged sub-mode of either the inter-mode or the bi-prediction mode.

[0096] 8 shows an example diagram of a video decoder (810). The video decoder (810) is configured to receive coded pictures that are part of a coded video sequence and decode the coded pictures to generate reconstructed pictures. In one example, the video decoder (810) is used in place of the video decoder (410) in the example of FIG. 4.

[0097] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872), coupled to each other as shown in FIG. 8.

[0098] The entropy decoder (871) may be configured to reconstruct from the coded picture certain symbols that represent syntax elements that make up the coded picture. Such symbols may include, for example, the mode in which the block is coded (e.g., intra mode, inter mode, bi-predictive mode, etc., where inter mode and bi-predictive mode are in merged submode or other submode), as well as prediction information (e.g., intra prediction information or inter prediction information, etc.) that may identify certain samples or metadata used for prediction by the intra decoder (872) or the inter decoder (880), respectively. The symbols may also include, for example, residual information in the form of quantized transform coefficients, etc. In one example, when the prediction mode is an inter mode or bi-predictive mode, the inter prediction information is provided to the inter decoder (880); when the prediction type is an intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information may undergo inverse quantization and is provided to the residual decoder (873).

[0099] The inter decoder (880) is configured to receive the inter prediction information and to generate inter prediction results based on the inter prediction information.

[0100] The intra decoder (872) is configured to receive the intra prediction information and to generate a prediction result based on the intra prediction information.

[0101] The residual decoder (873) is configured to perform inverse quantization to extract inverse quantized transform coefficients, and to process the inverse quantized transform coefficients to transform the residual information from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (to include quantizer parameters (QP)), which may be provided by the entropy decoder (871) (data path not shown as this may be only low volume control information).

[0102] The reconstruction module (874) is configured to combine, in the spatial domain, the residual information output by the residual decoder (873) and the prediction result (output by the inter prediction module or the intra prediction module, as the case may be) to form a reconstructed block, which may be part of a reconstructed picture, which may be part of a reconstructed video. It should be noted that other suitable operations, such as a deblocking operation, may be performed to improve visual quality.

[0103] It should be noted that the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using any suitable technology. In one embodiment, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using one or more integrated circuits. In other embodiments, the video encoders (403), (603), and (603) and the video decoders (410), (510), and (810) may be implemented using one or more processors executing software instructions.

[0104] This disclosure includes embodiments related to an affine coding mode, in which affine motion parameters can be derived based on bilateral matching instead of signaling.

[0105] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). In 2015, the two standardization bodies jointly formed the Joint Video Research Team (JVET) to explore the possibility of developing the next video coding standard beyond HEVC. In October 2017, the two standardization bodies announced a Joint Call for Proposals on Video Compression with Capability beyond HEVC (CfP). By February 15, 2018, 22 CfP responses had been submitted for standard dynamic range (SDR), 12 for high dynamic range (HDR), and 12 for the 360 ​​video category. In April 2018, all received CfP responses were evaluated at the 122 MPEG / 10th JVET meeting. As a result of the meeting, JVET formally launched the standardization process for next-generation video coding beyond HEVC. This new standard was named Versatile Video Coding (VVC) and JVET was renamed the Joint Video Experts Team. In 2020, ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the VVC video coding standard (version 1).

[0106] In inter prediction, motion parameters are needed for each inter predicted coding unit (CU), e.g., to code the VVC features used for inter predicted sample generation. The motion parameters may include motion vectors, reference picture indexes, reference picture list usage indexes, and / or additional information. The motion parameters may be signaled in an explicit or implicit manner. If a CU is coded in skip mode, it may be associated with one PU, and significant residual coefficients, coded motion vector deltas, and / or reference picture indexes may not be needed. If a CU is coded in merge mode, the motion parameters of the CU may be obtained from neighboring CUs. The neighboring CUs may include spatial and temporal candidates, as well as additional schedules (or additional candidates) as introduced in VVC. The merge mode may be applied to any inter predicted CU, not just for skip mode. An alternative to the merge mode is explicit transmission of motion parameters, where the motion vectors, the corresponding reference picture indexes of each reference picture list, the reference picture list usage flag, and / or other necessary information may be explicitly signaled for each CU.

[0107] In VVC, the VVC Test Model (VTM) reference software can include several new and improved inter-predictive coding tools, which can include one or more of the following: (1) Enhanced Merge Prediction (2) Merged Motion Vector Differential (MMVD) (3) AMVP mode with symmetric MVD signaling (4) Affine motion compensation prediction (5) Sub-block based temporal motion vector prediction (SbTMVP) (6) Adaptive Motion Vector Resolution (AMVR) (7) Motion field storage: 1 / 16 luma sample MV storage and 8x8 motion field compression (8) Bi-prediction with CU-level weights (BCW) (9) Bidirectional Optical Flow (BDOF) (10) Decoder-side Motion Vector Refinement (DMVR) (11) Combined Inter- and Intra-Prediction (CIIP) (12) Geometric Partition Mode (GPM)

[0108] In HEVC, a translational motion model is applied to motion compensated prediction (MCP). In the real world, many types of motion can exist, such as zoom in / out, rotation, perspective motion, and other irregular motions. Block-based affine transform motion compensated prediction can be applied to VTM, etc. Figure 9A shows an affine motion field of a block (902) described by motion information of two control points (4 parameters). Figure 9B shows an affine motion field of a block (904) described by a three control point motion vector (6 parameters).

[0109] As shown in FIG. 9A, in a four-parameter affine motion model, the motion vector at a sample location (x,y) in a block (902) can be derived in equation (1) as follows:

number

number

[0110] As shown in FIG. 9B, in a six-parameter affine motion model, the motion vector at a sample location (x,y) within a block (904) can be derived in equation (3) as follows:

number

number

[0111] As shown in FIG. 10, to simplify the motion compensation prediction, a block-based affine transformation prediction may be applied. To derive a motion vector for each 4×4 luma sub-block, the motion vector of a central sample (e.g., (1002)) of each sub-block (e.g., (1004)) in a current block (1000) may be calculated according to Equations (1)-(4) and rounded to 1 / 16 fractional precision. A motion compensation interpolation filter may be applied to generate a prediction for each sub-block using the derived motion vector. The sub-block size of the chroma components may also be set as 4×4. The MV of a 4×4 chroma sub-block may be calculated as the average of the MVs of four corresponding 4×4 luma sub-blocks.

[0112] In affine merge prediction, the affine merge (AF_MERGE) mode can be applied to CUs whose width and height are both 8 or more. The CPMV of the current CU can be generated based on the motion information of spatially neighboring CUs. Up to five CPMVP candidates can be applied to affine merge prediction, and an index can be signaled to indicate which of the five CPMVP candidates can be used for the current CU. In affine merge prediction, three types of CPMV candidates can be used to form the affine merge candidate list: (1) inherited affine merge candidates extrapolated from the CPMV of neighboring CUs, (2) affine merge candidates constructed with CPMVP derived using the translation MV of neighboring CUs, and (3) zero MV.

[0113] In VTM3, up to two inherited affine candidates can be applied. The two inherited affine candidates can be derived from the affine motion models of the neighboring blocks. For example, one inherited affine candidate can be derived from the left neighboring CU, and the other inherited affine candidate can be derived from the upper neighboring CU. An exemplary candidate block can be shown in FIG. 11. As shown in FIG. 11, for the left predictor (or the left inherited affine candidate), the scanning order can be A0→A1, and for the upper predictor (or the upper inherited affine candidate), the scanning order can be B0→B1→B2. Therefore, only the first available inherited candidate from each side can be selected. No pruning check may be performed between the two inherited candidates. Once the neighboring affine CU is identified, the control point motion vector of the neighboring affine CU can be used to derive the CPMVP candidate in the affine merge list of the current CU. As shown in FIG. 12, when block A adjacent to the lower left of the current block (1204) is coded in affine mode, motion vectors v2, v3, and v4 of the upper left corner, upper right corner, and lower left corner of the CU (1202) including block A can be achieved. If block A is coded with a four-parameter affine model, two CPMVs of the current CU (1204) can be calculated according to v2 and v3 of the CU (1202). If block A is coded with a six-parameter affine model, three CPMVs of the current CU (1204) can be calculated according to v2, v3, and v4 of the CU (1202).

[0114] The constructed affine candidate of the current block may be a candidate constructed by combining the neighboring translational motion information of each control point of the current block. The motion information of the control points may be derived from specific spatial and temporal neighborhoods, which may be shown in FIG. 13. As shown in FIG. 13, the CPMV k(k=1,2,3,4) represents the kth control point of the current block (1302). For CPMV1, the B2→B3→A2 block can be checked and the MV of the first available block can be used. For CPMV2, the B1→B0 block can be checked. For CPMV3, the A1→A0 block can be checked. If CPM4 is not available, TMVP can be used as CPMV4.

[0115] After the MVs of the four control points are achieved, affine merge candidates for the current block (1302) can be constructed based on the motion information of the four control points. For example, affine merge candidates can be constructed based on the combinations of the MVs of the four control points in the following order: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}.

[0116] A combination of three CPMVs can construct a six-parameter affine merge candidate, and a combination of two CPMVs can construct a four-parameter affine merge candidate. To avoid motion scaling processing, if the reference indexes of the control points are different, the associated combination of control point MVs can be discarded.

[0117] After the inherited and constructed affine merge candidates have been checked, if the list is not already full, a zero MV may be inserted at the end of the list.

[0118] In affine AMVP prediction, affine AMVP mode can be applied to CUs with both width and height of 16 or more. A CU-level affine flag can be signaled in the bitstream to indicate whether affine AMVP mode is used, and then another flag can be signaled to indicate whether 4-parameter affine or 6-parameter affine is applied. In affine AMVP prediction, the difference between the CPMV of the current CU and the predictor of the CPMVP of the current CU can be signaled in the bitstream. The size of the affine AMVP candidate list can be 2, and the affine AMVP candidate list can be generated by using the four types of CPMV candidates in the following order: (1) Hereditary affine AMVP candidates extrapolated from the CPMVs of nearby CUs, (2) A constructed affine AMVP candidate with CPMVP derived using the translational MV of nearby CUs; (3) translational MVs from nearby CUs, and (4) Zero MV.

[0119] The check order of the inherited affine AMVP candidates may be the same as the check order of the inherited affine merge candidates. To determine the AVMP candidates, only affine CUs with the same reference picture as the current block may be considered. When an inherited affine motion predictor is inserted into the candidate list, the pruning process may not be applied.

[0120] The constructed AMVP candidate may be derived from the specified spatial neighborhood. As shown in FIG. 13, the same check order as in the affine merge candidate construction may be applied. In addition, the reference picture index of the neighboring block may also be checked. The first block in the check order may be inter-coded and have the same reference picture as the current CU (1302). If the current CU (1302) is coded in a four-parameter affine mode and both mv0 and mv1 are available, one constructed AMVP candidate may be determined. The constructed AMVP candidate may be further added to the affine AMVP list. If the current CU (1302) is coded in a six-parameter affine mode and all three CPMVs are available, the constructed AMVP candidate may be added as one candidate in the affine AMVP list. Otherwise, the constructed AMVP candidate may be set as unavailable.

[0121] After the inherited affine AMVP candidates and constructed AMVP candidates are checked, if there are still less than two candidates in the affine AMVP list, mv0, mv1, and mv2 can be added in order. mv0, mv1, and mv2 can serve as translation MVs to predict all control point MVs of the current CU (e.g., (1302)), if available. Finally, if the affine AMVP is not yet full, a zero MV can be used to fill the affine AMVP list.

[0122] Sub-block-based affine motion compensation can save memory access bandwidth and reduce computational complexity compared to pixel-based motion compensation, at the expense of a penalty in prediction accuracy. To achieve finer granularity of motion compensation, prediction refinement by optical flow (PROF) can be used to refine sub-block-based affine motion compensation prediction without increasing memory access bandwidth for motion compensation. In VVC, after sub-block-based affine motion compensation is performed, the luma prediction samples can be refined by adding the difference derived by the optical flow equation. PROF can be described in the following four steps:

[0123] Step (1): Sub-block-based affine motion compensation can be performed to generate a sub-block prediction I(i,j).

[0124] Step (2): Spatial gradient of subblock prediction g x (i,j) and g y (i,j) can be calculated at each sample position using a 3-tap filter [-1,0,1]. The gradient calculation can be the same as the gradient calculation in BDOF. For example, the spatial gradient g x (i,j) and g y (i,j) can be calculated based on equations (5) and (6), respectively. g x (i,j)=(I(i+1,j)>>shift1)-(I(i-1,j)>>shift1) Equation (5) g y (i,j)=(I(i,j+1)>>shift1)-(I(i,j-1)>>shift1) Equation (6) As shown in equations (5) and (6), shift1 can be used to control the accuracy of the gradient. Sub-block (e.g., 4×4) predictions can be extended by one sample on each side for gradient calculation. To avoid additional memory bandwidth and additional interpolation calculations, the extended samples on the extension boundary can be copied from the nearest integer pixel location in the reference picture.

[0125] Step (3): The luma prediction accuracy can be calculated by the optical flow formula as shown in Equation (7). ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) Equation (7) where Δv(i,j) is the sum of the sample MV (denoted by v(i,j)) calculated for sample position (i,j) and the sub-block MV (v SB MV (represented by v(i,j)) that corresponds to the reference pixel (1406). The subblock (1402) may be a difference between the sample MV and the subblock MV. FIG. 14 shows an example diagram of the difference between the sample MV and the subblock MV. As shown in FIG. 14, the subblock (1402) may be included in the current block (1400), and the sample (1404) may be included in the subblock (1402). The sample (1404) may include a sample motion vector v(i,j) that corresponds to the reference pixel (1406). The subblock (1402) may include a subblock motion vector v(i,j) that corresponds to the reference pixel (1406). SB The sub-block motion vector v SB Based on, the sample (1404) may correspond to the reference pixel (1408). The difference between the sample MV and the sub-block MV, represented by Δv(i,j), may be indicated by the difference between the reference pixel (1406) and the reference pixel (1408). Δv(i,j) may be quantized in units of 1 / 32 luma sample precision.

[0126] Since the affine model parameters and sample positions relative to the sub-block center may not change from sub-block to another sub-block, Δv(i,j) may be calculated for a first sub-block (e.g., (1402)) and reused for other sub-blocks (e.g., (1410)) in the same CU (e.g., (1400)). Let dx(i,j) be the horizontal offset and dy(i,j) be the distance from the sample position (i,j) to the center of the sub-block (x SB ,y SB ), Δv(x,y) can be derived by the following equations (8) and (9):

number

[0127] To maintain accuracy, the center of the subblock (x SB ,y SB ) is ((W SB -1) / 2,(H SB -1) / 2), where W SB and H SB are the width and height of the sub-block, respectively.

[0128] Once Δv(x,y) is obtained, the parameters of the affine model can be obtained. For example, in the case of a four-parameter affine model, the parameters of the affine model can be shown in Equation (10).

number

number

[0129] Step (4): Finally, the luma prediction accuracy ΔI(i,j) may be added to the sub-block prediction I(i,j). The final prediction I′ may be generated as shown in equation (12). I'(i,j)=I(i,j)+ΔI(i,j) Equation (12)

[0130] PROF may not be applied to an affine coded CU in two cases: (1) when all control points MV are the same, indicating that the CU has only translational motion, and (2) when the affine motion parameters are larger than the specified limit since sub-block-based affine MC has been downgraded to CU-based MC to avoid large memory access bandwidth requirements.

[0131] Affine motion estimation (ME), such as the VVC reference software VTM, can be operated for both uni-prediction and bi-prediction. Uni-prediction can be performed for either reference list L0 or reference list L1, and bi-prediction can be performed for both reference list L0 and reference list L1.

[0132] FIG. 15 shows a schematic diagram of affine ME (1500). As shown in FIG. 15, in affine ME (1500), affine uni-prediction (S1502) can be performed on reference list L0 to obtain a prediction P0 of the current block based on an initial reference block in reference list L0. Affine uni-prediction (S1504) can also be performed on reference list L1 to obtain a prediction P1 of the current block based on an initial reference block in reference list L1. In (S1506), affine bi-prediction can be performed. Affine bi-prediction (S1506) can start with an initial prediction residual (2I-P0)-P1, where I can be the initial value of the current block. Affine bi-prediction (S1506) can search candidates in reference list L1 around the initial reference block in reference list L1 to find the best (or selected) reference block with the smallest prediction residual (2I-P0)-Px, where Px is the prediction of the current block based on the selected reference block.

[0133] Using the reference picture, for the current coding block, the affine ME process can first choose a set of control point motion vectors (CPMVs) as a base. An iterative method can be used to generate a prediction output of the current affine model corresponding to the set of CPMVs, calculate the gradient of the prediction sample, and then solve a linear equation to determine the delta CPMVs and optimize the affine prediction. The iteration can be stopped when all delta CPMVs are 0 or the maximum number of iterations is reached. The CPMV obtained from the iteration can be the final CPMV of the reference picture.

[0134] After the best affine CPVM of both reference lists L0 and L1 is determined for affine uni-prediction, an affine bi-prediction search can be performed using the best uni-prediction CPMV and the reference list on one side to search for the best CPMV of the other reference list to optimize the affine bi-prediction output. The affine bi-prediction search can be performed iteratively on the two reference lists to obtain the optimal result.

[0135] 16 shows an example affine ME process (1600) in which a final CPMV associated with a reference picture may be calculated. The affine ME process (1600) may begin at (S1602). At (S1602), a base CPMV of a current block may be determined. The base CPMV may be determined based on one of a merge index, an advanced motion vector prediction (AMVP) predictor index, an affine merge index, etc.

[0136] In (S1604), an initial affine prediction of the current block may be obtained based on the base CPMV. For example, according to the base CPMV, a four-parameter affine motion model or a six-parameter affine motion model may be applied to generate the initial affine prediction.

[0137] In (S1606), a gradient of the initial affine prediction may be obtained. For example, the gradient of the initial affine prediction may be obtained based on Equation (5) and Equation (6).

[0138] At (S1608), a delta CPMV may be determined. In some embodiments, the delta CPMV may be associated with a displacement between an initial affine prediction and a subsequent affine prediction, such as a first affine prediction. Based on a gradient of the initial affine prediction and the delta CPMV, a first affine prediction may be obtained. The first affine prediction may correspond to the first CPMV.

[0139] At (S1610), a decision may be made to check whether the delta CPMV is 0 or the number of iterations is greater than or equal to a threshold. If the delta CPMV is 0 or the number of iterations is greater than or equal to a threshold, at (S1612), a final (or selected) CPMV may be determined. The final (or selected) CPMV may be a first CPMV determined based on the gradient of the initial affine prediction and the delta CPMV.

[0140] Further referring to (S1610), if the delta CPMV is not 0 or the number of iterations is less than a threshold, a new iteration may be started. In the new iteration, an updated CPMV (e.g., the first CPMV) may be provided to (S1604) to generate an updated affine prediction. The affine ME process (1600) may then proceed to (S1606) where a gradient of the updated affine prediction may be calculated. The affine ME process (1600) may then proceed to (S1608) to continue the new iteration.

[0141] In the affine motion model, the four-parameter affine motion model can be further described by an equation including the rotation and zoom motions. For example, the four-parameter affine motion model can be rewritten in equation (13) as follows:

number

number

number

[0142] Bidirectional Optical Flow (BDOF) in VVC was previously called BIO in JEM. Compared to the JEM version, BDOF in VVC can be a simpler version that requires less computation, especially in terms of the number of multiplications and the size of the multipliers.

[0143] BDOF may be used to refine the bi-predictive signal of a CU at the 4×4 sub-block level. BDOF may be applied to a CU if the CU satisfies the following conditions: (1) The CU is coded using a “true” bi-prediction mode, i.e., one of the two reference pictures is before the current picture in display order and the other is after the current picture in display order; (2) The distances (e.g., POC differences) from the two reference pictures to the current picture are the same; (3) both reference pictures are short-term reference pictures; (4) the CU is not coded using affine mode or SbTMVP merge mode; (5) the CU has more than 64 luma samples; (6) both the CU height and the CU width are greater than or equal to 8 luma samples; (7) the BCW Weight Index shows equal weights; (8) No weighting position (WP) is available for the current CU; and (9) CIIP mode is not used for the current CU.

[0144] BDOF may be applied only to the luma component. As the name BDOF suggests, the BDOF mode may be based on the concept of optical flow, which assumes that object motion is smooth. For each 4 × 4 sub-block, the motion accuracy (v x ,v y ) may be calculated. The motion accuracy may then be used to adjust the bi-predictive sample values ​​in the 4×4 sub-block. BDOF may include the following steps:

[0145] First, the horizontal gradients of the two predicted signals from reference list L0 and reference list L1

number

number

number

[0146] Then, the autocorrelation and cross-correlation of the gradients S1, S2, S3, S5, and S6 can be calculated according to the following equations (18) to (22):

number

number

[0147] Next, the motion accuracy (v x ,v y ) can be derived using cross-correlation and auto-correlation terms using the following equations (26) and (27):

number

number

number

number

number

number

[0148] Finally, the BDOF samples of the CU can be calculated by adjusting the bi-predictive samples in equation (29) as follows: pred BDOF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+o offset )>>shift Eq.(29) The values ​​can be selected such that the multipliers in the BDOF process do not exceed 15 bits and the maximum bit width of the intermediate parameters in the BDOF process can be kept within 32 bits.

[0149] To derive the gradient value, we select some predicted samples I in list k (k=0, 1) outside the current CU boundary. (k) (i,j) needs to be generated. As shown in FIG. 17, BDOF in VVC can use one extended row / column (1702) around the boundary (1706) of the CU (1704). To control the computational complexity of generating out-of-bounds predicted samples, the predicted samples in the extended region (e.g., the unshaded region in FIG. 17) can be generated by directly taking reference samples of nearby integer positions (e.g., using floor() operation on the coordinates) without interpolation, and a normal 8-tap motion compensation interpolation filter can be used to generate predicted samples in the CU (e.g., the shaded region in FIG. 17). The extended sample values ​​can be used only in gradient calculation. In the remaining steps of the BDOF process, if samples and gradient values ​​outside the CU boundary are required, the samples and gradient values ​​can be padded (e.g., repeated) from the nearest neighbors of the samples and gradient values.

[0150] If the width and / or height of a CU is greater than 16 luma samples, the CU may be divided into sub-blocks with width and / or height equal to 16 luma samples, and the sub-block boundaries may be treated as CU boundaries in the BDOF process. The maximum unit size of the BDOF process may be limited to 16×16. For each sub-block, the BDOF process may be skipped. If the sum of absolute differences (SAD) between the initial L0 predicted sample and the initial L1 predicted sample is less than a threshold, the BDOF process may not be applied to the sub-block. The threshold may be set equal to (8*W*(H>>1), where W may indicate the width of the sub-block and H may indicate the height of the sub-block. To avoid further complexity of the SAD calculation, the SAD between the initial L0 predicted sample and the initial L1 predicted sample calculated in the DMVR process may be reused in the BBOF process.

[0151] If BCW is enabled for the current block, i.e., if the BCW weight index indicates unequal weights, bidirectional optical flow may be disabled. Similarly, if WP is enabled for the current block, i.e., if the luma weight flag (e.g., luma_weight_lx_flag) for either of the two reference pictures is 1, BDOF may also be disabled. If the CU is coded in symmetric MVD mode or CIIP mode, BDOF may also be disabled.

[0152] To improve the accuracy of the MV of the merge mode, a bilateral matching (BM)-based decoder-side motion vector refinement in VVC, etc., can be applied. In bi-predictive operation, a refined MV can be searched around the initial MV in the reference picture list L0 and the reference picture list L1. The BM method can calculate the distortion between two candidate blocks in the reference picture list L0 and the reference picture list L1.

[0153] Figure 18 shows an example schematic diagram of BM-based decoder-side motion vector refinement. As shown in Figure 18, a current picture (1802) may include a current block (1808). The current picture may include a reference picture list L0 (1804) and a reference picture list L1 (1806). The current block (1808) may include an initial reference block (1812) in reference picture list L0 (1804) with an initial motion vector MV0 and an initial reference block (1814) in reference picture list L1 (1806) with an initial motion vector MV1. A search process may be performed around an initial MV0 in reference picture list L0 (1804) and an initial MV1 in reference picture list L1 (1806). For example, a first candidate reference block (1810) may be identified in reference picture list L0 (1804), and a first candidate reference block (1816) may be identified in reference picture list L1 (1806). The SAD between the candidate reference blocks (e.g., (1810) and (1816)) based on each MV candidate (e.g., MV0' and MV1') around the initial MV (e.g., MV0 and MV1) may be calculated. The MV candidate with the lowest SAD becomes the refined MV and may be used to generate a bi-predictive signal for predicting the current block (1808).

[0154] The application of DMVR may be restricted and may only be applied to CUs coded based on mode and feature, such as VVC, as follows: (1) CU-level merge mode with bi-predictive MVs, (2) For a current picture, one reference picture is in the past and another reference picture is in the future; (3) The distances (e.g., POC differences) from the two reference pictures to the current picture are the same; (4) both reference pictures are short-term reference pictures; (5) The CU has more than 64 luma samples. (6) Both the CU height and the CU width are equal to or greater than 8 luma samples; (7) The BCW weight index indicates equal weights; (8) WP is not enabled for the current block, and (9) CIIP mode is not used for the current block.

[0155] The refined MVs derived by the DMVR process can be used to generate inter-prediction samples and can be used in temporal motion vector prediction for future picture coding, while the original MVs can be used in the deblocking process and can be used in spatial motion vector prediction for future CU coding.

[0156] In DVMR, the search points can surround the initial MV, and the MV offset can follow the MV difference mirroring rule. In other words, any point checked by DMVR, indicated by the candidate MV pair (MV0, MV1), can follow the MV difference mirroring rule shown in equations (30) and (31): MV0'=MV0+MV_offset Equation (30) MV1'=MV1-MV_offset formula (31) Here, MV_offset may represent a refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range may be two integer luma samples from the initial MV. The search may include an integer sample offset search stage and a fractional sample refinement stage.

[0157] For example, a full search of 25 points can be applied to the integer sample offset search. The SAD of the initial MV pair can be calculated first. If the SAD of the initial MV pair is less than a threshold, the integer sample stage of the DMVR can be terminated. Otherwise, the SAD of the remaining 24 points can be calculated and checked in a scan order, such as a raster scan order. The point with the smallest SAD can be selected as the output of the integer sample offset search stage. In order to reduce the penalty of uncertainty in DMVR refinement, the original MVs in the DMVR process can have a priority to be selected. The SAD between the reference blocks referenced by the initial MV candidates can be reduced by ¼ of the SAD value.

[0158] The integer sample search can be followed by fractional sample refinement. To reduce computational complexity, the fractional sample refinement can be derived by using a parametric error surface equation instead of an additional search with SAD comparison. The fractional sample refinement can be conditionally invoked based on the output of the integer sample search stage. If the integer sample search stage ends at the center with the smallest SAD in either the first iteration search or the second iteration search, the fractional sample refinement can be further applied.

[0159] In parametric error surface based sub-pixel offset estimation, the central location cost and the costs at the four neighboring locations from the center are used to fit a 2D parabolic error surface equation based on Eq. (32): E(x,y)=A(xx min ) 2 +B(yy min ) 2 +C formula (32) Here, (x min ,y min ) can correspond to the minimum cost fractional position, and C can correspond to the minimum cost value. By solving equation (32) using the cost values ​​of the five search points, (x min ,ymin ) can be calculated as follows:

number

[0160] In VVC, etc., quadratic interpolation and sample padding can be applied. The resolution of the MV can be, for example, 1 / 16 luma samples. The fractional position samples can be interpolated using an 8-tap interpolation filter. In DMVR, the search points can surround the initial fractional pel MV with integer sample offsets, so the fractional position samples need to be interpolated for the DMVR search process. To reduce the computational complexity, a bilinear interpolation filter can be used to generate fractional samples for the search process in DMVR. In another important effect, by using a bilinear filter with a 2-sample search range, DVMR does not access more reference samples compared to a normal motion compensation process. After the refined MV is achieved using the DMVR search process, a normal 8-tap interpolation filter can be applied to generate the final prediction. Due to the lack of access to more reference samples compared to a normal MC process, samples that may not be needed for the original MV-based interpolation process but may be needed for the refined MV-based interpolation process can be padded from the available samples.

[0161] If the width and / or height of a CU is greater than 16 luma samples, the CU may be further divided into sub-blocks having widths and / or heights equal to 16 luma samples. The maximum unit size of the DMVR search process may be limited to 16×16.

[0162] In related bilateral matching processes such as BDOF and DMVR, only the translational motion model is applied, which may not capture complex motion information.

[0163] In this disclosure, bilateral matching can be provided to derive / refine affine motion parameters instead of signaling. The affine parameters can include translation, rotation, zoom, and / or other motions. Bilateral matching can be performed to find the best (or selected) match between two blocks (e.g., a first reference block and a second reference block) in a reference frame (e.g., a first reference frame and a second reference frame) with respect to a minimum matching error such as SAD, SSE, etc. The two blocks in the reference frame can be constrained by an affine motion model.

[0164] In some embodiments, the affine motion model may be a four-parameter affine motion model or a six-parameter affine motion model. In some embodiments, the constraints applied to the two blocks in the reference frame may be associated with parameters of the affine motion model. The parameters may include a zoom factor, a rotation angle, a translation portion, etc.

[0165] Let affine motion 0 (or affMv0) and affine motion 1 (or affMv1) be the affine motions from the current block to the first reference block and the second reference block, respectively, and the two blocks (or reference blocks) can be located (or determined) by affMv0 and affMv1 under the constraints of this disclosure. In one embodiment, affMV0 or affMV1 can be described by a four-parameter affine motion model shown in equation (1). In another embodiment, affMV0 or affMV1 can be described by a six-parameter affine motion model shown in equation (3). In another embodiment, affMv0 and affMv1 can represent delta affine motions with respect to the initial motions of the two reference blocks as indicated by the merge index or predictor index. In one example, temporal distance 0 (or dPoc0) and temporal distance 1 (or dPoc1) are the temporal distances between the current frame and each of the two reference frames.

[0166] If Poc_Cur represents the picture order count (POC) of the current frame, RefPoc_L0 represents the POC of the reference frame on list L0 (or the first reference frame), and RefPoc_L1 represents the POC of the reference frame on list L1 (or the second reference frame), then dPoc0 and dPoc1 can be written as the following Equations (35) and (36): dPoc0=Poc_Cur-RefPoc_L0 Equation (35) dPoc1=Poc_Cur-RefPoc_L1 Equation (36) dPoc0 and dPoc1 may have different signs (e.g., positive and negative signs), which may indicate that the two reference frames are on different temporal sides of the current frame.

[0167] In the present disclosure, a temporal distance-based constraint may be applied to the affine motions (e.g., affMV0 and affMV1). In one embodiment, the translational portions of affMv0 and affMv1 may be related to dPoc0 / dPoc1. For example, the translational portions of affMv0 and affMv1 may be proportional to dPoc0 / dPoc1. The translational portions may be represented by c in a first direction (e.g., X direction) and f in a second direction (e.g., Y direction) as shown in equation (13).

[0168] In one embodiment, the zoom factor portions of affMv0 and affMv1 may be related to dPoc0 / dPoc1. The zoom factor portions of affMv0 and affMv1 may be proportional to dPoc0 / dPoc1, such as exponentially proportional. The zoom factor may be r as shown in equation (13). For example, if the zoom factor of affMv0 for the first reference block is r0, the zoom factor r1 of affMv1 may be shown in equation (37):

number

[0169] In one embodiment, α0 and α1 can be the delta zoom portions of affMv0 and affMv1, respectively, relative to 1. α0=r0-1, α1=r1-1. Constraints can be applied to α0 and α1 such that they can be proportional to dPoc0 / dPoc1, such as linearly proportional to dPoc0 / dPoc1 in equation (38):

number

[0170] In one embodiment, the rotation angles θ0 and θ1 of affMv0 and affMv1 (such as in equation (13)) may be related to dPoc0 / dPoc1, such as being proportional to dPoc0 / dPoc1. For example, as shown in equation (39), the rotation angles θ0 and θ1 of affMv0 and affMv1 may be linearly proportional to dPoc0 / dPoc1.

number

[0171] In one embodiment, when different weighting factors are applied to two reference lists (e.g., reference list L0 and reference list L1) of affine bi-prediction (e.g., BCW or picture-level weighted bi-prediction), an additional weighting factor can be applied to the temporal distance-based constraint. In one example, if the weighting factor derived from the bi-prediction weights is w, the linear ratio dPoc1 / dPoc0 used in the above equations (37)-(39) can be replaced by:

number

[0172] In this disclosure, a bilateral matching process can be applied to predict a current block in a current frame.

[0173] In one embodiment, the bilateral matching process can be performed by searching for the minimum matching error within a certain search range. In one example, the search can be performed similarly to DMVR. Unlike DMVR, the search range can include not only the translation part but also the rotation factor and the zoom factor. The search range can be signaled in the bitstream or can be predefined. Furthermore, different search ranges can be used in different cases based on the temporal distance between the current frame and the reference frame, the frame type, and / or the temporal level, etc.

[0174] According to the bilateral matching process, a plurality of candidate reference block pairs may be determined according to one or more constraints of an affine motion model provided in at least one of Equations (37)-(39) based on a search range. Each candidate reference block pair of the plurality of candidate reference block pairs may include a respective candidate reference block in a first reference frame and a respective candidate reference block in a second reference frame. A respective cost value may be determined for each candidate reference block pair of the plurality of candidate reference block pairs. The cost value may indicate a difference between the candidate reference block in the first reference frame and the candidate reference block in the second reference frame. The cost value may be determined based on one of a mean square error (MSE), a mean absolute difference (MAD), a SAD, a sum of absolute differences after transformation (SATD), and the like. A reference block pair in the plurality of candidate reference block pairs associated with a minimum cost value may be selected to predict a current block. For example, an affine motion model may be determined based on a CPMV associated with the selected reference block pair. A block level, a sub-block level, or a pixel level prediction of the current block may be further performed based on the determined affine motion model.

[0175] In one embodiment, the bilateral matching process can be operated by minimizing the distortion model of two reference blocks of the current block. The distortion model can be based on a first-order Taylor expansion, which can be similar to the BDOF of Equation (7) or the affine ME of the VTM software provided in Figures 15-16.

[0176] The distortion model (or distortion value) between two reference blocks in the first iteration of the bilateral matching process can be given by the following Equations (40)-(42): D1(i,j)=P 1,L0 (i,j)-P 1,L1 (i,j) Equation (40) P 1,L0 (i,j)=P 0,L0 (i,j)+g x0,L0(i,j)*Δv x0,L0 (i,j)+g y0,L0 (i,j)*Δv y0,L0 (i,j) Equation (41) P 1,L1 (i,j)=P 0,L1 (i,j)+g x0,L1 (i,j)*Δv x0,L1 (i,j)+g y0,L1 (i,j)*Δv y0,L1 (i,j) Equation (42) As shown in equation (40), P 1,L0 (i, j) may be the first predictor of the current block based on the first reference block in the reference list L0. 1,L1 (i,j) may be the first predictor of the current block based on the first reference block in the reference list L1. (i,j) may be the location of a pixel (or sample). D1(i,j) is the first predictor of the current block based on the first reference block in the reference list L1. 1,L0 (i,j) and P 1,L1 (i,j) indicates the pixel difference between the current block and the first reference block in reference list L0. The first reference block in reference list L1 can be determined by affine motion 1 (e.g., affMV1) from the current block to the first reference block in reference list L1. Affine motion 0 and affine motion 1 can be constrained by at least one of equations (37) to (39).

[0177] P 1,L0 (i,j) and P 1,L1 (i,j) can be determined according to the affine ME search process shown in Figure 16. As shown in equation (41), P 0,L0 (i,j) may be the initial predictor of the current block based on the initial reference block (or base CPMV) in the reference list L0. x0,L0 (i,j) is the initial predictor P 0,L0 It can be taken as the gradient in the x direction of (i,j). y0,L0 (i,j) is the initial predictor P 0,L0It can be taken as the gradient in the y direction of (i,j). Δv x0,L0 (i,j) may be the difference or displacement along the x direction of two reference blocks (or sub-blocks), such as the initial reference block and the first reference block in the reference list L0. y0,L0 (i,j) may be the difference or displacement along the y direction of two reference blocks (or sub-blocks), such as the initial reference block and the first reference block in reference list L0.

[0178] Similarly, as shown in equation (42), P 0,L1 (i,j) may be the initial predictor of the current block based on the initial reference block (or base CPMV) in the reference list L1. x0,L1 (i,j) is the initial predictor P 0,L1 It can be taken as the gradient in the x direction of (i,j). y0,L1 (i,j) is the initial predictor P 0,L1 It can be said that the gradient of (i,j) in the y direction is Δv x0,L1 (i,j) may be the difference or displacement along the x direction of two reference blocks (or sub-blocks), such as the initial reference block and the first reference block in the reference list L1. y0,L1 (i,j) may be the difference or displacement along the y direction of two reference blocks (or sub-blocks), such as the initial reference block and the first reference block in reference list L1.

[0179] Δv x0,L0 (i,j), Δv y0,L0 (i,j), Δv x0,L1 (i,j), and Δv y0,L1 In response to at least one of (i,j) being non-zero, the bilateral matching may proceed to a second iteration, in which a second predictor P 2,L0 (i,j), and a second predictor P of the current block based on the second reference block in the reference list L1. 2,L1(i,j) can be determined according to the following equations (43)-(44): P 2,L0 (i,j)=P 1,L0 (i,j)+g x1,L0 (i,j)*Δv x1,L0 (i,j)+g y1,L0 (i,j)*Δv y1,L0 (i,j) Equation (43) P 2,L1 (i,j)=P 1,L1 (i,j)+g x1,L1 (i,j)*Δv x1,L1 (i,j)+g y1,L1 (i,j)*Δv y1,L1 (i,j) Equation (44) As shown in equation (43), g x1,L0 (i,j) is the first predictor P 1,L0 It can be taken as the gradient in the x direction of (i,j). y1,L0 (i,j) is the first predictor P 1,L0 It can be said that the gradient of (i,j) in the y direction is Δv x1,L0 (i,j) may be the difference or displacement along the x direction between the first and second reference blocks in the reference list L0. y1,L0 (i,j) can be the difference or displacement between the first and second reference blocks along the y direction. As shown in equation (44), g x1,L1 (i,j) is the first predictor P 1,L1 It can be taken as the gradient in the x direction of (i,j). y1,L1 (i,j) is the first predictor P 1,L1 It can be said that the gradient of (i,j) in the y direction is Δv x1,L1 (i,j) can be the difference or displacement along the x direction between the first and second reference blocks in the reference list L1. y1,L1 (i,j) may be the difference or displacement along the y direction between the first and second reference blocks in reference list L1.

[0180] The second reference block in reference list L0 may be represented by an affine motion 0' (e.g., affMV0') from the current block to the second reference block in reference list L0. The second reference block in reference list L1 may be represented by an affine motion 1' (e.g., affMV1') from the current block to the second reference block in reference list L1. Affine motion 0' and affine motion 1' may be constrained by at least one of equations (37) to (39).

[0181] Furthermore, P 2,L0 (i,j) and P 2,L1 The pixel difference (or cost value) for (i,j) can be calculated according to equation (45) below: D2(i,j)=P 2,L0 (i,j)-P 2,L1 (i,j) Equation (45)

[0182] The bilateral matching process is repeated until the number of iterations N is equal to or greater than a threshold, or until the displacement Δv N+1,L0 (i, j) is 0 or the displacement Δv between the Nth reference block in the reference list L1 and the (N+1)th reference block in the reference list L1 N+1,L1 The bilateral matching process may be terminated when (i,j) is 0. Thus, N reference block pairs may be generated based on the bilateral matching process. Each of the N reference block pairs may include a respective reference block in reference list L0 and a respective reference block in reference list L1. Each of the N reference block pairs may also include a respective distortion value indicating a difference between a corresponding reference block in reference list L0 and a corresponding reference block in reference list L1.

[0183] According to the distortion values ​​(or cost values) of the N reference block pairs, the current block can be predicted based on the best (or selected) reference block pair having the smallest distortion value. For example, an affine motion model can be determined based on the CPMV associated with the selected reference block pair. Block-level, sub-block-level, or pixel-level prediction of the current block can be further performed based on the determined affine motion model.

[0184] It should be noted that each of the N reference block pairs may include a respective reference block in the first reference list L0 and a respective reference block in the second reference list L1. Each reference block in the first reference list L0 and the second reference list L1 may have a regular or irregular shape. A regular shape may be a shape with all sides equal and all interior angles equal. For example, a reference block may have a square shape. An irregular shape may not have equal sides or equal angles.

[0185] The number of iterations of the bilateral matching process can be signaled in the bitstream or predefined based on one or more conditions. One or more additional constraints can be applied during the bilateral matching process. For example, the delta affine motion can be further required to be within a certain range, such as within 2 pixels. In one embodiment, such a range can be signaled in the bitstream or predefined. The delta affine motion can be the difference between the current motion vector and the previous motion vector during the bilateral matching process. The delta affine motion can be at the sub-block level, the control point level, or the pixel level.

[0186] In this disclosure, the starting point of the bilateral matching process (e.g., P 0,L0 (i,j) and P 0,L1(i,j)) can be indicated by a predictor such as a merge index, an AMVP predictor index, or an affine motion indicated by an affine merge index.

[0187] In some embodiments, after bilateral matching, the derived affine motion can be used directly as the motion of the current block, or the affine motion can be used as a motion predictor for the current block.

[0188] In this disclosure, the bilateral matching process can be combined with existing coding modes such as merge, affine / sub-block merge, MMVD, or GPM.

[0189] In one embodiment, an additional flag can be signaled in the bitstream to indicate whether bilateral matching processing is used.

[0190] In one embodiment, the bilateral matching process may be always on whenever a block (or the current block) is bi-predicted.

[0191] In one embodiment, the bilateral matching process is always on whenever a block (or the current block) is bi-predicted and coded in affine mode.

[0192] In one embodiment, the bilateral matching process can be added to a candidate list and identified by an index. When the bilateral matching process is added to a candidate list, such as an affine candidate list, the CPMV associated with the N reference block pairs and the starting point can serve as a candidate in the affine candidate list. The candidate can be inserted after SbTmvp, after an inherited affine candidate, after a constructed affine candidate, or after a history-based candidate, for example.

[0193] In some embodiments, the bilateral matching process can be used as an alternative or general form of other bilateral processes such as BDOF.

[0194] In the present disclosure, a three-parameter affine model or a four-parameter affine model can be applied in the affine bilateral matching process. In one example, a three-parameter scaling (or zoom) model can be applied in the affine bilateral matching process. The three-parameter scaling model associated with the reference list L0 can be simplified in equation (46), where (1+α) can be a scaling coefficient and the parameters (c, f) can represent translational motion.

number

number

[0195] In one example, the four-parameter affine model can include a rotation transformation and a translational motion (e.g., a scaling factor r=1). Thus, the four-parameter affine model associated with the reference list L0 can be simplified to Equation (48).

number

number

[0196] Affine bilateral matching can be achieved based on the above affine model provided in equations (46)-(49).

[0197] For example, the current iteration of affine bilateral matching can be shown in Equation (50) and Equation (51). As shown in Equation (50) and Equation (51), p0 and p1 can represent predictions on reference list L0 and reference list L1, respectively. p0 and p1 can be obtained from the base CPMV or from the previous iteration. p'0 and p'1 can be the refined predictions after the current iteration with ΔMV0 and ΔMV1 applied to each sample of the current block. ΔMV0 is associated with reference list L0, and the component ΔMV in the x direction is x0 and the y-component ΔMV y0 ΔMV0 may represent an affine motion vector difference between the current iteration and the previous iteration. x and g y can be the horizontal and vertical gradients of the prediction (e.g., p0 or p1). p'0=p0+g x0 ΔMV x0 +g y0 ΔMV y0 Formula (50) p'1=p1+g x1 ΔMV x1 +g y1 ΔMV y1 Formula (51) An affine bilateral matching process can be performed. The mean squared error (MSE) of the two refined predictions after each iteration is Σ(p'0-p'1) 2 Such as, the distortion between the refinement predictions p'0 and p'1 can be minimized. Based on the symmetric model linear functions shown in Eqs. (46)-(49), the affine bilateral matching can be the same optimization process shown in the affine ME of FIG.

[0198] As shown in FIG. 16, the affine bilateral matching process may start from an initial CPMV (or base CPMV) of an affine candidate. The initial CPMV may be refined in a first iteration to generate an updated CPMV value. In each iteration, an affine prediction may be generated based on the updated CPMV value. Similar to the affine motion estimation method provided in the VTM software, a gradient-based affine equation solving method may be applied in each iteration. The corresponding affine model may be used for affine parameter derivation. The new delta affine parameters (e.g., delta CPMV) generated in each iteration may be applied to each reference list based on a symmetric model, such as the symmetric model of Equations (46)-(49), to generate updated CPMVs for both reference lists. The updated CPMV may correspond to an updated reference block in a reference list (e.g., reference list L0 or L1).

[0199] In some embodiments, the iterations can end when the delta CPMV (e.g., ΔMV0 or ΔMV1) becomes zero or the iterations reach a predetermined number of iterations. In some embodiments, the constraints stated in equations (37-39) can be applied to the symmetric model of equations (46)-(49).

[0200] FIG. 19 shows a flowchart outlining an exemplary decoding process (1900) according to some embodiments of the present disclosure. FIG. 20 shows a flowchart outlining an exemplary encoding process (2000) according to some embodiments of the present disclosure. The proposed processes may be used separately or combined in any order. Furthermore, each of the processes (or embodiments), the encoder, and the decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored in a non-transitory computer-readable medium.

[0201] The operations of the processes (e.g., (1900) and (2000)) can be combined or arranged in any quantity or order as desired. In embodiments, two or more of the operations of the processes (e.g., (1900) and (2000)) may be performed in parallel.

[0202] The processes (e.g., (1900) and (2000)) can be used in the reconstruction and / or encoding of a block to generate a prediction block for a block being reconstructed. In various embodiments, the processes (e.g., (1900) and (2000)) are performed by processing circuitry of terminal devices (310), (320), (330), and (340), processing circuitry performing the functions of a video encoder (403), processing circuitry performing the functions of a video decoder (410), processing circuitry performing the functions of a video decoder (510), processing circuitry performing the functions of a video encoder (603), etc. In some embodiments, the processes (e.g., (1900) and (2000)) are implemented with software instructions, such that the processing circuitry performs the processes (e.g., (1900) and (2000)) when the processing circuitry executes the software instructions.

[0203] As shown in Figure 19, the process (1900) may begin at (S1901) and proceed to (S1910), where prediction information for a current block in a current picture may be decoded from a coded video bitstream, and the prediction information may indicate that the current block should be predicted based on an affine model.

[0204] In (S1920), the affine motion parameters of the affine model may be derived by affine bilateral matching, in which the affine model is derived based on reference blocks in a first reference picture and a second reference picture of the current picture. The affine motion parameters may not be included in the coded video bitstream.

[0205] At (S1930), control point motion vectors of the affine model may be determined based on the derived affine motion parameters.

[0206] In (S1940), the current block can be reconstructed based on the derived affine model.

[0207] A plurality of affine motion parameters of the affine model can be derived based on a reference block pair from a plurality of candidate reference block pairs of reference blocks in a first reference picture and a second reference picture. The reference block pair can include a first reference block in the first reference picture and a second reference block in the second reference picture based on the affine model and a cost value constraint. The constraint can be associated with a temporal distance ratio based on (i) a first temporal distance between the current picture and the first reference picture, and (ii) a second temporal distance between the current picture and the second reference picture. The cost value can be based on a difference between the first reference block and the second reference block.

[0208] In some embodiments, the temporal distance ratio may be equal to the product of a weighting factor and a ratio of (i) a first temporal distance between the current picture and the first reference picture and (ii) a second temporal distance between the current picture and the second reference picture, where the weighting factor may be a positive integer.

[0209] In some embodiments, the affine model constraint may indicate that a first translation coefficient of a first affine motion vector from the current block to the first reference block is proportional to a temporal distance ratio. The affine model constraint may indicate that a second translation coefficient of a second affine motion vector from the current block to the second reference block is proportional to a temporal distance ratio.

[0210] In some embodiments, the constraints of the affine model may further indicate that a second zoom factor of a second affine motion vector from the current block to the second reference block is equal to a first zoom factor of a first affine motion vector from the current block to the first reference block relative to a power of the temporal distance ratio.

[0211] In some embodiments, the constraint of the affine model may indicate that a ratio of (i) a first delta zoom factor of a first affine motion vector from the current block to the first reference block and (ii) a second delta zoom factor of a second affine motion vector from the current block to the second reference block is equal to a temporal distance ratio. The first delta zoom factor may be equal to the first zoom factor -1, and the second delta zoom factor may be equal to the second zoom factor -1.

[0212] In some embodiments, the constraints of the affine model may further indicate that a ratio between (i) a first rotation angle of a first affine motion vector from the current block to the first reference block and (ii) a second rotation angle of a second affine motion vector from the current block to the second reference block is equal to a temporal distance ratio.

[0213] To determine the reference block pair, a plurality of candidate reference block pairs can be determined according to the constraints of the affine model. Each candidate reference block pair of the plurality of candidate reference block pairs can include a respective candidate reference block in a first reference picture and a respective candidate reference block in a second reference picture. A respective cost value can be determined for each candidate reference block pair of the plurality of candidate reference block pairs. The reference block pair associated with the minimum cost value can be determined as the candidate reference block pair of the plurality of candidate reference block pairs.

[0214] In some embodiments, the plurality of candidate reference block pairs may include a first candidate reference block pair, and the first candidate reference block pair may include a first candidate reference block in a first reference picture and a first candidate reference block in a second reference picture. To determine the reference block pair, an initial predictor associated with the first reference picture of the current block may be determined based on the initial reference block in the first reference picture. An initial predictor associated with the second reference picture of the current block may be determined based on the initial reference block in the second reference picture. The first predictor associated with the first reference picture of the current block may be determined based on the initial predictor associated with the first reference picture of the current block, and the first predictor associated with the first reference picture of the current block may be associated with the first candidate reference block in the first reference picture. The first predictor associated with the second reference picture of the current block can be determined based on an initial predictor associated with the second reference picture of the current block, and the first predictor associated with the second reference picture of the current block can be associated with a first candidate reference block in the second reference picture. The first cost value can be determined based on a difference between the first predictor associated with the first reference picture of the current block and the first predictor associated with the second reference picture of the current block.

[0215] In the process (1900), the initial predictor associated with the first reference picture of the current block may be indicated by one of a merge index, an advanced motion vector prediction (AMVP) predictor index, and an affine merge index.

[0216] To determine a first predictor associated with the first reference picture of the current block, a first component in a first direction of a gradient value of an initial predictor associated with the first reference picture of the current block may be determined. A second component in a second direction of a gradient value of an initial predictor associated with the first reference picture of the current block may be determined. The second direction may be perpendicular to the first direction. A first component in a first direction of a displacement between an initial reference block in the first reference picture and a first candidate reference block in the first reference picture may be determined. A second component in a second direction of a displacement between an initial reference block in the first reference picture and a first candidate reference block in the first reference picture may be determined. A first predictor associated with the first reference picture of the current block may be determined to be equal to the sum of (i) an initial predictor associated with the first reference picture of the current block, (ii) a product of a first component of a gradient value of the initial predictor and a first component of the displacement, and (iii) a product of a second component of the gradient value of the initial predictor and a second component of the displacement.

[0217] The plurality of candidate reference block pairs may include an Nth candidate reference block pair, and the Nth candidate reference block pair may include an Nth candidate reference block in a first reference picture and an Nth candidate reference block in a second reference picture. To determine the reference block pair, an Nth predictor associated with the Nth candidate reference block in a first reference picture of the current block may be determined based on an (N-1)th predictor associated with the (N-1)th candidate reference block in a first reference picture of the current block. An Nth predictor associated with the Nth candidate reference block in a second reference picture of the current block may be determined based on an (N-1)th predictor associated with the (N-1)th candidate reference block in a second reference picture of the current block. An Nth cost value may be determined based on a difference between the Nth predictor associated with the first reference picture of the current block and the Nth predictor associated with the second reference picture of the current block.

[0218] To determine the N-th predictor associated with the first reference picture of the current block, a first component in a first direction of a gradient value of the (N-1)-th predictor associated with the first reference picture of the current block may be determined. A second component in a second direction of a gradient value of the (N-1)-th predictor associated with the first reference picture of the current block may be determined. The second direction may be perpendicular to the first direction. A first component of a displacement may be determined. The first component of the displacement may be a difference in a first direction between the N-th candidate reference block in the first reference picture and the (N-1)-th candidate reference block in the first reference picture. A second component of the displacement may be a difference in a second direction between the N-th candidate reference block in the first reference picture and the (N-1)-th candidate reference block in the first reference picture. The Nth predictor may be determined based on the first reference picture of the current block as being equal to the sum of (i) the (N-1)th predictor associated with the first reference picture of the current block, (ii) the product of a first component of the gradient value of the (N-1)th predictor and a first component of the displacement, and (iii) the product of a second component of the gradient value of the (N-1)th predictor and a second component of the displacement.

[0219] In some embodiments, the multiple candidate reference block pairs may include N candidate reference block pairs based on one of: (i) N is equal to an upper limit value; and (ii) a displacement between the Nth candidate reference block in the first reference picture and the (N+1)th candidate reference block in the first reference picture is 0.

[0220] In some embodiments, the delta motion vector of each sub-block in the current block may be less than or equal to a threshold.

[0221] The prediction information may include a flag indicating whether the current block is predicted based on affine bilateral matching using an affine model.

[0222] In the process (1900), a list of candidate affine model motion vectors associated with the current block may be determined. The candidate motion vector list may include control point motion vectors associated with the reference block pair.

[0223] After (S1940), the process proceeds to (S1999) and ends.

[0224] The process (1900) may be adapted as appropriate. Step(s) of the process (1900) may be modified and / or omitted. Additional step(s) may be added. Any suitable order of implementation may be used.

[0225] As shown in Figure 20, the process (2000) can start at (S2001) and proceed to (S2010). At (S2010), a constraint of an affine model of a current block in a current picture can be determined. The constraint can be associated with a temporal distance ratio based on (i) a first temporal distance between the current picture and a first reference picture of the current picture, and (ii) a second temporal distance between the current picture and a second reference picture of the current picture.

[0226] In (S2020), a plurality of affine motion parameters of the affine model can be determined based on a reference block pair from a candidate reference block pair in a first reference picture and a second reference picture. The reference block pair can include a first reference block in the first reference picture and a second reference block in the second reference picture. The reference block pair can be determined from the plurality of candidate reference block pairs based on a constraint of the affine model and a cost value. The cost value can be associated with a difference between the first reference block and the second reference block.

[0227] At (S2030), control point motion vectors of the affine model may be determined based on the determined plurality of affine motion parameters.

[0228] In (S2040), prediction information for the current block can be generated based on the determined affine model.

[0229] The process then proceeds to (S2099) and ends.

[0230] The process (2000) may be adapted as appropriate. Steps in the process (2000) may be modified and / or omitted. Additional step(s) may be added. Any suitable order of implementation may be used.

[0231] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 21 illustrates a computer system (2100) suitable for implementing certain embodiments of the disclosed subject matter.

[0232] Computer software can be coded using any suitable machine code or computer language that is amenable to mechanisms such as assembly, compilation, linking, etc., to create code that includes instructions that can be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., directly, or via interpretation, microcode execution, etc.

[0233] The instructions may be executed on various types of computers or computer components including, for example, personal computers, tablet computers, servers, smartphones, gaming consoles, Internet of Things devices, and the like.

[0234] 21 for the computer system (2100) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The arrangement of components should not be interpreted as having any dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of the computer system (2100).

[0235] The computer system (2100) may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, via tactile input (such as keystrokes, swipes, data glove movements, etc.), audio input (such as voice, clapping, etc.), visual input (such as gestures), or olfactory input (not shown). Human interface devices may also be used to capture certain media that are not necessarily directly associated with conscious human input, such as audio (such as voice, music, ambient sounds, etc.), images (such as scanned images, photographic images obtained from a still image camera, etc.), and video (such as two-dimensional video, three-dimensional video including stereoscopic video, etc.).

[0236] The input human interface devices may include one or more of a keyboard (2101), a mouse (2102), a trackpad (2103), a touch screen (2110), a data glove (not shown), a joystick (2105), a microphone (2106), a scanner (2107), and a camera (2108) (only one of each is shown).

[0237] The computer system (2100) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through haptic output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touch screen (2110), data gloves (not shown), or joystick (2105), although there may also be haptic feedback devices that do not function as input devices), audio output devices (such as speakers (2109), headphones (not shown)), visual output devices (such as screens (2110), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without haptic feedback capabilities, some of which may be capable of outputting two-dimensional visual output or three- or more-dimensional output via means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0238] The computer system (2100) may also include human-accessible storage devices and their associated media, such as optical media, including CD / DVD ROM / RW (2120) with CD / DVD or similar media (2121), thumb drives (2122), removable hard drives or solid state drives (2123), legacy magnetic media such as tapes and floppy disks (not shown), and specialized ROM / ASIC / PLD based devices (not shown) such as security dongles.

[0239] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.

[0240] The computer system (2100) may also include an interface (2154) to one or more communication networks (2155). The networks may be, for example, wireless, wired, optical. The networks may further be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial including CANBus, etc. Certain networks typically require an external network interface adapter attached to a specific general-purpose data port (e.g., a USB port of the computer system (2100)) or peripheral bus (2149), while other networks are typically integrated into the core of the computer system (2100) by attaching to a system bus described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (2100) can communicate with other entities. Such communications can be one-way receive only (e.g., broadcast TV), one-way transmit only (e.g., CANbus to a specific CANbus device), or bidirectional, for example, with other computer systems using local or wide area digital networks. Specific protocols and protocol stacks can be used with each of these networks and network interfaces, as described above.

[0241] The aforementioned human interface devices, human accessible storage devices, and network interfaces may be attached to the core (2140) of the computer system (2100).

[0242] The cores (2140) may include specialized programmable processing devices in the form of one or more central processing units (CPUs) (2141), graphics processing units (GPUs) (2142), field programmable gate areas (FPGAs) (2143), hardware accelerators for specific tasks (2144), graphics adapters (2150), and the like. These devices may be connected via a system bus (2148), along with read-only memory (ROM) (2145), random access memory (2146), and internal mass storage (2147), such as an internal hard drive or SSD that is not accessible to the user. In some computer systems, the system bus (2148) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, and the like. Peripheral devices may be attached directly to the core's system bus (2148) or via a peripheral bus (2149). In one example, a screen (2110) may be connected to a graphics adapter (2150). Architectures for peripheral buses include PCI, USB, and the like.

[0243] The CPU (2141), GPU (2142), FPGA (2143), and accelerator (2144) can execute certain instructions that can combine to constitute the aforementioned computer code. That computer code can be stored in ROM (2145) or RAM (2146). Persistent data can be stored, for example, in internal mass storage (2147), while transitory data can also be stored in RAM (2146). Rapid storage and retrieval from any of the memory devices can be enabled using cache memory, which can be closely associated with one or more of the CPU (2141), GPU (2142), mass storage (2147), ROM (2145), RAM (2146), etc.

[0244] The computer-readable medium can bear computer code for performing various computer-implemented operations. The medium and computer code can be those specially designed and constructed for the purposes of the present disclosure, or they can be of the available kind well known to those skilled in the computer software arts.

[0245] As an example, but not by way of limitation, a computer system (2100) having an architecture, specifically a core (2140), can provide functionality as a result of a processor(s) (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be the user-accessible mass storage introduced above, as well as media associated with specific storage of the core (2140) of a non-transitory nature, such as the core internal mass storage (2147) or ROM (2145). Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (2140). The computer-readable media can include one or more memory devices or chips, depending on the particular needs. The software can cause the core (2140), and specifically the processors (including CPUs, GPUs, FPGAs, etc.) therein, to perform certain processes or certain portions of certain processes described herein, including defining data structures stored in RAM (2146) and modifying such data structures according to the software-defined processes. Additionally, or alternatively, the computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator (2144)) that may operate in place of or together with software to perform certain processes or certain portions of certain processes described herein. Where appropriate, references to software may encompass logic and vice versa. Where appropriate, references to computer-readable media may encompass circuitry (such as integrated circuits (ICs)) that store software for execution, circuitry that embodies logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0246] Appendix A: Acronyms JEM: Joint exploration model VVC: versatile video coding BMS: benchmark set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplementary Enhancement Information VUI: Video Usability Information GOP: Group of Pictures TU: Transform Unit PU: Prediction Unit CTU: Coding Tree Unit CTB: Coding Tree Block PB: Prediction Block HRD: Hypothetical Reference Decoder SNR: Signal to Noise Ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: Cathode Ray Tube LCD: Liquid Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile communications LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Area SSD: Solid-state drive IC: Integrated Circuit CU: Coding Unit

[0247] While this disclosure describes several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods not explicitly shown or described herein, but which embody the principles of this disclosure and are therefore within the spirit and scope of the present disclosure. [Explanation of symbols]

[0248] 101 sample, point, 102 arrow, 103 arrow, 104 block, square block, 110 schematic diagram, 201 block, 300 communication system, 310 terminal device, 320 terminal device, 330 terminal device, 340 terminal device, 350 network, 350 communication network, 400 communication system, 401 video source, 402 stream, 403 video encoder, 404 video data, 405 streaming server, 406 client subsystem, 407 input copy, 410 video decoder, 411 output stream, 412 display, 413 capture subsystem, 420 electronic device, 430 electronic device, 501 channel, 510 video decoder, 512 rendering device, 515 buffer memory, 520 parser, 521 symbol, 530 electronic device, 531 receiver, 551 inverse transformation unit, 552 Intra picture prediction unit, Intra prediction unit, 553 Motion compensated prediction unit, 555 Aggregator, 556 Loop filter unit, 557 Reference picture memory, 558 Current picture buffer, 601 Video source, 603 Video encoder, 620 Electronic device, 630 Source coder, 632 Coding engine, 633 Decoder, Local decoder, Local video decoder, 634 Reference picture memory, 635 Predictor, 640 Transmitter, 643 Video sequence, 645 Entropy coder, 650 Controller, 660 Communication channel, 703 Video encoder, 721 Generic controller, 722 Intra encoder, 723 Residual calculator, 724 Residual encoder, 725 Entropy encoder, 726 Switch, 728 Residual decoder, 730 Inter encoder, 810 Video decoder, 871 Entropy decoder, 872 Intra decoder, 873 Residual decoder, 874 Reconstruction module, 880 Interdecoder, 1000 Current block, 1002 Center sample, 1004 Sub-block, 1202 CU, 1204 Current block, current CU, 1302 Current block, current CU, 1400 Current block, CU, 1402 Sub-block, 1404 Sample, 1406Reference pixel, 1408 Reference pixel, 1500 Affine ME, 1600 Affine ME process, 1702 Expanded row / column, 1704 CU, 1706 Boundary, 1802 Current picture, 1804 Reference picture list L0, 1806 Reference picture list L1, 1808 Current block, 1812 Initial reference block, 1814 Initial reference block, 1816 First candidate reference block, 1900 Decoding process, 2000 Encoding process, 2100 Computer system, 2101 Keyboard, 2102 Mouse, 2103 Trackpad, 2105 Joystick, 2106 Microphone, 2107 Scanner, 2108 Camera, 2109 Speaker, 2110 Touch screen, 2120 CD / DVD ROM / RW, 2121 Media, 2122 Thumb drive, 2123 solid state drive, 2140 core, 2141 central processing unit, 2142 graphics processing unit, 2143 field programmable gate area, 2144 hardware accelerator, 2145 read only memory, 2146 random access memory, 2147 core internal mass storage, 2148 system bus, 2149 peripheral bus, 2150 graphics adapter, 2154 interface, 2155 communication network, L0 reference list, L1 reference list, MV0 initial motion vector, MV1 initial motion vector, P0 prediction of current block, P1 prediction of current block, S1502 affine uni-prediction, S1504 affine uni-prediction, S1506 affine bi-prediction

Claims

1. 1. A method of video decoding performed by a video decoder, comprising: decoding prediction information for a current block in a current picture from a coded video bitstream, the prediction information indicating that the current block should be predicted based on an affine model; deriving a plurality of affine motion parameters of the affine model via affine bilateral matching, where the affine model is derived based on reference blocks in a first reference picture and a second reference picture of the current picture, and the plurality of affine motion parameters are not included in the coded video bitstream, the step further comprising: deriving the plurality of affine motion parameters of the affine model based on a reference block pair from a plurality of candidate reference block pairs of the reference blocks in the first reference picture and the second reference picture, the reference block pair including a first reference block in the first reference picture and a second reference block in the second reference picture, based on a constraint of the affine model and a cost value, the constraint being associated with a temporal distance ratio based on (i) a first temporal distance between the current picture and the first reference picture and (ii) a second temporal distance between the current picture and the second reference picture, the cost value being based on a difference between the first reference block and the second reference block; determining control point motion vectors of the affine model based on the derived affine motion parameters; reconstructing the current block based on the derived affine model; Including, the temporal distance ratio is equal to a product of a weighting factor and a ratio of (i) the first temporal distance between the current picture and the first reference picture and (ii) the second temporal distance between the current picture and the second reference picture, the weighting factor being a positive integer.

2. The constraints of the affine model are: a first translation coefficient of a first affine motion vector from the current block to the first reference block is proportional to the temporal distance ratio; and a second translation coefficient of a second affine motion vector from the current block to the second reference block is proportional to the temporal distance ratio; The method of claim 1, further comprising:

3. The constraints of the affine model are: a second zoom factor of a second affine motion vector from the current block to the second reference block is equal to a first zoom factor of a first affine motion vector from the current block to the first reference block to a power of the temporal distance ratio; The method of claim 1, further comprising:

4. The constraints of the affine model are: a ratio of (i) a first delta zoom factor of a first affine motion vector from the current block to the first reference block and (ii) a second delta zoom factor of a second affine motion vector from the current block to the second reference block is equal to the temporal distance ratio; the first delta zoom factor is equal to the first zoom factor minus 1; and the second delta zoom factor is equal to the second zoom factor minus 1; The method of claim 1, further comprising:

5. The constraints of the affine model are: a ratio of (i) a first rotation angle of a first affine motion vector from the current block to the first reference block and (ii) a second rotation angle of a second affine motion vector from the current block to the second reference block is equal to the temporal distance ratio. The method of claim 1, further comprising:

6. The step of deriving the plurality of affine motion parameters comprises: determining the plurality of candidate reference block pairs according to the constraints of the affine model, each candidate reference block pair of the plurality of candidate reference block pairs including a respective candidate reference block in the first reference picture and a respective candidate reference block in the second reference picture; determining a respective cost value for each candidate reference block pair of the plurality of candidate reference block pairs; determining the reference block pair as a candidate reference block pair from among the plurality of candidate reference block pairs associated with a minimum cost value; The method of claim 1, further comprising:

7. the plurality of candidate reference block pairs includes a first candidate reference block pair, the first candidate reference block pair including a first candidate reference block in the first reference picture and a first candidate reference block in the second reference picture; The step of deriving the plurality of affine motion parameters comprises: determining an initial predictor associated with the first reference picture of the current block based on an initial reference block in the first reference picture; determining an initial predictor associated with the second reference picture of the current block based on an initial reference block in the second reference picture; determining a first predictor associated with the first reference picture of the current block based on the initial predictor associated with the first reference picture of the current block, wherein the first predictor associated with the first reference picture of the current block is associated with the first candidate reference block in the first reference picture; determining a first predictor associated with the second reference picture of the current block based on the initial predictor associated with the second reference picture of the current block, wherein the first predictor associated with the second reference picture of the current block is associated with the first candidate reference block in the second reference picture; determining a first cost value based on a difference between the first predictor associated with the first reference picture of the current block and the first predictor associated with the second reference picture of the current block; 7. The method of claim 6, further comprising:

8. 8. The method of claim 7, wherein the initial predictor associated with the first reference picture of the current block is indicated by one of a merge index, an advanced motion vector prediction (AMVP) predictor index, and an affine merge index.

9. The step of determining the first predictor associated with the first reference picture of the current block comprises: determining a first component in a first direction of a gradient value of the initial predictor associated with the first reference picture of the current block; determining a second component in a second direction of the gradient value of the initial predictor associated with the first reference picture of the current block, the second direction being perpendicular to the first direction; determining a first component of a displacement in the first direction between the initial reference block in the first reference picture and the first candidate reference block in the first reference picture; determining a second component in the second direction of the displacement between the initial reference block in the first reference picture and the first candidate reference block in the first reference picture; determining the first predictor associated with the first reference picture of the current block as equal to the sum of (i) the initial predictor associated with the first reference picture of the current block, (ii) the product of the first component of the gradient value of the initial predictor and the first component of the displacement, and (iii) the product of the second component of the gradient value of the initial predictor and the second component of the displacement; 8. The method of claim 7, further comprising:

10. the plurality of candidate reference block pairs includes an Nth candidate reference block pair including an Nth candidate reference block in the first reference picture and an Nth candidate reference block in the second reference picture; The step of deriving the plurality of affine motion parameters comprises: determining an N-th predictor associated with the N-th candidate reference block in the first reference picture for the current block based on an (N-1)-th predictor associated with an (N-1)-th candidate reference block in the first reference picture for the current block; determining an N-th predictor associated with the N-th candidate reference block in the second reference picture of the current block based on an (N-1)-th predictor associated with an (N-1)-th candidate reference block in the second reference picture of the current block; determining an Nth cost value based on a difference between the Nth predictor associated with the first reference picture of the current block and the Nth predictor associated with the second reference picture of the current block; 7. The method of claim 6, further comprising:

11. The step of determining the Nth predictor associated with the first reference picture of the current block comprises: determining a first component in a first direction of a gradient value of the (N-1)th predictor associated with the first reference picture of the current block; determining a second component in a second direction of the gradient value of the (N-1)th predictor associated with the first reference picture of the current block, the second direction being perpendicular to the first direction; determining a first component of a displacement in the first direction between the Nth candidate reference block in the first reference picture and the (N-1)th candidate reference block in the first reference picture; determining a second component in the second direction of the displacement between the Nth candidate reference block in the first reference picture and the (N-1)th candidate reference block in the first reference picture; determining the Nth predictor based on the first reference picture of the current block as being equal to the sum of (i) the (N-1)th predictor associated with the first reference picture of the current block, (ii) a product of the first component of the gradient value of the (N-1)th predictor and the first component of the displacement, and (iii) a product of the second component of the gradient value of the (N-1)th predictor and the second component of the displacement; 11. The method of claim 10, further comprising:

12. The plurality of candidate reference block pairs include: (i) N is equal to the upper limit; and (ii) a displacement between the Nth candidate reference block in the first reference picture and an (N+1)th candidate reference block in the first reference picture is 0; The method of claim 11 , further comprising:

13. The method of claim 11 , wherein the delta motion vector of each sub-block in the current block is less than or equal to a threshold.

14. The method of claim 1 , wherein the prediction information includes a flag indicating whether the current block is predicted based on the affine bilateral matching using the affine model.

15. determining a list of candidate motion vectors for the affine model associated with the current block, the list of candidate motion vectors including the control point motion vectors associated with the reference block pair; The method of claim 1, further comprising:

16. Apparatus comprising processing circuitry configured to perform the method of any one of claims 1 to 15.

17. A computer program for causing a computer to carry out the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Improvements to frame rate upconversion coding mode

    JP2019534622A

  • Video picture prediction method and apparatus

    JP2022511637A

  • Coding device, decoding device, coding method, and decoding method

    WO2019049912A1

  • Symmetric merge mode motion vector coding

    WO2020185925A1

Cited By

  • Using Affine Models in Affine Bilateral Matching

    JP2025513971A