Using Affine Models in Affine Bilateral Matching

The use of affine models for motion prediction in video coding addresses the challenge of encoding complex motion patterns, resulting in improved coding efficiency and reduced bitrates.

JP2025513971APending Publication Date: 2025-05-02TENCENT AMERICA LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024515115
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-29
Filing Date
2022-09-30
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

Current video coding techniques face challenges in efficiently predicting and encoding complex motion patterns in videos, particularly with the increasing number of possible directions for intra-prediction, which can lead to increased bitrates and reduced compression efficiency.

Method used

The proposed method employs affine models, specifically three-parameter and four-parameter zooming models, and three-parameter and four-parameter rotation models, for affine bilateral matching. These models are used to predict the motion of blocks between reference pictures, allowing for more accurate and efficient encoding of complex motion patterns.

Benefits of technology

By using affine models for motion prediction, the method achieves improved coding efficiency and reduced bitrates, effectively handling complex motion patterns that previous techniques struggled with, while maintaining high video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025513971000001_ABST
    Figure 2025513971000001_ABST
Patent Text Reader

Abstract

A first affine parameter and a second affine parameter of each of the one or more affine models associated with the current block are determined. The first affine parameter of each affine model is associated with a first motion vector of a first reference picture of the current block. The second affine parameter of each affine model is associated with a second motion vector of a second reference picture of the current block. The first affine parameter and the second affine parameter of each affine model have opposite signs. A control point motion vector (CPMV) of the affine motion of the current block is determined by performing affine bilateral matching based on the determined first affine parameter and the determined second affine parameter of the one or more affine models. The current block is reconstructed based on the CPMV of the affine motion of the current block.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Patent Application No. 17 / 955,927, entitled "AFFINE MODELS USE IN AFFINE BILATERAL MATCHING," filed September 29, 2022, which claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 334,558, entitled "Affine Models Use in Affine Bilateral Matching," filed April 25, 2022. The disclosures of the prior applications are incorporated herein by reference in their entireties.

[0002] This disclosure describes embodiments that relate generally to video coding. [Background technology]

[0003] The discussion of the background art provided herein is intended to generally present the context of the present disclosure. The work of the presently named inventors is not expressly or impliedly admitted as prior art to the present disclosure, to the extent that that work is described in this background section, as well as aspects of the description that are not admitted as prior art at the time of filing.

[0004] Uncompressed digital images and / or videos may include a sequence of pictures, each having spatial dimensions of, for example, 1920x1080 luminance samples and associated chrominance samples. The sequence of pictures may have a fixed or variable picture rate (also informally known as frame rate), for example, 60 pictures per second or 60 Hz. Uncompressed images and / or videos have specific bitrate requirements. For example, 1080p60 4:2:0 video with 8 bits per sample (1920x1080 luminance sample resolution at 60 Hz frame rate) requires a bandwidth approaching 1.5 Gbit / s. One hour of such video requires more than 600 gigabytes of storage space.

[0005] One objective of image and / or video coding and decoding may be the reduction of redundancy in the input image and / or video signal through compression. Compression may help reduce the aforementioned bandwidth and / or storage space requirements, in some cases by more than one order of magnitude. The description herein uses video encoding / decoding as an illustrative example, but the same techniques may be applied to image encoding / decoding in a similar manner without departing from the spirit of this disclosure. Both lossless and lossy compression, as well as combinations thereof, may be used. Lossless compression refers to techniques where an exact copy of the original signal can be reconstructed from a compressed original signal. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough to make the reconstructed signal useful for the intended application. For video, lossy compression has been widely adopted. The amount of distortion tolerated depends on the application, e.g., users of a particular consumer streaming application may tolerate higher distortion than users of a television distribution application. The achievable compression ratio may reflect that a higher tolerable / acceptable distortion can result in a higher compression ratio.

[0006] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transform processing, quantization, and entropy coding.

[0007] Video codec techniques may include a technique known as intra-coding. In intra-coding, sample values ​​are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into blocks of samples. When all blocks of samples are coded in intra mode, the picture may be an intra picture. Intra pictures and their derivatives, such as independent decoder refresh pictures, may be used to reset the decoder state and thus may be used as the first picture in a coded video bitstream and video session or as a still image. Samples of an intra block undergo a transform and the transform coefficients may be quantized before entropy coding. Intra prediction may be a technique that minimizes sample values ​​in the pre-transform domain. In some cases, the smaller the DC value after the transform and the smaller the AC coefficients, the fewer bits are needed at a given quantization step size to represent the block after entropy coding.

[0008] For example, conventional intra-coding used in MPEG-2 generation coding techniques does not use intra-prediction. However, some newer video compression techniques include techniques that attempt to perform prediction based on surrounding sample data and / or metadata obtained during encoding and / or decoding of a block of data. Such techniques are hereinafter referred to as "intra-prediction" techniques. It should be noted that, at least in some cases, intra-prediction uses reference data only from the current picture being reconstructed, and not from reference pictures.

[0009] Intra prediction may take many different forms. When more than one of such techniques may be used in a given video coding technique, the particular technique in use may be coded as a particular intra prediction mode that uses that particular technique. In certain cases, an intra prediction mode may have sub-modes and / or parameters that may be coded individually or included in a mode codeword that defines the prediction mode being used. Which codeword to use for a given mode, sub-mode, and / or parameter combination may affect the coding efficiency gain from intra prediction and may also affect the entropy coding technique used to convert the codeword into a bitstream.

[0010] A specific mode of intra prediction was introduced with H.264, elaborated in H.265, and further elaborated in newer coding techniques such as the Joint Search Model (JEM), Generic Video Coding (VVC), and Benchmark Set (BMS). A predictor block may be formed using neighboring sample values ​​of already available samples. The sample values ​​of the neighboring samples are copied to the predictor block according to the direction. The reference to the direction in use may be coded in the bitstream or may itself be predicted.

[0011] Referring to FIG. 1A, at the bottom right, a subset of 9 known predictor directions is shown from the 33 possible predictor directions defined in H.265 (corresponding to the 33 angle modes of the 35 intra modes). The point (101) where the arrows converge represents the sample being predicted. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples to the upper right at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples to the lower left of sample (101) at an angle of 22.5 degrees from the horizontal.

[0012] Still referring to FIG. 1A, at the top left, a square block (104) of 4×4 samples (indicated by a dashed bold line) is depicted. The square block (104) includes 16 samples, each labeled with "S", its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample (from the top) in the Y dimension and the first sample (from the left) in the X dimension. Similarly, sample S44 is the fourth sample in the block (104) in both the Y and X dimensions. Since the block is 4×4 samples in size, S44 is at the bottom right. Further shown are reference samples that follow a similar numbering scheme. The reference sample is labeled with R, its Y position (e.g., row index) and X position (column index) relative to the block (104). In both H.264 and H.265, the predicted samples are close to the block being reconstructed, and therefore negative values ​​do not need to be used.

[0013] Intra-picture prediction can work by copying reference sample values ​​from neighboring samples indicated by the signaled prediction direction. For example, assume that the coded video bitstream includes signaling for this block indicating a prediction direction that coincides with the arrow (102), i.e., the sample is predicted from the sample to the upper right at an angle of 45 degrees from the horizontal. In that case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.

[0014] In certain cases, especially when the orientation is not evenly divisible by 45 degrees, the values ​​of multiple reference samples may be combined, for example via interpolation, to calculate the reference sample.

[0015] The number of possible directions is increasing as video coding techniques develop. In H.264 (2003), nine different directions could be represented. That increased to 33 in H.265 (2013). Currently, JEM / VVC / BMS can support up to 65 directions. Experiments are being carried out to identify the most likely directions, using certain techniques in entropy coding to represent those more likely directions with fewer bits, accepting certain penalties for less likely directions. Furthermore, the direction itself can possibly be predicted from neighboring directions used in neighboring, already decoded blocks.

[0016] FIG. 1B shows a schematic diagram (110) of 65 intra prediction directions with JEM to illustrate the increasing number of prediction directions over time.

[0017] The mapping of intra-prediction direction bits representing directions in the coded video bitstream may vary from one video coding technique to another. Such mappings may range, for example, from simple direct mappings to codewords, complex adaptive schemes with most probable modes, and similar techniques. However, in most cases, there may be certain directions that are statistically less likely to occur in the video content than certain other directions. Since the goal of video compression is to reduce redundancy, these less likely directions are represented by a larger number of bits than more likely directions in well-performing video coding techniques.

[0018] Image and / or video coding and decoding may be performed using inter-picture prediction with motion compensation. Motion compensation may be a lossy compression technique and may refer to a technique in which blocks of sample data from a previously reconstructed picture or part thereof (reference picture) are used for prediction of a newly reconstructed picture or picture part after being spatially shifted in a direction indicated by a motion vector (hereinafter MV). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions X and Y, or three dimensions, the third dimension being an indication of the reference picture in use (the latter may indirectly be a temporal dimension).

[0019] In some video compression techniques, the MV applicable to a particular area of ​​sample data may be predicted from other MVs, e.g., from an MV related to another area of ​​sample data that is spatially adjacent to the area being reconstructed and precedes that MV in decoding order. Doing so may substantially reduce the amount of data required to code the MV, thereby removing redundancy and increasing compression. MV prediction may work effectively, for example, when coding an input video signal derived from a camera (known as natural video), because there is a statistical likelihood that areas larger than the area to which a single MV is applicable move in similar directions and therefore, in some cases, may be predicted using similar motion vectors derived from MVs of nearby areas. As a result, the MV found for a given region will be similar or the same as the MV predicted from the surrounding MVs, and after entropy coding, it may be represented with fewer bits than would be used when coding the MV directly. In some cases, MV prediction may be an example of lossless compression of a signal (i.e., an MV) derived from the original signal (i.e., a sample stream). In other cases, the MV prediction itself may be lossy, for example, due to rounding errors when computing a predictor from multiple surrounding MVs.

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms offered by H.265, one that will be described with reference to Figure 2 is a technique hereafter referred to as "spatial merging".

[0021] Referring to Figure 2, a current block (201) contains samples that were discovered by the encoder during the motion search process to be predictable from a previous block of the same size but spatially shifted. Instead of coding its MV directly, the MV may be derived from metadata related to one or more reference pictures, e.g., from the nearest reference picture (in decoding order), using MVs related to any one of five surrounding samples, denoted A0, A1, and B0, B1, B2 (202-206, respectively). In H.265, MV prediction may use predictors from the same reference picture that neighboring blocks are using. Summary of the Invention [Means for solving the problem]

[0022] Aspects of the present disclosure provide a method and apparatus for video encoding / decoding. In some examples, the apparatus for video decoding includes a processing circuit.

[0023] According to one aspect of the present disclosure, a method of video decoding implemented in a video decoder is provided. In the method, prediction information of a current block in a current picture may be received from a coded video bitstream. The prediction information may indicate that the current block is predicted based on affine bilateral matching, in which an affine motion of the current block is derived based on reference blocks in a first reference picture and a second reference picture associated with the current picture. The current picture may be disposed between the first reference picture and the second reference picture. A first affine parameter and a second affine parameter for each of one or more affine models associated with the affine bilateral matching may be determined. The first affine parameter of each affine model may be associated with a first motion vector of the first reference picture. The second affine parameter of each affine model may be associated with a second motion vector of the second reference picture. The first affine parameter and the second affine parameter may have opposite signs. A control point motion vector (CPMV) of the affine motion of the current block may be determined by performing affine bilateral matching based on the determined first affine parameters and the determined second affine parameters of the one or more affine models. The current block may be reconstructed based on the CPMV of the affine motion of the current block.

[0024] To determine the CPMV, a first CPMV of the affine motion of the current block may be determined based on a first affine model of the one or more affine models. A refined CPMV of the affine motion of the current block may be determined based on the first CPMV and a second affine model of the one or more affine models. The second affine model may be the same as or different from the first affine model. Thus, the current block may be reconstructed based on the refined CPMV of the affine motion of the current block.

[0025] The one or more affine models may include a three-parameter zooming model. The three-parameter zooming model may further include a first component of a first motion vector of a sample in the current block associated with a first reference picture along a first direction. The first component of the first motion vector may be equal to a first component of a translation coefficient along the first direction and a product of a position of the sample along the first direction and a zoom factor. The three-parameter zooming model may include a second component of a first motion vector of a sample in the current block associated with a first reference picture along a second direction. The second component of the first motion vector may be equal to a second component of a translation coefficient along the second direction and a product of a position of the sample along the second direction and a zoom factor, the second direction may be perpendicular to the first direction. The three-parameter zooming model may include a first component of a second motion vector of a sample in the current block associated with a second reference picture along the first direction. The first component of the second motion vector may be equal to a first component of an opposite translation coefficient along the first direction plus a product of the position of the sample along the first direction and the opposite zoom coefficient. The three-parameter zooming model may further include a second component of a second motion vector of the sample in the current block associated with the second reference picture along the second direction. The second component of the second motion vector may be equal to a second component of an opposite translation coefficient along the second direction plus a product of the position of the sample along the second direction and the opposite zoom coefficient.

[0026] The one or more affine models may include a three-parameter rotation model. The three-parameter rotation model may further include a first component of a first motion vector of a sample in the current block associated with a first reference picture along a first direction. The first component of the first motion vector may be equal to a sum of a first component of a translation coefficient along the first direction and a product of a position of the sample along a second direction and a rotation coefficient. The second direction may be perpendicular to the first direction. The three-parameter rotation model may include a second component of a first motion vector of a sample in the current block associated with a first reference picture along a second direction. The second component of the first motion vector may be equal to a sum of a second component of a translation coefficient along the second direction and a product of a position of the sample along the first direction and an inverse rotation coefficient. The three-parameter rotation model may include a first component of a second motion vector of a sample in the current block associated with a second reference picture along the first direction. The first component of the second motion vector may be equal to a first component of an opposite translation coefficient in the first direction plus a product of the position of the sample along the second direction and the inverse rotation coefficient. The three-parameter rotation model may include a second component of a second motion vector of a sample in the current block associated with a second reference picture along the second direction. The second component of the second motion vector may be equal to a second component of an opposite translation coefficient along the second direction plus a product of the position of the sample along the first direction and the rotation coefficient.

[0027] The one or more affine models may include a four-parameter zooming model. The four-parameter zooming model may further include a first component of a first motion vector of a sample in the current block associated with a first reference picture along a first direction. The first component of the first motion vector may be equal to a sum of a first component of a translation coefficient along the first direction and a product of a position of the sample along the first direction and a first component of a zoom factor in the first direction. The four-parameter zooming model may include a second component of a first motion vector of a sample in the current block associated with a first reference picture along a second direction. The second component of the first motion vector may be equal to a sum of a second component of a translation coefficient along the second direction and a product of a position of the sample along the second direction and a second component of a zoom factor along the second direction. The second direction may be perpendicular to the first direction. The four-parameter zooming model may include a first component of a second motion vector of a sample in the current block associated with a second reference picture along the first direction. The first component of the second motion vector may be equal to the sum of a first component of an opposite translation coefficient along the first direction and a product of a position of the sample along the first direction and a first component of an opposite zoom coefficient along the first direction. The four-parameter zooming model may include a second component of a second motion vector of a sample in the current block associated with a second reference picture along the second direction. The second component of the second motion vector may be equal to the sum of a second component of an opposite translation coefficient along the second direction and a product of a position of the sample along the second direction and a second component of an opposite zoom coefficient along the second direction.

[0028] The one or more affine models may include a four-parameter rotation model. The four-parameter rotation model may further include a first component of a first motion vector of a sample in the current block associated with a first reference picture along a first direction. The first component of the first motion vector may be equal to the sum of (i) a first component of a translation coefficient along the first direction, (ii) a product of a position of the sample along the first direction and a first coefficient associated with the rotation coefficient, and (iii) a product of a position of the sample along a second direction and a second coefficient associated with the rotation coefficient. The first direction may be perpendicular to the second direction. The four-parameter rotation model may include a second component of a first motion vector of a sample in the current block associated with a first reference picture along the second direction. The second component of the first motion vector may be equal to the sum of (i) a second component of a translation coefficient along the second direction, (ii) a product of the position of the sample along the first direction and an opposite second coefficient associated with the rotation coefficient, and (iii) a product of the position of the sample along the second direction and a first coefficient associated with the rotation coefficient. The four-parameter rotation model may include a first component of a second motion vector of a sample in the current block associated with a second reference picture along the first direction. The first component of the second motion vector may be equal to the sum of (i) a first component of an opposite translation coefficient along the first direction, (ii) a product of the position of the sample along the first direction and a first coefficient associated with the rotation coefficient, and (iii) a product of the position of the sample along the second direction and an opposite second coefficient associated with the rotation coefficient. The four-parameter rotation model may include a second component of a second motion vector of a sample in the current block associated with a second reference picture along the second direction. A second component of the second motion vector may be equal to the sum of (i) a second component of the opposite translation coefficient along the second direction, (ii) a product of the sample's position along the first direction and a second coefficient associated with the rotation coefficient, and (iii) a product of the sample's position along the second direction and a first coefficient associated with the rotation coefficient.

[0029] To determine a first CPMV of the affine motion of the current block, a first component of an initial predictor associated with a first reference picture of the current block may be determined. A second component of an initial predictor associated with a second reference picture of the current block may be determined. The first component of the first predictor associated with the first reference picture of the current block may be determined based on the first component of the initial predictor associated with the first reference picture of the current block. The second component of the first predictor associated with the second reference picture of the current block may be determined based on the second component of the initial predictor associated with the second reference picture of the current block. The first CPMV of the affine motion may be determined based on the first component of the first predictor associated with the first reference picture and the second component of the first predictor associated with the second reference picture.

[0030] In some embodiments, a first component of the initial predictor associated with the first reference picture of the current block may be determined based on one of a merge candidate, an advanced motion vector prediction (AMPV) candidate, and an affine merge candidate.

[0031] To determine a first component of the first predictor associated with the first reference picture of the current block, a first component of a gradient value of the first component of the initial predictor along a first direction may be determined. A second component of a gradient value of the first component of the initial predictor along a second direction may be determined. The second direction may be perpendicular to the first direction. A first component of the initial predictor along the first direction and a first component of a displacement associated with the first component of the first predictor may be determined according to a first affine model of the one or more affine models. A first component of the initial predictor along the second direction and a second component of a displacement associated with the first component of the first predictor may be determined according to a first affine model of the one or more affine models. The first component of the first predictor may be determined based on a sum of (i) the first component of the initial predictor, (ii) a product of the first component of the gradient value and the first component of the displacement, and (iii) a product of the second component of the gradient value and the second component of the displacement.

[0032] To determine a refined CPMV of the affine motion, the refined CPMV of the affine motion may be determined based on an Nth predictor associated with the first reference picture and an Nth predictor associated with the second reference picture in response to one of: (i) N is equal to an upper limit value of the iterative process; and (ii) a displacement based on the Nth predictor associated with the first reference picture and the (N+1)th predictor associated with the first reference picture is zero.

[0033] To determine the defined CPMV of the affine motion based on the N-th predictor associated with the first reference picture of the current block, a first component of a gradient value of the (N-1)-th predictor associated with the first reference picture of the current block along a first direction may be determined. A second component of a gradient value of the (N-1)-th predictor associated with the first reference picture of the current block along a second direction may be determined. A first component of a displacement associated with the N-th predictor associated with the first reference picture along the first direction and the (N-1)-th predictor associated with the first reference picture may be determined according to a second affine model of the one or more affine models. A second component of a displacement associated with the N-th predictor associated with the first reference picture along the second direction and the (N-1)-th predictor associated with the first reference picture may be determined according to a second affine model of the one or more affine models. The Nth predictor associated with the first reference picture of the current block may then be determined based on the sum of (i) the (N-1)th predictor associated with the first reference picture of the current block, (ii) the product of a first component of the gradient value of the (N-1)th predictor and a first component of the displacement, and (iii) the product of a second component of the gradient value of the (N-1)th predictor and a second component of the displacement.

[0034] In some embodiments, to determine the first CPMV of the affine motion, the first CPMV of the affine motion may be determined based on an Nth predictor associated with the first reference picture and an Nth predictor associated with the second reference picture.

[0035] In some embodiments, to determine the defined CPMV of the affine motion, the defined CPMV of the affine motion may be determined based on (i) an Mth predictor associated with the first reference picture derived from an Nth predictor associated with the first reference picture, and (ii) an Mth predictor associated with the second reference picture derived from an Nth predictor associated with the second reference picture, in response to one of: (i) M being equal to an upper limit value of the iterative process, and (ii) a displacement based on the Mth predictor associated with the first reference picture and the (M+1)th predictor associated with the first reference picture is zero.

[0036] In some embodiments, the first and second affine parameters for each of the one or more affine models may be determined based on a syntax element included in the prediction information. The syntax element may be included in one of a sequence parameter set, a picture parameter set, and a slice header.

[0037] According to another aspect of the present disclosure, an apparatus is provided, the apparatus including a processing circuit, the processing circuit may be configured to perform any of the methods for video encoding / decoding.

[0038] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to perform any of the methods for video encoding / decoding.

[0039] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Brief description of the drawings]

[0040] [Figure 1A] FIG. 2 is a schematic diagram of an example subset of intra-prediction modes. [Figure 1B] FIG. 2 is a diagram of an example intra-prediction direction. [Diagram 2] FIG. 2 is a schematic diagram of a current block and its surrounding spatial merge candidates in one example. [Diagram 3] FIG. 3 is a schematic diagram of a simplified block diagram of a communication system (300) according to one embodiment. [Figure 4] FIG. 4 is a schematic diagram of a simplified block diagram of a communication system (400) according to one embodiment. [Diagram 5] FIG. 2 is a schematic diagram of a simplified block diagram of a decoder according to one embodiment. [Figure 6] FIG. 2 is a schematic diagram of a simplified block diagram of an encoder according to one embodiment. [Figure 7] 4 shows a block diagram of an encoder according to another embodiment; [Figure 8] 4 shows a block diagram of a decoder according to another embodiment; [Figure 9A] 4 shows a schematic diagram of a four-parameter affine model according to another embodiment. [Figure 9B] 1 shows a schematic diagram of a six-parameter affine model according to another embodiment. [Figure 10] FIG. 4 is a schematic diagram of an affine motion vector field associated with sub-blocks within a block according to another embodiment; [Figure 11] 13 shows a schematic diagram of exemplary locations of spatial merging candidates according to another embodiment; [Figure 12] 4 shows a schematic diagram of control point motion vector inheritance according to another embodiment; [Figure 13] 13 shows a schematic diagram of candidate positions for constructing an affine merge mode according to another embodiment; [Figure 14] 1 shows a schematic diagram of Prediction Refinement Using Optical Flow (PROF) according to another embodiment. [Figure 15]4 shows a schematic diagram of an affine motion estimation process according to another embodiment; [Figure 16] 4 shows a flowchart of an affine motion estimation search according to another embodiment. [Figure 17] 1 shows a schematic diagram of an extended coding unit (CU) region for bidirectional optical flow (BDOF) according to another embodiment. [Figure 18] 4 shows a schematic diagram of a decoding side motion vector refinement according to another embodiment; [Figure 19] 1 shows a flowchart outlining an exemplary decoding process according to some embodiments of the present disclosure. [Figure 20] 1 shows a flowchart outlining an exemplary encoding process according to some embodiments of the present disclosure. [Figure 21] FIG. 1 is a schematic diagram of a computer system according to one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0041] FIG. 3 illustrates an example block diagram of a communication system (300). The communication system (300) includes a plurality of terminal devices capable of communicating with each other, for example, via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) perform unidirectional transmission of data. For example, the terminal device (310) may code video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to the other terminal device (320) via the network (350). The encoded video data may be transmitted in the form of one or more coded video bitstreams. The terminal device (320) may receive the coded video data from the network (350), decode the coded video data to reconstruct the video pictures, and display the video pictures according to the reconstructed video data. Unidirectional data transmission may be common, such as in media serving applications.

[0042] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) implementing bidirectional transmission of coded video data, for example during a video conference. For the bidirectional transmission of data, in one example, each of the terminal devices (330) and (340) may code video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) over the network (350). Each of the terminal devices (330) and (340) may also receive coded video data transmitted by the other of the terminal devices (330) and (340), decode the coded video data to reconstruct the video pictures, and display the video pictures on an accessible display device according to the reconstructed video data.

[0043] In the example of FIG. 3, terminal devices (310), (320), (330), and (340) are shown as a server, a personal computer, and a smartphone, respectively, although the principles of the present disclosure may not be so limited. Embodiments of the present disclosure apply to laptop computers, tablet computers, media players, and / or dedicated videoconferencing equipment. Network (350) represents any number of networks that convey coded video data between terminal devices (310), (320), (330), and (340), including, for example, wired and / or wireless communication networks. Communication network (350) may exchange data over circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of network (350) may not be important to the operation of the present disclosure unless described herein below.

[0044] 4 shows a video encoder and video decoder in a streaming environment as an example of an application for the disclosed subject matter. The disclosed subject matter is equally applicable to other video-enabled applications including, for example, videoconferencing, digital TV, streaming services, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0045] The streaming system may include a capture subsystem (413) that may include a video source (401), such as a digital camera, that creates a stream of uncompressed video pictures (402). In one example, the stream of video pictures (402) includes samples taken by a digital camera. The stream of video pictures (402), shown as a thick line to emphasize the high data volume compared to the encoded video data (404) (or coded video bitstream), may be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (404) (or coded video bitstream), shown as a thin line to emphasize the lower data volume compared to the stream of video pictures (402), may be stored on a streaming server (405) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include, for example, a video decoder (410) within an electronic device (430). The video decoder (410) decodes the incoming copy of the encoded video data (407) and creates an outgoing stream of video pictures (411) that can be rendered on a display (412) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., a video bitstream) can be encoded according to a particular video coding / compression standard. Examples of those standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC).The disclosed subject matter may be used in the context of a VVC.

[0046] It should be noted that the electronic devices (420) and (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).

[0047] 5 shows an example block diagram of a video decoder (510). The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). The video decoder (510) can be used in place of the video decoder (410) of the example of FIG. 4.

[0048] The receiver (531) may receive one or more coded video sequences to be decoded by the video decoder (510). In one embodiment, one coded video sequence is received at a time, and the decoding of each coded video sequence is independent of the decoding of the other coded video sequences. The coded video sequences may be received from a channel (501), which may be a hardware / software link to a storage device that stores the coded video data. The receiver (531) may receive the coded video data along with other data, e.g., coded audio data and / or auxiliary data streams, which may be forwarded to their respective using entities (not shown). The receiver (531) may separate the coded video sequences from the other data. To combat network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter, "parser (520)"). In certain applications, the buffer memory (515) is part of the video decoder (510). In other applications, the buffer memory (515) can be external to the video decoder (510) (not shown). In still other applications, there can be buffer memories (not shown) external to the video decoder (510), e.g., to combat network jitter, plus other buffer memories (515) internal to the video decoder (510), e.g., to handle playout timing. When the receiver (531) is receiving data from a storage / forwarding device of sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory (515) may not be needed or can be small. For use over a best-effort packet network such as the Internet, the buffer memory (515) may be needed and can be relatively large, advantageously adaptively sized, and implemented at least in part within an operating system or similar element (not shown) external to the video decoder (510).

[0049] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the coded video sequence. The categories of symbols include information used to manage the operation of the video decoder (510) and, in some cases, information for controlling a rendering device such as a rendering device (512) (e.g., a display screen) that is not an integral part of the electronic device (530) but may be coupled to the electronic device (530) as shown in FIG. 5. The control information for the rendering device(s) may be in the form of a supplemental enhancement information (SEI) message or a video usability information (VUI) parameter set fragment (not shown). The parser (520) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles including variable length coding with or without context dependency, Huffman coding, arithmetic coding, etc. The parser (520) may extract, from the coded video sequence, a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on the at least one parameter corresponding to the group. The subgroups may include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (520) may also extract information from the coded video sequence, such as transform coefficients, quantization parameter values, motion vectors, etc.

[0050] The parser (520) may generate symbols (521) by performing an entropy decoding / parsing operation on the video sequence received from the buffer memory (515).

[0051] The reconstruction of the symbols (521) can involve several different units, depending on the type of coded video picture or portion thereof (inter-picture and intra-picture, inter-block and intra-block, etc.), as well as other factors. Which units are involved and how can be controlled by subgroup control information parsed by the parser (520) from the coded video sequence. The flow of such subgroup control information between the parser (520) and the following units is not depicted for the sake of clarity.

[0052] Beyond the functional blocks already mentioned, the video decoder (510) may be conceptually subdivided into a number of functional units, as described below. In an actual implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate:

[0053] The first unit is a scalar / inverse transform unit (551), which receives quantized transform coefficients as well as control information from the parser (520) including which transform to use, block size, quantization coefficients, quantization scaling matrices, etc. as symbol(s) (521). The scalar / inverse transform unit (551) can output blocks comprising sample values ​​that can be input to an aggregator (555).

[0054] In some cases, the output samples of the scaler / inverse transform unit (551) may relate to intra-coded blocks. Intra-coded blocks are blocks that do not use prediction information from a previously reconstructed picture, but may use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates a block of the same size and shape as the block being reconstructed using surrounding already reconstructed information fetched from a current picture buffer (558). The current picture buffer (558) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. The aggregator (555) may add, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).

[0055] In other cases, the output samples of the scalar / inverse transform unit (551) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion compensated prediction unit (553) may access the reference picture memory (557) to fetch samples used for prediction. After motion compensating the fetched samples according to the symbols (521) related to the block, these samples may be added by the aggregator (555) to the output of the scalar / inverse transform unit (551) (in this case referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (557) from which the motion compensated prediction unit (553) fetches the prediction samples may be controlled by motion vectors available to the motion compensated prediction unit (553), for example, in the form of symbols (521) that may have X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​fetched from the reference picture memory (557) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, and the like.

[0056] The output samples of the aggregator (555) may be subjected to various loop filtering techniques in the loop filter unit (556). Video compression techniques may include in-loop filter techniques controlled by parameters contained in the coded video sequence (also called the coded video bitstream) and available to the loop filter unit (556) as symbols (521) from the parser (520). Video compression may also be performed in response to meta-information obtained during decoding of a previous portion (in decoding order) of the coded picture or coded video sequence, or in response to previously reconstructed and loop filtered sample values.

[0057] The output of the loop filter unit (556) may be a sample stream that can be output to a rendering device (512) and also stored in a reference picture memory (557) for use in future inter-picture prediction.

[0058] Once a particular coded picture is fully reconstructed, it may be used as a reference picture for future prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) may become part of the reference picture memory (557), and a new current picture buffer may be reallocated before beginning reconstruction of the next coded picture.

[0059] The video decoder (510) may perform decoding operations according to a given video compression technique or standard, such as ITU-T Recommendation H.265. The coded video sequence may comply with the syntax specified by the video compression technique or standard being used, in the sense that the coded video sequence adheres to both the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. Specifically, the profile may select certain tools from among all tools available in the video compression technique or standard as tools that are only available to them under the profile. Also, compliance may require that the complexity of the coded video sequence be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may be further limited in some cases by a hypothetical reference decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.

[0060] In one embodiment, the receiver (531) can receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence(s). The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0061] 6 shows an example block diagram of a video encoder (603). The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmitting circuit). The video encoder (603) can be used in place of the video encoder (403) of the example of FIG.

[0062] The video encoder (603) may receive video samples from a video source (601) (which is not part of the electronic device (620) in the example of FIG. 6) that may capture the video image(s) to be coded by the video encoder (603). In other examples, the video source (601) is part of the electronic device (620).

[0063] The video source (601) may provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream that may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...) and suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (601) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a number of separate pictures that give motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each pixel may contain one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0064] According to one embodiment, the video encoder (603) may code and compress pictures of a source video sequence into a coded video sequence (643) in real time or under any other time constraints required. Enforcing an appropriate coding rate is one function of the controller (650). In some embodiments, the controller (650) controls and is operatively coupled to other functional units described below. This coupling is not depicted for clarity. Parameters set by the controller (650) may include rate control related parameters (picture skip, quantizer, lambda value for rate distortion optimization techniques, ...), picture size, Group of Pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other appropriate functions for the video encoder (603) optimized for a particular system design.

[0065] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As an oversimplified explanation, in one example, the coding loop can include a source coder (630) (responsible for creating symbols, such as a symbol stream, based on an input picture to be coded and reference picture(s)) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a similar manner as the (remote) decoder would create them. The reconstructed sample stream (sample data) is input to a reference picture memory (634). Since the decoding of the symbol stream results in bit-exact results regardless of the location of the decoder (local or remote), the contents of the reference picture memory (634) are also bit-exact between the local and remote encoders. In other words, the predictive part of the encoder "sees" exactly the same sample values ​​as the decoder would "see" when using prediction during decoding as reference picture samples. This basic principle of reference picture synchrony (and the drift that occurs when synchrony cannot be maintained, for example due to channel errors) is also used in several related techniques.

[0066] The operation of the "local" decoder (633) may be the same as the operation of a "remote" decoder, such as the video decoder (510), already described in detail above in conjunction with Figure 5. Referring also briefly to Figure 5, however, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder (645) and parser (520) may be lossless, the entropy decoding portion of the video decoder (510), including the buffer memory (515), and the parser (520), may not be fully implemented in the local decoder (633).

[0067] In one embodiment, the decoder techniques, except for parsing / entropy decoding, present in the decoder, are present in the corresponding encoder in the same or substantially the same functional form. Thus, the disclosed subject matter focuses on the operation of the decoder. The description of the encoder techniques can be omitted, since they are the inverse of the decoder techniques described generically. In certain areas, more detailed descriptions are provided below.

[0068] In some examples, during operation, the source coder (630) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as “reference pictures.” In this manner, the coding engine (632) codes differences between pixel blocks of the input picture and pixel blocks of reference picture(s) that may be selected as predictive reference(s) to the input picture.

[0069] The local video decoder (633) may decode the coded video data of pictures that may be designated as reference pictures based on the symbols created by the source coder (630). The operation of the coding engine (632) may advantageously be a lossy process. When the coded video data may be decoded in a video decoder (not shown in FIG. 6), the reconstructed video sequence may usually be a copy of the source video sequence with some errors. The local video decoder (633) may replicate the decoding process that may be performed by the video decoder on the reference pictures and store the reconstructed reference pictures in the reference picture memory (634). In this way, the video encoder (603) may locally store copies of reconstructed reference pictures that have common content as reconstructed reference pictures that would be obtained by the far-end video decoder (without transmission errors).

[0070] The predictor (635) may perform a predictive search for the coding engine (632). That is, for a new picture to be coded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc., that can serve as suitable predictive references for the new picture. The predictor (635) may operate on sample blocks for each pixel block to find a suitable predictive reference. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have predictive references drawn from multiple reference pictures stored in the reference picture memory (634).

[0071] The controller (650) may manage the coding operations of the source coder (630), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0072] The output of all the aforementioned functional units may undergo entropy coding in an entropy coder (645), which converts the symbols produced by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0073] The transmitter (640) may buffer the coded video sequence(s) created by the entropy coder (645) to prepare them for transmission over a communication channel (660), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (640) may merge the coded video data from the video encoder (603) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0074] The controller (650) may manage the operation of the video encoder (603). During coding, the controller (650) may assign a particular coded picture type to each coded picture, which may affect the coding technique that may be applied to the respective picture. For example, pictures are often assigned as one of the following picture types:

[0075] An intra picture (I-picture) may be one that can be coded and decoded without using any other picture in a sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh ("IDR") pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.

[0076] A predictive picture (P picture) may be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict the sample values ​​of each block.

[0077] Bidirectionally predicted pictures (B-pictures) may be coded and decoded using intra- or inter-prediction, which uses up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predicted pictures may use more than two reference pictures and associated metadata for the reconstruction of a block.

[0078] A source picture may generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and coded block by block. A block may be predictively coded with reference to other (already coded) blocks determined by the coding assignment applied to the block's respective picture. For example, a block of an I picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial or intra prediction). A pixel block of a P picture may be predictively coded with reference to one previously coded reference picture, via spatial prediction, or via temporal prediction. A block of a B picture may be predictively coded with reference to one or two previously coded reference pictures, via spatial prediction, or via temporal prediction.

[0079] The video encoder (603) may perform coding operations in accordance with a given video coding technique or standard, such as ITU-T Recommendation H.265. In its operations, the video encoder (603) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified in the video coding technique or standard being used.

[0080] In one embodiment, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0081] A video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) uses spatial correlation within a given picture, while inter-picture prediction uses correlation (temporal or other) between pictures. In one example, a particular picture being coded / decoded, called a current picture, is divided into blocks. When a block in the current picture is similar to a reference block in a reference picture that was previously coded and is still buffered in the video, the block in the current picture may be coded by a vector called a motion vector. A motion vector points to a reference block in a reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0082] In some embodiments, bi-prediction techniques may be used for inter-picture prediction. According to bi-prediction techniques, two reference pictures, such as a first reference picture and a second reference picture, are used, both of which are before the decoding order of the current picture in the video (but the display order may be past and future, respectively). A block in the current picture may be coded by a first motion vector pointing to a first reference block in the first reference picture and by a second motion vector pointing to a second reference block in the second reference picture. A block may be predicted by a combination of the first reference block and the second reference block.

[0083] Furthermore, to improve coding efficiency, merge mode techniques can be used for inter-picture prediction.

[0084] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed on a block-by-block basis. For example, according to the HEVC standard, a picture in a sequence of video pictures is divided into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, 16×16 pixels, etc. In general, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU can be recursively quad-tree partitioned into one or more coding units (CUs). For example, a CTU of 64×64 pixels can be partitioned into one CU of 64×64 pixels, or into four CUs of 32×32 pixels, or into 16 CUs of 16×16 pixels. In one example, each CU is analyzed to determine the prediction type of the CU, such as an inter prediction type or an intra prediction type. A CU is divided into one or more prediction units (PUs) according to temporal predictability and / or spatial predictability. Generally, each PU includes one luma prediction block (PB) and two chroma PBs. In one embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Taking a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values ​​(e.g., luma values) of 8×8 pixels, 16×16 pixels, 8×16 pixels, and 16×8 pixels, etc.

[0085] 7 shows an example diagram of a video encoder (703). The video encoder (703) is configured to receive a processed block of sample values ​​(e.g., a predictive block) in a current video picture in a sequence of video pictures and to encode the processed block into a coded picture that is part of a coded video sequence. In one example, the video encoder (703) is used in place of the video encoder (403) in the example of FIG. 4.

[0086] In an HEVC example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as a predictive block of 8×8 samples. The video encoder (703) determines whether the processing block is best coded using intra-mode, inter-mode, or bi-predictive mode, for example using rate-distortion optimization. When the processing block is to be coded in intra-mode, the video encoder (703) may encode the processing block into a coded picture using intra-prediction techniques, and when the processing block is to be coded in inter-mode or bi-predictive mode, the video encoder (703) may encode the processing block into a coded picture using inter-prediction techniques or bi-prediction techniques, respectively. In certain video coding techniques, the merge mode may be an inter-picture prediction sub-mode in which a motion vector is derived from one or more motion vector predictors without the aid of coded motion vector components outside the predictors. In certain other video coding techniques, there may be motion vector components applicable to the current block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the mode of the processing blocks.

[0087] In the example of FIG. 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725), coupled to each other as shown in FIG. 7.

[0088] The inter-encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block to one or more reference blocks in reference pictures (e.g., blocks in previous and subsequent pictures), generate inter-prediction information (e.g., inter-coding techniques, motion vectors, description of redundant information through merge mode information), and calculate an inter-prediction result (e.g., a predictive block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture that is decoded based on the encoded video information.

[0089] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), possibly compare the block with already coded blocks in the same picture, generate transformed quantized coefficients, and possibly also generate intra prediction information (e.g., intra prediction direction information according to one or more intra encoding techniques). In one example, the intra encoder (722) also calculates intra prediction results (e.g., prediction blocks) based on the intra prediction information and reference blocks in the same picture.

[0090] The generic controller (721) is configured to determine generic control data and control other components of the video encoder (703) based on the generic control data. In one example, the generic controller (721) determines the mode of the block and provides a control signal to the switch (726) based on the mode. For example, when the mode is an intra mode, the generic controller (721) controls the switch (726) to select an intra mode result to be used by the residual calculator (723) and controls the entropy encoder (725) to select intra prediction information and include the intra prediction information in the bitstream, and when the mode is an inter mode, the generic controller (721) controls the switch (726) to select an inter prediction result to be used by the residual calculator (723) and controls the entropy encoder (725) to select inter prediction information and include the inter prediction information in the bitstream.

[0091] The residual calculator (723) is configured to calculate a difference (residual data) between the received block and a prediction result selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) is configured to operate based on the residual data to encode the residual data to generate transform coefficients. In one example, the residual encoder (724) is configured to transform the residual data from the spatial domain to the frequency domain to generate transform coefficients. The transform coefficients then undergo a quantization process to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be used by the intra-encoder (722) and the inter-encoder (730) as appropriate. For example, the inter-encoder (730) may generate decoded blocks based on the decoded residual data and the inter-prediction information, and the intra-encoder (722) may generate decoded blocks based on the decoded residual data and the intra-prediction information. In some examples, the decoded blocks may be appropriately processed to generate a decoded picture, which may be buffered in a memory circuit (not shown) and used as a reference picture.

[0092] The entropy encoder (725) is configured to format a bitstream to include the encoded block. The entropy encoder (725) is configured to include various information in the bitstream according to an appropriate standard, such as the HEVC standard. In one example, the entropy encoder (725) is configured to include in the bitstream general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information. It should be noted that, according to the disclosed subject matter, when coding a block in a merged sub-mode of either an inter mode or a bi-prediction mode, no residual information is present.

[0093] 8 shows an example diagram of a video decoder (810). The video decoder (810) is configured to receive coded pictures that are part of a coded video sequence and decode the coded pictures to generate reconstructed pictures. In one example, the video decoder (810) is used in place of the video decoder (410) of the example of FIG. 4.

[0094] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter-decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-decoder (872), coupled to each other as shown in FIG.

[0095] The entropy decoder (871) may be configured to reconstruct from the coded picture certain symbols that represent syntax elements of which the coded picture is composed. Such symbols may include, for example, prediction information (e.g., intra-mode, inter-mode, bi-predictive mode, etc., where inter-mode and bi-predictive mode are in merged submode or other submode) that may identify the mode in which the block is coded, as well as certain samples or metadata used for prediction by the intra-decoder (872) or inter-decoder (880), respectively (e.g., intra-predictive information or inter-predictive information, etc.). The symbols may also include, for example, residual information in the form of quantized transform coefficients, etc. In one example, when the prediction mode is an inter-mode or bi-predictive mode, the inter-predictive information is provided to the inter-decoder (880); when the prediction type is an intra-predictive type, the intra-predictive information is provided to the intra-decoder (872). The residual information may undergo inverse quantization and be provided to the residual decoder (873).

[0096] The inter decoder (880) is configured to receive the inter prediction information and to generate inter prediction results based on the inter prediction information.

[0097] The intra decoder (872) is configured to receive the intra prediction information and to generate a prediction result based on the intra prediction information.

[0098] The residual decoder (873) is configured to perform inverse quantization to extract inverse quantized transform coefficients, and to process the inverse quantized transform coefficients to transform the residual information from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (to include quantization parameters (QPs)), which may be provided by the entropy decoder (871) (this may be only a small amount of control information, so a data path is not depicted).

[0099] The reconstruction module (874) is configured to combine, in the spatial domain, the residual information output by the residual decoder (873) and the prediction result (possibly output by the inter prediction module or the intra prediction module) to form a reconstructed block that may become part of a reconstructed picture, which may become part of a reconstructed video. It should be noted that other suitable operations, such as a deblocking operation, may be performed to improve visual quality.

[0100] It should be noted that the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) can be implemented using any suitable technology. In one embodiment, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (603) and the video decoders (410), (510), and (810) can be implemented using one or more processors executing software instructions.

[0101] The present disclosure includes embodiments related to affine models used in affine coding modes, including affine bilateral matching. In one embodiment, the affine bilateral matching can apply an affine model including at least one of a three-parameter zooming model, a four-parameter zooming model, a three-parameter rotation model, or a four-parameter rotation model.

[0102] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). In 2015, the two standardization bodies jointly formed the Joint Video Research Team (JVET) to explore the possibility of developing the next video coding standard beyond HEVC. In October 2017, the two standardization bodies announced a Joint Call for Proposals on Video Compression with Capability beyond HEVC (CfP). By February 15, 2018, 22 CfP responses had been submitted for standard dynamic range (SDR), 12 for high dynamic range (HDR), and 12 for the 360 ​​video category. In April 2018, all received CfP responses were evaluated at the 122 MPEG / 10th JVET meeting. As a result of the meeting, JVET formally launched the standardization process for next-generation video coding beyond HEVC. This new standard was named Versatile Video Coding (VVC) and JVET was renamed the Joint Video Experts Team. In 2020, ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the VVC video coding standard (version 1).

[0103] In inter prediction, motion parameters are needed for each inter predicted coding unit (CU), e.g., to code the VVC features used for inter predicted sample generation. The motion parameters may include motion vectors, reference picture indexes, reference picture list usage indexes, and / or additional information. The motion parameters may be signaled in an explicit or implicit manner. If a CU is coded in skip mode, it may be associated with one PU, and significant residual coefficients, coded motion vector deltas, and / or reference picture indexes may not be needed. If a CU is coded in merge mode, the motion parameters of the CU may be obtained from neighboring CUs. The neighboring CUs may include spatial and temporal candidates, as well as additional schedules (or additional candidates) as introduced in VVC. The merge mode may be applied to any inter predicted CU, not just for skip mode. An alternative to the merge mode is explicit transmission of motion parameters, where the motion vectors, the corresponding reference picture indexes of each reference picture list, the reference picture list usage flag, and / or other necessary information may be explicitly signaled for each CU.

[0104] In VVC, the VVC Test Model (VTM) reference software may include several new refined inter-predictive coding tools, which may include one or more of the following: (1) Enhanced Merge Prediction (2) Merged Motion Vector Differential (MMVD) (3) AMVP mode with symmetric MVD signaling (4) Affine motion compensation prediction (5) Sub-block-based Temporal Motion Vector Prediction (SbTMVP) (6) Adaptive Motion Vector Resolution (AMVR) (7) Motion field storage: 1 / 16 luma sample MV storage and 8x8 motion field compression (8) Bi-prediction with CU-level weights (BCW) (9) Bidirectional Optical Flow (BDOF) (10) Decoder-side Motion Vector Refinement (DMVR) (11) Combined Inter- and Intra-Prediction (CIIP) (12) Geometric Partition Mode (GPM)

[0105] In HEVC, a translational motion model is applied to motion compensated prediction (MCP). In the real world, many types of motion can exist, such as zoom in / out, rotation, perspective motion, and other irregular motion. In VTM, etc., block-based affine transform motion compensated prediction can be applied. Figure 9A shows an affine motion field of a block (902) described by motion information of two control points (4 parameters). Figure 9B shows an affine motion field of a block (904) described by three control point motion vectors (6 parameters).

[0106] As shown in FIG. 9A, in a four-parameter affine motion model, the motion vector at a sample location (x,y) in a block (902) can be derived in equation (1) as follows:

number

number

[0107] As shown in FIG. 9B, in a six-parameter affine motion model, the motion vector at a sample location (x,y) within a block (904) can be derived in equation (3) as follows:

number

number

[0108] As shown in FIG. 10, to simplify the motion compensation prediction, a block-based affine transformation prediction can be applied. To derive a motion vector for each 4×4 luma sub-block, the motion vector (e.g., (1002)) of the center sample of each sub-block (e.g., (1004)) in the current block (1000) can be calculated according to equations (1)-(4) and rounded to 1 / 16 fractional precision. Then, a motion compensation interpolation filter can be applied to generate a prediction of each sub-block using the derived motion vector. The sub-block size of the chroma components can also be set as 4×4. The MV of a 4×4 chroma sub-block can be calculated as the average of the MVs of the four corresponding 4×4 luma sub-blocks.

[0109] In affine merge prediction, the affine merge (AF_MERGE) mode can be applied to CUs whose width and height are both 8 or more. The CPMV of the current CU can be generated based on the motion information of spatially neighboring CUs. Up to five CPMVP candidates can be applied to affine merge prediction, and an index can be signaled to indicate which of the five CPMVP candidates can be used for the current CU. In affine merge prediction, three types of CPMV candidates can be used to form the affine merge candidate list: (1) inherited affine merge candidates extrapolated from the CPMV of neighboring CUs, (2) affine merge candidates constructed with CPMVP derived using the translation MV of neighboring CUs, and (3) zero MV.

[0110] In VTM3, up to two inherited affine candidates can be applied. The two inherited affine candidates can be derived from the affine motion models of the neighboring blocks. For example, one inherited affine candidate can be derived from the left neighboring CU, and the other inherited affine candidate can be derived from the upper neighboring CU. An exemplary candidate block can be shown in FIG. 11. As shown in FIG. 11, for the left predictor (or the left inherited affine candidate), the scanning order can be A0→A1, and for the upper predictor (or the upper inherited affine candidate), the scanning order can be B0→B1→B2. Therefore, only the first available inherited candidate from each side can be selected. No pruning check may be performed between the two inherited candidates. Once the neighboring affine CU is identified, the control point motion vector of the neighboring affine CU can be used to derive the CPMVP candidate in the affine merge list of the current CU. As shown in FIG. 12, when block A, which is adjacent to the lower left of the current block (1204), is coded in affine mode, motion vectors v2, v3, and v4 of the upper left corner, upper right corner, and lower left corner of the CU (1202) including block A can be achieved. If block A is coded with a four-parameter affine model, two CPMVs of the current CU (1204) can be calculated according to v2 and v3 of the CU (1202). If block A is coded with a six-parameter affine model, three CPMVs of the current CU (1204) can be calculated according to v2, v3, and v4 of the CU (1202).

[0111] The constructed affine candidate of the current block may be a candidate constructed by combining the neighboring translational motion information of each control point of the current block. The motion information of the control points may be derived from specific spatial and temporal neighborhoods, which may be shown in FIG. 13. As shown in FIG. 13, the CPMV k(k=1,2,3,4) represents the kth control point of the current block (1302). For CPMV1, the B2→B3→A2 block can be checked and the MV of the first available block can be used. For CPMV2, the B1→B0 block can be checked. For CPMV3, the A1→A0 block can be checked. If CPM4 is not available, TMVP can be used as CPMV4.

[0112] After the MVs of the four control points are achieved, affine merge candidates for the current block (1302) can be constructed based on the motion information of the four control points. For example, the affine merge candidates can be constructed based on the combination of the MVs of the four control points in the following order: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}.

[0113] A combination of three CPMVs can be used to construct a six-parameter affine merge candidate, and a combination of two CPMVs can be used to construct a four-parameter affine merge candidate. To avoid motion scaling processing, if the reference indexes of the control points are different, the associated combination of control point MVs can be discarded.

[0114] After the inherited and constructed affine merge candidates have been checked, if the list is not already full, a zero MV may be inserted at the end of the list.

[0115] In affine AMVP prediction, affine AMVP mode can be applied to CUs with both width and height of 16 or more. A CU-level affine flag can be signaled in the bitstream to indicate whether affine AMVP mode is used, and then another flag can be signaled to indicate whether 4-parameter affine or 6-parameter affine is applied. In affine AMVP prediction, the difference between the CPMV of the current CU and the predictor of the CPMVP of the current CU can be signaled in the bitstream. The size of the affine AMVP candidate list can be 2, and the affine AMVP candidate list can be generated by using the four types of CPMV candidates in the following order: (1) Hereditary affine AMVP candidates extrapolated from the CPMVs of nearby CUs, (2) A constructed affine AMVP candidate with CPMVP derived using the translational MV of nearby CUs; (3) translational MVs from neighboring CUs, and (4) Zero MV.

[0116] The check order of the inherited affine AMVP candidates may be the same as the check order of the inherited affine merge candidates. To determine the AVMP candidates, in one example, only affine CUs with the same reference picture as the current block are considered. When an inherited affine motion predictor is inserted into the candidate list, the pruning process may not be applied.

[0117] The constructed AMVP candidate may be derived from the specified spatial neighborhood. As shown in FIG. 13, the same check order as in the affine merge candidate construction may be applied. In addition, the reference picture index of the neighboring block may also be checked. The first block in the check order may be inter-coded and have the same reference picture as the current CU (1302). If the current CU (1302) is coded in a four-parameter affine mode and both mv0 and mv1 are available, one constructed AMVP candidate may be determined. The constructed AMVP candidate may be further added to the affine AMVP list. If the current CU (1302) is coded in a six-parameter affine mode and all three CPMVs are available, the constructed AMVP candidate may be added as one candidate in the affine AMVP list. Otherwise, the constructed AMVP candidate may be set as unavailable.

[0118] After the inherited affine AMVP candidates and constructed AMVP candidates are checked, if there are still less than two candidates in the affine AMVP list, mv0, mv1, and mv2 can be added in order. mv0, mv1, and mv2 can serve as translation MVs to predict all control point MVs of the current CU (e.g., (1302)), if available. Finally, if the affine AMVP is not yet filled, the zero MV can be used to fill the affine AMVP list.

[0119] Sub-block-based affine motion compensation can save memory access bandwidth and reduce computational complexity compared to pixel-based motion compensation, at the expense of a penalty in prediction accuracy. To achieve finer granularity of motion compensation, prediction refinement by optical flow (PROF) can be used to refine sub-block-based affine motion compensation prediction without increasing memory access bandwidth for motion compensation. After sub-block-based affine motion compensation is implemented, the luma prediction samples can be refined by adding the difference derived by the optical flow equation as in VVC. PROF can be described in the following four steps:

[0120] Step (1): Sub-block-based affine motion compensation may be performed to generate a sub-block prediction I(i,j).

[0121] Step (2): Spatial gradient of subblock prediction g x (i,j) and g y (i,j) can be calculated at each sample position using a 3-tap filter [-1,0,1]. The gradient calculation may be the same as the gradient calculation in BDOF. For example, the spatial gradient g x (i,j) and g y (i,j) can be calculated based on equations (5) and (6), respectively. g x (i,j)=(I(i+1,j)≫shift1)-(I(i-1,j)≫shift1) Equation (5) g y (i,j)=(I(i,j+1)≫shift1)-(I(i,j-1)≫shift1) Equation (6) As shown in equations (5) and (6), shift1 can be used to control the accuracy of the gradient. Sub-block (e.g., 4×4) predictions can be extended by one sample on each side for gradient calculation. To avoid additional memory bandwidth and additional interpolation calculations, the extended samples on the extension boundary can be copied from the nearest integer pixel location in the reference picture.

[0122] Step (3): The luma prediction refinement can be calculated by the optical flow equation as shown in equation (7). ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) Equation (7) where Δv(i,j) is the sum of the sample MV calculated for sample position (i,j), denoted as v(i,j), and the MV of the sub-block to which sample (i,j) belongs, v SB The subblock MV may be a difference between the sample MV and the subblock MV, represented as v(i,j). FIG. 14 shows an example diagram of the difference between the sample MV and the subblock MV. As shown in FIG. 14, the subblock (1402) may be included in the current block (1400), and the sample (1404) may be included in the subblock (1402). The sample (1404) may include a sample motion vector v(i,j) corresponding to the reference pixel (1406). The subblock (1402) may include a subblock motion vector v(i,j). SB The sub-block motion vector v SB Based on, the sample (1404) may correspond to the reference pixel (1408). The difference between the sample MV and the sub-block MV, denoted by Δv(i,j), may be denoted by the difference between the reference pixel (1406) and the reference pixel (1408). Δv(i,j) may be quantized in units of 1 / 32 luma sample precision.

[0123] Since the affine model parameters and sample positions relative to the sub-block center may not change from sub-block to another sub-block, Δv(i,j) may be calculated for a first sub-block (e.g., (1402)) and reused for other sub-blocks (e.g., (1410)) in the same CU (e.g., (1400)). dx(i,j) is calculated from the sample position (i,j) to the sub-block center (x SB ,y SB ), and dy(i,j) be the vertical offset, Δv(x,y) can be derived by equations (8) and (9) as follows:

number

number

[0124] To maintain accuracy, the subblock centers (x SB ,y SB ) is ((W SB -1) / 2,(H SB -1) / 2), where W SB and H SB are the width and height of the sub-block, respectively.

[0125] Once Δv(x,y) is obtained, the parameters of the affine model can be obtained. For example, in the case of a four-parameter affine model, the parameters of the affine model can be shown in Equation (10).

number

number

[0126] Step (4): Finally, the luma prediction refinement ΔI(i,j) may be added to the sub-block prediction (i,j). The final prediction I′ may be generated as shown in equation (12). I'(i,j)=I(i,j)+ΔI(i,j) Equation (12)

[0127] PROF may not be applied to an affine coded CU in two cases: (1) when all control point MVs are the same, indicating that the CU has only translational motion, and (2) when the affine motion parameters are larger than the specified limit because sub-block-based affine MC has been downgraded to CU-based MC to avoid large memory access bandwidth requirements.

[0128] Affine motion estimation (ME), such as the VVC reference software VTM, can be operated for both uni-prediction and bi-prediction. Uni-prediction can be performed for either reference list L0 or reference list L1, and bi-prediction can be performed for both reference list L0 and reference list L1.

[0129] FIG. 15 shows a schematic diagram of affine ME (1500). As shown in FIG. 15, affine uni-prediction (S1502) can be performed on reference list L0 in affine ME (1500) to obtain a prediction P0 of the current block based on an initial reference block in reference list L0. Affine uni-prediction (S1504) can also be performed on reference list L1 to obtain a prediction P1 of the current block based on an initial reference block in reference list L1. In (S1506), affine bi-prediction can be performed. Affine bi-prediction (S1506) can start with an initial prediction residual (2I-P0)-P1, where I can be the initial value of the current block. Affine bi-prediction (S1506) can search candidates in reference list L1 around the initial reference block in reference list L1 to find the best (or selected) reference block with the smallest prediction residual (2I-P0)-Px, where Px is the prediction of the current block based on the selected reference block.

[0130] Using the reference picture, for the current coding block, the affine ME process can first choose a set of control point motion vectors (CPMVs) as a base. An iterative method can be used to generate a prediction output of the current affine model corresponding to the set of CPMVs, calculate the gradient of the prediction sample, and then solve a linear equation to determine the delta CPMVs and optimize the affine prediction. The iteration can be stopped when all delta CPMVs are 0 or the maximum number of iterations is reached. The CPMVs obtained from the iterations can be the final CPMVs of the reference picture.

[0131] After the best affine CPVM of both reference lists L0 and L1 is determined for affine uni-prediction, an affine bi-prediction search can be performed using the best uni-prediction CPMV and the reference list on one side to search for the best CPMV of the other reference list to optimize the affine bi-prediction output. The affine bi-prediction search can be performed iteratively on the two reference lists to obtain the optimal result.

[0132] 16 shows an example affine ME process (1600) that can calculate a final CPMV associated with a reference picture. The affine ME process (1600) can begin at (S1602). At (S1602), a base CPMV of a current block can be determined. The base CPMV can be determined based on one of a merge index, an advanced motion vector prediction (AMVP) predictor index, an affine merge index, etc.

[0133] In (S1604), an initial affine prediction of the current block can be obtained based on the base CPMV. For example, according to the base CPMV, a four-parameter affine motion model of a six-parameter affine motion model can be applied to generate the initial affine prediction.

[0134] In (S1606), a gradient of the initial affine prediction may be obtained. For example, the gradient of the initial affine prediction may be obtained based on Equation (5) and Equation (6).

[0135] At (S1608), a delta CPMV may be determined. In some embodiments, the delta CPMV may be associated with a displacement between an initial affine prediction and a subsequent affine prediction, such as the first affine prediction. Based on a gradient of the initial affine prediction and the delta CPMV, a first affine prediction may be obtained. The first affine prediction may correspond to the first CPMV.

[0136] At (S1610), a determination may be made to check whether the delta CPMV is zero or the number of iterations is greater than or equal to a threshold. If the delta CPMV is zero or the number of iterations is greater than or equal to a threshold, at (S1612), a final (or selected) CPMV may be determined. The final (or selected) CPMV may be a first CPMV determined based on the gradient of the initial affine prediction and the delta CPMV.

[0137] Further referring to (S1610), if the delta CPMV is not zero or the number of iterations is less than a threshold, a new iteration may be started. In the new iteration, an updated CPMV (e.g., the first CPMV) may be provided to (S1604) to generate an updated affine prediction. The affine ME process (1600) may then proceed to (S1606) where a gradient of the updated affine prediction may be calculated. The affine ME process (1600) may then proceed to (S1608) to continue the new iteration.

[0138] In the affine motion model, the four-parameter affine motion model can be further described by an equation including the rotation and zoom motions. For example, the four-parameter affine motion model can be rewritten in equation (13) as follows:

number

number

number

[0139] BDOF in VVC was previously called BIO in JEM. Compared to the JEM version, BDOF in VVC can be a simpler version that requires fewer calculations, especially in terms of the number of multiplications and the size of the multipliers.

[0140] BDOF can be used to refine the bi-predictive signal of a CU at the 4×4 sub-block level. BDOF can be applied to a CU if the CU satisfies the following conditions: (1) The CU is coded using a “true” bi-prediction mode, i.e., one of the two reference pictures is before the current picture in display order and the other is after the current picture in display order; (2) The distances (e.g., POC differences) from the two reference pictures to the current picture are the same; (3) both reference pictures are short-term reference pictures; (4) the CU is not coded using affine mode or SbTMVP merge mode; (5) the CU has more than 64 luma samples; (6) both the CU height and the CU width are greater than or equal to 8 luma samples; (7) the BCW Weight Index shows equal weights; (8) No weighting position (WP) is available for the current CU; and (9) CIIP mode is not used for the current CU.

[0141] BDOF may be applied only to the luma component. As the name BDOF suggests, the BDOF mode may be based on the concept of optical flow, which assumes that object motion is smooth. For each 4 × 4 sub-block, motion refinement (v x ,v y) can be calculated. Then, motion refinement can be used to adjust the bi-predictive sample values ​​in the 4×4 sub-block. BDOF can include the following steps:

[0142] First, the horizontal gradients of the two predicted signals from reference list L0 and reference list L1

number

number

number

number

[0143] Then, the autocorrelation and cross-correlation of the gradients S1, S2, S3, S5, and S6 can be calculated according to equations (18)-(22) as follows: S1=Σ (i,j)∈Ω Abs(ψ x (i,j)), equation (18) S2=Σ (i,j)∈Ω ψ x (i,j) Sign(ψ y (i,j)) Equation (19) S3=Σ (i,j)∈Ω θ(i,j)·Sign(ψ x (i,j)) Equation (20) S5=Σ(i,j)∈Ω Abs(ψ y (i,j)) Equation (21) S6=Σ (i,j)∈Ω θ(i,j)·Sign(ψ y (i,j)) Equation (22) Here, ψ x (i,j), ψ y (i,j) and θ(i,j) can be given by equations (23) to (25), respectively.

number

number

[0144] Then, motion refinement (v x ,v y ) can be derived using cross-correlation and auto-correlation terms using the following equations (26) and (27):

number

number

number

number

number

number

[0145] Finally, the BDOF samples of the CU can be calculated by adjusting the bi-predictive samples in equation (29) as follows: pred BDOF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+ο offset )≫shift expression (29) The values ​​can be selected such that the multipliers in the BDOF process do not exceed 15 bits and the maximum bit width of the intermediate parameters in the BDOF process can be kept within 32 bits.

[0146] To derive the gradient value, we select some predicted samples I in list k (k=0, 1) outside the current CU boundary. (k) (i,j) needs to be generated. As shown in FIG. 17, BDOF in VVC can use one extended row / column (1702) around the boundary (1706) of the CU (1704). To control the computational complexity of generating out-of-bounds predicted samples, the predicted samples in the extended region (e.g., the unshaded region in FIG. 17) can be generated by directly taking reference samples of nearby integer positions (e.g., using floor() operation on the coordinates) without interpolation, and a normal 8-tap motion compensation interpolation filter can be used to generate predicted samples in the CU (e.g., the shaded region in FIG. 17). The extended sample values ​​can be used only in gradient calculation. In the remaining steps of the BDOF process, if samples and gradient values ​​outside the CU boundary are required, the samples and gradient values ​​can be padded (e.g., repeated) from the nearest neighbors of the samples and gradient values.

[0147] If the width and / or height of a CU is greater than 16 luma samples, the CU may be divided into sub-blocks with width and / or height equal to 16 luma samples, and the sub-block boundaries may be treated as CU boundaries in the BDOF process. The maximum unit size of the BDOF process may be limited to 16×16. For each sub-block, the BDOF process may be skipped. If the sum of absolute differences (SAD) between the initial L0 predicted sample and the initial L1 predicted sample is less than a threshold, the BDOF process may not be applied to the sub-block. The threshold may be set equal to (8*W*(H≫1), where W may indicate the width of the sub-block and H may indicate the height of the sub-block. To avoid further complexity of the SAD calculation, the SAD between the initial L0 predicted sample and the initial L1 predicted sample calculated in the DMVR process may be reused in the BDOF process.

[0148] If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, bidirectional optical flow may be disabled. Similarly, if WP is enabled for the current block, i.e., the luma weight flag (e.g., luma_weight_lx_flag) for either of the two reference pictures is 1, BDOF may be disabled. If the CU is coded in symmetric MVD mode or CIIP mode, BDOF may also be disabled.

[0149] To improve the accuracy of the MV of the merge mode, a bilateral matching (BM)-based decoder-side motion vector refinement such as in VVC can be applied. In bi-predictive operation, the refined MV can be searched around the initial MV in the reference picture list L0 and the reference picture list L1. The BM method can calculate the distortion between two candidate blocks in the reference picture list L0 and the reference picture list L1.

[0150] Figure 18 shows an example schematic diagram of BM-based decoder-side motion vector refinement. As shown in Figure 18, a current picture (1802) may include a current block (1808). The current picture may include a reference picture list L0 (1804) and a reference picture list L1 (1806). The current block (1808) may include an initial reference block (1812) in reference picture list L0 (1804) with an initial motion vector MV0 and an initial reference block (1814) in reference picture list L1 (1806) with an initial motion vector MV1. A search process may be performed around an initial MV0 in reference picture list L0 (1804) and an initial MV1 in reference picture list L1 (1806). For example, a first candidate reference block (1810) may be identified in reference picture list L0 (1804), and a first candidate reference block (1816) may be identified in reference picture list L1 (1806). The SAD between the candidate reference blocks (e.g., (1810) and (1816)) based on each MV candidate (e.g., MV0′ and MV1′) around the initial MV (e.g., MV0 and MV1) may be calculated. The MV candidate with the lowest SAD becomes the refined MV and may be used to generate a bi-predictive signal for predicting the current block (1808).

[0151] The application of DMVR may be restricted and may only be applied to CUs coded based on mode and feature, such as VVC, as follows: (1) CU-level merge mode with bi-predictive MVs, (2) For a current picture, one reference picture is in the past and another reference picture is in the future. (3) The distances (e.g., POC differences) from the two reference pictures to the current picture are the same; (4) both reference pictures are short-term reference pictures; (5) The CU has more than 64 luma samples. (6) Both the CU height and the CU width are equal to or greater than 8 luma samples; (7) The BCW weight index indicates equal weights; (8) WP is not currently enabled for blocking, and (9) CIIP mode is not currently used for blocks.

[0152] The refined MVs derived by the DMVR process can be used to generate inter-prediction samples and can be used in temporal motion vector prediction for future picture coding, while the original MVs can be used in the deblocking process and can be used in spatial motion vector prediction for future CU coding.

[0153] In DVMR, the search points can surround the initial MV, and the MV offset can follow the MV difference mirroring rule. In other words, any point checked by DMVR, represented by the candidate MV pair (MV0, MV1), can follow the MV difference mirroring rule shown in Equation (30) and Equation (31). MV0 ’ =MV0+MV_offset formula (30) MV1 ’ =MV1-MV_offset formula (31) Here, MV_offset may represent a refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range may be two integer luma samples from the initial MV. The search may include an integer sample offset search stage and a fractional sample refinement stage.

[0154] For example, a full search of 25 points can be applied to the integer sample offset search. The SAD of the initial MV pair can be calculated first. If the SAD of the initial MV pair is less than a threshold, the integer sample stage of the DMVR can be terminated. Otherwise, the SAD of the remaining 24 points can be calculated and checked in a scan order, such as a raster scan order. The point with the smallest SAD can be selected as the output of the integer sample offset search stage. In order to reduce the penalty of uncertainty in DMVR refinement, the original MVs in the DMVR process can have a priority to be selected. The SAD between the reference blocks referenced by the initial MV candidates can be reduced by ¼ of the SAD value.

[0155] The integer sample search can be followed by fractional sample refinement. To reduce computational complexity, the fractional sample refinement can be derived by using a parametric error surface equation instead of an additional search with SAD comparison. The fractional sample refinement can be conditionally invoked based on the output of the integer sample search stage. If the integer sample search stage ends at the center with the smallest SAD in either the first iteration search or the second iteration search, the fractional sample refinement can be further applied.

[0156] In parametric error surface based sub-pixel offset estimation, the center location cost and the costs at the four neighboring locations from the center are used to fit a 2D parabolic error surface equation based on Eq. (32): E(x,y)=A(xx min ) 2 +B(yy min ) 2 +C formula (32) Here, (x min ,y min ) may correspond to the minimum cost fractional position, and C may correspond to the minimum cost value. By solving equation (32) using the cost values ​​of the five search points, (x min ,y min) can be calculated using equations (33) and (34). x min =(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) Equation (33) y min =(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0))) Equation (34) x min and y min The value of x can be automatically constrained to be between -8 and 8 since all cost values ​​are positive and the minimum value is E(0,0). min and y min The value constraint of can correspond to a half pel (or pixel) offset with 1 / 16 pel MV accuracy in VVC. To get a sub-pixel accurate refinement delta MV, the calculated fraction (x min ,y min ) can be added to the refinement MV, which is an integer distance.

[0157] In VVC, etc., quadratic interpolation and sample padding can be applied. The resolution of the MV can be, for example, 1 / 16 luma samples. The fractional position samples can be interpolated using an 8-tap interpolation filter. In DMVR, the search points can surround the initial fractional pel MV with integer sample offsets, so the fractional position samples need to be interpolated for the DMVR search process. To reduce the computational complexity, a bilinear interpolation filter can be used to generate fractional samples for the search process in DMVR. In another important effect, by using a bilinear filter with a 2-sample search range, DVMR does not access more reference samples compared to a normal motion compensation process. After the refined MV is achieved using the DMVR search process, a normal 8-tap interpolation filter can be applied to generate the final prediction. Due to the lack of access to more reference samples compared to a normal MC process, samples that may not be needed for the original MV-based interpolation process but may be needed for the refined MV-based interpolation process can be padded from the available samples.

[0158] If the width and / or height of a CU is greater than 16 luma samples, the CU may be further divided into sub-blocks having widths and / or heights equal to 16 luma samples. The maximum unit size of the DMVR search process may be limited to 16×16.

[0159] In VVC, etc., an affine bilateral matching mode can be applied. In the affine bilateral matching mode, a four-parameter affine model can be applied to find the best (or selected) affine parameters based on bilateral matching. However, since the regular four-parameter affine model is a nonlinear function, it may be difficult to directly solve the affine parameters from the regular four-parameter affine model such as the four-parameter affine model in equations (1), (2), and (13).

[0160] In the present disclosure, a simplified four-parameter affine model may be provided. In some embodiments, the simplified four-parameter affine model may be derived, for example, from Equation (13).

[0161] In one embodiment, a three-parameter (or 3p) zooming model may be given in equations (35) and (36) as follows, where α, c, and f are three affine parameters:

number

number

[0162] As shown in equations (35) and (36), V x0 V may be a first motion vector of a sample in the current block relative to a first reference picture in reference list L0 along a first direction (e.g., the X direction). y0 V may be a first motion vector of a sample in the current block associated with a first reference picture in reference list L0 along a second direction (e.g., the Y direction). x1 V may be a second motion vector of the sample in the current block associated with the second reference picture in the reference list L1 along the first direction. y1 may be a second motion vector of the sample in the current block associated with a second reference picture in the reference list L1 along the second direction, c may be a translation coefficient (or a first component of the translation coefficient) in the first direction, and f may be a translation coefficient (or a second component of the translation coefficient) in the second direction. (x,y) may be a position of the sample in the current block. α=r−1, where r may be the zoom factor shown in equation (13).

[0163] In one example, equation (35) can be derived from equation (13) when θ is small, cosθ=1, and sinθ=0. Furthermore, as shown in equations (35) and (36), the first motion vector (V x0 ,V y0 ) is the affine parameter of the second motion vector (V x1 ,V y1 ) has the opposite sign value compared to the affine parameter

[0164] In one embodiment, a four-parameter (or 4p) zooming model may be given in equations (37) and (38) as follows, where α 1 , α 2 , c, and f are the four affine parameters.

number

number

[0165] As shown in equations (37) and (38), V x0 V may be a first motion vector of a sample in the current block associated with a first reference picture in reference list L0 along a first direction (e.g., the X direction). y0 V may be a first motion vector of a sample in the current block associated with a first reference picture in reference list L0 along a second direction (e.g., the Y direction). x1 V may be a second motion vector of the sample in the current block associated with the second reference picture in the reference list L1 along the first direction. y1may be a second motion vector of a sample in a current block associated with a second reference picture in the reference list L1 along a second direction, c may be a translation coefficient (or a first component of the translation coefficient) in a first direction, f may be a translation coefficient (or a second component of the translation coefficient) in a second direction, and (x,y) may be a position of the sample in a current block. α1=r1-1, where r1 may be a zoom factor (or a first component of the zoom factor) in a first direction, and α2=r2-1, where r2 may be a zoom factor (or a second component of the zoom factor) in a second direction. In one example, α1 is not equal to α2, which indicates that the first motion vector may have a zoom factor in a first direction (e.g., α1) that is different from the zoom factor in a second direction (e.g., α2).

[0166] In one example, equation (37) can be derived from equation (13) when θ is small, cosθ=1, and sinθ=0. Also, as shown in equations (37) and (38), the first motion vector (V x0 ,V y0 ) affine parameters and the second motion vector (V x1 ,V y1 ) are the inverses of the affine parameters.

[0167] In one embodiment, a three-parameter rotation model can be provided in equations (39) and (40) as follows, where θ, c, and f are the three affine parameters:

number

number

[0168] As shown in equations (39) and (40), V x0V may be a first motion vector of a sample in the current block relative to a first reference picture in reference list L0 along a first direction (e.g., the X direction). y0 V may be a first motion vector of a sample in the current block associated with a first reference picture in reference list L0 along a second direction (e.g., the Y direction). x1 V may be a second motion vector of the sample in the current block associated with the second reference picture in the reference list L1 along the first direction. y1 may be a second motion vector of the sample in the current block associated with a second reference picture in the reference list L1 along the second direction, c may be a translation coefficient (or a first component of the translation coefficient) in the first direction, f may be a translation coefficient (or a second component of the translation coefficient) in the second direction. (x,y) may be a position of the sample in the current block, and θ may be a rotation coefficient according to equation (13).

[0169] In one example, equation (39) can be derived from equation (13) when θ is small, r=1, cos θ=1, and sin θ=0. Also, as shown in equations (39) and (40), the first motion vector (V x0 ,V y0 ) affine parameters and the second motion vector (V x1 ,V y1 ) have an inverse relationship with the affine parameters

[0170] In one embodiment, a four-parameter rotation model can be provided in equations (41) and (42) as follows, where a, b, c, and f are the four affine parameters:

number

number

[0171] As shown in equations (41) and (42), V x0 V may be a first motion vector of a sample in the current block relative to a first reference picture in reference list L0 along a first direction (e.g., the X direction). y0 V may be a first motion vector of a sample in the current block associated with a first reference picture in reference list L0 along a second direction (e.g., the Y direction). x1 V may be a second motion vector of the sample in the current block associated with the second reference picture in the reference list L1 along the first direction. y1 may be a second motion vector of the sample in the current block associated with a second reference picture in the reference list L1 along the second direction, c may be a translation coefficient (or a first component of the translation coefficient) in the first direction, f may be a translation coefficient (or a second component of the translation coefficient) in the second direction, and (x,y) may be a position of the sample in the current block.

[0172] In one example, equation (41) may be the same as equation (13), where the rotation factor may be θ and the zoom factor r may be 1. Thus, α=cosθ−1 and b=sinθ. α may be a first coefficient related to the rotation factor θ, b may be a second coefficient related to the rotation factor, and −b may be an opposite second coefficient. Equation (42) may be derived based on equation (14), where the rotation factor may be −θ and the zoom factor r may be 1.

[0173] In this disclosure, the affine models provided in Equations (35)-(42) may be used separately or in combination in affine bilateral matching. In affine bilateral matching, the affine motion of a current block may be derived (or refined) via an iterative search based on reference blocks in a first reference picture and a second reference picture associated with the current picture, where the current picture may be located between the first reference picture and the second reference picture. The affine models provided in Equations (35)-(42), or a subset of the affine models, may also be used in combination with a predetermined number of iterations to derive a final CPMV (refined affine merge candidate).

[0174] In one embodiment, 3p zooming model based affine bilateral matching can be applied to refine the affine merge candidates.

[0175] In one example, the affine bilateral matching may be performed based on an iterative process (or iterative search), such as the iterative process of affine motion estimation (ME) shown in Figures 15-16. Thus, the 3p zooming model may be applied in the iterative process to derive a final CPMV (or refine an affine merge candidate) according to the base CPMV.

[0176] As shown in FIG. 16, in a first iteration of the iterative process, at (S1602), a base CPMV may be determined for a current block. The base CPMV may be associated with an initial reference block in reference list L0 and an initial reference block in reference list L1, respectively. In one embodiment, the base CPMV may be determined based on an affine merge candidate. At (S1604), a starting point (or initial predictor) P 0,L0 (i,j) may be determined based on an initial reference block (or base CPMV) in the reference list L0, and a starting point (or initial predictor) P 0,L1 (i,j) may be determined based on the initial reference block (or base CPMV) in the reference list L1. 0,L0The gradient of (i,j) and P 0,L1 The gradient of (i,j) can be calculated based on equations (5) and (6), etc. For example, g x0,L0 (i,j) is the initial predictor P 0,L0 is the gradient of (i,j) in the x direction, and g y0,L0 (i,j) is the initial predictor P 0,L0 is the gradient of (i,j) in the y direction, and g x1,L1 (i,j) is the first predictor P 1,L1 is the gradient of (i,j) in the x direction, and g y1,L1 (i,j) is the first predictor P 1,L1 is the gradient of (i,j) in the y direction.

[0177] At (S1608), a delta CPMV associated with two reference blocks (or sub-blocks) in reference list L0, such as the initial reference block and the first reference block in reference list L0, may be calculated based on one of the affine models given in equations (35)-(42). In one example, the affine model may be the 3p zooming model shown in equations (35) and (36). The delta CPMV is calculated based on Δv x0,L0 (i,j) and Δv y0,L0 (i,j) can be expressed as Δv x0,L0 (i,j) may be the difference or displacement along the x direction of two reference blocks (or sub-blocks), such as the initial reference block and the first reference block in the reference list L0. y0,L0 (i,j) may be the difference or displacement along the y direction of two reference blocks (or sub-blocks), such as the initial reference block and the first reference block in reference list L0. Similarly, the delta CPMV associated with two reference blocks (or sub-blocks) in reference list L1, such as the initial reference block and the first reference block in reference list L1, may be calculated based on one of the affine models given in equations (35)-(42). In one example, the affine model may be the 3p zooming model shown in equations (35) and (36). The delta CPMV is calculated based on Δv x0,L1 (i,j) and Δv y0,L1It can be expressed as (i,j). Δv x0,L1 (i,j) may be the difference or displacement of two reference blocks (or sub-blocks), such as the initial reference block and the first reference block in the reference list L1 along the x direction. y0,L1 (i,j) may be the difference or displacement of two reference blocks (or sub-blocks), such as the initial reference block and the first reference block in reference list L1, along the y direction.

[0178] At (S1610), a first predictor P of the current block based on the first reference block in the reference list L0 is calculated. 1,L0 (i,j), and the first predictor P of the current block based on the first reference block in the reference list L1. 1,L1 (i,j) may be determined according to equations (43) and (44). P 1,L0 (i,j)=P 0,L0 (i,j)+g x0,L0 (i,j)*Δv x0,L0 (i,j)+g y0,L0 (i,j)*Δv y0,L0 (i,j) Equation (43) P 1,L1 (i,j)=P 0,L1 (i,j)+g x0,L1 (i,j)*Δv x0,L1 (i,j)+g y0,L1 (i,j)*Δv y0,L1 (i,j) Equation (44) Here, (i,j) may be the position of a pixel (or sample) in the current block.

[0179] Δv x0,L0 (i,j), Δv y0,L0 (i,j), Δv x0,L1 (i,j), and Δv y0,L1 In response to at least one of (i,j) being non-zero, the affine bilateral matching may proceed to a second iteration according to (S1614). At (S1614), an updated CPMV (e.g., P 1,L0 (i,j) and P 1,L1The first CPMV associated with (i,j) may be provided to (S1604), where an updated CPMV (or an updated affine prediction) may be generated. The affine bilateral matching may then proceed to (S1606), where a gradient of the updated CPMV may be calculated. The affine bilateral matching may then proceed to (S1608) to continue with a new iteration (e.g., a second iteration). In the second iteration, a second predictor P of the current block based on a second reference block in the reference list L0 may be provided to (S1605), where an updated CPMV (or an updated affine prediction) may be generated. 2,L0 (i,j), and a second predictor P of the current block based on the second reference block in the reference list L1. 2,L1 (i,j) can be determined according to the following equations (45)-(46): P 2,L0 (i,j)=P 1,L0 (i,j)+g x1,L0 (i,j)*Δv x1,L0 (i,j)+g y1,L0 (i,j)*Δv y1,L0 (i,j) Equation (45) P 2,L1 (i,j)=P 1,L1 (i,j)+g x1,L1 (i,j)*Δv x1,L1 (i,j)+g y1,L1 (i,j)*Δv y1,L1 (i,j) Equation (46) As shown in equation (45), g x1,L0 (i,j) is the first predictor P 1,L0 It can be taken as the gradient in the x direction of (i,j). y1,L0 (i,j) is the first predictor P 1,L0 It can be said that the gradient of (i,j) in the y direction is Δv x1,L0 (i,j) may be the difference or displacement along the x direction between the first and second reference blocks in the reference list L0. y1,L0 (i,j) can be the difference or displacement along the y direction between the first reference block and the second reference block. As shown in equation (46), g x1,L1 (i,j) is the first predictor P 1,L1 It can be taken as the gradient in the x direction of (i,j).y1,L1 (i,j) is the first predictor P 1,L1 It can be said that the gradient of (i,j) in the y direction is Δv x1,L1 (i,j) may be the difference or displacement between the first and second reference blocks in the reference list L1 along the x direction. y1,L1 (i,j) may be the difference or displacement between the first and second reference blocks in reference list L1 along the y direction.

[0180] The bilateral matching iteration is started when the iteration number N is equal to or greater than the iteration threshold (or the maximum iteration number) or when the displacement between the Nth reference block in the reference list L0 and the N+1th reference block in the reference list L1 (e.g., Δv xN,L0 (i,j),Δv yN,L0 (i,j)) is zero, or the displacement between the Nth reference block in the reference list L1 and the N+1th reference block in the reference list L1 (e.g., Δv xN,L1 (i,j),Δv yN,L1 16, the process may terminate when all of the Nth reference blocks in the reference list L0 and the Nth reference block in the reference list L1 are equal to or greater than the Nth reference block in the reference list L1.

[0181] In one example, the final CPMV may be an affine merge candidate of the current block. In one example, the final CPMV may be directly applied to generate prediction information of the current block. In one example, the final CPMV may also be applied to perform bi-prediction, such as affine bi-prediction shown in FIG. 15. For example, as shown in FIG. 15, in (S1502), a prediction P0 of the current block may be determined based on the final CPMV associated with the N-th reference block in the reference list L0. In (S1504), a prediction P1 of the current block may be determined based on the final CPMV associated with the N-th reference block in the reference list L1. In (S1506), affine bi-prediction may be performed. The affine bi-prediction (1506) may start with an initial prediction residual (2I-P0)-P1, where I may be an initial value of the current block or an initial predictor of the current block. The affine bi-prediction (1506) may search candidates in the reference list L1 around P1 in the reference list L1 to find a best (or selected) reference block with a minimum prediction residual (2I-P0)-Px. Px may be a prediction of the current block based on the selected reference block in the reference list L1. In some embodiments, the affine bi-prediction may further start from the initial prediction residual (2I-P0)-Px and search for a best (or selected) reference block around P0 in the reference list L0 with a minimum prediction residual (2I-Py)-Px. Py may be a prediction of the current block based on the selected reference block in the reference list L0.

[0182] In another embodiment, both the 3p zooming model and the 4p rotation model may be applied in affine bilateral matching, where the starting point (or base CPMV) may be generated based on translational motion, such as based on a merge candidate indicated by a merge index or an AMPV candidate indicated by an AMVP predictor index.

[0183] In one example, P 1,L0 (i,j) and P 1,L1A first CPMV, denoted by (i,j), may be calculated in a first iteration based on a first affine model of the affine models in equations (35)-(42), such as a 3p zooming model. Subsequent iterations may then proceed by applying a second affine model of the affine models in equations (35)-(42), such as a 4p rotation model, to calculate P N,L0 (i,j) and P N,L1 The final CPMV, denoted by (i,j), can be derived.

[0184] In one example, N iterations may be performed to generate a final CPMV0 based on a first affine model of the affine models in equations (35)-(42), such as a 3p zooming model. The final CPMV0 may then be applied as an initial point (or base CPMV) to generate a final CPMV1 in M ​​iterations based on a second affine model of the affine models in equations (35)-(42), such as a 4p rotation model. N and M may be any positive integers.

[0185] In one example, some affine models may be applied, such as M (M=1 to 4) affine models among the affine models in equations (35) to (42). For each affine model, N i (i=1 to M) iterations can be applied. i (i=1 to M) iterations may be performed sequentially to generate a final CPMV. For example, a 3p zooming model, a 3p rotation model, or a 4p zooming model may be selected first. Based on the 3p zooming model, N1 iterations may be performed to derive a final CPMV0. Based on the 3p rotation model, the final CPMV0 may be used as a starting point, and N2 iterations may be performed to generate a final CPMV1. Based on the 4p zooming model, the final CPMV1 may be used as a starting point, and N3 iterations may be performed to generate a final CPMV2. The final CPMV2 may be a refined affine merge candidate for the current block.

[0186] In one embodiment, the type of affine model used in affine bilateral matching (e.g., which one of the affine models) may be signaled by a high-level syntax, such as an SPS at the sequence level or a slice header at the slice level.

[0187] In one example, the type of affine model used for bilateral matching to derive affine merge candidates and the type of affine model for deriving translational merge candidates may be signaled separately.

[0188] FIG. 19 shows a flowchart outlining an exemplary decoding process (1900) according to some embodiments of the present disclosure. FIG. 20 shows a flowchart outlining an exemplary encoding process (2000) according to some embodiments of the present disclosure. The proposed processes may be used separately or combined in any order. Furthermore, each of the processes (or embodiments), the encoder, and the decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium.

[0189] The operations of the processes (e.g., (1900) and (2000)) can be combined or arranged in any quantity or order as desired. In embodiments, two or more of the operations of the processes (e.g., (1900) and (2000)) may be performed in parallel.

[0190] The processes (e.g., (1900) and (2000)) can be used in the reconstruction and / or encoding of a block to generate a prediction block for a block being reconstructed. In various embodiments, the processes (e.g., (1900) and (2000)) are performed by processing circuitry of terminal devices (310), (320), (330), and (340), processing circuitry performing the functions of a video encoder (403), processing circuitry performing the functions of a video decoder (410), processing circuitry performing the functions of a video decoder (510), processing circuitry performing the functions of a video encoder (603), etc. In some embodiments, the processes (e.g., (1900) and (2000)) are implemented with software instructions, such that the processing circuitry performs the processes (e.g., (1900) and (2000)) when the processing circuitry executes the software instructions.

[0191] As shown in FIG. 19, the process (1900) can start at (S1901) and proceed to (S1910). At (S1910), prediction information of a current block in a current picture can be received from a coded video bitstream. The prediction information can indicate that the current block is predicted based on affine bilateral matching, in which an affine motion of the current block is derived based on reference blocks in a first reference picture and a second reference picture associated with the current picture. The current picture can be located between the first reference picture and the second reference picture.

[0192] At (S1920), a first affine parameter and a second affine parameter for each of one or more affine models associated with the affine bilateral matching may be determined. The first affine parameter of each affine model may be associated with a first motion vector of a first reference picture. The second affine parameter of each affine model may be associated with a second motion vector of a second reference picture. The first affine parameter and the second affine parameter may have opposite signs.

[0193] In (S1930), a control point motion vector (CPMV) of an affine motion of the current block may be determined by performing affine bilateral matching based on the determined first affine parameter and the determined second affine parameter of one or more affine models.

[0194] At (S1940), the current block may be reconstructed based on the CPMV of the affine motion of the current block.

[0195] To determine the CPMV, a first CPMV of the affine motion of the current block may be determined based on a first affine model of the one or more affine models. A refined CPMV of the affine motion of the current block may be determined based on the first CPMV and a second affine model of the one or more affine models. The second affine model may be the same as or different from the first affine model. Thus, the current block may be reconstructed based on the refined CPMV of the affine motion of the current block.

[0196] The one or more affine models may include a three-parameter zooming model. The three-parameter zooming model may further include a first component of a first motion vector of a sample in the current block associated with a first reference picture along a first direction. The first component of the first motion vector may be equal to a first component of a translation coefficient along the first direction and a product of a position of the sample along the first direction and a zoom factor. The three-parameter zooming model may include a second component of a first motion vector of a sample in the current block associated with a first reference picture along a second direction. The second component of the first motion vector may be equal to a second component of a translation coefficient along the second direction and a product of a position of the sample along the second direction and a zoom factor, the second direction may be perpendicular to the first direction. The three-parameter zooming model may include a first component of a second motion vector of a sample in the current block associated with a second reference picture along the first direction. The first component of the second motion vector may be equal to a first component of an opposite translation coefficient along the first direction plus a product of the position of the sample along the first direction and the opposite zoom coefficient. The three-parameter zooming model may further include a second component of a second motion vector of the sample in the current block associated with the second reference picture along the second direction. The second component of the second motion vector may be equal to a second component of an opposite translation coefficient along the second direction plus a product of the position of the sample along the second direction and the opposite zoom coefficient.

[0197] The one or more affine models may include a three-parameter rotation model. The three-parameter rotation model may further include a first component of a first motion vector of a sample in the current block associated with a first reference picture along a first direction. The first component of the first motion vector may be equal to a sum of a first component of a translation coefficient along the first direction and a product of a position of the sample along a second direction and a rotation coefficient. The second direction may be perpendicular to the first direction. The three-parameter rotation model may include a second component of a first motion vector of a sample in the current block associated with a first reference picture along a second direction. The second component of the first motion vector may be equal to a sum of a second component of a translation coefficient along the second direction and a product of a position of the sample along the first direction and an inverse rotation coefficient. The three-parameter rotation model may include a first component of a second motion vector of a sample in the current block associated with a second reference picture along the first direction. The first component of the second motion vector may be equal to a first component of an opposite translation coefficient in the first direction plus a product of the position of the sample along the second direction and the inverse rotation coefficient. The three-parameter rotation model may include a second component of a second motion vector of a sample in the current block associated with a second reference picture along the second direction. The second component of the second motion vector may be equal to a second component of an opposite translation coefficient along the second direction plus a product of the position of the sample along the first direction and the rotation coefficient.

[0198] The one or more affine models may include a four-parameter zooming model. The four-parameter zooming model may further include a first component of a first motion vector of a sample in the current block associated with a first reference picture along a first direction. The first component of the first motion vector may be equal to a sum of a first component of a translation coefficient along the first direction and a product of a position of the sample along the first direction and a first component of a zoom factor in the first direction. The four-parameter zooming model may include a second component of a first motion vector of a sample in the current block associated with a first reference picture along a second direction. The second component of the first motion vector may be equal to a sum of a second component of a translation coefficient along the second direction and a product of a position of the sample along the second direction and a second component of a zoom factor along the second direction. The second direction may be perpendicular to the first direction. The four-parameter zooming model may include a first component of a second motion vector of a sample in the current block associated with a second reference picture along the first direction. The first component of the second motion vector may be equal to the sum of a first component of an opposite translation coefficient along the first direction and a product of a position of the sample along the first direction and a first component of an opposite zoom coefficient along the first direction. The four-parameter zooming model may include a second component of a second motion vector of a sample in the current block associated with a second reference picture along the second direction. The second component of the second motion vector may be equal to the sum of a second component of an opposite translation coefficient along the second direction and a product of a position of the sample along the second direction and a second component of an opposite zoom coefficient along the second direction.

[0199] The one or more affine models may include a four-parameter rotation model. The four-parameter rotation model may further include a first component of a first motion vector of a sample in the current block associated with a first reference picture along a first direction. The first component of the first motion vector may be equal to the sum of (i) a first component of a translation coefficient along the first direction, (ii) a product of a position of the sample along the first direction and a first coefficient associated with the rotation coefficient, and (iii) a product of a position of the sample along a second direction and a second coefficient associated with the rotation coefficient. The first direction may be perpendicular to the second direction. The four-parameter rotation model may include a second component of a first motion vector of a sample in the current block associated with a first reference picture along the second direction. The second component of the first motion vector may be equal to the sum of (i) a second component of a translation coefficient along the second direction, (ii) a product of the position of the sample along the first direction and an opposite second coefficient associated with the rotation coefficient, and (iii) a product of the position of the sample along the second direction and a first coefficient associated with the rotation coefficient. The four-parameter rotation model may include a first component of a second motion vector of a sample in the current block associated with a second reference picture along the first direction. The first component of the second motion vector may be equal to the sum of (i) a first component of an opposite translation coefficient along the first direction, (ii) a product of the position of the sample along the first direction and a first coefficient associated with the rotation coefficient, and (iii) a product of the position of the sample along the second direction and an opposite second coefficient associated with the rotation coefficient. The four-parameter rotation model may include a second component of a second motion vector of a sample in the current block associated with a second reference picture along the second direction. A second component of the second motion vector may be equal to the sum of (i) a second component of the opposite translation coefficient along the second direction, (ii) a product of the sample's position along the first direction and a second coefficient associated with the rotation coefficient, and (iii) a product of the sample's position along the second direction and a first coefficient associated with the rotation coefficient.

[0200] To determine a first CPMV of the affine motion of the current block, a first component of an initial predictor associated with a first reference picture of the current block may be determined. A second component of an initial predictor associated with a second reference picture of the current block may be determined. The first component of the first predictor associated with the first reference picture of the current block may be determined based on the first component of the initial predictor associated with the first reference picture of the current block. The second component of the first predictor associated with the second reference picture of the current block may be determined based on the second component of the initial predictor associated with the second reference picture of the current block. The first CPMV of the affine motion may be determined based on the first component of the first predictor associated with the first reference picture and the second component of the first predictor associated with the second reference picture.

[0201] In some embodiments, a first component of the initial predictor associated with the first reference picture of the current block may be determined based on one of a merge candidate, an advanced motion vector prediction (AMPV) candidate, and an affine merge candidate.

[0202] To determine a first component of the first predictor associated with the first reference picture of the current block, a first component of a gradient value of the first component of the initial predictor along a first direction may be determined. A second component of a gradient value of the first component of the initial predictor along a second direction may be determined. The second direction may be perpendicular to the first direction. A first component of the initial predictor along the first direction and a first component of a displacement associated with the first component of the first predictor may be determined according to a first affine model of the one or more affine models. A first component of the initial predictor along the second direction and a second component of a displacement associated with the first component of the first predictor may be determined according to a first affine model of the one or more affine models. The first component of the first predictor may be determined based on a sum of (i) the first component of the initial predictor, (ii) a product of the first component of the gradient value and the first component of the displacement, and (iii) a product of the second component of the gradient value and the second component of the displacement.

[0203] To determine a refined CPMV of the affine motion, the refined CPMV of the affine motion may be determined based on an Nth predictor associated with the first reference picture and an Nth predictor associated with the second reference picture in response to one of: (i) N is equal to an upper limit value of the iterative process; and (ii) a displacement based on the Nth predictor associated with the first reference picture and the (N+1)th predictor associated with the first reference picture is zero.

[0204] To determine the defined CPMV of the affine motion based on the N-th predictor associated with the first reference picture of the current block, a first component of a gradient value of the (N-1)-th predictor associated with the first reference picture of the current block along a first direction may be determined. A second component of a gradient value of the (N-1)-th predictor associated with the first reference picture of the current block along a second direction may be determined. A first component of a displacement associated with the N-th predictor associated with the first reference picture along the first direction and the (N-1)-th predictor associated with the first reference picture may be determined according to a second affine model of the one or more affine models. A second component of a displacement associated with the N-th predictor associated with the first reference picture along the second direction and the (N-1)-th predictor associated with the first reference picture may be determined according to a second affine model of the one or more affine models. The Nth predictor associated with the first reference picture of the current block may then be determined based on the sum of (i) the (N-1)th predictor associated with the first reference picture of the current block, (ii) the product of a first component of the gradient value of the (N-1)th predictor and a first component of the displacement, and (iii) the product of a second component of the gradient value of the (N-1)th predictor and a second component of the displacement.

[0205] In some embodiments, to determine the first CPMV of the affine motion, the first CPMV of the affine motion may be determined based on an Nth predictor associated with the first reference picture and an Nth predictor associated with the second reference picture.

[0206] In some embodiments, to determine the defined CPMV of the affine motion, the defined CPMV of the affine motion may be determined based on (i) an Mth predictor associated with the first reference picture derived from an Nth predictor associated with the first reference picture, and (ii) an Mth predictor associated with the second reference picture derived from an Nth predictor associated with the second reference picture, in response to one of: (i) M being equal to an upper limit value of the iterative process, and (ii) a displacement based on the Mth predictor associated with the first reference picture and the (M+1)th predictor associated with the first reference picture is zero.

[0207] In some embodiments, the first and second affine parameters for each of the one or more affine models may be determined based on a syntax element included in the prediction information. The syntax element may be included in one of a sequence parameter set, a picture parameter set, and a slice header.

[0208] After (S1940), the process proceeds to (S1999) and ends.

[0209] The process (1900) may be adapted as appropriate. Step(s) of the process (1900) may be modified and / or omitted. Additional step(s) may be added. Any suitable order of implementation may be used.

[0210] As shown in FIG. 20, the process (2000) can start at (S2001) and proceed to (S2010). In (S2010), a first affine parameter and a second affine parameter for each of one or more affine models associated with the affine bilateral matching can be determined. The affine bilateral matching can be applied to a current block in a current picture. The first affine parameter of each affine model can be associated with a first motion vector of a first reference picture of the current block. The second affine parameter of each affine model can be associated with a second motion vector for a second reference picture of the current block. The first affine parameter and the second affine parameter can have opposite signs.

[0211] In (S2020), a CPMV of the current block can be determined by performing affine bilateral matching based on the determined first affine parameters and the determined second affine parameters of the one or more affine models.

[0212] In step S2030, prediction information for the current block can be generated based on the determined CPMV of the current block.

[0213] The process then proceeds to (S2099) and ends.

[0214] The process (2000) may be adapted as appropriate. Step(s) in the process (2000) may be modified and / or omitted. Additional step(s) may be added. Any suitable order of implementation may be used.

[0215] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 21 illustrates a computer system (2100) suitable for implementing certain embodiments of the disclosed subject matter.

[0216] The computer software may be coded using any suitable machine code or computer language that may undergo assembly, compilation, linking, or similar mechanisms to generate code including instructions that may be executed by one or more computer central processing units (CPUs) and / or graphics processing units (GPUs), either directly, or through interpretation and execution of microcode, etc.

[0217] The instructions may be executed in various types of computers or components thereof including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, and the like.

[0218] 21 for the computer system (2100) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The arrangement of components should not be interpreted as having any dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of the computer system (2100).

[0219] The computer system (2100) may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, via tactile input (such as keystrokes, swipes, data glove movements), audio input (such as voice, clapping), visual input (such as gestures), or olfactory input (not shown). Human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as sound (such as speech, music, ambient sounds, etc.), images (such as scanned images, photographic images obtained from a still image camera, etc.), and video (such as two-dimensional video, three-dimensional video including stereoscopic video, etc.).

[0220] The input human interface devices may include one or more of a keyboard (2101), a mouse (2102), a trackpad (2103), a touch screen (2110), a data glove (not shown), a joystick (2105), a microphone (2106), a scanner (2107), and a camera (2108) (only one of each is shown).

[0221] The computer system (2100) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through haptic output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touch screen (2110), data gloves (not shown), or joystick (2105), although there may also be haptic feedback devices that do not function as input devices), audio output devices (such as speakers (2109), headphones (not shown)), visual output devices (such as screens (2110), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capability, each with or without haptic feedback capability, some of which may be capable of outputting two-dimensional visual output or three- or more-dimensional output via means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0222] The computer system (2100) may also include human-accessible storage devices and their associated media, such as optical media, including CD / DVD ROM / RW (2120) with CD / DVD or similar media (2121), thumb drives (2122), removable hard drives or solid state drives (2123), legacy magnetic media such as tapes and floppy disks (not shown), and specialized ROM / ASIC / PLD based devices (not shown) such as security dongles.

[0223] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.

[0224] The computer system (2100) may also include an interface (2154) to one or more communication networks (2155). The networks may be, for example, wireless, wired, optical. The networks may further be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial including CANBus, etc. Certain networks typically require an external network interface adapter attached to a specific general-purpose data port (e.g., a USB port of the computer system (2100)) or peripheral bus (2149), while other networks are typically integrated into the core of the computer system (2100) by attaching to a system bus described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (2100) can communicate with other entities. Such communications can be one-way receive only (e.g., broadcast TV), one-way transmit only (e.g., CANbus to a specific CANbus device), or two-way with other computer systems using, for example, local or wide area digital networks. Specific protocols and protocol stacks may be used with each of these networks and network interfaces, as described above.

[0225] The aforementioned human interface devices, human accessible storage devices, and network interfaces may be attached to the core (2140) of the computer system (2100).

[0226] The cores (2140) may include specialized programmable processing devices in the form of one or more central processing units (CPUs) (2141), graphics processing units (GPUs) (2142), field programmable gate areas (FPGAs) (2143), hardware accelerators for specific tasks (2144), graphics adapters (2150), and the like. These devices may be connected via a system bus (2148), along with read-only memory (ROM) (2145), random access memory (2146), and internal mass storage (2147), such as an internal hard drive or SSD that is not accessible to the user. In some computer systems, the system bus (2148) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, and the like. Peripheral devices may be attached directly to the core's system bus (2148) or via a peripheral bus (2149). In one example, a screen (2110) may be connected to a graphics adapter (2150). Architectures for peripheral buses include PCI, USB, etc.

[0227] The CPU (2141), GPU (2142), FPGA (2143), and accelerator (2144) can execute certain instructions that can combine to constitute the aforementioned computer code. The computer code can be stored in ROM (2145) or RAM (2146). Transient data can also be stored in RAM (2146), and persistent data can be stored, for example, in internal mass storage (2147). Rapid storage and retrieval from any of the memory devices can be enabled using cache memory, which can be closely associated with one or more of the CPU (2141), GPU (2142), mass storage (2147), ROM (2145), RAM (2146), etc.

[0228] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.

[0229] As an example, but not by way of limitation, a computer system (2100) having an architecture, specifically a core (2140), can provide functionality as a result of a processor(s) (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be the user-accessible mass storage introduced above, as well as media associated with specific storage of the core (2140) of a non-transitory nature, such as the core internal mass storage (2147) or ROM (2145). Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (2140). The computer-readable media can include one or more memory devices or chips, depending on the particular needs. The software can cause the core (2140), and specifically the processors (including CPUs, GPUs, FPGAs, etc.) therein, to perform certain processes or certain portions of certain processes described herein, including defining data structures stored in RAM (2146) and modifying such data structures according to the software-defined processes. Additionally, or alternatively, the computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator (2144)) that may operate in place of or together with software to perform certain processes or certain portions of certain processes described herein. Where appropriate, references to software may encompass logic and vice versa. Where appropriate, references to computer-readable media may encompass circuitry (such as integrated circuits (ICs)) that store software for execution, circuitry that embodies logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software. Appendix A: Acronyms JEM: Joint Exploration Model VVC: Versatile Video Coding BMS: Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplemental Extended Information VUI: Video Usability Information GOP: Groups of Pictures TU: conversion unit PU: Prediction unit CTU: Coding Tree Unit CTB: coding tree block PBs: Predicted blocks HRD: Hypothetical Reference Decoder SNR: Signal to Noise Ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: cathode ray tube LCD: Liquid crystal display device OLED: Organic Light Emitting Diode CD:Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE: Long Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Area SSD: Solid State Drive IC: Integrated Circuit CU: coding unit

[0230] While this disclosure has described several exemplary embodiments, there are modifications, substitutions, and various substitute equivalents which are within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods which, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure. [Explanation of symbols]

[0231] 101 sample, 102 arrow, 103 arrow, 104 square block, 110 schematic diagram, 201 current block, 300 communication system, 310 terminal device, 320 terminal device, 330 terminal device, 340 terminal device, 350 network, 401 video source, 402 stream, 403 video encoder, 404 video data, 405 streaming server, 406 client subsystem, 407 copy, 408 client subsystem, 409 copy, 410 video decoder, 413 capture subsystem, 420 electronic device, 430 electronic device, 501 channel, 510 video decoder, 512 rendering device, 515 buffer memory, 520 parser, 521 symbol, 530 electronic device, 531 receiver, 551 scaler / inverse transform unit, 552 intra-picture prediction unit, 553 motion compensation prediction unit, 555 aggregator, 556 loop filter unit, 557 reference picture memory, 558 current picture buffer, 603 video encoder, 620 electronic device, 630 source coder, 632 coding engine, 633 decoder, 634 reference picture memory, 635 predictor, 640 transmitter, 643 video sequence, 645 entropy coder, 650 controller, 660 communication channel, 703 video encoder, 721 controller, 722 intra encoder, 723 residual calculator, 724 residual encoder, 725 entropy encoder, 726 switch, 728 residual decoder, 730 inter encoder, 810 video decoder, 871 entropy decoder, 872 intra decoder, 873 residual decoder, 874 reconstruction, 880 inter decoder, 902 block, 904 block, 1000 current block, 1202 CU, 1204 current block, 1302 current block, 1400 current block, 1402 subblock, 1404 sample, 1406 reference pixel, 1408 reference pixel, 1500 affine ME, 1506 affine bi-prediction, 1600 affine ME process, 1702 row / column, 1704 CU, 1802 current picture, 1804Reference picture list L0, 1806 Reference picture list L1, 1808 Current block, 1810 Candidate reference block, 1812 Initial reference block, 1814 Initial reference block, 1816 First candidate reference block, 1900 Process, 2000 Process, 2100 Computer system, 2101 Keyboard, 2102 Mouse, 2103 Track pad, 2105 Joystick, 2106 Microphone, 2107 Scanner, 2108 Camera, 2109 Speaker, 2110 Touch screen, 2121 Media, 2122 Thumb drive, 2123 Solid state drive, 2140 Core, 2141 Central processing unit, 2142 Graphics processing unit, 2143 Field programmable gate area, 2144 Hardware accelerator, 2145 Read only memory, 2146 Random access memory, 2147 Internal mass storage, 2148 system bus, 2149 peripheral bus, 2150 graphics adapter, 2154 interface, 2155 communications network

Claims

1. 1. A method of video decoding implemented in a video decoder, the method comprising: receiving prediction information for a current block in a current picture from a coded video bitstream, the prediction information indicating that the current block is predicted based on affine bilateral matching, in which an affine motion of the current block is derived based on a plurality of reference blocks in a first reference picture and a second reference picture associated with the current picture, the current picture being between the first reference picture and the second reference picture; determining a first affine parameter and a second affine parameter for each of one or more affine models associated with the affine bilateral matching, the first affine parameter of the respective affine model being associated with a first motion vector for the first reference picture and the second affine parameter of the respective affine model being associated with a second motion vector for the second reference picture, the first affine parameter and the second affine parameter having opposite signs; determining a control point motion vector (CPMV) of the affine motion of the current block by performing the affine bilateral matching based on the determined first affine parameters and the determined second affine parameters of the one or more affine models; reconstructing the current block based on the CPMV of the affine motion of the current block; A method comprising:

2. The step of determining the CPMV comprises: determining a first CPMV of the affine motion of the current block based on a first affine model of the one or more affine models; determining a refined CPMV of the affine motion of the current block through an iterative process based on the first CPMV and a second affine model of the one or more affine models, the second affine model being the same as or different from the first affine model; Further comprising: The step of reconstructing comprises: reconstructing the current block based on the refined CPMV of the affine motion of the current block. The method of claim 1, further comprising:

3. The one or more affine models include a three-parameter zooming model, the three-parameter zooming model being: a first component of a first motion vector of a sample in the current block associated with the first reference picture along a first direction, the first component of the first motion vector being equal to a first component of a translation coefficient along the first direction plus a product of a position of the sample along the first direction and a zoom factor; a second component of the first motion vector of the sample in the current block associated with the first reference picture along a second direction, the second component of the first motion vector being equal to a sum of a second component of the translation coefficient along the second direction and a product of the position of the sample along the second direction and the zoom factor, the second direction being perpendicular to the first direction; a first component of a second motion vector of the sample in the current block associated with the second reference picture along the first direction, the first component of the second motion vector being equal to a sum of a first component of an opposite translation coefficient along the first direction and a product of the position of the sample along the first direction and an opposite zoom coefficient; a second component of the second motion vector of the sample in the current block associated with the second reference picture along the second direction, the second component of the second motion vector being equal to a sum of a second component of the opposite translation coefficient along the second direction and a product of the position of the sample along the second direction and the opposite zoom coefficient; The method of claim 1, further comprising:

4. The one or more affine models include a three-parameter rotation model, the three-parameter rotation model being: a first component of a first motion vector of a sample in the current block associated with the first reference picture along a first direction, the first component of the first motion vector being equal to a first component of a translation coefficient along the first direction plus a product of a position of the sample along a second direction and a rotation coefficient, the second direction being perpendicular to the first direction; a second component of the first motion vector of the sample in the current block associated with the first reference picture along the second direction, the second component of the first motion vector being equal to a sum of a second component of the translation coefficient along the second direction and a product of the position of the sample along the first direction and an inverse rotation coefficient; a first component of a second motion vector of the sample in the current block associated with the second reference picture along the first direction, the first component of the second motion vector being equal to a first component of an inverse translation coefficient in the first direction plus a product of the position of the sample along the second direction and the inverse rotation coefficient; a second component of the second motion vector of the sample in the current block associated with the second reference picture along the second direction, the second component of the second motion vector being equal to a sum of a second component of the opposite translation coefficient along the second direction and a product of the position of the sample along the first direction and the rotation coefficient; The method of claim 1, further comprising:

5. The one or more affine models include a four-parameter zooming model, the four-parameter zooming model being: a first component of a first motion vector of a sample in the current block associated with the first reference picture along a first direction, the first component of the first motion vector being equal to a sum of a first component of a translation coefficient along the first direction and a product of a position of the sample along the first direction and a first component of a zoom coefficient in the first direction; a second component of the first motion vector of the sample in the current block associated with the first reference picture along a second direction, the second component of the first motion vector being equal to the sum of a second component of the translation coefficient along the second direction and a product of the position of the sample along the second direction and a second component of the zoom factor along the second direction, the second direction being perpendicular to the first direction; a first component of a second motion vector of the sample in the current block associated with the second reference picture along a first direction, the first component of the second motion vector being equal to a sum of a first component of an opposite translation coefficient along the first direction and a product of the position of the sample along the first direction and a first component of an opposite zoom coefficient along the first direction; a second component of the second motion vector of the sample in the current block associated with the second reference picture along the second direction, the second component of the second motion vector being equal to the sum of a second component of the opposite translation coefficient along the second direction and a product of the position of the sample along the second direction and a second component of the opposite zoom coefficient along the second direction; The method of claim 1, further comprising:

6. The one or more affine models include a four-parameter rotation model, the four-parameter rotation model being: a first component of a first motion vector of a sample in the current block associated with a first reference picture along a first direction, the first component of the first motion vector being equal to the sum of (i) a first component of a translation coefficient along the first direction, (ii) a product of a position of the sample along the first direction and a first coefficient associated with a rotation coefficient, and (iii) a product of the position of the sample along a second direction and a second coefficient associated with the rotation coefficient, the first direction being perpendicular to the second direction; a second component of the first motion vector of the sample in the current block associated with the first reference picture along the second direction, the second component of the first motion vector being equal to the sum of (i) a second component of the translation coefficient along the second direction, (ii) a product of the position of the sample along the first direction and an opposite second coefficient associated with the rotation coefficient, and (iii) a product of the position of the sample along the second direction and the first coefficient associated with the rotation coefficient; a first component of a second motion vector of the sample in the current block associated with the second reference picture along the first direction, the first component of the second motion vector being equal to the sum of (i) a first component of an opposite translation coefficient along the first direction, (ii) a product of the position of the sample along the first direction and the first coefficient associated with the rotation coefficient, and (iii) a product of the position of the sample along the second direction and the opposite second coefficient associated with the rotation coefficient; a second component of the second motion vector of the sample in the current block associated with the second reference picture along the second direction, the second component of the second motion vector being equal to the sum of (i) a second component of an opposite translation coefficient along the second direction, (ii) a product of the position of the sample along the first direction and the second coefficient associated with the rotation coefficient, and (iii) a product of the position of the sample along the second direction and the first coefficient associated with the rotation coefficient; The method of claim 1, further comprising:

7. The step of determining the first CPMV of the affine motion of the current block comprises: determining a first component of an initial predictor associated with the first reference picture of the current block; determining a second component of the initial predictor associated with the second reference picture of the current block; determining a first component of a first predictor associated with the first reference picture of the current block based on the first component of the initial predictor associated with the first reference picture of the current block; determining a second component of the first predictor associated with the second reference picture of the current block based on the second component of the initial predictor associated with the second reference picture of the current block; determining the first CPMV of the affine motion based on the first component of the first predictor associated with the first reference picture and the second component of the first predictor associated with the second reference picture; The method of claim 2, further comprising:

8. 8. The method of claim 7, wherein the first component of the initial predictor associated with the first reference picture of the current block is determined based on one of a merge candidate, an advanced motion vector prediction (AMPV) candidate, and an affine merge candidate.

9. The step of determining a first component of the first predictor associated with the first reference picture of the current block comprises: determining a first component of a gradient value of the first component of the initial predictor along a first direction; determining a second component of the gradient value of the first component of the initial predictor along a second direction, the second direction being perpendicular to the first direction; determining a first component of a displacement associated with the first component of the initial predictor and the first component of the first predictor along the first direction according to the first affine model of the one or more affine models; determining a second component of the displacement associated with the first component of the initial predictor and the first component of the first predictor along the second direction according to the first affine model of the one or more affine models; determining the first component of the first predictor based on the sum of (i) the first component of the initial predictor, (ii) the product of the first component of the gradient value and the first component of the displacement, and (iii) the product of the second component of the gradient value and the second component of the displacement; 9. The method of claim 8, further comprising:

10. The step of determining the refined CPMV of the affine motion comprises: Based on an N-th predictor associated with the first reference picture and the N-th predictor associated with the second reference picture, (i) N is equal to an upper limit of the iterative process; and (ii) a displacement based on the Nth predictor associated with the first reference picture and an (N+1)th predictor associated with the first reference picture is zero. determining the refined CPMV of the affine motion in response to one of 10. The method of claim 9, further comprising:

11. The step of determining the refined CPMV of the affine motion based on the N-th predictor associated with the first reference picture of the current block comprises: determining a first component along the first direction of a gradient value of an (N-1)th predictor associated with the first reference picture of the current block; determining a second component along the second direction of the gradient value of the (N-1)th predictor associated with the first reference picture of the current block; determining a first component of a displacement associated with the Nth predictor associated with the first reference picture and the (N-1)th predictor associated with the first reference picture along the first direction according to the second affine model of the one or more affine models; determining a second component of the displacement associated with the Nth predictor associated with the first reference picture and the (N-1)th predictor associated with the first reference picture along the second direction according to a second affine model of the one or more affine models; determining the N-th predictor associated with the first reference picture of the current block based on a sum of: (i) the (N-1)-th predictor associated with the first reference picture of the current block; (ii) a product of the first component of the gradient value of the (N-1)-th predictor and the first component of the displacement; and (iii) a product of the second component of the gradient value of the (N-1)-th predictor and the second component of the displacement.

11. The method of claim 10, further comprising:

12. The step of determining the first CPMV of the affine motion comprises: determining the first CPMV of the affine motion based on the N-th predictor associated with the first reference picture and the N-th predictor associated with the second reference picture; 12. The method of claim 11, further comprising:

13. The step of determining the refined CPMV of the affine motion comprises: based on (i) an Mth predictor associated with the first reference picture, the Mth predictor being derived from the Nth predictor associated with the first reference picture, and (ii) an Mth predictor associated with the second reference picture, the Mth predictor being derived from the Nth predictor associated with the second reference picture; (i) M is equal to an upper limit for said iterative process; and (ii) a displacement based on the Mth predictor associated with the first reference picture and an (M+1)th predictor associated with the first reference picture is zero. determining the refined CPMV of the affine motion in response to one of 13. The method of claim 12, further comprising:

14. The step of determining the first affine parameters and the second affine parameters comprises: determining the first affine parameters and the second affine parameters for each of the one or more affine models based on a syntax element included in the prediction information, the syntax element being included in one of a sequence parameter set, a picture parameter set, and a slice header; The method of claim 1.

15. An apparatus comprising: A processing circuit, receiving prediction information for a current block in a current picture from a coded video bitstream, the prediction information indicating that the current block is predicted based on affine bilateral matching, in which an affine motion of the current block is derived based on a plurality of reference blocks in a first reference picture and a second reference picture associated with the current picture, the current picture being between the first reference picture and the second reference picture; determining a first affine parameter and a second affine parameter for each of one or more affine models associated with the affine bilateral matching, the first affine parameter of the respective affine model being associated with a first motion vector for the first reference picture and the second affine parameter of the respective affine model being associated with a second motion vector for the second reference picture, the first affine parameter and the second affine parameter having opposite signs; determining a control point motion vector (CPMV) of the affine motion of the current block by performing the affine bilateral matching based on the determined first affine parameters and the determined second affine parameters of the one or more affine models; reconstructing the current block based on the CPMV of the affine motion of the current block A processing circuit configured to An apparatus comprising:

16. The processing circuitry includes: determining a first CPMV of the affine motion of the current block based on a first affine model of the one or more affine models; determining a refined CPMV of the affine motion of the current block through an iterative process based on the first CPMV and a second affine model of the one or more affine models, the second affine model being the same as or different from the first affine model; reconstructing the current block based on the refined CPMV of the affine motion of the current block; The apparatus of claim 15, configured to:

17. The one or more affine models include a three-parameter zooming model, the three-parameter zooming model being: a first component of a first motion vector of a sample in the current block associated with the first reference picture along a first direction, the first component of the first motion vector being equal to a first component of a translation coefficient along the first direction plus a product of a position of the sample along the first direction and a zoom factor; a second component of the first motion vector of the sample in the current block associated with the first reference picture along a second direction, the second component of the first motion vector being equal to a sum of a second component of the translation coefficient along the second direction and a product of the position of the sample along the second direction and the zoom factor, the second direction being perpendicular to the first direction; a first component of a second motion vector of the sample in the current block associated with the second reference picture along the first direction, the first component of the second motion vector being equal to a sum of a first component of an opposite translation coefficient along the first direction and a product of the position of the sample along the first direction and an opposite zoom coefficient; a second component of the second motion vector of the sample in the current block associated with the second reference picture along the second direction, the second component of the second motion vector being equal to a sum of a second component of the opposite translation coefficient along the second direction and a product of the position of the sample along the second direction and the opposite zoom coefficient; 16. The apparatus of claim 15, further comprising:

18. The one or more affine models include a three-parameter rotation model, the three-parameter rotation model being: a first component of a first motion vector of a sample in the current block associated with the first reference picture along a first direction, the first component of the first motion vector being equal to a first component of a translation coefficient along the first direction plus a product of a position of the sample along a second direction and a rotation coefficient, the second direction being perpendicular to the first direction; a second component of the first motion vector of the sample in the current block associated with the first reference picture along the second direction, the second component of the first motion vector being equal to a sum of a second component of the translation coefficient along the second direction and a product of the position of the sample along the first direction and an inverse rotation coefficient; a first component of a second motion vector of the sample in the current block associated with the second reference picture along the first direction, the first component of the second motion vector being equal to a first component of an inverse translation coefficient in the first direction plus a product of the position of the sample along the second direction and the inverse rotation coefficient; a second component of the second motion vector of the sample in the current block associated with the second reference picture along the second direction, the second component of the second motion vector being equal to a sum of a second component of the opposite translation coefficient along the second direction and a product of the position of the sample along the first direction and the rotation coefficient; 16. The apparatus of claim 15, further comprising:

19. The one or more affine models include a four-parameter zooming model, the four-parameter zooming model being: a first component of a first motion vector of a sample in the current block associated with the first reference picture along a first direction, the first component of the first motion vector being equal to a sum of a first component of a translation coefficient along the first direction and a product of a position of the sample along the first direction and a first component of a zoom coefficient in the first direction; a second component of the first motion vector of the sample in the current block associated with the first reference picture along a second direction, the second component of the first motion vector being equal to the sum of a second component of the translation coefficient along the second direction and a product of the position of the sample along the second direction and a second component of the zoom factor along the second direction, the second direction being perpendicular to the first direction; a first component of a second motion vector of the sample in the current block associated with the second reference picture along a first direction, the first component of the second motion vector being equal to a sum of a first component of an opposite translation coefficient along the first direction and a product of the position of the sample along the first direction and a first component of an opposite zoom coefficient along the first direction; a second component of the second motion vector of the sample in the current block associated with the second reference picture along the second direction, the second component of the second motion vector being equal to the sum of a second component of the opposite translation coefficient along the second direction and a product of the position of the sample along the second direction and a second component of the opposite zoom coefficient along the second direction; 16. The apparatus of claim 15, further comprising:

20. The one or more affine models include a four-parameter rotation model, the four-parameter rotation model being: a first component of a first motion vector of a sample in the current block associated with a first reference picture along a first direction, the first component of the first motion vector being equal to the sum of (i) a first component of a translation coefficient along the first direction, (ii) a product of a position of the sample along the first direction and a first coefficient associated with a rotation coefficient, and (iii) a product of the position of the sample along a second direction and a second coefficient associated with the rotation coefficient, the first direction being perpendicular to the second direction; a second component of the first motion vector of the sample in the current block associated with the first reference picture along the second direction, the second component of the first motion vector being equal to the sum of (i) a second component of the translation coefficient along the second direction, (ii) a product of the position of the sample along the first direction and an opposite second coefficient associated with the rotation coefficient, and (iii) a product of the position of the sample along the second direction and the first coefficient associated with the rotation coefficient; a first component of a second motion vector of the sample in the current block associated with the second reference picture along the first direction, the first component of the second motion vector being equal to the sum of (i) a first component of an opposite translation coefficient along the first direction, (ii) a product of the position of the sample along the first direction and the first coefficient associated with the rotation coefficient, and (iii) a product of the position of the sample along the second direction and the opposite second coefficient associated with the rotation coefficient; a second component of the second motion vector of the sample in the current block associated with the second reference picture along the second direction, the second component of the second motion vector being equal to the sum of (i) a second component of an opposite translation coefficient along the second direction, (ii) a product of the position of the sample along the first direction and the second coefficient associated with the rotation coefficient, and (iii) a product of the position of the sample along the second direction and the first coefficient associated with the rotation coefficient; 16. The apparatus of claim 15, further comprising:

Citation Information

Patent Citations

  • Improvements to frame rate upconversion coding mode

    JP2019534622A

  • Bilateral matching using affine motion

    JP7675206B2

  • Symmetric merge mode motion vector coding

    WO2020185925A1

  • Motion refinement with bilateral matching for affine motion compensation in video coding

    WO2022266328A1