Symmetric affine modes

The symmetric affine mode in video coding improves compression efficiency by optimizing intra-prediction and motion compensation through a four-parameter affine model, addressing inefficiencies in existing technologies.

JP7807155B2Active Publication Date: 2026-01-27TENCENT AMERICA LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023560169
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-10-07
Filing Date
2022-10-11
Publication Date
2026-01-27
Estimated Expiration
2042-10-11

AI Technical Summary

Technical Problem

Existing video coding technologies face inefficiencies in intra-prediction and motion compensation, particularly in handling complex motion vectors and affine transformations, leading to suboptimal compression ratios and increased data requirements.

Method used

The introduction of a symmetric affine mode for video encoding/decoding, utilizing a four-parameter affine model with specific parameter relationships to determine control point motion vectors, allowing for more efficient reconstruction of video blocks based on multiple reference pictures.

Benefits of technology

Enhances video compression efficiency by reducing redundancy in intra-prediction and motion compensation, leading to improved compression ratios and reduced data requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007807155000017
    Figure 0007807155000017
  • Figure 0007807155000018
    Figure 0007807155000018
  • Figure 0007807155000019
    Figure 0007807155000019
Patent Text Reader

Abstract

A symmetric affine mode is applied to the current block. A first affine parameter of a first affine model of the symmetric affine mode is determined. The first affine model is associated with the current block and a first reference block of the current block in a first reference picture. A second affine parameter of a second affine model of the symmetric affine mode is derived based on the first affine parameter of the first affine model. The second affine model is associated with the current block and a second reference block of the current block in a second reference picture. The first affine parameter and the second affine parameter have one of opposite signs, inverse values, and a proportional relationship. A control point motion vector (CPMV) of the current block is determined based on the first and second affine models. The current block is reconstructed based on the determined CPMV.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Application No. 63 / 328,589, entitled "Symmetric Affine Mode," filed April 7, 2022, which claims the benefit of priority to U.S. Patent Application No. 17 / 962,422, entitled "SYMMETRIC AFFINE MODE," filed October 7, 2022. The disclosures of the prior applications are incorporated herein by reference in their entireties.

[0002] This disclosure generally describes embodiments related to video coding. [Background technology]

[0003] The background art discussion provided herein is intended to generally present the context for the present disclosure. The inventors' work, to the extent that it is described in this background art section, and aspects of the description that may not be admitted as prior art at the time of filing, are not admitted expressly or implicitly as prior art to the present disclosure.

[0004] Uncompressed digital images and / or video can include a series of pictures, each with spatial dimensions of, for example, 1920 x 1080 luma samples and associated chroma samples. The series of pictures can have a fixed or variable picture rate (informally known as a frame rate) of, for example, 60 pictures per second, or 60 Hz. Uncompressed images and / or video have specific bitrate requirements. For example, 1080p60 4:2:0 video (1920 x 1080 luma sample resolution at a 60 Hz frame rate) with 8 bits per sample requires a bandwidth approaching 1.5 Gbit / s. One hour of such video requires more than 600 GBytes of storage space.

[0005] One goal of image and / or video coding and decoding may be reducing redundancy in the input image and / or video signal through compression. Compression can help reduce the aforementioned bandwidth and / or storage space requirements, in some cases by two or more orders of magnitude. While the description herein uses video encoding / decoding as an illustrative example, the same techniques can be applied to image encoding / decoding in a similar manner without departing from the spirit of this disclosure. Both lossless and lossy compression, as well as combinations thereof, can be employed. Lossless compression refers to techniques in which an exact copy of the original signal can be reconstructed from a compressed version of the original signal. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signal is small enough to make the reconstructed signal useful for its intended use. For video, lossy compression is widely adopted. The amount of acceptable distortion depends on the application; for example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect that the higher the tolerable / acceptable distortion, the higher the compression ratio.

[0006] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transform processing, quantization, and entropy coding.

[0007] Video codec technology can include a technique known as intra-coding. In intra-coding, sample values ​​are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, pictures are spatially subdivided into blocks of samples. When all blocks of samples are coded in intra mode, the picture may be an intra-picture. Intra-pictures and their derivatives, such as independent decoder refresh pictures, are used to reset the decoder state and can therefore be used as the first picture in a coded video bitstream and video session, or as a still image. Samples in intra-blocks undergo a transform, and the transform coefficients may be quantized before entropy coding. Intra-prediction can be a technique that minimizes sample values ​​in the pre-transform domain. In some cases, the smaller the DC value and AC coefficients after the transform, the fewer bits required for a given quantization step size to represent the block after entropy coding.

[0008] Conventional intra-coding, for example, as used in MPEG-2 generation coding techniques, does not use intra-prediction. However, some newer video compression techniques include techniques that attempt to perform prediction based on surrounding sample data and / or metadata obtained during encoding / decoding of a block of data. Such techniques are hereinafter referred to as "intra-prediction" techniques. Note that, at least in some cases, intra-prediction uses only reference data from the current picture being reconstructed, and not from a reference picture.

[0009] Intra-prediction can take many different forms. When two or more of such techniques can be used in a given video coding technique, the particular technique in use can be coded as a particular intra-prediction mode that uses the particular technique. In certain cases, an intra-prediction mode can have sub-modes and / or parameters, which can be coded separately or included in a mode codeword that defines the prediction mode used. The codeword used for a given mode, sub-mode, and / or parameter combination can affect coding efficiency via intra-prediction, and therefore can also affect the entropy coding technique used to convert the codeword into a bitstream.

[0010] Certain modes of intra prediction were introduced in H.264, improved in H.265, and further refined in new coding techniques such as the Joint Search Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). Neighboring sample values ​​of already available samples can be used to form a predictor block. The sample values ​​of the neighboring samples are copied to the predictor block according to their direction. A reference to the direction in use can be coded in the bitstream, or it can be predicted itself.

[0011] Referring to FIG. 1A, depicted at the bottom right is a subset of nine known predictor directions from the 33 possible predictor directions defined in H.265 (corresponding to the 33 angular modes out of the 35 intra modes). The point where the arrows converge (101) represents the sample being predicted. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples to the upper right at an angle of 45 degrees from horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples to the lower left of sample (101) at an angle of 22.5 degrees from horizontal.

[0012] 1A , a square block (104) of 4×4 samples (indicated by a thick dashed line) is shown in the upper left. The square block (104) contains 16 samples, each labeled with “S,” its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample (from the top) in the Y dimension and the first sample (from the left) in the X dimension. Similarly, sample S44 is the fourth sample in both the Y and X dimensions within the block (104). Because the block is 4×4 samples in size, S44 is located in the lower right. Reference samples, which follow a similar numbering scheme, are also shown. The reference samples are labeled R, their Y position (e.g., row index), and their X position (column index) relative to the block (104). In both H.264 and H.265, predicted samples are neighbors of the block being reconstructed, so negative values ​​need not be used.

[0013] Intra-picture prediction can work by copying reference sample values ​​from neighboring samples indicated by the signaled prediction direction. For example, assume that the coded video bitstream includes signaling indicating a prediction direction consistent with the arrow (102) for this block, i.e., the sample is predicted from the sample to the upper right at a 45-degree angle from the horizontal. In this case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.

[0014] In certain cases, to calculate a reference sample, especially when the direction is not evenly divisible by 45 degrees, the values ​​of multiple reference samples may be combined, for example by interpolation.

[0015] The number of possible directions has increased as video coding technology has evolved. H.264 (2003) allowed for nine different directions to be represented. This increased to 33 in H.265 (2013). Currently, JEM / VVC / BMS can support up to 65 directions. Experiments have been conducted to identify the most likely directions, and specific entropy coding techniques are used to represent these likely directions with a small number of bits, accepting a certain penalty for less likely directions. Furthermore, the direction itself may be predictable from nearby directions used in previously decoded blocks.

[0016] Figure 1B shows a schematic diagram (110) showing 65 intra-prediction directions with JEM to illustrate the increasing number of prediction directions over time.

[0017] The mapping of intra-prediction direction bits, which represent directions within a coded video bitstream, may vary depending on the video coding technique. Such mappings may range from simple direct mappings to complex adaptive schemes including codewords, most probable modes, and similar techniques. However, in most cases, there may be certain directions that are statistically less likely to occur within the video content than certain other directions. Because the goal of video compression is to reduce redundancy, these less likely directions are represented with more bits than more likely directions in well-performing video coding techniques.

[0018] Image and / or video coding and decoding can be performed using inter-picture prediction with motion compensation. Motion compensation may be a lossy compression technique and may refer to a technique used to predict a newly reconstructed picture or portion of a picture after blocks of sample data from a previously reconstructed picture or portion thereof (reference picture) are spatially shifted in a direction indicated by a motion vector (hereinafter, MV). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or three dimensions, with the third dimension being an indication of the reference picture in use (the third dimension may indirectly be a temporal dimension).

[0019] In some video compression techniques, the MV applicable to a particular area of ​​sample data can be predicted from other MVs, e.g., from MVs associated with other areas of sample data that are spatially adjacent to the area being reconstructed and precede that MV in decoding order. Doing so can significantly reduce the amount of data required to code the MV, thereby eliminating redundancy and increasing compression ratios. For example, when coding an input video signal derived from a camera (known as natural video), MV prediction can work effectively because there is a statistical likelihood that areas larger than the area to which a single MV is applicable move in similar directions and, therefore, in some cases, can be predicted using similar motion vectors derived from MVs of nearby areas. As a result, the detected MV for a given area is similar or identical to the MV predicted from surrounding MVs, which, after entropy coding, can be represented with fewer bits than would be used if the MV were coded directly. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., an MV) derived from an original signal (i.e., a sample stream). In other cases, MV prediction itself can be lossy, for example, due to rounding errors when calculating a predictor from several surrounding MVs.

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Recommendation H.265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms provided by H.265, the one described with reference to Figure 2 is a technique hereafter referred to as "spatial merging".

[0021] Referring to Figure 2, a current block (201) contains samples that the encoder discovers during the motion search process are predictable from a spatially shifted previous block of the same size. Instead of directly coding its MV, the MV can be derived from metadata associated with one or more reference pictures, e.g., from the most recent reference picture (in decoding order), using the MV associated with any one of five surrounding samples, denoted A0, A1, and B0, B1, B2 (202-206, respectively). In H.265, MV prediction can use predictors from the same reference picture used by neighboring blocks. Summary of the Invention [Means for solving the problem]

[0022]

[0006] Aspects of the present disclosure provide a method and apparatus for video encoding / decoding. In some examples, an apparatus for video decoding includes a processing circuit.

[0023] According to one aspect of the present disclosure, a method of video decoding performed in a video decoder is provided. The method may receive coded information for a current block in a current picture from a coded video bitstream. The coded information may include a flag indicating whether a symmetric affine mode applies to the current block. In response to the flag indicating that the symmetric affine mode applies to the current block, first affine parameters of a first affine model for the symmetric affine mode may be determined from the received coded information. The first affine model may be associated with the current block and a first reference block for the current block in a first reference picture of the current picture. Second affine parameters of a second affine model for the symmetric affine mode may be derived based on the first affine parameters of the first affine model. The second affine model may be associated with the current block and a second reference block for the current block in a second reference picture of the current picture. The first and second affine parameters may have one of opposite signs, inverse values, and a proportional relationship based on a first temporal distance between the first reference picture and the current picture and a second temporal distance between the second reference picture and the current picture. A control point motion vector (CPMV) of the current block may be determined based on the first and second affine models. The current block may be reconstructed based on the determined CPMV of the current block.

[0024] In some embodiments, the flag may be coded via one of a context-adaptive binary arithmetic coding (CABAC) context and a bypass code.

[0025] In some embodiments, the symmetric affine mode may be determined to be associated with a four-parameter affine model in response to a flag indicating that the symmetric affine mode applies to the current block.

[0026] In some embodiments, the flag may indicate that symmetric affine mode is applied to the current block based on a first temporal distance between the current picture and a first reference picture being equal to a second temporal distance between the current picture and a second reference picture.

[0027] In response to the flag indicating that the symmetric affine mode is applied to the current block, reference index information can be derived. The reference index information can indicate which reference picture in the first reference list is the first reference picture and which reference picture in the second reference list is the second reference picture.

[0028] The first affine parameters can include a first translation factor and at least one of a first zoom factor or a first rotation factor, and the second affine parameters can include a second translation factor and at least one of a second zoom factor or a second rotation factor.

[0029] In one example, the sum of the first rotation factor and the second rotation factor can be zero, the sum of the first translation factor and the second translation factor can be zero, and the product of the first zoom factor and the second zoom factor can be one.

[0030] In one example, the ratio of the first rotation factor to the second rotation factor can be linearly proportional to the ratio of the first temporal distance to the second temporal distance, and the ratio of the first zoom factor to the second zoom factor can be exponentially proportional to the ratio of the first temporal distance to the second temporal distance.

[0031] According to another aspect of the present disclosure, a method of video encoding performed in a video encoder may be provided. In the method, first affine parameters of a first affine model of a current block in a current picture may be determined. The first affine model may be associated with the current block and a first reference block of the current block in a first reference picture. An initial CPMV of the current block associated with a second reference picture may be determined based on a second affine model derived from the first affine model. The second affine model may be associated with the current block and a second reference block of the current block in the second reference picture. The second affine parameters of the second affine model may be symmetric with respect to the first affine parameters of the first affine model. An improved CPMV of the current block associated with the second reference picture may be determined based on the initial CPMV of the current block associated with the second reference picture and the first affine motion search. The improved CPMV of the current block associated with the first reference picture may be determined based on the initial CPMV of the current block associated with the first reference picture and the second affine motion search. The initial CPMV of the current block associated with the first reference picture may be derived from the refined CPMV of the current block associated with the second reference picture and may be symmetric with respect to the refined CPMV. The prediction information of the current block may be determined based on the refined CPMV of the current block associated with the first reference picture and the refined CPMV of the current block associated with the second reference picture.

[0032] To determine an improved CPMV of the current block associated with the second reference picture, an initial predictor of the current block may be determined based on the initial CPMV of the current block associated with the second reference picture. A first predictor of the current block may be determined based on the initial predictor. The first predictor may be equal to the sum of (i) the initial predictor of the current block, (ii) a product of a first component of a gradient value of the initial predictor and a first component of a motion vector difference associated with the initial predictor and the first predictor, and (iii) a product of a second component of the gradient value of the initial predictor and a second component of the motion vector difference.

[0033] To determine an improved CPMV of the current block associated with the second reference picture, the improved CPMV of the current block can be determined based on the Nth predictor associated with the second reference picture in response to one of (i) N being equal to an upper iteration value of the first affine motion search, and (ii) a motion vector difference associated with the Nth predictor and the (N+1)th predictor is zero.

[0034] In some embodiments, the first affine parameters may include a first translation factor and at least one of a first zoom factor or a first rotation factor, and the second affine parameters may include a second translation factor and at least one of a second zoom factor or a second rotation factor.

[0035] In one example, the sum of the first rotation factor and the second rotation factor can be zero, the sum of the first translation factor and the second translation factor can be zero, and the product of the first zoom factor and the second zoom factor can be one.

[0036] In one example, the ratio between the first rotation factor and the second rotation factor may be linearly proportional to the ratio between the first temporal distance between the first reference picture and the current picture and the second temporal distance between the second reference picture and the current picture, and the ratio between the first zoom factor and the second zoom factor may be exponentially proportional to the ratio between the first temporal distance and the second temporal distance.

[0037] According to another aspect of the present disclosure, there is provided an apparatus, the apparatus including a processing circuit, the processing circuit being configured to perform any of the methods for video encoding / decoding.

[0038] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to perform any of the methods for video encoding / decoding.

[0039] Further features, nature and various advantages of the disclosed subject matter will become apparent from the following detailed description and accompanying drawings. [Brief explanation of the drawings]

[0040] [Figure 1A] FIG. 2 is a schematic diagram of an example subset of intra-prediction modes. [Figure 1B] FIG. 1 is a diagram of an exemplary intra-prediction direction. [Figure 2] FIG. 1 is a schematic diagram of a current block and its surrounding spatial merge candidates in one example. [Figure 3] FIG. 3 is a simplified block diagram schematic of a communication system (300) according to one embodiment. [Figure 4] FIG. 4 is a simplified block diagram schematic of a communication system (400) according to one embodiment. [Figure 5] FIG. 2 is a schematic diagram of a simplified block diagram of a decoder according to one embodiment. [Figure 6] FIG. 2 is a schematic diagram of a simplified block diagram of an encoder according to one embodiment. [Figure 7] FIG. 4 is a block diagram of an encoder according to another embodiment. [Figure 8] FIG. 10 is a block diagram of a decoder according to another embodiment. [Figure 9A] FIG. 1 is an exemplary schematic diagram of a four-parameter affine model. [Figure 9B] FIG. 1 is an exemplary schematic diagram of a six-parameter affine model. [Figure 10] 1 is an exemplary schematic diagram of an affine motion vector field associated with sub-blocks within a block; [Figure 11] FIG. 1 is a schematic diagram of exemplary locations of spatial merge candidates. [Figure 12] FIG. 10 is an exemplary schematic diagram of control point motion vector inheritance; [Figure 13] FIG. 10 is an exemplary schematic diagram of candidate positions for constructing an affine merge mode. [Figure 14] FIG. 1 is an exemplary schematic diagram of prediction refinement using optical flow (PROF). [Figure 15] FIG. 2 is an exemplary schematic diagram of an affine motion estimation process. [Figure 16] 1 is an exemplary flowchart of an affine motion estimation search. [Figure 17] FIG. 1 is an exemplary schematic diagram of a symmetric motion vector difference (MVD) mode. [Figure 18] 1 is a flowchart outlining an exemplary decoding process according to some embodiments of the present disclosure. [Figure 19] 1 is a flowchart outlining an exemplary encoding process according to some embodiments of the present disclosure. [Figure 20] 1 is a flowchart outlining an exemplary encoding process according to some embodiments of the present disclosure. [Figure 21] FIG. 1 is a schematic diagram of a computer system, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0041] Figure 3 shows an exemplary block diagram of a communication system (300). The communication system (300) includes multiple terminal devices that can communicate with each other, for example, via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of Figure 3, the first pair of terminal devices (310) and (320) implement unidirectional transmission of data. For example, the terminal device (310) may code video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to the other terminal device (320) via the network (350). The encoded video data may be transmitted in the form of one or more coded video bitstreams. The terminal device (320) may receive the coded video data from the network (350), decode the coded video data to reconstruct the video pictures, and display the video pictures according to the reconstructed video data. Unidirectional data transmission may be common, such as in media serving applications.

[0042] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that implement bidirectional transmission of coded video data, for example, during a video conference. In the case of bidirectional transmission of data, in one example, each of the terminal devices (330) and (340) can code video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) over the network (350). Each of the terminal devices (330) and (340) can also receive coded video data transmitted by the other of the terminal devices (330) and (340), decode the coded video data to recover the video pictures, and display the video pictures on an accessible display device according to the recovered video data.

[0043] In the example of FIG. 3 , terminal devices 310, 320, 330, and 340 are illustrated as a server, a personal computer, and a smartphone, respectively, although the principles of the present disclosure need not be so limited. Embodiments of the present disclosure apply with laptop computers, tablet computers, media players, and / or dedicated videoconferencing equipment. Network 350 represents any number of networks that convey coded video data between terminal devices 310, 320, 330, and 340, including, for example, wired and / or wireless communication networks. Communication network 350 may exchange data over circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of network 350 may not be important to the operation of the present disclosure, unless otherwise described herein.

[0044] 4 shows a video encoder and video decoder in a streaming environment as an example of an application for the disclosed subject matter, which may be equally applicable to other video-enabled applications including, for example, video conferencing, digital television, streaming services, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0045] The streaming system may include a capture subsystem (413) that may include a video source (401), such as a digital camera, that creates a stream of uncompressed video pictures (402). In one example, the stream of video pictures (402) includes samples taken by the digital camera. The stream of video pictures (402) is depicted with a bold line to emphasize its large amount of data compared to the encoded video data (404) (or coded video bitstream) and may be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (404) (or coded video bitstream) is depicted with a thin line to emphasize its small amount of data compared to the stream of video pictures (402) and may be stored on a streaming server (405) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include a video decoder (410), for example, within an electronic device (430). The video decoder (410) decodes the input copy (407) of the encoded video data and creates an output stream (411) of video pictures that can be rendered on a display (412) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., a video bitstream) may be encoded according to a particular video coding / compression standard. Examples of such standards include ITU-T Recommendation H.265.In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC), and the disclosed subject matter may be used in the context of VVC.

[0046] It should be noted that the electronic devices (420) and (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).

[0047] 5 shows an example block diagram of a video decoder (510). The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., receiving circuitry). The video decoder (510) can be used in place of the video decoder (410) in the example of FIG.

[0048] The receiver (531) may receive one or more coded video sequences to be decoded by the video decoder (510). In one embodiment, one coded video sequence is received at a time, and the decoding of each coded video sequence is independent of the decoding of other coded video sequences. The coded video sequences are received from a channel (501), which may be a hardware / software link to a storage device that stores the coded video data. The receiver (531) receives the coded video data along with other data, such as coded audio data and / or auxiliary data streams, which may be forwarded to their respective using entities (not shown). The receiver (531) may separate the coded video sequences from other data. To combat network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter, "parser (520)"). In certain applications, the buffer memory (515) is part of the video decoder (510). In other applications, the buffer memory (515) may be external to the video decoder (510) (not shown). In still other applications, there may be a buffer memory (not shown) external to the video decoder (510), for example, to combat network jitter, plus another buffer memory (515) internal to the video decoder (510), for example, to handle playout timing. When the receiver (531) is receiving data from a storage / forwarding device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory (515) may be unnecessary or may be small. For use with best-effort packet networks such as the Internet, the buffer memory (515) may be required and may be relatively large, advantageously adaptively sized, and at least partially implemented in an operating system or similar element (not shown) external to the video decoder (510).

[0049] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (510) and information for potentially controlling a rendering device, such as a render device (512) (e.g., a display screen) that is not an integral part of the electronic device (530) but may be coupled to the electronic device (530), as shown in FIG. 5. The control information for the rendering device(s) may be in the form of a supplemental enhancement information (SEI) message or a video usability information (VUI) parameter set fragment (not shown). The parser (520) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles, including variable length coding with or without context dependency, Huffman coding, arithmetic coding, etc. The parser (520) may extract, from the coded video sequence, a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroup may include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser (520) may also extract information from the coded video sequence, such as transform coefficients, quantization parameter values, motion vectors, etc.

[0050] The parser (520) may generate symbols (521) by performing entropy decoding / parsing operations on the video sequence received from the buffer memory (515).

[0051] The reconstruction of the symbols (521) can involve several different units, depending on the type of coded video picture or portion thereof (inter-picture and intra-picture, inter-block and intra-block, etc.), as well as other factors. Which units are involved and how can be controlled by the parser (520) through subgroup control information parsed from the coded video sequence. The flow of such subgroup control information between the parser (520) and the following units is not depicted for clarity.

[0052] Beyond the functional blocks already mentioned, the video decoder (510) may be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate:

[0053] The first unit is a scalar / inverse transform unit (551), which receives quantized transform coefficients as well as control information from the parser (520) including which transform to use, block size, quantization factors, quantization scaling matrices, etc. as symbol(s) (521). The scalar / inverse transform unit (551) can output blocks comprising sample values ​​that can be input to an aggregator (555).

[0054] In some cases, the output samples of the scaler / inverse transform unit (551) may relate to intra-coded blocks. Intra-coded blocks are blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates a block of the same size and shape as the block being reconstructed using surrounding already reconstructed information fetched from the current picture buffer (558). The current picture buffer (558), for example, buffers a partially reconstructed and / or fully reconstructed current picture. The aggregator (555) may add, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).

[0055] In other cases, the output samples of the scalar / inverse transform unit (551) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion-compensated prediction unit (553) may access a reference picture memory (557) to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols (521) related to the block, these samples may be added by the aggregator (555) to the output of the scalar / inverse transform unit (551) (in this case, referred to as residual samples or a residual signal) to generate output sample information. The addresses in the reference picture memory (557) from which the motion-compensated prediction unit (553) fetches prediction samples may be controlled by, for example, motion vectors available to the motion-compensated prediction unit (553) in the form of symbols (521), which may have X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​fetched from the reference picture memory (557) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, and the like.

[0056] The output samples of the aggregator (555) can undergo various loop filtering techniques in a loop filter unit (556). Video compression techniques can include in-loop filtering techniques controlled by parameters contained in the coded video sequence (also called the coded video bitstream) and provided to the loop filter unit (556) as symbols (521) from the parser (520). Video compression can also be performed in response to meta-information obtained during decoding of a coded picture or previous portion (in decoding order) of the coded video sequence, or in response to previously reconstructed and loop-filtered sample values.

[0057] The output of the loop filter unit (556) may be a sample stream that can be output to a render device (512) and stored in a reference picture memory (557) for use in future inter-picture prediction.

[0058] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a new current picture buffer can be reallocated before beginning reconstruction of the next coded picture.

[0059] The video decoder (510) may perform decoding operations according to a given video compression technology or standard, such as ITU-T Recommendation H.265. The coded video sequence may conform to the syntax specified by the video compression technology or standard being used, in the sense that the coded video sequence adheres to both the syntax of the video compression technology or standard and the profile documented in the video compression technology or standard. Specifically, a profile may select certain tools from among all tools available in the video compression technology or standard as the only tools available under that profile. Compliance may also require that the complexity of the coded video sequence be within a range defined by the level of the video compression technology or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further constrained by a hypothetical reference decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.

[0060] In one embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0061] 6 shows an example block diagram of a video encoder (603). The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmitting circuit). The video encoder (603) can be used in place of the video encoder (403) in the example of FIG.

[0062] The video encoder (603) may receive video samples from a video source (601) (which is not part of the electronic device (620) in the example of FIG. 6) that may capture video image(s) to be coded by the video encoder (603). In other examples, the video source (601) is part of the electronic device (620).

[0063] The video source (601) may provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream, which may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 YCrCB, RGB, etc.), and an appropriate sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (601) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. Video data may be provided as multiple individual pictures that convey motion when viewed sequentially. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following discussion focuses on samples.

[0064] According to one embodiment, the video encoder (603) may code and compress pictures of a source video sequence into a coded video sequence (643) in real time or under any other time constraints as required. Enforcing an appropriate coding rate is one function of the controller (650). In some embodiments, the controller (650) controls and is operatively coupled to other functional units described below. Coupling is not shown for clarity. Parameters set by the controller (650) may include rate control-related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured with other appropriate functions for the video encoder (603) optimized for a particular system design.

[0065] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As an overly simplified explanation, in one example, the coding loop may include a source coder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be coded and one or more reference pictures) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to that which a (remote) decoder would also create. The reconstructed sample stream (sample data) is input to a reference picture memory (634). Because decoding of the symbol stream yields bit-exact results regardless of the location (local or remote) of the decoder, the contents of the reference picture memory (634) are also bit-exact between the local and remote encoders. In other words, the predictive portion of the encoder "sees" the exact same sample values ​​as the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchronism (and the drift that occurs when synchronism cannot be maintained, for example due to channel errors) is also used in several related techniques.

[0066] The operation of the "local" decoder (633) may be the same as the operation of a "remote" decoder, such as the video decoder (510), already described in detail above in conjunction with Figure 5. Referring also briefly to Figure 5, however, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder (645) and parser (520) may be lossless, the entropy decoding portion of the video decoder (510), including the buffer memory (515), and the parser (520), may not be fully implemented in the local decoder (633).

[0067] In one embodiment, decoder technology, excluding parsing / entropy decoding, present in the decoder is present in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on the operation of the decoder. A description of the encoder technology can be omitted, as it is the reverse of the decoder technology described generically. In certain areas, more detailed descriptions are provided below.

[0068] In some examples, during operation, the source coder (630) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as “reference pictures.” In this manner, the coding engine (632) codes differences between pixel blocks of the input picture and pixel blocks of the reference picture(s) that may be selected as the predictive reference(s) for the input picture.

[0069] The local video decoder (633) may decode coded video data of pictures that may be designated as reference pictures based on symbols created by the source coder (630). The operation of the coding engine (632) may advantageously be a lossy process. When the coded video data is decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (633) may replicate the decoding process that may be performed by the video decoder on the reference pictures and store the reconstructed reference pictures in the reference picture memory (634). In this way, the video encoder (603) may locally store copies of reconstructed reference pictures that have common content as reconstructed reference pictures that would be obtained by the far-end video decoder (without transmission errors).

[0070] The predictor (635) may perform the predictive search for the coding engine (632). That is, for a new picture to be coded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors, block shapes, etc., that can serve as appropriate predictive references for the new picture. The predictor (635) may operate on sample blocks, pixel blocks at a time, to find an appropriate predictive reference. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have predictive references drawn from multiple reference pictures stored in the reference picture memory (634).

[0071] The controller (650) may manage the coding operations of the source coder (630), including, for example, setting parameters and subgroup parameters used to encode the video data.

[0072] The output of all of the aforementioned functional units may undergo entropy coding in an entropy coder (645), which converts the symbols produced by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0073] The transmitter (640) may buffer the coded video sequence(s) created by the entropy coder (645) and prepare them for transmission over a communication channel (660), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (640) may merge the coded video data from the video encoder (603) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0074] The controller (650) may manage the operation of the video encoder (603). During coding, the controller (650) may assign a particular coded picture type to each coded picture, which may affect the coding technique that may be applied to the respective picture. For example, pictures are often assigned as one of the following picture types:

[0075] An intra picture (I-picture) may be one that can be coded and decoded without using any other picture in a sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh ("IDR") pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.

[0076] A predictive picture (P picture) may be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict the sample values ​​of each block.

[0077] Bidirectionally predicted pictures (B-pictures) may be coded and decoded using intra- or inter-prediction, which uses up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predicted pictures may use more than two reference pictures and associated metadata for the reconstruction of a block.

[0078] A source picture is typically spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture may be predictively coded via spatial prediction with reference to one previously coded reference picture or via temporal prediction. Blocks of a B-picture may be predictively coded with reference to one or two previously coded reference pictures via spatial prediction or via temporal prediction.

[0079] The video encoder (603) may perform coding operations in accordance with a predetermined video coding technique or standard, such as ITU-T Recommendation H.265. In doing so, the video encoder (603) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancies in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.

[0080] In one embodiment, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0081] Video may be captured as a time sequence of multiple source pictures (video pictures). Intra-picture prediction (often abbreviated as intra-prediction) uses spatial correlation within a given picture, while inter-picture prediction uses correlation (temporal or other) between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. When a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture can be coded by a vector called a motion vector. The motion vector points to a reference block in the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0082] In some embodiments, bi-prediction techniques may be used for inter-picture prediction. According to bi-prediction techniques, two reference pictures, such as a first reference picture and a second reference picture, are used, both of which are before the decoding order of the current picture in the video (but their display orders may be past and future, respectively). A block in the current picture may be coded by a first motion vector pointing to a first reference block in the first reference picture and by a second motion vector pointing to a second reference block in the second reference picture. A block may be predicted by a combination of the first reference block and the second reference block.

[0083] Furthermore, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.

[0084] According to some embodiments of the present disclosure, prediction, such as inter-picture prediction or intra-picture prediction, is performed on a block-by-block basis. For example, according to the HEVC standard, pictures in a sequence of video pictures are divided into coding tree units (CTUs) for compression, and the CTUs within a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU can be recursively quadtree-decomposed into one or more coding units (CUs). For example, a 64x64 pixel CTU can be divided into one CU of 64x64 pixels, four CUs of 32x32 pixels, or 16 CUs of 16x16 pixels. In one example, each CU is analyzed to determine the prediction type of the CU, such as an inter-prediction type or an intra-prediction type. A CU is divided into one or more prediction units (PUs) according to temporal predictability and / or spatial predictability. Generally, each PU includes one luma prediction block (PB) and two chroma PBs. In one embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Taking a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values ​​(e.g., luma values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, and 16x8 pixels.

[0085] 7 shows an example diagram of a video encoder (703). The video encoder (703) is configured to receive a processed block of sample values ​​(e.g., a predictive block) in a current video picture in a sequence of video pictures and encode the processed block into a coded picture that is part of a coded video sequence. In one example, the video encoder (703) is used in place of the video encoder (403) of the example of FIG. 4.

[0086] In an HEVC example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as a predictive block of 8x8 samples. The video encoder (703) determines whether the processing block is optimally coded using intra-mode, inter-mode, or bi-predictive mode, for example, using rate-distortion optimization. When the processing block is to be coded in intra-mode, the video encoder (703) may encode the processing block into a coded picture using intra-prediction techniques, and when the processing block is to be coded in inter-mode or bi-predictive mode, the video encoder (703) may encode the processing block into a coded picture using inter-prediction techniques or bi-prediction techniques, respectively. In certain video coding techniques, the merge mode may be an inter-picture prediction sub-mode in which a motion vector is derived from one or more motion vector predictors without the aid of coded motion vector components outside the predictors. In certain other video coding techniques, there may be motion vector components applicable to the current block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the mode of the processing block.

[0087] In the example of Figure 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725), coupled to each other as shown in Figure 7.

[0088] The inter-encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in reference pictures (e.g., blocks in previous and subsequent pictures), generate inter-prediction information (e.g., inter-coding techniques, motion vectors, description of redundant information through merge mode information), and calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture that is decoded based on the coded video information.

[0089] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), possibly compare the block with previously coded blocks in the same picture, generate transformed quantized coefficients, and possibly also generate intra prediction information (e.g., intra prediction direction information according to one or more intra encoding techniques). In one example, the intra encoder (722) also calculates intra prediction results (e.g., prediction blocks) based on the intra prediction information and reference blocks in the same picture.

[0090] The general-purpose controller (721) is configured to determine general-purpose control data and control other components of the video encoder (703) based on the general-purpose control data. In one example, the general-purpose controller (721) determines the mode of the block and provides a control signal to the switch (726) based on the mode. For example, when the mode is intra mode, the general-purpose controller (721) controls the switch (726) to select the intra-mode result used by the residual calculator (723) and controls the entropy encoder (725) to select intra-prediction information and include the intra-prediction information in the bitstream. When the mode is inter mode, the general-purpose controller (721) controls the switch (726) to select the inter-prediction result used by the residual calculator (723) and controls the entropy encoder (725) to select inter-prediction information and include the inter-prediction information in the bitstream.

[0091] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and a prediction result selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) is configured to operate based on the residual data to encode the residual data to generate transform coefficients. In one example, the residual encoder (724) is configured to transform the residual data from the spatial domain to the frequency domain to generate transform coefficients. The transform coefficients then undergo a quantization process to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data may be used by the intra-encoder (722) and the inter-encoder (730) as appropriate. For example, the inter-encoder (730) may generate decoded blocks based on the decoded residual data and inter-prediction information, and the intra-encoder (722) may generate decoded blocks based on the decoded residual data and intra-prediction information. In some examples, the decoded blocks are processed appropriately to generate decoded pictures, which may be buffered in a memory circuit (not shown) and used as reference pictures.

[0092] The entropy encoder (725) is configured to format a bitstream to include the encoded blocks. The entropy encoder (725) is configured to include various information in the bitstream according to an appropriate standard, such as the HEVC standard. In one example, the entropy encoder (725) is configured to include in the bitstream general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information. It should be noted that, according to the disclosed subject matter, when coding a block in a merged sub-mode of either an inter mode or a bi-prediction mode, no residual information is present.

[0093] 8 shows an example diagram of a video decoder (810). The video decoder (810) is configured to receive coded pictures that are part of a coded video sequence and decode the coded pictures to generate reconstructed pictures. In one example, the video decoder (810) is used in place of the video decoder (410) of the example of FIG. 4.

[0094] In the example of Figure 8, the video decoder (810) includes an entropy decoder (871), an inter-decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-decoder (872), which are coupled together as shown in Figure 8.

[0095] The entropy decoder (871) can be configured to reconstruct, from a coded picture, specific symbols that represent the syntax elements of which the coded picture is composed. Such symbols can include, for example, prediction information (e.g., intra-prediction information or inter-prediction information) that can identify the mode in which the block is coded (e.g., intra-mode, inter-mode, bi-prediction mode, inter-mode and bi-prediction mode of merged or other submodes, etc.) and specific samples or metadata used for prediction by the intra decoder (872) or inter decoder (880), respectively. The symbols can also include, for example, residual information in the form of quantized transform coefficients. In one example, when the prediction mode is an inter-mode or bi-prediction mode, the inter-prediction information is provided to the inter decoder (880), and when the prediction type is an intra-prediction type, the intra-prediction information is provided to the intra decoder (872). The residual information can undergo inverse quantization and be provided to the residual decoder (873).

[0096] The inter decoder (880) is configured to receive the inter prediction information and generate an inter prediction result based on the inter prediction information.

[0097] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0098] The residual decoder (873) is configured to perform inverse quantization to extract inverse quantized transform coefficients and process the inverse quantized transform coefficients to transform the residual information from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (to include quantization parameters (QPs)), which may be provided by the entropy decoder (871) (this may be only a small amount of control information, so a data path is not depicted).

[0099] The reconstruction module (874) is configured to combine, in the spatial domain, the residual information output by the residual decoder (873) and the prediction results (possibly output by the inter-prediction module or the intra-prediction module) to form reconstructed blocks that may become part of a reconstructed picture, which may become part of a reconstructed video. It should be noted that other appropriate operations, such as a deblocking operation, may be performed to improve visual quality.

[0100] It should be noted that the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) can be implemented using any suitable technology. In one embodiment, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (603) and the video decoders (410), (510), and (810) can be implemented using one or more processors executing software instructions.

[0101] This disclosure includes embodiments related to affine coding modes, for example, in which both the references in future and past reference frames and the affine models applied to the future and past reference frames may be symmetric.

[0102] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (Version 1), 2014 (Version 2), 2015 (Version 3), and 2016 (Version 4). In 2015, the two standards organizations jointly formed the Joint Video Exploration Team (JVET) to explore the possibility of developing a next-generation video coding standard beyond HEVC. In October 2017, the two standards organizations announced a Joint Call for Proposals on Video Compression with Capability beyond HEVC (CfP). By February 15, 2018, 22 CfP responses had been submitted for standard dynamic range (SDR), 12 for high dynamic range (HDR), and 12 for the 360 ​​video category. In April 2018, all CfP responses received were evaluated at the 122 MPEG / 10th JVET meeting. As a result of this meeting, JVET formally launched the standardization process for next-generation video coding beyond HEVC. This new standard was named Versatile Video Coding (VVC), and JVET was renamed the Joint Video Experts Team. In 2020, ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the VVC video coding standard (Version 1).

[0103] In inter prediction, motion parameters are required for each inter-predicted coding unit (CU), e.g., to code VVC features used in generating inter-predicted samples. The motion parameters may include a motion vector, a reference picture index, a reference picture list usage index, and / or additional information. The motion parameters may be signaled in an explicit or implicit manner. If a CU is coded in skip mode, the CU may be associated with one PU, and significant residual coefficients, coded motion vector deltas, and / or reference picture indexes may not be required. If a CU is coded in merge mode, the motion parameters of the CU may be obtained from neighboring CUs. The neighboring CUs may include spatial and temporal candidates, as well as additional schedules (or additional candidates) such as those introduced in VVC. Merge mode can be applied to any inter-predicted CU, not just skip mode. An alternative to merge mode is explicit transmission of motion parameters; a motion vector, a corresponding reference picture index for each reference picture list, a reference picture list usage flag, and / or other necessary information may be explicitly signaled for each CU.

[0104] In VVC, the VVC Test Model (VTM) reference software may include several new and improved inter-predictive coding tools, which may include one or more of the following: (1) Enhanced Merge Prediction (2) Merge Motion Vector Difference (MMVD) (3) Advanced Motion Vector Prediction (AMVP) mode with symmetric MVD signaling (4) Affine motion compensation prediction (5) Sub-block-based temporal motion vector prediction (SbTMVP) (6) Adaptive Motion Vector Resolution (AMVR) (7) Motion field storage: 1 / 16 luminance sample MV storage and 8x8 motion field compression (8) Bi-prediction with CU-level weights (BCW) (9) Bidirectional Optical Flow (BDOF) (10) Decoder-side Motion Vector Improvement (DMVR) (11) Combined Inter- and Intra-Prediction (CIIP) (12) Geometric Partition Mode (GPM)

[0105] In HEVC, a translational motion model is applied to motion compensated prediction (MCP). In the real world, many types of motion may exist, such as zoom in / out, rotation, perspective motion, and other irregular motions. Block-based affine transform motion compensated prediction can be applied to VTM, etc. Figure 9A shows the affine motion field of a block (902) described by motion information of two control points (four parameters). Figure 9B shows the affine motion field of a block (904) described by three control point motion vectors (six parameters).

[0106] As shown in FIG. 9A, in a four-parameter affine motion model, the motion vector at a sample position (x, y) within a block (902) can be derived in equation (1) as follows:

number

number

[0107] As shown in FIG. 9B, in a six-parameter affine motion model, the motion vector at sample position (x, y) within block (904) can be derived in equation (3) as follows:

number

number

[0108] As shown in FIG. 10, block-based affine transformation prediction can be applied to simplify motion compensation prediction. To derive a motion vector for each 4×4 luma subblock, the motion vector (e.g., (1002)) of the center sample of each subblock (e.g., (1004)) in the current block (1000) can be calculated according to equations (1) to (4) and rounded to 1 / 16 fractional precision. A motion compensation interpolation filter can then be applied to generate a prediction for each subblock with the derived motion vector. The subblock size of the chroma component can also be set to 4×4. The motion vector (MV) of a 4×4 luma subblock can be calculated as the average of the motion vectors of the four corresponding 4×4 luma subblocks.

[0109] In affine merge prediction, the affine merge (AF_MERGE) mode can be applied to CUs whose width and height are both 8 or greater. The CPMV of the current CU can be generated based on the motion information of spatially neighboring CUs. Up to five CPMVP candidates can be applied to affine merge prediction, and an index can be signaled to indicate which of the five CPMVP candidates can be used for the current CU. In affine merge prediction, three types of CPMV candidates can be used to derive the affine merge candidate list: (1) inherited affine merge candidates extrapolated from the CPMVs of neighboring CUs, (2) constructed affine merge candidates with CPMVPs derived using the translational MVs of neighboring CUs (e.g., MVs that include only translation factors such as merge MVs), and (3) zero MVs.

[0110] In VTM3, up to two inherited affine candidates can be applied. The two inherited affine candidates can be derived from the affine motion models of neighboring blocks. For example, one inherited affine candidate can be derived from a left neighboring CU, and the other inherited affine candidate can be derived from an upper neighboring CU. An example candidate block can be shown in FIG. 11. As shown in FIG. 11, for the left predictor (or left inherited affine candidate), the scanning order can be A0->A1, and for the upper predictor (or upper inherited affine candidate), the scanning order can be B0->B1->B2. Therefore, only the first available inherited candidate from each side can be selected. No pruning check needs to be performed between the two inherited candidates. Once the neighboring affine CUs are identified, the control point motion vectors of the neighboring affine CUs can be used to derive CPMVP candidates in the affine merge list of the current CU. As shown in Figure 12, when block A, the lower left neighbor of the current block (1204), is coded in affine mode, motion vectors v2, v3, and v4 can be achieved for the upper left, upper right, and lower left corners of the CU (1202) containing block A. If block A is coded using a four-parameter affine model, two CPMVs for the current CU (1204) can be calculated according to v2 and v3 of the CU (1202). If block A is coded using a six-parameter affine model, three CPMVs for the current CU (1204) can be calculated according to v2, v3, and v4 of the CU (1202).

[0111] The constructed affine candidate for the current block can be constructed by combining the local translational motion information of each control point of the current block. The motion information of the control points can be derived from specific spatial and temporal neighborhoods, which can be shown in Figure 13. As shown in Figure 13, the CPMV k(k=1,2,3,4) represents the kth control point of the current block (1302). For CPMV1, the B2->B3->A2 block can be checked and the MV of the first available block can be used. For CPMV2, the B1->B0 block can be checked. For CPMV3, the A1->A0 block can be checked. If CPM4 is not available, TMVP can be used as CPMV4.

[0112] After the MVs of the four control points are achieved, affine merge candidates for the current block (1302) can be constructed based on the motion information of the four control points. For example, affine merge candidates can be constructed based on the MV combinations of the four control points in the following order: {CPMV1,CPMV2,CPMV3}, {CPMV1,CPMV2,CPMV4}, {CPMV1,CPMV3,CPMV4}, {CPMV2,CPMV3,CPMV4}, {CPMV1,CPMV2}, {CPMV1,CPMV3}.

[0113] A combination of three CPMVs can construct a six-parameter affine merge candidate, and a combination of two CPMVs can construct a four-parameter affine merge candidate. To avoid the motion scaling process, if the reference indices of the control points are different, the associated combination of control point MVs can be discarded.

[0114] After one or more inherited and constructed affine merge candidates have been checked, if the list is not already full, a zero MV may be inserted at the end of the list.

[0115] In affine AMVP prediction, affine AMVP mode can be applied to CUs whose width and height are both 16 or greater. A CU-level affine flag can be signaled in the bitstream to indicate whether affine AMVP mode is used, and then another flag can be signaled to indicate whether 4-parameter affine or 6-parameter affine is applied. In affine AMVP prediction, the difference between the CPMV of the current CU and the predictor of the CPMVP of the current CU can be signaled in the bitstream. The size of the affine AMVP candidate list can be 2, and the affine AMVP candidate list can be generated by using four types of CPMV candidates in the following order: (1) Inherited affine AMVP candidates extrapolated from the CPMVs of neighboring CUs; (2) constructed affine AMVP candidates with CPMVPs derived using the translational MVs of nearby CUs; (3) translational MVs from nearby CUs, and (4) Zero MV.

[0116] The check order of inherited affine AMVP candidates may be the same as the check order of inherited affine merge candidates. To determine AVMP candidates, only affine CUs with the same reference picture as the current block may be considered. When an inherited affine motion predictor is inserted into the candidate list, the pruning process may not be applied.

[0117] The constructed AMVP candidate can be derived from a specified spatial neighborhood. As shown in Figure 13, the same check order can be applied as the check order in affine merge candidate construction. The reference picture indexes of neighboring blocks can also be checked. The first block in the check order can be inter-coded and have the same reference picture as the current CU (1302). If the current CU (1302) is coded in a four-parameter affine mode and both mv0 and mv1 are available, the constructed AMVP candidate can be determined. The constructed AMVP candidate can be further added to the affine AMVP list. If the current CU (1302) is coded in a six-parameter affine mode and all three CPMVs are available, the constructed AMVP candidate can be added as a candidate in the affine AMVP list. Otherwise, the constructed AMVP candidate can be set as unavailable.

[0118] After the inherited and constructed affine AMVP candidates are checked, if there are still less than two candidates in the affine AMVP list, mv0, mv1, and mv2 can be added in order. mv0, mv1, and mv2, if available, can serve as translation MVs to predict all control point MVs for the current CU (e.g., (1302)). Finally, if the affine AMVP is not yet filled, a zero MV can be used to fill the affine AMVP list.

[0119] Subblock-based affine motion compensation can save memory access bandwidth and reduce computational complexity compared to pixel-based motion compensation, at the expense of a prediction accuracy penalty. To achieve finer granularity in motion compensation, prediction refinement by optical flow (PROF) can be used to refine subblock-based affine motion compensation predictions without increasing memory access bandwidth for motion compensation. In VVC, after subblock-based affine motion compensation is performed, luma prediction samples can be refined by adding the difference derived by the optical flow formula. PROF can be explained in the following four steps:

[0120] Step (1): Sub-block-based affine motion compensation can be performed to generate a sub-block prediction I(i,j).

[0121] Step (2): Spatial gradient of sub-block prediction g x (i,j) and g y (i,j) can be calculated at each sample position using a 3-tap filter [-1,0,1]. The gradient calculation can be the same as the gradient calculation in BDOF. For example, the spatial gradient g x (i,j) and g y (i,j) can be calculated based on equations (5) and (6), respectively. g x (i,j)=(I(i+1,j)≫shift1)-(I(i-1,j)≫shift1) Equation (5) g y (i,j)=(I(i,j+1)≫shift1)-(I(i,j-1)≫shift1) Equation (6) As shown in Equation (5) and Equation (6), shift1 can be used to control the accuracy of the gradient. Sub-block (e.g., 4x4) prediction can be extended by one sample on each side for gradient calculation. To avoid additional memory bandwidth and additional interpolation calculations, the extended samples on the extension boundary can be copied from the nearest integer-pixel position in the reference picture.

[0122] Step (3): The luminance prediction refinement can be calculated by the optical flow formula as shown in Equation (7). ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) Equation (7) where Δv(i,j) is the ratio between the sample MV calculated for sample position (i,j) represented by v(i,j) and the v of the sub-block to which sample (i,j) belongs. SB The motion vector v(i,j) may be a difference between the current block (1400) and the sub-block motion vector v(i,j). Figure 14 shows an example diagram of the difference between the sample motion vector and the sub-block motion vector v. As shown in Figure 14, the current block (1400) may include a sub-block (1402), which may include a sample (1404). The sample (1404) may include a sample motion vector v(i,j) corresponding to the reference pixel (1406). The sub-block (1402) may include a sub-block motion vector v(i,j). SB The sub-block motion vector v SB Based on this, the sample (1404) can correspond to the reference pixel (1408). The difference between the sample MV and the sub-block MV, represented by Δv(i,j), can be indicated by the difference between the reference pixel (1406) and the reference pixel (1408). Δv(i,j) can be quantized in units of 1 / 32 luma sample precision.

[0123] Since the affine model parameters and sample positions relative to the sub-block center cannot change from sub-block to another, Δv(i,j) can be calculated for the first sub-block (e.g., (1402)) and reused for other sub-blocks (e.g., (1410)) within the same CU (e.g., (1400)). Let dx(i,j) be the horizontal offset and dy(i,j) be the distance from sample position (i,j) to the sub-block (x SB ,y SB), Δv(x,y) can be derived by the following equations (8) and (9):

number

number

[0124] To maintain accuracy, the sub-block centers (x SB ,y SB ) is ((W SB -1) / 2,(H SB -1) / 2), where W SB and H SB are the width and height of the sub-block, respectively.

[0125] Once Δv(x,y) is obtained, the parameters of the affine model can be obtained. For example, in the case of a four-parameter affine model, the parameters of the affine model can be shown in Equation (10).

number

number

[0126] Step (4): Finally, the luma prediction refinement ΔI(i,j) can be added to the sub-block prediction (i,j). The final prediction I′ can be generated as shown in equation (12). I'(i,j)=I(i,j)+ΔI(i,j) Equation (12)

[0127] In some embodiments, PROF may not be applied to an affine-coded CU in two cases: (1) when all control point MVs are the same, indicating that the CU has only translational motion, and (2) when the affine motion parameters are larger than a specified limit, since sub-block-based affine MC degrades to CU-based MC to avoid large memory access bandwidth requirements.

[0128] Affine motion estimation (ME), such as the VVC reference software VTM, can be operated for both uni-prediction and bi-prediction: uni-prediction can be performed for either reference list L0 or reference list L1, and bi-prediction can be performed for both reference list L0 and reference list L1.

[0129] FIG. 15 shows a schematic diagram of affine ME (1500). As shown in FIG. 15, in affine ME (1500), affine uni-prediction (S1502) can be performed on reference list L0 to obtain a prediction P0 of the current block based on an initial reference block in reference list L0. Affine uni-prediction (S1504) can be performed on reference list L1 to obtain a prediction P1 of the current block based on the initial reference block in reference list L1. In (S1506), affine bi-prediction can be performed. The affine bi-prediction (S1506) can start with an initial prediction residual (2I-P0)-P1, where I can be the initial value of the current block. The affine bi-prediction (S1506) can search candidates around the initial reference block in reference list L1 to find the best (or selected) reference block with the smallest prediction residual (2I-P0)-Px, where Px is the prediction of the current block based on the selected reference block.

[0130] Using the reference picture, for the current coding block, the affine ME process can first select a set of control point motion vectors (CPMVs) as a base. An iterative method can be used to generate a prediction output of the current affine model corresponding to the set of CPMVs, calculate the gradient of the prediction sample, and then solve a linear equation to determine the delta CPMVs and optimize the affine prediction. The iteration can be stopped when all delta CPMVs are 0 or the maximum number of iterations is reached. The CPMVs obtained from the iterations can be the final CPMVs of the reference picture.

[0131] After the best affine CPVMs of both reference lists L0 and L1 are determined for affine uni-prediction, an affine bi-predictive search can be performed by using the best uni-predictive CPMV and the reference list on one side to search for the best CPMV of the other reference list to optimize the affine bi-predictive output. The affine bi-predictive search can be performed iteratively on the two reference lists to obtain the optimal result.

[0132] 16 shows an example affine ME process (1600) that can calculate a final CPMV associated with a reference picture. The affine ME process (1600) can begin at (S1602). At (S1602), a base CPMV for the current block can be determined. The base CPMV can be determined based on one of a merge index, an advanced motion vector prediction (AMVP) predictor index, an affine merge index, etc.

[0133] In (S1604), an initial affine prediction of the current block can be obtained based on the base CPMV. For example, according to the base CPMV, a four-parameter affine motion model of a six-parameter affine motion model can be applied to generate the initial affine prediction.

[0134] In (S1606), the gradient of the initial affine prediction can be obtained. For example, the gradient of the initial affine prediction can be obtained based on Equation (5) and Equation (6).

[0135] At (S1608), a delta CPMV may be determined. In some embodiments, the delta CPMV may be associated with a displacement between an initial affine prediction and a subsequent affine prediction, such as the first affine prediction. The first affine prediction may be obtained based on a gradient of the initial affine prediction and the delta CPMV. The first affine prediction may correspond to the first CPMV.

[0136] At (S1610), a determination may be made to check whether the delta CPMV is zero or the number of iterations is greater than or equal to a threshold. If the delta CPMV is zero or the number of iterations is greater than or equal to a threshold, a final (or selected) CPMV may be determined at (S1612). The final (or selected) CPMV may be a first CPMV determined based on the gradient of the initial affine prediction and the delta CPMV.

[0137] Further referring to (S1610), if the delta CPMV is not zero or the number of iterations is less than a threshold, a new iteration can be initiated. In the new iteration, an updated CPMV (e.g., the first CPMV) can be provided to (S1604) to generate an updated affine prediction. The affine ME process (1600) can then proceed to (S1606) and calculate the gradient of the updated affine prediction. The affine ME process (1600) can then proceed to (S1608) to continue the new iteration.

[0138] In the affine motion model, the four-parameter affine motion model can be further described by an equation including the rotation and zoom motion. For example, the four-parameter affine motion model can be rewritten in equation (13) as follows:

number

number

number

[0139] Symmetric motion vector difference (MVD) coding can be applied in VVC, etc. For example, in addition to unidirectionally predictive MVD signaling and bidirectionally predictive MVD signaling, a symmetric MVD (SMVD) mode for bidirectionally predictive MVD signaling can be applied. In the symmetric MVD mode, motion information including reference picture indexes of both list 0 and list 1 and the MVD of list 1 can be derived without being signaled.

[0140] The decoding process for the symmetric MVD mode can be provided as follows: (1) At the slice level, variables such as BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 can be derived as follows: a) If mvd_l1_zero_flag is 1, then BiDirPredFlag is set equal to 0. b) Otherwise, if the closest reference picture in list 0 and the closest reference picture in list 1 form a backward-forward pair of reference pictures of a forward-forward pair of reference pictures, then BiDirPredFlag is set to 1 and the reference pictures in both list 0 and list 1 are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. (2) At the CU level, if the CU is bi-predictively coded and BiDirPredFlag is equal to 1, a symmetric mode flag indicating whether symmetric mode is used or not can be explicitly signaled.

[0141] If the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly signaled. The reference indexes of list 0 and list 1 may be set equal to a pair of reference pictures, respectively. MVD1 may be set equal to (-MVD0). The final motion vector may be shown in equation (16) below.

number

[0142] Figure 17 shows an example diagram of the symmetric MVD mode. As shown in Figure 17, a current block (1708) may be included in a current picture (1702). The current block (1708) may have a reference block (1710) in a first reference picture (1704). The first reference picture (1704) may be included in a first reference list (or reference picture list) L0 and may correspond to a motion vector predictor of the current block associated with reference list L0. The current block may have a reference block (1716) in a second reference picture (1706). The second reference picture (1706) may be included in a second reference list (or second reference picture list) L1 and may correspond to a motion vector predictor of the current block (1708) associated with reference list L1. In some embodiments, an updated first reference block (1712) in the first reference picture (1704) may be determined. The updated first reference block (1712) and the first reference block (1710) may correspond to MVD0. According to a symmetric MVD mode, an updated second reference block (1714) in the second reference picture (1706) may be determined. The updated second reference block (1714) and the second reference block (1716) may correspond to MVD1. MVD1 may be symmetric with respect to MVD0 such that MVD1=-MVD0.

[0143] In the encoder, symmetric MVD motion estimation (or search) can start with initial MV evaluation. The set of initial MV candidates can include MVs obtained from unipredictive search, MVs obtained from bipredictive search, and MVs from the AMVP list. The MV with the lowest rate-distortion cost can be selected as the initial MV for symmetric MVD motion search.

[0144] A bi-predictive affine symmetric MVD may also be provided, for example, the bi-predictive affine mode symmetric MVD may be provided based on a process of bi-predicted translational motion symmetric MVD coding.

[0145] When a symmetric mode (e.g., bi-predictive affine symmetric MVD) is used, the MVP index flag (e.g., mvp_l0_flag and mvp_l1_flag) and the MVD in list 0 (e.g., MVD0) may be explicitly signaled. The reference indices of list 0 and list 1 may be set equal to a pair of reference pictures, which may be processed in the same manner as symmetric MVD coding. When the affine mode flag is true, the MVD of the top-left control point in list 1 may be set equal to the negative value of the MVD of the top-left control point in list 0. The MVDs of the other control points in list 1 may be set to 0. The final control point motion vector may be derived in equations (17) and (18) as follows: For the top-left control point:

number

number

[0146] The signaling cost of affine motion parameters in an affine motion model can be much higher than the signaling cost of translational motion. Although symmetric MVD coding can be applied to affine motion, the overall coding efficiency may not be sufficient.

[0147] The present disclosure may provide a symmetric affine mode. Based on the symmetric affine mode, affine motion information of a first reference list (e.g., reference list L0) may be signaled. Affine motion information of another reference list (e.g., reference list L1) may be derived based on the affine motion information of the first reference list. The affine motion information may include a type of affine model (e.g., a four-parameter affine model or a six-parameter affine model), affine motion parameters of the affine model, etc. According to the symmetric affine mode, in one example, the first affine parameter and the second affine parameter may have opposite signs. In one example, the first affine parameter and the second affine parameter may have opposite values. In one example, the first affine parameter and the second affine parameter may have a proportional relationship based on a first temporal distance between the first reference picture and the current picture and a second temporal distance between the second reference picture and the current picture.

[0148] In one embodiment, the Symmetric Affine mode can be indicated by signaling information in the bitstream. The signaling information in the bitstream can include a flag. The flag can be, for example, a Symmetric Affine Flag (SAFF). In one example, the flag (e.g., SAFF) can be coded in a context-adaptive binary arithmetic coding (CABAC) context. In one example, the flag can be bypass coded.

[0149] In one embodiment, the symmetric affine mode may be used only for a specific affine type. For example, the symmetric affine mode may be applicable only when a four-parameter affine model is used. In one example, if the symmetric affine flag (or SAFF flag) is signaled as true, the affine type (e.g., four-parameter affine model or six-parameter affine mode) is not signaled but may be derived as a four-parameter affine model. In one example, if the affine type is signaled as a six-parameter affine model, the SAFF flag is not signaled and may be derived as false.

[0150] In one embodiment, the symmetric affine mode can be used when a condition related to MVD is met (or SAFF is true). For example, the symmetric affine mode can be used when a symmetric MVD (SMVD) is met. Thus, if a current frame has a future reference frame and a past reference frame, the current frame can be at a temporal intermediate position of the further reference frame and the past reference frame.

[0151] In one example, the picture order count (POC) of a reference picture in list L0 (or the first reference list L0) may be represented as Ref_POC_L0. The POC of a reference picture in list L1 (or the second reference list L1) may be represented as Ref_POC_L1. The POC of the current picture may be represented as Curr_POC. Therefore, the symmetric affine mode may be used when the SMVD is satisfied, such as by satisfying the condition shown in equation (19). Ref_POC_L0-Curr_POC=Curr_POC-Ref_POC_L1 Equation (19) Here, Ref_POC_L0-Curr_POC may represent a first temporal distance between the current picture and a reference picture in a first reference list (e.g., list L0), and Curr_POC-Ref_POC_L1 may represent a second temporal distance between the current picture and a reference picture in a second reference list (e.g., list L1).

[0152] In one embodiment, when SAFF is on or true, reference index information may not be signaled. Instead, reference index information may be derived. In one example, the reference index information may be derived in the same manner as MVD is derived in symmetric MVD (SMVD) according to equation (16). The reference index information may indicate which reference picture in a reference list is a reference picture for the current picture. For example, once a reference picture for the current picture in a first reference list is determined, reference index information for a second reference picture for the current picture in a second reference list may be derived based on the reference picture for the current picture in the first reference list. The derived reference index information may convey which reference picture in the second reference list is the second reference picture for the current picture based on the symmetric affine mode.

[0153] In one embodiment, when SAFF is on or true, affine motion information (e.g., affine motion parameters or affine type) may be signaled only for a list (e.g., the first reference list L0). Affine motion information may be derived for another list (e.g., the second reference list L1) based on the symmetric affine mode described in equations (17) and (18). According to the symmetric affine mode, affine motion parameters in the two reference lists, such as a rotation factor (e.g., θ), a zoom factor (e.g., r), and a translational motion vector (e.g., c and f), may be symmetric. In one example, the sum of the rotation factors in the two reference lists, such as a first rotation factor in the first reference list and a second rotation factor in the second reference list, may be zero. In one example, the sum of the translational motion vectors (or translation factors) of the two lists may be zero. The zoom factors of the two lists may be inverse (i.e., reciprocal). In one example, the product of the zoom factors of the two lists may be one. In one example, the zoom factor associated with the first reference list may be r, and the zoom factor associated with the second reference list may be 1 / r. Thus, once the affine motion parameters of the affine model associated with the current block in the current picture and the reference blocks of the reference pictures in the first reference list are obtained, the affine motion parameters of the affine model associated with the current block in the current picture and the reference blocks of the reference pictures in the second reference list can be derived based on the symmetric affine mode.

[0154] It should be noted that MV derivation based on the symmetric affine mode can be performed using a transformation between control points, affine motion parameters (e.g., a, b, c, and d), or zoom factors and rotation factors. For example, the transformation between control points and affine motion parameters can be processed based on Equations (1) and (2). The transformation between affine motion parameters (e.g., a, b, c, and d) and zoom factors and rotation factors can be processed based on Equations (13) and (14).

[0155] In one embodiment, when SAFF is on or true, affine motion information (e.g., affine motion parameters or affine type) can be signaled for a list (e.g., a first reference list). Reference indices (e.g., indices for indicating which reference picture is used) of both reference lists (e.g., a first reference list and a second reference list) can also be signaled. Control point motion vectors for another list (e.g., a second reference list) can be derived based on the ratio between the POC distances (or temporal distances) of the current picture and reference pictures on list L0 (or first reference list L0) and list L1 (or second reference list L1) using an affine model such as the four-parameter affine model described in equation (2). Thus, the rotation parameter may be linearly proportional to the temporal distance, while the zoom factor may be exponentially proportional to the temporal distance.

[0156] In one example, the first temporal distance dPoc0 between the current picture and the first reference picture in list L0, and the second temporal distance dPoc1 between the current picture and the second reference picture in list L1 can be described by the following equations (20) and (21): dPoc0=Poc_Cur-RefPoc_L0 Equation (20) dPoc1=Poc_Cur-RefPoc_L1 Equation (21)

[0157] In one example, as shown in equation (22), the ratio between the first rotation factor θ0 and the second rotation factor θ1 can be linearly proportional to the ratio between the first temporal distance dPoc0 between the first reference picture and the current picture and the second temporal distance dPoc1 between the second reference picture and the current picture.

number

[0158] In one example, as shown in equation (23), the ratio between the first zoom factor r0 and the second zoom factor r1 can be exponentially proportional to the ratio between the first temporal distance dPoc0 and the second temporal distance dPoc1.

number

[0159] In the present disclosure, a symmetric affine mode may be applied to an encoder. Based on the symmetric affine mode, a starting point of a first reference list (e.g., a base CPMV or an initial CPMV) in affine motion estimation may be first determined. A starting point of a second reference list may be derived based on the starting point of the first reference list. In some embodiments, the symmetric affine mode may be used when a specific affine model is applied. For example, the symmetric affine mode may be used when a four-parameter affine model is applied in affine motion estimation.

[0160] In the symmetric affine mode, an iterative search can be used. The iterative search can be processed, for example, according to FIGS. 15 and 16. In the iterative search, an affine motion (or base CPMV), such as a four-parameter affine motion, can be first determined for the first reference list. In some embodiments, the affine motion can be determined based on one of a merge candidate, an AMVP candidate, and an affine merge candidate. A starting point (or base CPMV) of the second reference list can be derived based on the symmetric affine mode. For example, the starting point can be determined based on the affine motion of the first reference list such that the affine parameters of the starting point of the second reference list can be symmetric with respect to the affine parameters of the affine motion of the first reference list.

[0161] As shown in FIG. 16, an iterative search can be performed based on a starting point in the second reference list. At (S1602), in a first iteration of the iterative process, a base CPMV of the current block relative to the second reference list can be determined from the affine motion of the first reference list based on the symmetric affine mode. The base CPMV can be associated with an initial reference block in the second reference list L1. At (S1604), a starting point (or initial predictor) P 0,L1 (i,j) can be determined based on the initial reference block (or base CPMV) in the reference list L1 (or second reference list L1). In (S1606), P 0,L1 We can calculate the gradient of (i,j). For example, g x1,L1 (i,j) is the first predictor P in the x direction 1,L1 It can be the gradient of (i,j). g y1,L1 (i,j) is the first predictor P 1,L1 It can be the gradient in the y direction of (i,j).

[0162] In (S1608), the delta CPMV associated with two reference blocks (or sub-blocks) in the reference list L1, such as the initial reference block and the first reference block in the reference list L1, can be calculated based on an affine type, such as the four-parameter affine model shown in Equation (1) and Equation (2). The delta CPMV is calculated based on Δv x0,L1 (i,j) and Δv y0,L1 (i,j) can be expressed as Δv x0,L1 (i,j) can be the difference or displacement of two reference blocks (or sub-blocks), such as the initial reference block and the first reference block in the reference list L1 along the x direction. y0,L1 (i,j) may be the difference or displacement of two reference blocks (or sub-blocks), such as the initial reference block and the first reference block in reference list L1, along the y direction.

[0163] At (S1610), a first predictor P of the current block based on the first reference block in the reference list L1 is calculated. 1,L1 (i,j) can be determined according to equation (24). P 1,L1 (i,j)=P 0,L1 (i,j)+g x0,L1 (i,j)*Δv x0,L1 (i,j)+g y0,L1 (i,j)*Δv y0,L1 (i,j) Equation (24) Here, (i,j) can be the position of the pixel (or sample) within the current block.

[0164] Δv x0,L1 (i,j) and Δv y0,L1 In response to at least one of (i,j) being non-zero, the iterative search may then proceed to a second iteration according to (S1614). At (S1614), an updated CPMV (e.g., P 1,L1 The first CPMV associated with (i,j) may be provided to (S1604), where an updated CPMV (or updated affine prediction) may be generated. The iterative search may then proceed to (S1606), where a gradient of the updated CPMV may be calculated. The iterative search may then proceed to (S1608) to continue with a new iteration (e.g., a second iteration). In the second iteration, a second predictor P for the current block based on a second reference block in the reference list L1 is calculated. 2,L1 (i,j) can be determined according to equation (25). P 2,L1 (i,j)=P 1,L1 (i,j)+g x1,L1 (i,j)*Δv x1,L1 (i,j)+g y1,L1 (i,j)*Δv y1,L1 (i,j) Equation (25) As shown in equation (25), g x1,L1 (i,j) is the first predictor P in the x direction 1,L1 It can be the gradient of (i,j). g y1,L1 (i,j) is the first predictor P1,L1 It can be the gradient of (i,j) in the y direction. Δv x1,L1 (i,j) may be the difference or displacement between the first and second reference blocks in the reference list L1 along the x direction. y1,L1 (i,j) may be the difference or displacement between the first and second reference blocks in the reference list L1 along the y direction.

[0165] An iteration is initiated when the iteration number N is equal to or greater than a threshold (or maximum iteration number) for the iteration process, or when the displacement (e.g., Δv xN,L1 (i,j), Δv yN,L1 (i,j)) is zero, the process can be terminated. This allows the final CPMV (or refined CPMV) to be determined based on the Nth reference block in reference list L1, as shown in (S1612) of FIG.

[0166] After the affine motion (or affine motion parameters) of the second reference list L1 is refined, the affine motion (or affine motion parameters) of the first reference list L0 can be derived based on the refined affine motion of the second reference list L1 using the symmetric affine mode. For example, the starting point (or base CPMV) of the first reference list L0 can be derived based on the final CPMV of the second reference list L1 according to the symmetric affine mode. The refinement of the affine motion (or affine motion parameters) for the first reference list L0 can again proceed according to the iterative search shown in FIG. 16. Such a process can be performed iteratively until one or more conditions are met. For example, the process can be performed iteratively until a certain number of iterations is reached or until the rate-distortion cost of the affine motion reaches a certain threshold.

[0167] FIG. 18 shows a flowchart outlining an exemplary decoding process (1800) according to some embodiments of the present disclosure. FIG. 19 shows a flowchart outlining an exemplary encoding process (1900) according to some embodiments of the present disclosure. FIG. 20 shows a flowchart outlining an exemplary encoding process (2000) according to some embodiments of the present disclosure. The proposed processes may be used separately or combined in any order. Furthermore, each of the processes (or embodiments), the encoder, and the decoder, may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium.

[0168] The operations of the processes (e.g., 1800, 1900, and 2000) may be combined or arranged in any quantity or order as desired. In embodiments, two or more of the operations of the processes (e.g., 1800, 1900, and 2000) may be performed in parallel.

[0169] The processes (e.g., (1800), (1900), and (2000)) can be used in block reconstruction and / or encoding to generate a predictive block for a block being reconstructed. In various embodiments, the processes (e.g., (1800), (1900), and (2000)) are performed by processing circuits such as the processing circuits of the terminal devices (310), (320), (330), and (340), the processing circuits that perform the functions of the video encoder (403), the processing circuits that perform the functions of the video decoder (410), the processing circuits that perform the functions of the video decoder (510), the processing circuits that perform the functions of the video encoder (603), etc. In some embodiments, the processes (e.g., (1800), (1900), and (2000)) are implemented with software instructions, and thus, when the processing circuits execute the software instructions, the processing circuits perform the processes (e.g., (1800), (1900), and (2000)).

[0170] As shown in Figure 18, the process (1800) may start at (S1801) and proceed to (S1810). At (S1810), coded information for a current block in a current picture may be received from a coded video bitstream. The coded information may include a flag indicating whether a symmetric affine mode is applied to the current block.

[0171] At (S1820), in response to a flag indicating that the symmetric affine mode is applied to the current block, first affine parameters of a first affine model of the symmetric affine mode can be determined from the received coded information. The first affine model can be associated with the current block and a first reference block of the current block in a first reference picture of the current picture.

[0172] At (S1830), second affine parameters of a second affine model in a symmetric affine mode may be derived based on the first affine parameters of the first affine model. The second affine model may be associated with a current block in a second reference picture of the current picture and a second reference block of the current block. The first affine parameters and the second affine parameters may have one of opposite signs, inverse values, and a proportional relationship based on a first temporal distance between the first reference picture and the current picture and a second temporal distance between the second reference picture and the current picture.

[0173] At (S1840), the CPMV of the current block may be determined based on the first affine model and the second affine model.

[0174] At (S1850), the current block can be reconstructed based on the determined CPMV of the current block.

[0175] In some embodiments, the flag may be coded via one of a CABAC context and a bypass code.

[0176] In some embodiments, the symmetric affine mode may be determined to be associated with a four-parameter affine model in response to a flag indicating that the symmetric affine mode applies to the current block.

[0177] In some embodiments, the flag may indicate that symmetric affine mode is applied to the current block based on a first temporal distance between the current picture and a first reference picture being equal to a second temporal distance between the current picture and a second reference picture.

[0178] In response to the flag indicating that the symmetric affine mode is applied to the current block, reference index information can be derived. The reference index information can indicate which reference picture in the first reference list is the first reference picture and which reference picture in the second reference list is the second reference picture.

[0179] The first affine parameters can include a first translation factor and at least one of a first zoom factor or a first rotation factor, and the second affine parameters can include a second translation factor and at least one of a second zoom factor or a second rotation factor.

[0180] In one example, the sum of the first rotation factor and the second rotation factor can be zero, the sum of the first translation factor and the second translation factor can be zero, and the product of the first zoom factor and the second zoom factor can be one.

[0181] In one example, the ratio of the first rotation factor to the second rotation factor can be linearly proportional to the ratio of the first temporal distance to the second temporal distance, and the ratio of the first zoom factor to the second zoom factor can be exponentially proportional to the ratio of the first temporal distance to the second temporal distance.

[0182] The process then proceeds to (S1899) and ends.

[0183] The process 1800 may be adapted as appropriate. Step(s) of the process 1800 may be modified and / or omitted. Additional step(s) may be added. Any suitable order of implementation may be used.

[0184] As shown in Figure 19, the process (1900) may start at (S1901) and proceed to (S1910). At (S1910), first affine parameters of a first affine model of a current block in a current picture may be determined. The first affine model may be associated with the current block and a first reference block of the current block in a first reference picture.

[0185] At (S1920), an initial CPMV of a current block associated with a second reference picture may be determined based on a second affine model derived from the first affine model. The second affine model may be associated with the current block and a second reference block of the current block in the second reference picture. Second affine parameters of the second affine model may be symmetric with respect to the first affine parameters of the first affine model.

[0186] At (S1930), an improved CPMV of the current block associated with the second reference picture can be determined based on the initial CPMV of the current block associated with the second reference picture and the first affine motion search.

[0187] At (S1940), an improved CPMV of the current block associated with the first reference picture may be determined based on the initial CPMV of the current block associated with the first reference picture and the second affine motion search. The initial CPMV of the current block associated with the first reference picture may be derived from the improved CPMV of the current block associated with the second reference picture and may be symmetric with respect to the improved CPMV.

[0188] At (S1950), prediction information of the current block can be determined based on an improved CPMV of the current block associated with the first reference picture and an improved CPMV of the current block associated with the second reference picture.

[0189] To determine an improved CPMV of the current block associated with the second reference picture, an initial predictor of the current block may be determined based on the initial CPMV of the current block associated with the second reference picture. A first predictor of the current block may be determined based on the initial predictor. The first predictor may be equal to the sum of (i) the initial predictor of the current block, (ii) a product of a first component of a gradient value of the initial predictor and a first component of a motion vector difference associated with the initial predictor and the first predictor, and (iii) a product of a second component of the gradient value of the initial predictor and a second component of the motion vector difference.

[0190] To determine an improved CPMV of the current block associated with the second reference picture, the improved CPMV of the current block can be determined based on the Nth predictor associated with the second reference picture in response to one of (i) N being equal to an upper iteration value of the first affine motion search, and (ii) a motion vector difference associated with the Nth predictor and the (N+1)th predictor is zero.

[0191] In some embodiments, the first affine parameters may include a first translation factor and at least one of a first zoom factor or a first rotation factor, and the second affine parameters may include a second translation factor and at least one of a second zoom factor or a second rotation factor.

[0192] In one example, the sum of the first rotation factor and the second rotation factor can be zero, the sum of the first translation factor and the second translation factor can be zero, and the product of the first zoom factor and the second zoom factor can be one.

[0193] In one example, the ratio between the first rotation factor and the second rotation factor may be linearly proportional to the ratio between the first temporal distance between the first reference picture and the current picture and the second temporal distance between the second reference picture and the current picture, and the ratio between the first zoom factor and the second zoom factor may be exponentially proportional to the ratio between the first temporal distance and the second temporal distance.

[0194] The process then proceeds to (S1999) and ends.

[0195] The process 1900 may be adapted as appropriate. Step(s) of the process 1900 may be modified and / or omitted. Additional step(s) may be added. Any suitable order of implementation may be used.

[0196] As shown in Figure 20, the process (2000) can start at (S2001) and proceed to (S2010). At (S2010), first affine parameters of a first affine model of a symmetric affine mode to be applied to a current block in a current picture can be determined. The first affine model can be associated with the current block and a first reference block in a first reference picture of the current picture.

[0197] At (S2020), second affine parameters of a second affine model in a symmetric affine mode may be determined based on the first affine parameters of the first affine model. The second affine model may be associated with a current block in a second reference picture and a second reference block for the current block. The first affine parameters and the second affine parameters may have one of opposite signs, inverse values, and a proportional relationship based on a first temporal distance between the first reference picture and the current picture and a second temporal distance between the second reference picture and the current picture.

[0198] At (S2030), a control point motion vector (CPMV) of the current block can be determined based on the first affine model and the second affine model.

[0199] In (S2040), prediction information for the current block may be generated based on the determined CPMV of the current block and a flag indicating whether the symmetric affine mode is applied to the current block.

[0200] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 21 illustrates a computer system (2100) suitable for implementing certain embodiments of the disclosed subject matter.

[0201] Computer software can be coded using any suitable machine or computer language that can be subject to assembly, compilation, linking, or similar mechanisms to create code containing instructions that can be executed directly or via interpretation, microcode execution, etc. by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.

[0202] The instructions may be executed on various types of computers or computer components including, for example, personal computers, tablet computers, servers, smartphones, gaming consoles, Internet of Things devices, and the like.

[0203] The components of the computer system (2100) illustrated in Figure 21 are exemplary in nature and are not intended to impose any limitations on the scope or functionality of use of the computer software for implementing embodiments of the present disclosure. The arrangement of components should not be interpreted as having any dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of the computer system (2100).

[0204] The computer system (2100) may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, via tactile input (such as keystrokes, swipes, or data glove movements), audio input (such as voice or claps), visual input (such as gestures), or olfactory input (not depicted). Human interface devices may also be used to capture certain media not necessarily directly associated with conscious human input, such as audio (such as voice, music, or ambient sounds), images (such as scanned images, photographic images obtained from a still image camera), or video (such as two-dimensional video, three-dimensional video, including stereoscopic video).

[0205] The input human interface devices may include one or more of a keyboard (2101), a mouse (2102), a trackpad (2103), a touchscreen (2110), a data glove (not shown), a joystick (2105), a microphone (2106), a scanner (2107), and a camera (2108) (only one of each is depicted).

[0206] The computer system (2100) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen (2110), data gloves (not shown), or joystick (2105), although haptic feedback devices that do not function as input devices may also exist), audio output devices (such as speakers (2109) or headphones (not depicted)), visual output devices (such as screens (2110), including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touchscreen input capability and each with or without haptic feedback capability, some of which may be capable of outputting two-dimensional visual output or three-dimensional or higher-dimensional output via means such as stereographic output, virtual reality glasses (not depicted), holographic displays, and smoke tanks (not depicted)), and printers (not depicted).

[0207] The computer system (2100) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (2120) with CD / DVD or similar media (2121), thumb drives (2122), removable hard drives or solid state drives (2123), legacy magnetic media (not depicted) such as tape and floppy disks, and specialized ROM / ASIC / PLD-based devices (not depicted) such as security dongles.

[0208] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.

[0209] The computer system (2100) may also include an interface (2154) to one or more communication networks (2155). The networks may be, for example, wireless, wired, or optical. The networks may further be local, wide-area, metropolitan, vehicular, and industrial, real-time, delay-tolerant, and the like. Examples of networks include local area networks such as Ethernet and wireless LAN; cellular networks including GSM, 3G, 4G, 5G, LTE, and the like; wired or wireless wide-area digital TV networks including cable TV, satellite TV, and terrestrial broadcast TV; and vehicular and industrial networks including CAN Bus. Certain networks typically require an external network interface adapter attached to a particular general-purpose data port (e.g., a USB port on the computer system (2100)) or peripheral bus (2149); other networks are typically integrated into the core of the computer system (2100) by attaching to a system bus (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (2100) can communicate with other entities. Such communication may be one-way receive only (e.g., broadcast TV), one-way transmit only (e.g., CANbus to a specific CANbus device), or two-way with other computer systems using, for example, local or wide-area digital networks. Specific protocols and protocol stacks may be used with each of these networks and network interfaces, as described above.

[0210] The aforementioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the core (2140) of the computer system (2100).

[0211] The cores (2140) may include specialized programmable processing devices in the form of one or more central processing units (CPUs) (2141), graphics processing units (GPUs) (2142), field programmable gate arrays (FPGAs) (2143), task-specific hardware accelerators (2144), graphics adapters (2150), etc. These devices may be connected via a system bus (2148), along with read-only memory (ROM) (2145), random access memory (2146), and internal mass storage (2147) such as a non-user-accessible internal hard drive or SSD. In some computer systems, the system bus (2148) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (2148) or via a peripheral bus (2149). In one example, a screen (2110) may be connected to the graphics adapter (2150). Architectures for peripheral buses include PCI, USB, etc.

[0212] The CPU (2141), GPU (2142), FPGA (2143), and accelerator (2144) can combine to execute specific instructions that may constitute the aforementioned computer code. The computer code can be stored in ROM (2145) or RAM (2146). Transient data can also be stored in RAM (2146), and persistent data can be stored, for example, in internal mass storage (2147). Rapid storage and retrieval from any of the memory devices can be enabled using cache memory, which can be closely associated with one or more of the CPU (2141), GPU (2142), mass storage (2147), ROM (2145), RAM (2146), etc.

[0213] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.

[0214] By way of example, and not limitation, a computer system (2100) having an architecture, and specifically a core (2140), can provide functionality as a result of processor(s) (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be the user-accessible mass storage introduced above, as well as media associated with specific storage of the core (2140) that is non-transitory in nature, such as the core's internal mass storage (2147) or ROM (2145). Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (2140). The computer-readable media can include one or more memory devices or chips, depending on particular needs. The software can cause the core (2140), and specifically the processor(s) therein (including a CPU, GPU, FPGA, etc.), to perform particular processes or portions of particular processes described herein, including defining data structures stored in RAM (2146) and modifying such data structures according to software-defined processes. Additionally, or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator (2144)) that can operate in place of or together with software to perform particular processes or portions of particular processes described herein. Where appropriate, references to software can encompass logic, and vice versa. Where appropriate, references to computer-readable media can encompass circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software. Appendix A: Acronyms JEM: Joint Exploration Model VVC: Versatile Video Coding BMS: Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplemental Extended Information VUI: Video Usability Information GOP: Group of Pictures TU: Conversion unit PU: Prediction Unit CTU: Coding Tree Unit CTB: coding tree block PB: Predicted Block HRD: Hypothetical Reference Decoder SNR: Signal-to-Noise Ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: cathode ray tube LCD: Liquid crystal display OLED: Organic Light Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE: Long Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Area SSD: Solid State Drive IC: Integrated Circuit CU: Coding Unit

[0215] While this disclosure describes several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure. [Explanation of symbols]

[0216] 0 Reference list, 1 Reference list, 101 Point where arrows converge, Sample, 102 Arrow, 103 Arrow, 104 Square block, 110 Schematic diagram, 201 Current block, 202 Sample, 203 Sample, 204 Sample, 205 Sample, 206 Sample, 300 Communication system, 310 Terminal device, 320 Terminal device, 330 Terminal device, 340 Terminal device, 350 Communication network, 400 Communication system, 401 Video source, 402 Stream of video pictures, 403 Video encoder, 404 Encoded video data, 405 Streaming server, 406 Client subsystem, 407 Input copy of encoded video data, 408 Client subsystem, 409 Encoded video data, 410 Video decoder, 411 Output stream of video pictures, 412 Display, 413 Capture subsystem, 420 Electronic device, 430 Electronic device, 501, channel, 510, video decoder, 512, render device, 515, buffer memory, 520, parser, 521, symbol, 530, electronic device, 531, receiver, 551, scaler / inverse transform, 552, intra-picture prediction unit, 553, motion compensation prediction unit, 555, aggregator, 556, loop filter unit, 557, reference picture memory, 558, current picture buffer, 601, video source, 603, video encoder, 620, electronic device, 630, source coder, 632, coding engine, 633, local video decoder, 634, reference picture memory, 635, predictor, 640, transmitter, 643, video sequence, 645, entropy coder, 650, controller, 660, communication channel, 703, video encoder, 721, general controller, 722, intra-encoder, 723, residual calculator, 724 Residual encoder, 725 entropy encoder, 726 switch, 728 residual decoder, 730 inter-encoder, 810 video decoder, 871 entropy decoder, 872 intra-decoder, 873 residual decoder, 874 reconstruction module, 880 inter-decoder, 902 block, 904 block, 1000 block, 1002Motion vector of center sample, 1004 Sub-block, 1202 CU, 1204 Current block, current CU, 1302 Current block, current CU, 1400 Current block, 1402 Sub-block, 1404 Sample, 1406 Reference pixel, 1408 Reference pixel, 1410 Sub-block, 1600 Affine ME process, 1702 Current picture, 1704 First reference picture, 1706 Second reference picture, 1708 Current block, 1710 First reference block, 1712 First reference block, 1714 Second reference block, 1716 Second reference block, 1800 Decoding process, 1900 Encoding process, 2000 Encoding process, 2100 Computer system, 2101 Keyboard, 2102 Mouse, 2103 Trackpad, 2105 Joystick, 2106 microphone, 2107 scanner, 2108 camera, 2109 speaker, 2110 touch screen, 2120 CD / DVD ROM / RW, 2121 CD / DVD or similar medium, 2122 thumb drive, 2123 removable hard drive or solid state drive, 2140 core, 2141 central processing unit, CPU, 2142 graphics processing unit, GPU, 2143 field programmable gate area, FPGA, 2144 hardware accelerator, 2145 read only memory, ROM, 2146 random access memory, RAM, 2147 core internal mass storage, 2148 system bus, 2149 specific general purpose data port or peripheral bus, 2150 graphics adapter, 2154 interface, 2155 communication network

Claims

1. 1. A method of video decoding performed by a video decoder, the method comprising: receiving coded information of a current block in a current picture from a coded video bitstream, the coded information including a flag indicating whether a Symmetric Affine mode is applied to the current block, and wherein first affine parameters of a first affine model of the symmetric affine mode, the first affine parameters including a first translation factor and at least one of a first zoom factor or a first rotation factor, and second affine parameters of a second affine model of the symmetric affine mode, the second affine parameters including a second translation factor and at least one of a second zoom factor or a second rotation factor, have a proportional relationship based on a first temporal distance between a first reference picture and the current picture and a second temporal distance between a second reference picture and the current picture; determining the first affine parameters from the received coded information in response to the flag indicating that the symmetric affine mode is to be applied to the current block, the first affine model being associated with the current block and a first reference block of the current block in a first reference picture of the current picture; deriving the second affine parameters based on the first affine parameters of the first affine model, the second affine model being associated with the current block and a second reference block of the current block in a second reference picture of the current picture, and each type of the first affine parameters and the second affine parameters having a proportional relationship based on a first temporal distance between the first reference picture and the current picture and a second temporal distance between the second reference picture and the current picture; determining a control point motion vector (CPMV) of the current block based on the first affine model and the second affine model; reconstructing the current block based on the determined CPMV of the current block; Including, If the absolute value of the first temporal distance and the absolute value of the second temporal distance are different, a ratio of the first rotation factor to the second rotation factor is linearly proportional to a ratio of the first temporal distance to the second temporal distance; a ratio between the first zoom factor and the second zoom factor that is exponentially proportional to the ratio between the first temporal distance and the second temporal distance; method.

2. The method of claim 1 , wherein the flag is coded via one of a context-adaptive binary arithmetic coding (CABAC) context and a bypass code.

3. determining that the symmetric affine mode is associated with a four-parameter affine model in response to the flag indicating that the symmetric affine mode is to be applied to the current block; The method of claim 1.

4. If the absolute value of the first temporal distance and the absolute value of the second temporal distance are equal, the flag indicates that the symmetric affine mode is applied to the current block based on an absolute value of the first temporal distance between the current picture and the first reference picture being equal to an absolute value of the second temporal distance between the current picture and the second reference picture. The method of claim 1.

5. in response to the flag indicating that the symmetric affine mode is to be applied to the current block, deriving reference index information indicating which reference picture in a first reference list is the first reference picture and which reference picture in a second reference list is the second reference picture; The method of claim 1 further comprising:

6. If the absolute value of the first temporal distance and the absolute value of the second temporal distance are equal, the first affine parameters include a first translation factor and at least one of a first zoom factor or a first rotation factor; the second affine parameters include a second translation factor and at least one of a second zoom factor or a second rotation factor; the sum of the first rotation factor and the second rotation factor is zero; the sum of the first translation factor and the second translation factor is zero; the product of the first zoom factor and the second zoom factor is 1; The method of claim 1.

7. Apparatus configured to carry out the method of any one of claims 1 to 6.

8. the coded information includes an affine type message; In response to the affine type message indicating that a six-parameter affine model is applied to the current block, the flag indicates that the symmetric affine mode is not applied to the current block.

8. The apparatus of claim 7.

Citation Information

Patent Citations

  • Affine motion prediction for video coding

    JP2019519980A

  • Improvements to frame rate upconversion coding mode

    JP2019534622A

  • Side motion refinement in video encoding / decoding systems.

    JP2022515875A