Affine Merge Mode Using Translation Vectors
By employing affine-translational merge candidates for video blocks, the patent addresses inefficiencies in existing video coding technologies, enhancing compression efficiency and reducing redundancy in video data.
Patent Information
- Application Number
- JP2024516533
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-08
- Filing Date
- 2022-11-09
- Publication Date
- 2025-07-03
AI Technical Summary
Existing video coding technologies face challenges in efficiently reducing redundancy and improving compression ratios, particularly in handling complex motions and directional predictions, which affect the bitrate and storage requirements of video data.
The implementation of affine-translational merge candidates for video blocks, which combine affine motion information from one reference picture and translational motion information from another, allowing for more accurate and efficient prediction by deriving motion information from multiple reference lists.
Enhances compression efficiency by reducing redundancy and improving prediction accuracy, leading to lower bitrates and reduced storage needs for video data.
Smart Images

Figure 2025520236000001_ABST
Abstract
Description
Technical Field
[0001]
[0001] Incorporation by Reference This application claims the benefit of priority to U.S. Patent Application No. 17 / 982,938, filed Nov. 8, 2022, "Affine Merge Mode Using Translation Motion Vectors," which claims the benefit of priority to U.S. Provisional Application No. 63 / 353,319, filed Jun. 17, 2022, "Affine Merge Mode Using Translation Motion Vectors." The disclosure of the prior application is incorporated herein by reference in its entirety.
[0002]
[0002] Technical Field This disclosure generally describes embodiments related to video coding.
Background Art
[0003]
[0003] Background The background description provided herein is for the purpose of generally presenting the context of the disclosure. The work done under the present inventor's name is not admitted as prior art to this disclosure, either expressly or by implication, to the extent that the work is described in a manner that does not otherwise render it eligible as prior art at the time of filing, not only in this background section but also in other respects.
[0004]
[0004] Uncompressed digital images and / or videos can include a series of pictures, each picture having spatial dimensions of, for example, 1920×1080 luminance samples and associated chrominance samples. The series of pictures can have a fixed or variable picture rate (informally known as the frame rate), for example, 60 pictures per second, or 60 Hz. Uncompressed images and / or videos have significant bitrate requirements. For example, 1080p60 4:2:0 video with 8 bits per sample (luminance sample resolution of 1920×1080 at a frame rate of 60 Hz) requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires a storage space of more than 600 GB.
[0005]
[0005] One of the purposes of encoding and decoding images and / or videos can be said to be the reduction of redundancy in the input image and / or video signal by compression. Compression can, in some cases, help reduce the aforementioned bandwidth and / or storage space requirements by more than an order of magnitude. The description in this case uses video encoding / decoding as an illustrative example, but the same techniques can be applied to image encoding / decoding in a similar manner without departing from the spirit of the present disclosure. Both lossless compression and lossy compression, as well as combinations thereof, can be used. Lossless compression (reversible compression) refers to a technique in which an exact copy of the original signal can be reconstructed from the compressed original signal. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough for the reconstructed signal to be useful for the intended application. In the case of video, lossy compression is widely used. The amount of acceptable distortion depends on the application; for example, users of a particular consumer streaming application may be able to tolerate higher distortion than users of a television distribution application. The achievable compression ratio can reflect the fact that higher acceptable / tolerable distortion can result in a higher compression ratio.
[0006]
[0006] Video encoders and decoders can utilize techniques in several broad categories, including, for example, motion compensation, transform processing, quantization, and entropy coding.
[0007]
[0007] Video codec technology can include techniques known as intra coding. In intra coding, sample values are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially divided into blocks of samples. If all blocks of samples are coded in the intra mode, that picture can be an intra picture. Intra pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and thus can be used as the first picture in a coded video bitstream and video session or as a still image. Samples of an intra block can be subjected to a transform, and the transform coefficients can be quantized prior to entropy coding. Intra prediction can be a technique for minimizing sample values in the pre-transform domain. In some cases, the smaller the post-transform DC value and the smaller the AC coefficients, the fewer bits are required to represent the block after entropy coding for a given quantization step size.
[0008]
[0008] For example, conventional intra-coding used in MPEG-2 generation coding technology does not use intra prediction. However, some new video compression technologies include techniques that attempt to perform prediction based on surrounding sample data and / or metadata obtained during encoding and / or decoding of data blocks. Such techniques are hereinafter referred to as "intra prediction" techniques. It should be noted that in at least some cases, intra prediction uses only reference data from the current picture being reconstructed and does not use reference data from reference pictures.
[0009]
[0009] There can be many different forms of intra prediction. If more than one such technique may be used in a given video coding technology, the particular technique used may be coded as a particular intra prediction mode that uses that particular technique. In certain cases, the intra prediction mode may have submodes and / or parameters, in which case the submodes and / or parameters may be individually coded or included in a mode codeword, which defines the prediction mode being used. Which codeword to use for a given combination of mode, submode, and / or parameter can affect the coding efficiency gain due to intra prediction and can also affect the entropy coding technology used to convert the codeword into the bitstream.
[0010]
[0010] The specific mode of intra prediction was introduced in H.264, improved in H.265, and further refined in new coding technologies such as JEM (joint exploration model), VVC (versatile video coding), and MBS (benchmark set). The prediction block can be formed using sample values of samples in the vicinity of already available samples. The sample values of the neighboring samples are copied to the predictor block according to a certain direction. The reference for the direction in use can be coded in the bitstream or can be predicted itself.
[0011]
[0011] Referring to FIG. 1A, what is depicted at the lower right is a subset of 9 predictor directions that can be seen from the 33 possible predictor directions defined in H.265 (corresponding to the 33 angular modes out of 35 intra - modes). The point (101) where the arrows converge represents the sample to be predicted. The arrows represent the directions from which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from the sample that goes diagonally up to the right at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from the sample that goes diagonally down to the left of sample (101) at an angle of 22.5 degrees from the horizontal.
[0012]
[0012] Referring further to FIG. 1A, a square block (104) of 4×4 samples is depicted in the top left (shown by the thick dashed line). The square block (104) contains 16 samples, each labeled with an “S”, its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample from the top in the Y dimension and the first sample from the left in the X dimension. Similarly, sample S44 is the fourth sample within block (104) in both the Y and X dimensions. Since the block size is 4×4 samples, S44 is in the bottom right. Further, reference samples following a similar numbering scheme are shown. The reference samples are labeled with the Y position (e.g., row index) and X position (column index) relative to block (104) and an R. In both H.264 and H.265, since the predicted samples are in the neighborhood of the block being reconstructed, there is no need to use negative values.
[0013]
[0013] Intra-picture prediction can be performed by copying the reference sample value from neighboring samples indicated by the signaled prediction direction. For example, assume that the coded video bitstream contains signaling indicating a prediction direction that matches arrow (102) for this block, i.e., assume that the samples are predicted from samples going diagonally up and to the right at a 45-degree angle from horizontal. In this case, samples S41, S32, S23, S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.
[0014]
[0014] In certain cases, the values of multiple reference samples can be combined, particularly for cases where the direction is not evenly divisible by 45 degrees, for example, by interpolation, to calculate the reference sample.
[0015]
[0015] As video coding technology develops, the number of possible directions is increasing. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013). Currently, JEM / VVC / BMS can support up to 65 directions. Experiments have been conducted to identify the most likely directions, and specific techniques in entropy coding are used to represent these likely directions with fewer bits while accepting a specific penalty for less likely directions. Furthermore, the directions themselves can sometimes be predicted from neighboring directions already decoded in neighboring blocks.
[0016]
[0016] FIG. 1B shows a schematic diagram (110) depicting 65 intra prediction directions according to JEM, indicating the increasing number of prediction directions over time.
[0017]
[0017] The mapping of intra prediction direction bits representing directions within a coded video bitstream can vary for each video coding technology. Such mapping can range from a simple direct mapping to complex adaptive schemes including codewords, the most likely modes, and similar techniques. However, in most cases, there may be specific directions (specific directions that statistically occur less likely in the video content than other specific directions). Since the goal of video compression is redundancy reduction, these less likely directions are represented with more bits than the likely directions in well - operating video coding technologies.
[0018]
[0018] The encoding and decoding of images and / or videos may be performed using inter-picture prediction with motion compensation. Motion compensation may be a lossless compression technique, and a block of sample data from a previously reconstructed image or a part thereof (reference image) may be spatially shifted in the direction indicated by a motion vector (hereinafter referred to as MV) and then used for the prediction of a newly reconstructed picture or picture part. In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions X and Y, or three dimensions, the third being an index of the reference picture used (the latter may indirectly be the temporal dimension).
[0019]
[0019] In some video compression techniques, the motion vectors (MVs) applicable to a given region of sample data can be predicted from other MVs, for example, from MVs associated with another area of sample data that is spatially adjacent to the area being reconstructed and that precedes that MV in decoding order. By doing so, the amount of data required to code the MVs can be significantly reduced, thereby eliminating redundancy and increasing the compression ratio. For example, when coding an input video signal derived from a camera (known as natural video), there is a statistical likelihood that areas larger than the area to which a single MV is applicable will move in a similar direction, and thus in some cases, it is possible to predict using a similar motion vector derived from the MVs of adjacent areas, so MV prediction can function effectively. This results in an MV that is found to be similar or identical to the MV predicted from surrounding MVs for a given area, and it can be represented with fewer bits than when directly coding the MV after entropy coding. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., the MV) derived from the original signal (i.e., the sample stream). In other cases, MV prediction itself can be lossy due to rounding errors, for example, when calculating a predictor from several surrounding MVs.
[0020]
[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, “High Efficiency Video Coding”, December 2016). Of the many MV prediction mechanisms provided by H.265, the one described with reference to FIG. 2 is a technique hereafter referred to as “spatial merge”.
[0021]
[0021] Referring to FIG. 2, the current block (201) includes samples discovered by the encoder during the motion search process such that it can be predicted from previous blocks of the same size that are spatially shifted. Instead of directly coding its MV, the MV can be derived from the metadata associated with one or more reference pictures, for example, using the MV associated with any of five surrounding samples (202 to 206 respectively) shown as A0, A1, and B0, B1, B2, from the latest reference picture (in the decoding order). In H.265, MV prediction can use predictors from the same reference picture as used by neighboring blocks.
Summary of the Invention
[0022]
[0022] Aspects of the present disclosure provide methods and apparatuses for video encoding / decoding. In some examples, an apparatus for video decoding includes a processing circuit. The processing circuit receives a coded video bitstream including a current block in a current picture. The processing circuit determines a first affine-translational merge candidate for prediction of the current block in the current picture from among a candidate list. The first affine-translational merge candidate provides affine motion information associated with a first reference picture in a first reference list and translational motion information associated with a second reference picture in a second reference list. The processing circuit generates a first prediction for samples in the current block according to the affine motion information associated with the first reference picture, and generates a second prediction for samples in the current block according to the translational motion information associated with the second reference picture. The processing circuit reconstructs the samples of the current block according to a combination of the first prediction and the second prediction.
[0023]
[0023] In some examples, the processing circuit derives the affine motion information and the translational motion information of the first affine-translational merge candidate, and inserts the first affine-translational merge candidate into the candidate list.
[0024]
[0024] In some examples, the processing circuit derives the affine motion information according to the first affine motion information of the first affine merge candidate by using uni-prediction. In another example, the processing circuit derives the affine motion information according to the second affine motion information of the second affine merge candidate by using bi-prediction, and the second affine motion information is associated with the first reference list.
[0025]
[0025] In one example, the processing circuit derives the translational motion information according to the first translational motion information of the first translational merge candidate by using uni-prediction. In another example, the processing circuit derives the translational motion information according to the second translational motion information of the second translational merge candidate by using bi-prediction, and the second translational motion information is associated with the second reference list.
[0026]
[0026] In one example, to derive the translational motion information, regarding the merge candidate with the merge index in the translational merge candidate list, the first list is selected from the first reference list and the second reference list based on the merge index according to a pre-set rule. The processing circuit derives the translational motion information based on the first translational motion information of the merge candidate associated with the first list in response to the presence of the first translational motion information; and derives the translational motion information based on the second translational motion information of the merge candidate associated with another list different from the first list in response to the absence of the first translational motion information.
[0027]
[0027] In some examples, the translational motion information can be derived from the affine merge candidates. In one example, the processing circuit derives the translational motion information based on the control point motion vectors of the affine merge candidates. In another example, the processing circuit derives the translational motion information based on the average of a plurality of control point motion vectors of the affine merge candidates. In another example, the processing circuit derives the translational motion information based on the motion vector at the center of the current block, and the motion vector is determined according to the affine model for the affine merge candidate.
[0028]
[0028] In one example, the processing circuit inserts the first affine-translational merge candidate after the last inherited affine merge candidate and into the existing sub-block-based merge candidate list. In another example, the processing circuit inserts the first affine-translational merge candidate after the last constructed affine merge candidate and into the existing sub-block-based merge candidate list. In another example, the processing circuit inserts the first affine-translational merge candidate after the last affine merge candidate and before the zero motion vector candidate and into the existing sub-block-based merge candidate list.
[0029]
[0029] In one example, the processing circuit inserts the first affine-translational merge candidate into the affine-translational merge list separately from the existing sub-block-based merge candidate list. In one example, in response to the decoded flag indicating the selection of the affine-translational merge list, the processing circuit selects the first affine-translational merge candidate from the affine-translational merge list according to the decoded index. In response to the decoded flag indicating the selection of the sub-block-based merge candidate list, the merge candidate is selected from the sub-block-based merge candidate list according to the decoded index.
[0030]
[0030] In some examples, candidates in the affine-translation merge list are reordered (or sorted) according to the candidate's template matching cost.
[0031]
[0031] In some examples, the processing circuit inserts a first affine-translation merge candidate into the candidate list in response to the count of affine-translation merge candidates in the candidate list being lower than a value. That value can be determined in advance or can be signaled at a high level from the encoder side to the decoder side.
[0032]
[0032] In some examples, the processing circuit derives affine motion information from a subset of affine merge candidates that are before other affine merge candidates in the affine merge candidate list, and the count of affine merge candidates in the subset is equal to a value. That value can be determined in advance or can be signaled at a high level from the encoder side to the decoder side.
[0033]
[0033] In some examples, the processing circuit derives translational motion information from a subset of translational merge candidates that are before other translational merge candidates in the translational merge candidate list, and the count of translational merge candidates in the subset is equal to a value. That value can be determined in advance or can be signaled at a high level from the encoder side to the decoder side.
[0034]
[0034] In some examples, the processing circuit determines whether bi-prediction is allowed in the current picture and, in response to that determination, determines a first affine-translation merge candidate for predicting the current block.
[0035]
[0035] In some examples, the processing circuit determines that the reference picture of the current picture is earlier in time than the current picture, and in response to that determination, determines a first affine-translation merge candidate for prediction of the current block.
[0036]
[0036] In some examples, a high-level syntax indicating allowance of the first affine-translation merge candidate for prediction of the current block is determined (e.g., decoded from a bitstream).
[0037]
[0037] In some examples, in response to local illumination compensation (LIC) being enabled but bi-prediction being prohibited, the processing circuit sets the LIC flag associated with the first affine-translation merge candidate to false.
[0038]
[0038] In one example, the processing circuit sets a bi-prediction with coding unit level weighting (BCW) index to a default value to combine the first prediction and the second prediction with equal weights.
[0039]
[0039] In another example, only one of the affine motion information and the translational motion information has a bi-prediction (BCW) index with coding unit level weighting. The first prediction and the second prediction are combined according to the BCW index.
[0040]
[0040] In another example, a first BCW index is associated with the affine motion information, and a second BCW index is associated with the translational motion information. The first prediction and the second prediction are combined according to the first BCW index.
[0041]
[0041] Aspects of the present disclosure also provide a non-transitory computer-readable storage medium storing instructions that, when executed by a computer for video decoding, cause the computer to perform a method for video decoding.
Brief Description of the Drawings
[0042]
[0042] Further features, characteristics, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.
Figure 1A
[0043] FIG. 1A is a schematic diagram of an exemplary subset of intra prediction modes.
Figure 1B
[0044] FIG. 1B is a diagram of exemplary intra prediction directions.
Figure 2
[0045] FIG. 2 is a schematic diagram of a current block and its surrounding spatial merge candidates in one example.
Figure 3
[0046] FIG. 3 is a schematic diagram of a simplified block diagram of a communication system (300) according to an embodiment.
Figure 4
[0047] FIG. 4 is a schematic diagram of a simplified block diagram of a communication system (400) according to an embodiment.
Figure 5
[0048] FIG. 5 is a schematic diagram of a simplified block diagram of a decoder according to an embodiment.
Figure 6
[0049] FIG. 6 is a schematic diagram of a simplified block diagram of an encoder according to an embodiment.
Figure 7
[0050] FIG. 7 shows a block diagram of an encoder according to another embodiment.
Figure 8
[0051] FIG. 8 shows a block diagram of a decoder according to another embodiment.
Figure 9A
[0052] FIG. 9A shows a schematic diagram of a four-parameter affine model according to another embodiment.
Figure 9B
[0053] Figure 9B shows a schematic diagram of a six-parameter affine model according to another embodiment.
Figure 10
[0054] Figure 10 shows a schematic diagram of an affine motion vector field associated with sub-blocks within a block according to another embodiment.
Figure 11
[0055] Figure 11 shows a schematic diagram of an exemplary position of a spatial merge candidate according to another embodiment.
Figure 12
[0056] Figure 12 shows a schematic diagram of control point motion vector inheritance according to another embodiment.
Figure 13
[0057] Figure 13 shows a schematic diagram of candidate positions for constructing an affine merge mode according to another embodiment.
Figure 14
[0058] Figure 14 shows a schematic diagram of prediction refinement with optical flow (PROF) according to another embodiment.
Figure 15
[0059] Figure 15 shows a schematic diagram of an affine motion estimation process according to another embodiment.
Figure 16
[0060] Figure 16 shows a flowchart of an affine motion estimation search according to another embodiment.
Figure 17
[0061] Figure 17 shows a schematic diagram of an extended coding unit (CU) region related to bi-directional optical flow according to another embodiment.
Figure 18
[0062] Figure 18 shows an exemplary schematic diagram of motion vector refinement on the decoder side.
Figure 19
[0063] Figure 19 shows an example of a search process in one example.
Figure 20
[0064] Figure 20 shows an example of search points in one example.
Figure 21
[0065] Figure 21 shows an example of template matching in some examples.
Figure 22
[0066] Figure 22 shows an example of template matching in affine merge mode in one example.
Figure 23
[0067] Figure 23 is a diagram showing a reference sample of a template of a current block for a dual-prediction merge candidate.
Figure 24
[0068] Figure 24 shows the derivation of a template and a reference sample of the template for a current block using sub-block-based merge candidates in some examples.
Figure 25
[0069] Figure 25 is a diagram showing a refinement direction in some examples.
Figure 26
[0070] Figure 26 is a diagram showing a first history parameter table and a second history parameter table in some examples.
Figure 27
[0071] Figure 27 is a diagram showing a history parameter table stored in a line buffer in some examples.
Figure 28A
[0072] Figure 28A shows generating additional inherited / constructed affine merge / AMVP candidates in some examples.
Figure 28B
[0072] Figure 28B shows generating additional inherited / constructed affine merge / AMVP candidates in some examples.
Figure 29
[0073] Figure 29 shows generating affine merge / AMVP candidates constructed from non-adjacent neighborhoods in some examples.
Figure 30
[0074] FIG. 30 is a flowchart showing an overview of a process according to some embodiments of the present disclosure.
Figure 31
[0075] FIG. 31 shows a flowchart showing an overview of another process according to some embodiments of the present disclosure.
Figure 32
[0076] FIG. 32 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0043]
[0077] FIG. 3 shows an exemplary block diagram of a communication system (300). The communication system (300) includes a plurality of terminal devices capable of communicating with each other via, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via a network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) perform unidirectional data transmission. For example, the terminal device (310) can code video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The coded video data can be transmitted in the form of one or more coded video bitstreams. The terminal device (320) can receive the coded video data from the network (350), decode the coded video data to restore the video picture, and display the video picture according to the restored video data. Unidirectional data transmission may be common in media serving applications and the like.
[0044]
[0078] In another example, the communication system (300) includes, for example, a second pair of terminal devices (330) and (340) that perform bidirectional transmission of coded video data, for example, during a video conference. Regarding the bidirectional transmission of data, for example, each of the terminal devices (330) and (340) can code video data (e.g., a stream of video pictures captured by the terminal device) to transmit to the other of the terminal devices (330) and (340) via the network (350). Each of the terminal devices (330) and (340) can also receive the coded video data transmitted by the other of the terminal devices (330) and (340), can decode the coded video data to restore the video picture, and can display the video picture on an accessible display device according to the restored video data.
[0045]
[0079] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) are shown as a server, a personal computer, and smartphones, but the principles of the present disclosure need not be so limited. Embodiments of the present disclosure have applications using laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network (350) represents any number of networks that carry coded video data between the terminal devices (310), (320), (330), and (340), including, for example, wired (wired) and / or wireless communication networks. The communication network (350) can exchange data over circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of the present disclosure, the architecture and topology of the network (350) may not be important for the operation of the present disclosure, unless otherwise described below.
[0046]
[0080] FIG. 4 shows a video encoder and a video decoder in a streaming environment as an application example of the disclosed subject matter. The disclosed subject matter is equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, streaming services, storage of compressed video on digital media (including CDs, DVDs, memory sticks, etc.).
[0047]
[0081] The streaming system can include a video source (401), such as a digital camera, and may include a capture subsystem (413) capable of generating a stream of, for example, uncompressed video pictures (402). In one example, the stream of video pictures (402) includes samples taken by a digital camera. The stream of video pictures (402), drawn as a thick line to emphasize the large amount of data when compared to the encoded video data (404) (or coded video bitstream), can be processed by an electronic device (420) including a video encoder (403) coupled to the video source (401). The video encoder (403) includes hardware, software, or a combination thereof and can be operable to implement or realize aspects of the disclosed subject matter as detailed below. The encoded video data (404) (or encoded video bitstream), drawn as a thin line to emphasize the smaller amount of data when compared to the stream of video pictures (402), can be stored in a streaming server (405) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include, for example, a video decoder (410) within an electronic device (430). The video decoder (410) decodes an incoming copy (407) of the encoded video data and generates an output stream of video pictures (411) that can be rendered on a display (412), such as a display screen, or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., video bitstreams) can be encoded according to a particular video coding / compression standard.Examples of these standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.
[0048]
[0082] Note that the electronic devices (420) and (430) can include other components (not shown). For example, the electronic device (420) can include a video decoder (not shown), and the electronic device (430) can also include a video encoder (not shown).
[0049]
[0083] FIG. 5 shows an exemplary block diagram of a video decoder (510). The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). The video decoder (510) can be used in place of the video decoder (410) in the example of FIG. 4.
[0050]
[0084] Receiver (531) is capable of receiving one or more coded video sequences to be decoded by video decoder (510). In an embodiment, if the decoding of each coded video sequence is independent of other coded video sequences, it is possible to receive one coded video sequence at a time. The coded video sequence can be received from channel (501), which may be a hardware / software link to a storage device storing the encoded video data. Receiver (531) can receive the encoded video data together with other data, such as coded audio data and / or auxiliary data streams, and these data can be transferred using respective entities (not shown). Receiver (531) can separate the coded video sequence from other data. To handle network jitter, buffer memory (515) may be coupled between receiver (531) and entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In a particular application, buffer memory (515) is part of video decoder (510). In other cases, it may be outside video decoder (510) (not shown). In yet another example, for example to handle network jitter, there may be buffer memory (not shown) outside video decoder (510), and furthermore, for example to handle playback timing, there may be another buffer memory (515) inside video decoder (510). If receiver (531) is receiving data from a store-and-forward device with sufficient bandwidth and controllability, or from a synchronous network, buffer memory (515) may not be required or can be made smaller.For use in a best-effort packet network such as the Internet, buffer memory (515) may be required, which may be relatively large and advantageously may be of an adaptable size and may be implemented at least partially in an operating system or similar element (not shown) outside the video decoder (510).
[0051]
[0085] The video decoder (510) can include a parser (520) to reconstruct symbols (521) from the coded video sequence. The categories of these symbols include information used to manage the operation of the video decoder (510) and potentially information for controlling a rendering device (512) (e.g., a display screen) that is not an essential part of the electronic device (530) but can be coupled to the electronic device (530) as shown in FIG. 5. The control information for the rendering device can be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (520) can parse / entropy decode the received coded video sequence. The coding of the video sequence to be coded can conform to a video coding technology or standard and can follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context influence, etc. The parser (520) can extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder based on at least one parameter corresponding to a group. The subgroup can include a Group of Pictures (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), etc. The parser (520) can also extract from the coded video sequence information such as transform coefficients, quantization parameter values, motion vectors, etc.
[0052]
[0086] The parser (520) can perform entropy decoding / analysis processing on the video sequence received from the buffer memory (515) to generate symbols (521).
[0053]
[0087] The reconstruction of the symbols (521) can include a plurality of different units depending on the type of the coded video picture or a part thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. How each unit is included can be controlled by subgroup control information analyzed by the parser (520) from the coded video sequence. Such a flow of subgroup control information between the parser (520) and the plurality of subsequent units is not depicted for clarity.
[0054]
[0088] The video decoder (510) can be conceptually subdivided into a plurality of functional units as described below, in addition to the functional blocks already described. In a practical implementation operating under commercial constraints, many of these units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, the conceptual subdivision into the following functional units is appropriate.
[0055]
[0089] The first unit is the scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives, as symbols (521) from the parser (520), not only the quantized transform coefficients but also control information (including the transform to be used, block size, quantization factor, quantization scaling matrix, etc.). The scaler / inverse transform unit (551) can output a block including sample values that can be input to the aggregator (555).
[0056]
[0090] In some cases, the output samples of the scaler / inverse transform unit (551) may be related to intra-coded blocks. An intra-coded block is a block that does not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed part of the current picture. Such prediction information can be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) uses the already reconstructed surrounding information fetched from the buffer (558) of the current picture to generate a block of the same size and shape as the block being reconstructed. The current picture buffer (558) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. The aggregator (555) may, in some cases, add, sample by sample, the prediction information generated by the intra prediction unit (552) to the output sample information as provided by the scaler / inverse transform unit (551).
[0057]
[0091] Otherwise, the output samples of the scaler / inverse transform unit (551) can be related to the inter-coded and potentially motion-compensated blocks. In such a case, the motion compensation prediction unit (553) can access the reference picture memory (557) to retrieve the samples to be used for prediction. According to the symbol (521) related to the block, after motion-compensating the retrieved samples, these samples are added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (in this case, called residual samples or residual signal), generating output sample information. The address in the reference picture memory (557) from which the motion compensation prediction unit (553) retrieves the prediction samples can be controlled by the motion vectors available to the motion compensation prediction unit (553) in the form of, for example, X, Y, and the symbol (521) which can have reference picture components. Also, motion compensation can include interpolation of sample values taken from the reference picture memory (557), motion vector prediction mechanisms, etc. when exact sub-sample motion vectors are used.
[0058]
[0092] The output samples of the aggregator (555) can be affected by various loop filtering techniques within the loop filter unit (556). Video compression techniques can include in-loop filtering techniques that are included in the coded video sequence (also called the coded video bitstream) and are controlled by parameters made available to the loop filter unit (556) as symbols (521) from the parser (520). Also, video compression can be responsive to meta information obtained during the decoding of the previously decoded parts of the coded picture or coded video sequence (in decoding order), and can also be responsive to previously reconstructed loop-filtered sample values.
[0059]
[0093] The output of the loop filter unit (556) can be not only output to the rendering device (512), but also a sample stream that can be stored in the reference picture memory (557) for future inter-picture prediction.
[0060]
[0094] Once a given coded picture is completely reconstructed, it can be used as a reference picture for future prediction. For example, when the coded picture corresponding to the current picture is completely reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the buffer (558) of the current picture can become part of the reference picture memory (557), and the buffer of the new current picture can be reallocated before starting the reconstruction of the subsequent coded pictures.
[0061]
[0095] The video decoder (510) is capable of performing a decoding operation according to a standard such as ITU-T Rec.H.265 or a predetermined video compression technology. The coded video sequence can comply with the syntax specified by the video compression technology or standard being used, in the sense that the coded video sequence complies with both the syntax of the video compression technology or standard and the profile as documented in the video compression technology or standard. Specifically, the profile can select specific tools as the only tools available under that profile from all the tools available in the video compression technology or standard. Also, for compliance, it is necessary that the complexity of the coded video sequence falls within the range defined by the level of the video compression technology or standard. In some cases, the level restricts the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further restricted by the Hypothetical Reference Decoder (HRD) specifications and metadata for HRD buffer management signaled in the coded video sequence.
[0062]
[0096] In an embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0063]
[0097] Figure 6 shows an exemplary block diagram of a video encoder (603). The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used instead of the video encoder (403) in the example of FIG. 4.
[0064]
[0098] The video encoder (603) can receive video samples from a video source (601) (not part of the electronic device (620) in the example of FIG. 6) that can capture video images to be coded by the video encoder (603). In another example, the video source (601) is part of the electronic device (620).
[0065]
[0099] The video source (601) can provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits,...), any color space (e.g., BT.601 YCrCb, RGB,...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (601) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that convey motion when viewed in sequence. Each picture itself can be organized as a spatial array of pixels, and each pixel can include one or more samples depending on the sampling structure, color space, etc. in use. One of ordinary skill in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.
[0066]
[0100] According to an embodiment, the video encoder (603) can code and compress pictures of a source video sequence into a coded video sequence (643) in real time or under any other required time constraints. Enforcing an appropriate coding speed is one function of the controller (650). In some embodiments, the controller (650) controls other functional units and is functionally coupled to other functional units as described below. The coupling is not depicted for clarity. The parameters set by the controller (650) can include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization techniques, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other appropriate functions related to the video encoder (603) optimized for a particular system design.
[0067]
[0101] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As a very simplified explanation, in one example, the coding loop can include a source coder (630) (which, for example, is responsible for generating symbols such as a symbol stream based on an input picture and reference pictures to be coded) and a (local) decoder (633) incorporated in the video encoder (603). The decoder (633) reconstructs the symbols to generate sample data in the same way as a (remote) decoder does. The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream results in a bit-exact result regardless of the position of the decoder (local or remote), the content in the reference picture memory (634) is also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" exactly the same sample values as the samples that the decoder would "see" when using prediction during decoding as the reference picture samples. This basic principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, for example, due to channel errors) is also used similarly in some related technologies.
[0068]
[0102] It is possible to assume that the operation of the "local" decoder (633) is the same as that of a "remote" decoder such as the video decoder (510) already described in detail above in relation to FIG. 5. However, referring briefly to FIG. 5, since it is possible to assume that the symbols are available and the encoding / decoding of the symbols into the coded video sequence by the entropy coder (646) and the parser (520) is lossless, the entropy decoding part of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully realized in the local decoder (633).
[0069]
[0103] In one embodiment, decoder techniques other than parsing / entropy decoding present in the decoder are present in the corresponding encoder in an ideal or substantially identical functional form. Accordingly, the disclosed subject matter focuses on the operation of the decoder. Since the description of encoder techniques is the reverse of the decoder techniques described comprehensively, it can be omitted. In certain areas, more detailed descriptions are given below.
[0070]
[0104] During operation, in some examples, the source coder (630) can perform motion-compensated predictive coding that predictive-codes an input picture by referring to one or more previously-coded pictures from a video sequence designated as a "reference picture". In this way, the coding engine (632) codes the difference between a pixel block of the input picture and a pixel block of a reference picture that can be selected as a prediction reference for the input picture.
[0071]
[0105] The local video decoder (633) can decode the coded video data of a picture that can be specified as a reference picture based on the symbols generated by the source coder (630). The operation of the coding engine (632) may advantageously be a lossless process. If the coded video data can be decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (633) can repeat the decoding process that can be performed by the video decoder in the reference picture, causing the reconstructed reference picture to be stored in the reference picture cache (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference picture having common content as the reconstructed reference picture obtained by the video decoder at the remote end (assuming no transmission errors).
[0072]
[0106] The predictor (635) can perform a prediction search for the coding engine (632). That is, for a new picture to be coded, the predictor (635) can search the reference picture memory (634) for sample data (as a candidate reference pixel block) or predetermined metadata (such as a reference picture motion vector, block shape, etc.), which may serve as an appropriate prediction reference for the new picture. The predictor (635) can operate on a sample block - pixel block basis to find an appropriate prediction reference. In some cases, the input picture may have a prediction reference drawn from a plurality of reference pictures stored in the reference picture memory (634) as determined by the search result obtained by the predictor (635).
[0073]
[0107] The controller (650) can manage the coding operation of the source coder (630), including, for example, the setting of parameters and subgroup parameters used for encoding video data.
[0074]
[0108] All outputs of the aforementioned functional units can be entropy-coded in the entropy coder (646). The entropy coder (645) converts the symbols generated by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable-length coding, arithmetic coding, etc.
[0075]
[0109] The transmitter (640) can buffer the coded video sequence as created by the entropy coder (645) and prepare for transmission via the communication channel (660), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (640) can merge the coded video data from the video coder (603) with other data to be transmitted, such as, for example, coded audio data and / or auxiliary data streams (sources not shown).
[0076]
[0110] The controller (650) can manage the operation of the video encoder (603). During coding, the controller (650) can assign a specific coded picture type to each of the coded pictures, which can affect the coding techniques applicable to each picture. For example, a picture may often be assigned as one of the following picture types.
[0077]
[0111] An Intra Picture (I Picture) can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of Intra Pictures, such as, for example, Independent Decoder Refresh (IDR) Pictures. Those skilled in the art are aware of these variations of I Pictures, as well as their respective uses and characteristics.
[0078]
[0112] A Predicted Picture (P Picture) can be encoded and decoded using Intra prediction or Inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.
[0079]
[0113] A Bi-Directionally Predicted Picture (B Picture) can be encoded and decoded using Intra prediction or Inter prediction that uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures can use more than two reference pictures and associated metadata for the reconstruction of one block.
[0080]
[0114] The source picture is typically spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and can be coded block by block. The blocks can be predictive coded with reference to other (already coded) blocks as determined by the coding assignment applied to each block of the picture. For example, blocks of an I picture may be non-predictively coded, or they may be predictive coded with reference to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be predictive coded by spatial or temporal prediction with reference to one previously coded reference picture. Blocks of a B picture may be predictive coded by spatial or temporal prediction with reference to one or two previously coded reference pictures.
[0081]
[0115] The video encoder (603) can perform coding operations according to a predetermined video coding technology or standard such as ITU-T Rec.H.266. In this operation, the video encoder (603) can execute various compression operations including predictive coding operations that utilize the temporal and spatial redundancies in the input video sequence. The coded video data can thus conform to the syntax specified by the video coding technology or standard being used.
[0082]
[0116] In an embodiment, the transmitter (640) can transmit additional data along with the coded video. The source coder (630) can include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data (redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.).
[0083]
[0117] Video can be captured as a plurality of source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra prediction) utilizes the spatial correlation in a given picture, while inter-picture prediction utilizes the (temporal or other) correlation between pictures. In one example, a particular picture under coding / decoding, called the current picture, is partitioned into blocks. If a block in the current picture is similar to a reference block in a reference picture that has been previously coded and is still buffered in the video, the block in the current picture can be coded by a vector called a motion vector. The motion vector points to the reference block in the reference picture and can have a third dimension to identify the reference picture when multiple reference pictures are used.
[0084]
[0118] In some embodiments, it is possible to use dual-prediction techniques for inter-picture prediction. According to the dual-prediction technique, two reference pictures such as a first reference picture and a second reference picture that both precede the current picture in decoding order within the video (however, they may be in the past and future respectively in display order) are used. A block in the current picture can be coded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.
[0085]
[0119] Further, in order to improve coding efficiency, it is possible to use merge-mode techniques for inter-picture prediction.
[0086]
[0120] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and the CTUs within a picture have the same size such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs) which are one luma CTB and two chroma CTBs. Each CTU can be recursively quadtree partitioned into one or more coding units (CUs). For example, a 64×64 pixel CTU can be partitioned into one 64×64 pixel CU, four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine a prediction type of the CU, such as an inter prediction type or an intra prediction type. The CU is partitioned into one or more prediction units (PUs) depending on temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In an embodiment, the prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of values (e.g., luma values) for pixels such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0087]
[0121] FIG. 7 shows an exemplary diagram of a video encoder (703). The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values within a current video picture in a sequence of video pictures and encode the processing block into a coded picture that is part of a coded video sequence. In one example, the video encoder (703) is used instead of the video encoder (403) of the example of FIG. 4.
[0088]
[0122] In an example of HEVC, a video encoder (703) receives a matrix of sample values of a processing block, such as a prediction block of 8×8 samples. The video encoder (703) uses an intra-mode, an inter-mode, or a bi-prediction mode to determine whether the processing block is best coded, for example, using rate distortion optimization. If the processing block is to be coded in the intra-mode, the video encoder (703) can use intra prediction techniques to code the processing block into the coded picture; if the processing block is to be coded in the inter-mode or the bi-prediction mode, the video encoder (703) can use inter prediction techniques or bi-prediction techniques respectively to code the processing block into the coded picture. In certain video coding techniques, the merge mode may be an inter-picture prediction sub-mode, in which case the motion vector is derived from one or more motion vector predictors without the benefit of coded motion vector components outside the predictor. In certain other video coding techniques, there may be motion vector components applicable to the target block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown), to determine the mode of the processing block.
[0089]
[0123] In the example of FIG. 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general-purpose controller (721), and an entropy encoder (725) coupled together as shown in FIG. 7.
[0090]
[0124] The inter-encoder (730) receives samples of a current block (e.g., a processing block), compares the block with one or more reference blocks in a reference picture (e.g., blocks of a previous picture and blocks in a subsequent picture), generates inter-prediction information (e.g., a description of redundant information by an inter-coding technique, a motion vector, merge mode information), and is configured to calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using some suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on the encoded video information.
[0091]
[0125] The intra-encoder (722) receives samples of a current block (e.g., a processing block), and in some cases, compares the block with blocks already coded within the same picture, generates quantized coefficients after transformation, and in some cases, is also configured to generate intra-prediction information (e.g., intra-prediction direction information according to one or more intra-coding techniques). In one example, the intra-encoder (722) also calculates an intra-prediction result (e.g., a predicted block) based on the intra-prediction information and reference blocks within the same picture.
[0092]
[0126] The general-purpose controller (721) is configured to determine general control data and control other components of the video encoder (703) based on the general control data. In one example, the general-purpose controller (721) determines the mode of a block and provides a control signal to the switch (726) based on that mode. For example, when the mode is the intra mode, the general-purpose controller (721) controls the switch (726) to select the intra mode result for use by the residual calculator (723), and controls the entropy encoder (725) to select the intra prediction information and include the intra prediction information in the bitstream; also, when the mode is the inter mode, the general-purpose controller (721) controls the switch (726) to select the inter prediction result for use by the residual calculator (723), and controls the entropy encoder (725) to select the inter prediction information and include the inter prediction information in the bitstream.
[0093]
[0127] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) is configured to operate based on the residual data to encode the residual data and generate transform coefficients. In one example, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain and generate transform coefficients. Next, the transform coefficients are subjected to quantization processing to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform inverse transformation and generate decoded residual data. The decoded residual data can be appropriately used by the intra-encoder (722) and the inter-encoder (730). For example, the inter-encoder (730) can generate a decoded block based on the decoded residual data and the inter-prediction information, and the intra-encoder (722) can generate a decoded block based on the decoded residual data and the intra-prediction information. The decoded block is appropriately processed to generate a decoded picture, and the decoded picture is buffered in a memory circuit (not shown) and can be used as a reference picture in some examples.
[0094]
[0128] The entropy encoder (725) is configured to format the bitstream to include the encoded blocks. The entropy encoder (725) is configured to include various information in the bitstream according to an appropriate standard such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other appropriate information in the bitstream. It should be noted that there is no residual information when coding a block in any merge submode of the inter mode or the bi-prediction mode according to the disclosed subject matter.
[0095]
[0129] FIG. 8 shows an exemplary diagram of a video decoder (810). The video decoder (810) is configured to receive a coded picture that is part of a coded video sequence and decode the coded picture to generate a reconstructed picture. In one example, the video decoder (810) is used in place of the video decoder (410) in the example of FIG. 4.
[0096]
[0130] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) coupled together as shown in FIG. 8.
[0097]
[0131] The entropy decoder (871) can be configured to reconstruct from the coded picture certain symbols representing the syntax elements that make up the coded picture. Such symbols can include, for example, the mode in which a block is coded (e.g., intra mode, inter mode, bi-prediction mode, merge sub-mode in the latter two, or another sub-mode), prediction information (e.g., intra prediction information or inter prediction information) that can identify specific samples or metadata used for prediction by the intra decoder (872) or the inter decoder (880) respectively. Also, the symbols can include, for example, residual information in the form of quantized transform coefficients. In one example, when the prediction mode is inter or bi-prediction mode, inter prediction information is provided to the inter decoder (880); when the prediction type is intra prediction type, intra prediction information is provided to the intra decoder (872). The residual information can be inverse quantized and provided to the residual decoder (873).
[0098]
[0132] The inter decoder (880) is configured to receive inter prediction information and generate an inter prediction result based on the inter prediction information.
[0099]
[0133] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.
[0100]
[0134] The residual decoder (873) is configured to perform inverse quantization to extract non-quantized transform coefficients, process the non-quantized transform coefficients, and convert the residual information from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (including quantization parameter (QP)), and that information may be provided by the entropy decoder (871) (since this may be only a small amount of control information, the data path is not depicted).
[0101]
[0135] The reconstruction module (874) is configured to combine, in the spatial domain, the residual information as the output by the residual decoder (873) and the prediction result (which may be output by an inter or intra prediction module in some cases) to form a reconstructed block, which is part of a reconstructed picture, and the picture may be part of a reconstructed video. It should be noted that other appropriate processes such as deblocking processing may be performed to improve visual quality.
[0102]
[0136] Note that the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) can be implemented using any appropriate technology. In an embodiment, the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) can be implemented using one or more processors that execute software instructions.
[0103]
[0137] Some aspects of the present disclosure provide techniques for an affine merge mode using translational motion vectors.
[0104]
[0138] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) have published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). In 2015, the two standardization organizations jointly established JVET (Joint Video Exploration Team) to explore the possibility of developing the next video coding standard beyond HEVC. In October 2017, the two standardization organizations issued a Joint Call for Proposals (CfP) on video compression with capabilities beyond HEVC. By February 15, 2018, 22 CfP responses regarding standard dynamic range (SDR), 12 CfP responses regarding high dynamic range (HDR), and 12 CfP responses regarding the 360-degree video category were respectively submitted. In April 2018, all the accepted CfP responses were evaluated at the 122nd MPEG / JVET meeting. As a result of the meeting, JVET officially started the standardization process for the next-generation video coding beyond HEVC. The new standard was named VVC (Versatile Video Coding), and JVET was renamed the Joint Video Experts Team. In 2020, ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the VVC video coding standard (version 1).
[0105]
[0139] In inter prediction, for each inter-predicted coding unit (CU), for example, motion parameters for the coding features of VVC used for generating inter-predicted samples are required. The motion parameters can include motion vectors, reference picture indices, reference picture list usage indices, and / or additional information. The motion parameters can be signaled in an explicit or implicit manner. When a CU is coded in skip mode, the CU can be associated with one PU, and significant residual coefficients, coded motion vector deltas, and / or reference picture indices may not be required. When a CU is coded in merge mode, the motion parameters for the CU can be obtained from neighboring CUs. The neighboring CUs can include spatial and temporal candidates, as well as additional schedules (or additional candidates), such as those introduced in VVC. The merge mode can be applied not only to skip mode but also to any inter-predicted CU. An alternative to the merge mode is the explicit transmission of motion parameters, in which case the motion vector, the corresponding reference picture index for each reference picture list, the reference picture list usage flag, and / or other necessary information can be explicitly signaled for each CU.
[0106]
[0140] In VVC, the VVC Test Model (VTM) reference software may include many new and refined inter-prediction coding tools, which may include one or more of the following: (1) Extended merge prediction (2) Merge motion vector difference (MMVD) (3) AMVP mode using symmetric MVD signaling (4) Affine motion compensation prediction (5) Subblock-based temporal motion vector prediction (SbTMVP) (6) Adaptive motion vector resolution (AMVR) (7) Motion field storage: 1 / 16 luma sample MV storage and 8x8 motion field compression (8) Bi-prediction with CU-level weights (9) Bi-directional optical flow (BDOF) (10) Decoder-side motion vector refinement (DMVR) (11) Combined inter and intra prediction (CIIP) (12) Geometric partitioning mode (GPM)
[0141] In HEVC, a translational motion model is applied to motion compensation prediction (MCP). In the real world, there may be many types of motions such as zoom in / out, rotation, diagonal motion, and other irregular motions. Block-based affine transform motion compensation prediction may be applied in VTM etc. Figure 9A shows an affine motion field of a block (902) described by motion information of two control points (4 parameters). Figure 9B shows an affine motion field of a block (904) described by three control point motion vectors (6 parameters).
[0107]
[0142] As shown in FIG. 9A, in the 4-parameter affine motion model, the motion vector at the sample position (x, y) within the block (902) can be derived by Equation (1) as follows:
[0108]
Equation
[0109]
Equation
[0143] As shown in FIG. 9B, in the 6-parameter affine motion model, the motion vector at the sample position (x, y) within the block (904) can be derived by Equation (3) as follows:
[0110]
Equation
[0111]
Equation
[0112] (mv 1x , mv 1y ) can be assumed to be the motion vector of the control point at the top-right corner.
[0113] (mv 2x ,mv 2y ) can be the motion vector of the control point at the bottom-left corner.
[0114]
[0144] As shown in FIG. 10, in order to simplify motion compensation prediction, it is possible to apply block-based affine transform prediction. To derive the motion vector for each 4×4 luma sub-block, the motion vector (e.g., (1002)) of the central sample of the sub-block (e.g., (1004)) within the current block (1000) is calculated according to equations (1)-(4) and can be rounded to 1 / 16 fractional precision. Then, in order to generate the prediction for each sub-block using the derived motion vector, it is possible to apply a motion compensation interpolation filter. Also, the sub-block size of the chroma component may also be set to 4×4. The MV of a 4×4 chroma sub-block can be calculated as the average of the MVs of four corresponding 4×4 luma sub-blocks.
[0115]
[0145] In affine merge prediction, for a CU where both the width and height are 8 or more, the affine merge (AF_MERGE) mode can be applied. The CPMVs of the current CU can be generated based on the motion information of spatially neighboring CUs. Up to 5 CPMVP candidates can be applied to affine merge prediction, and an index can be signaled to indicate which of the 5 CPMVP candidates can be used for the current CU. In affine merge prediction, there are three types of CPMV candidates: (1) Inherited affine merge candidates estimated from the CPMVs of neighboring CUs, (2) Affine merge candidates constructed using CPMVPs derived using the translational MVs of neighboring CUs, and (3) Zero MVs It is possible to form an affine merge candidate list using this.
[0116]
[0146] In VTM3, up to two inherited affine candidates can be applied. The two inherited affine candidates can be derived from the affine motion models of neighboring blocks. For example, one inherited affine candidate can be derived from the left neighboring CU, and the other inherited affine candidate can be derived from the above neighboring CU. It is possible to assume that the exemplary candidate blocks are those shown in FIG. 11. As shown in FIG. 11, in the case of the left predictor (or left inherited affine candidate), it is possible to assume that the scan order is A0→A1, and in the case of the above predictor (or above inherited affine candidate), it is possible to assume that the scan order is B0→B1→B2. Therefore, only the first available inherited candidate from each side can be selected. Between the two inherited candidates, a pruning check may not be executed. When a neighboring affine CU is identified, the control point motion vectors of the neighboring affine CU can be used to derive the CPMV candidates in the affine merge list of the current CU. As shown in FIG. 12, when the lower left neighboring block A of the current block (1204) is coded in affine mode, the motion vectors v2, v3, v4 of the top left corner, above right corner, and lower left corner of the CU (1202) including block A can be obtained. When block A is coded in the 4-parameter affine model, the two CPMVs of the current CU (1204) can be calculated according to v2, v3 of CU (1202). When block A is coded in the 6-parameter affine model, the three CPMVs of the current CU (1204) can be calculated according to v2, v3, v4 of CU (1202).
[0117]
[0147] The constructed affine candidates of the current block are candidates constructed by combining the translational motion information in the vicinity of each control point of the current block. It is possible to assume that the motion information of the control points can be derived from specific spatial and temporal vicinities that can be shown in FIG. 13. As shown in FIG. 13, CPMV k (k = 1, 2, 3, 4) represents the k-th control point of the current block (1302). In the case of CPMV1, it is possible to check the B2->B3->A2 blocks and use the MV of the first available block. In the case of CPMV2, it is possible to check the B1->B0 blocks. In the case of CPMV3, it is possible to check the A1->A0 blocks. If CPMV4 is not available, TMVP may be used as CPMV4.
[0118]
[0148] After the MVs of the four control points are obtained, based on the motion information of the four control points, an affine merge candidate for the current block (1302) can be constructed. For example, it is possible to construct an affine merge candidate based on the combination of the MVs of the four control points in the following order: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, and {CPMV1, CPMV3}.
[0119]
[0149] The combination of three CPMVs can construct a six-parameter affine merge candidate, and the combination of two CPMVs can construct a four-parameter affine merge candidate. In order to avoid the motion scaling process, if the reference indices of the control points are different, it is possible to discard the related combinations of the control point MVs.
[0120]
[0150] After the inherited affine merge candidates and the constructed affine merge candidates are checked, if the list is still not full, zero MVs may be inserted at the end of the list.
[0121]
[0151] In affine AMVP prediction, it is possible to apply the affine AMVP mode to a CU whose both width and height are 16 or more. The affine flag at the CU level can be signaled in the bitstream to indicate whether the affine AMVP mode is used, and then another flag can be signaled to indicate whether 4-parameter affine or 6-parameter affine is applied. In affine AMVP prediction, it is possible to signal the predictor of the CPMVPs of the current CU and the difference of the CPMVs of the current CU in the bitstream. It is possible to assume that the size of the affine AMVP candidate list is 2, and the affine AMVP candidate list can be generated by using four types of CPMV candidates in the following order: (1) Inherited affine AMVP candidates estimated from the CPMVs of neighboring CUs, (2) Affine AMVP candidates constructed using the CPMVPs derived using the translational MVs of neighboring CUs, (3) Translational MVs from neighboring CUs, and (4) Zero MVs
[0152] The order of checking the inherited affine AMVP candidates can be the same as the order of checking the inherited affine merge candidates. To determine the AVMP candidates, it is possible to consider only the affine CUs having the same reference picture as the current block. When the inherited affine motion predictor is inserted into the candidate list, the pruning process may not be applied.
[0122]
[0153] The constructed AMVP candidates can be derived from the specified spatial neighborhood. As shown in FIG. 13, it is possible to apply the same checking order as that during the construction of the affine merge candidates. Also, it is possible to check the reference picture indices of neighboring blocks. The first block in the checking order may be inter-coded and have the same reference picture as the current CU (1302). When the current CU (1302) is coded in the 4-parameter affine mode, one constructed AMVP candidate can be determined and both mv0 and mv1 can be used. The constructed AMPV candidates can be further added to the affine AMVP list. When the current CU (1302) is coded in the 6-parameter affine mode and all three CPMVs are available, the constructed AMVP candidates can be added as one candidate in the affine AMVP list. Otherwise, the constructed AMVP candidates can be set as unavailable.
[0123]
[0154] After the inherited affine AMVP candidates and the constructed AMVP candidates are checked, if the number of candidates in the affine AMVP list is still less than 2, mv0, mv1, and mv2 may be added in sequence. mv0, mv1, and mv2, if available, may function as translational MVs for predicting all control point MVs of the current CU (e.g., (1302)). Finally, if the affine AMVP is still not full, zero MVs can be used to fill the affine AMVP list.
[0124]
[0155] Sub-block-based affine motion compensation can save memory access bandwidth and reduce the computational complexity compared to pixel-based motion compensation at the cost of sacrificing prediction accuracy. To achieve a finer granularity of motion compensation, prediction refinement with optical flow (PROF) can be used to refine the sub-block-based affine motion compensation prediction without increasing the memory access bandwidth for motion compensation. In VVC, after sub-block-based affine motion compensation is performed, the luma prediction samples can be refined by adding the differences derived from the optical flow equations. PROF can be described in the following four steps:
[0156] Step (1): Sub-block-based affine motion compensation can be performed to generate a sub-block prediction I(i, j).
[0125]
[0157] Step (2): The spatial gradients g x (i, j) and g y (i, j) can be calculated at each sample position using a 3-tap filter [-1, 0, 1]. It is possible to assume that the gradient calculation is the same as that of BDOF. For example, the spatial gradients g x (i, j) and g y (i, j) can be calculated based on Equations (5) and (6) respectively.
[0126]
Equation
[0127]
[0158] Step (3): The luma prediction refinement can be calculated by the optical flow equation as shown in Equation (7).
[0128]
number
[0129]
[0159] Since the affine model parameters and the sample positions relative to the center of the sub-block may not change between sub-blocks, Δv(i,j) can be calculated for the first sub-block (e.g., (1402)) and reused for other sub-blocks (e.g., (1410)) within the same CU (e.g., (1400)). Let the horizontal offset from the sample position (i,j) to the center of the sub-block (x SB ,y SB ) be dx(i,j) and dy(i,j), then Δv(x,j) can be derived as follows by equations (8) and (9):
[0130]
Equation
[0160] To maintain accuracy, the center of the sub-block (x SB ,y SB ) can be calculated as ((W SB -1) / 2,(H SB -1) / 2) where W SB and H SB are the width and height of the sub-block respectively.
[0131]
[0161] Once Δv(x,y) is obtained, it is possible to obtain the parameters of the affine model. For example, in the case of a 4-parameter affine model, the parameters of the affine model can be expressed as in equation (10):
[0132]
Equation
[0133]
Equation
[0134]
[0162] Step (4): Finally, the luma prediction refinement ΔI(i, j) can be added to the sub-block prediction. The final prediction I' can be generated as shown in Equation (12):
[0135]
Equation
[0163] PROF may not be applicable in the following two cases for affine-coded CUs: (1) the case where all control point MVs are the same, indicating that the CU has only translational motion; and (2) the case where the affine motion parameters are larger than the specified limits, because the sub-block-based affine MC degrades with respect to the CU-based MC in order to avoid large memory access bandwidth requirements.
[0136]
[0164] It should be noted that when coding in the affine AMVP mode, each control point of the affine coding block has a motion vector difference (MVD). For each reference picture, the MVD of the control point is calculated from the actual CPMV value and the CPMV value of the affine AMVP predictor.
[0137]
[0165] In one example, in the case of four-parameter affine, two MVDs (denoted by MVD0 and MVD1) are coded for each reference list according to Equations (13) and (14):
[0138]
Number
[0139]
[0166] In another example, for the six-parameter affine case, three MVDs (denoted by MVD0, MVD1, and MVD2) are coded for each reference list according to Equations (15), (16), and (17):
[0140]
Number
[0141]
[0167] Affine motion estimation (ME) such as the VVC reference software VTM can operate for both single-prediction and dual-prediction. Single-prediction can be executed for one of reference list L0 and reference list L1, and dual-prediction can be executed for both reference list L0 and reference list L1.
[0142]
[0168] Affine motion estimation (ME) such as the VVC reference software VTM can operate for both single-prediction and dual-prediction. Single-prediction can be executed for one of reference list L0 and reference list L1, and dual-prediction can be executed for both reference list L0 and reference list L1.
[0143]
[0169] FIG. 15 shows a schematic diagram of the affine ME (1500). As shown in FIG. 15, in the affine ME (1500), an affine segment - prediction (S1502) is executed on the reference list L0, and based on the initial reference block of the reference list L0, it is possible to obtain the prediction P0 of the current block. An affine segment - prediction (S1504) is also executed on the reference list L1, and based on the initial reference block of the reference list L1, it is possible to obtain the prediction P1 of the current block. In (S1506), it is possible to execute an affine double - prediction. The affine double - prediction (S1506) can start with the initial prediction residual (2I - P0) - P1, where I can be assumed to be the initial value of the current block. The affine double - prediction (S1506) can search for candidates within the reference list L1 around the initial reference block in the reference list L1 and find the best (or selected) reference block having the minimum prediction residual (2I - P0) - Px, where Px is the prediction of the current block based on the selected reference block.
[0144]
[0170] Using the reference picture, for the current coding block, the affine ME process can first be selected based on a set of control point motion vectors (CPMVs). Using an iterative method, it is possible to generate the prediction output of the current affine model corresponding to the set of CPMVs, calculate the gradient of the prediction samples, and then solve a linear equation to determine the delta CPMVs for optimizing the affine prediction. The iteration can stop when all the delta CPMVs become 0, or when the maximum number of iterations is reached. The CPMVs obtained from the iteration may be the final CPMVs for the reference picture.
[0145]
[0171] After the best affine CPVMs for both reference lists L0 and L1 are determined with respect to affine segment - prediction, an affine dual - prediction search can be performed using the reference list on one side and the best segment - prediction CPMVs, and the best CPMVs for the other reference list can be searched to optimize the affine dual - prediction output. The affine dual - prediction search can be repeatedly performed for the two reference lists in order to obtain optimal results.
[0146]
[0172] FIG. 16 shows an exemplary affine ME process (1600) capable of calculating the final CPMVs related to a reference picture. The affine ME process (1600) can start from (S1602). In (S1602), it is possible to determine the base CPMVs of the current block. The base CPMVs can be determined based on any of a merge index, an advanced motion vector prediction (AMVP) predictor index, an affine merge index, etc.
[0147]
[0173] In (S1604), the initial affine prediction of the current block can be obtained based on the base CPMVs. For example, according to the base CPMVs, it is possible to apply a 4 - parameter affine motion model among the 6 - parameter affine motion models to generate the initial affine prediction.
[0148]
[0174] In (S1606), it is possible to obtain the gradient of the initial affine prediction. For example, the gradient of the initial affine prediction can be obtained based on equations (5) and (6).
[0149]
[0175] In (S1608), it is possible to determine ΔCPMVs. In some embodiments, ΔCPMVs can be associated with the displacement between an initial affinity prediction and a subsequent affinity prediction such as a first affinity prediction. Based on the initial affinity prediction and the gradient of the delta CPMVs, the first affinity prediction can be obtained. The first affinity prediction can correspond to the first CPMVs.
[0150]
[0176] In (S1610), it is possible to determine whether the delta CPMVs are zero or whether the number of iterations is greater than or equal to a threshold. If the delta CPMVs are zero or the number of iterations is greater than or equal to the threshold, then in (S1612), the final (or selected) CPMVs can be determined. The final (or selected) CPMVs can be the first CPMVs determined based on the initial affinity prediction and the gradient of the delta CPMVs.
[0151]
[0177] Continuing to refer to (S1610), if the delta CPMVs are not zero or the number of iterations is less than the threshold, it is possible to start a new iteration. In the new iteration, the updated CPMVs (e.g., the first CPMVs) can be provided to (S1604) to generate an updated affinity prediction. Then, the affine ME process (1600) can proceed to (S1606) to calculate the gradient of the updated affinity prediction. Then, the affine ME process (1600) can proceed to (S1608) to continue the new iteration.
[0152]
[0178] In the affine motion model, the four-parameter affine motion model can be further described by an equation including rotational and zoom motions. For example, the four-parameter affine motion model can be rewritten in Equation (18) as follows:
[0153]
Equation
[0154]
Number
[0155]
Number
[0179] The bidirectional optical flow (BDOF) in VVC was previously called BIO in JEM. Compared with the JEM version, the BDOF in VVC is a simple version that requires fewer operations, especially with respect to the number of multiplications and the size of the multiplier.
[0156]
[0180] BDOF can be used to refine the bi-predicted signals of the CU at the 4×4 sub-block level. BDOF can be applied to the CU when the CU satisfies the following conditions: (1) When the CU is coded using the "true" bi-prediction mode, i.e., when one of the two reference pictures is before the current picture in the display order and the other is after the current picture in the display order, (2) If the distance from the two reference pictures to the current picture (e.g., POC difference) is the same, (3) If both reference pictures are short-term reference pictures, (4) If the CU is not coded using affine mode or SbTMVP merge mode, (5) If the CU has more than 64 luma samples, (6) If both the CU height and the CU width are 8 luma samples or more, (7) If the BCW weight index indicates equal weights, (8) If the weighted position (WP) is not enabled for the current CU, and (9) If the CIIP mode is not used for the current CU.
[0157]
[0181] BDOF may only be applied to the luma component. As the name BDOF mode suggests, the BDOF mode can be based on the concept of optical flow that assumes smooth motion of objects. For each 4x4 sub-block, the motion refinement (v x , v y ) can be calculated by minimizing the difference between the L0 and L1 prediction samples. Then, the motion refinement can be used to adjust the bi-prediction sample values within the 4x4 sub-block. BDOF can include the following steps:
[0182] First, the horizontal and vertical gradients
[0158]
Equation
[0159]
Equation
[0160]
[0183] Next, the autocorrelation and cross-correlation S1, S2, S3, S5, S6 of the gradient can be calculated according to the following equations (23)-(27):
[0161]
Equation
[0162]
Equation
[0163]
[0184] The motion refinement (v x , v y ) can then be derived using cross-correlation and autocorrelation as follows using equations (31) and (32):
[0164]
Equation
[0165]
Num
[0166]
Num
[0167]
Num
[0185] Finally, the CU's BDOF samples can be calculated as follows by adjusting the dual-prediction samples of Equation (34):
[0168]
Num
[0169]
[0186] To derive the gradient value, several prediction samples I in list k (k = 0, 1) outside the current CU boundary (k)(i,j) needs to be generated. As shown in FIG. 17, for BDOF in VVC, one extended row / column (1702) can be used around the boundary (1706) of the CU (1704). To suppress the computational complexity of generating prediction samples outside the boundary, the prediction samples in the extended area (e.g., the non-shaded area in FIG. 17) can be directly generated by obtaining the reference samples at the neighboring integer positions (e.g., using the floor() operation for coordinates) without interpolation, and the prediction samples can be generated within the CU (e.g., the shaded area in FIG. 17) using the normal 8-tap motion compensation interpolation filter. The extended sample values can be used only for gradient calculation. For the remaining steps of the BDOF process, when samples and gradient values outside the CU boundary are required, those samples and gradient values can be padded (e.g., repeated) from the nearest neighbors of the samples and gradient values.
[0170]
[0187] If the width and / or height of the CU is greater than 16 luma samples, the CU can be divided into sub-blocks having a width and / or height equal to 16 luma samples, and the sub-block boundaries can be treated as the CU boundaries in the BDOF process. The maximum unit size of the BDOF process can be limited to 16x16. For each sub-block, the BDOF process can be skipped. If the sum of absolute differences (SAD) between the initial L0 and L1 prediction samples is less than the threshold, the BDOF process may not be applied to that sub-block. The threshold can be set equal to (8*W*(H>>1)), where W may indicate the width of the sub-block and H may indicate the height of the sub-block. To avoid the additional complexity of the SAD operation, the SAD between the initial L0 and L1 prediction samples calculated in the DMVR process can be reused in the BDOF process.
[0171]
[0188] When BCW is enabled for the current block, i.e., when the BCW weight index indicates unequal weighting, the bidirectional optical flow may be disabled. Similarly, when WP is enabled for the current block, i.e., when the luma weight flag (e.g., luma_weight_lx_flag) is 1 for either of the two reference pictures, BDOF may also be disabled. When the CU is coded in the symmetric MVD mode or the CIIP mode, BDOF may also be disabled.
[0172]
[0189] To improve the accuracy of the merge mode MV, decoder-side motion vector refinement based on bilateral matching (BM) can be applied as in VVC. In the bi-prediction operation, refined MVs can be searched around the initial MVs of the reference picture list L0 and the reference picture list L1. The BM method can calculate the distortion between two candidate blocks in the reference picture lists L0 and L1.
[0173]
[0190] Figure 18 shows an exemplary schematic diagram of BM-based decoder-side motion vector refinement. As shown in Figure 18, the current picture (1802) can include a current block (1808). The current picture can include a reference picture list L0 (1804) and a reference picture list L1 (1806). The current block (1808) can include an initial reference block (1812) of the reference picture list L0 (1804) corresponding to the initial motion vector MV0 and an initial reference block (1814) of the reference picture list L1 (1806) corresponding to the initial motion vector MV1. A search process can be executed around the initial MV0 of the reference picture list L0 (1804) and the initial MV1 of the reference picture list L1 (1806). For example, a first reference candidate block (1810) may be identified in the reference picture list L0 (1804), and a first reference candidate block (1816) may be identified in the reference picture list L1 (1806). The SAD between candidate reference blocks (e.g., (1810) and (1816)) based on each MV candidate (e.g., MV0' and MV1') around the initial MV (e.g., MV0 and MV1) can be calculated. The MV candidate having the lowest SAD becomes the refined MV and can be used to generate a bi-prediction signal to predict the current block (1808).
[0174]
[0191] The application of DMVR may be restricted and may be applied only to CUs coded based on modes and characteristics as in VVC as follows: (1) When it is a CU-level merge mode using a bi-prediction MV, (2) When there is one reference picture in the past and another reference picture in the future for the current picture, (3) When the distances (e.g., POC differences) from two reference pictures to the current picture are the same, (4) When both reference pictures are short-term reference pictures, (5) When the CU has more than 64 luma samples, (6) When both the CU height and the CU width are 8 luma samples or more, (7) When the BCW weight index indicates equal weights, (8) When WP is not enabled for the current CU, and (9) When the CIIP mode is not used for the current CU.
[0175]
[0192] The refined MV derived by the DMVR process can be used to generate inter-prediction samples and may be used in temporal motion vector prediction for future picture coding. The original MV, on the other hand, can be used in the deblocking process and may be used in spatial motion vector prediction for future CU coding.
[0176]
[0193] In DVMR, the search point can surround the initial MV, and the MV offset can follow the MV differential mirroring rule. In other words, any point checked by DMVR, as indicated by the candidate MV pair (MV0, MV1), can follow the MV differential mirroring rule shown in Equations (35) and (36):
[0177]
Equation
[0178]
[0194] For example, for integer sample offset search, a 25-point full search may be applied. First, the SAD of the initial MV pair can be calculated. If the SAD of the initial MV pair is smaller than the threshold, the integer sample stage of DMVR can be terminated. Otherwise, the SAD of the remaining 24 points can be calculated and checked in a scanning order such as raster scan order. The point with the minimum SAD can be selected as the output of the integer sample offset search stage. To reduce the penalty of the uncertainty of DMVR refinement, the original MV during the DMVR process may have a selected priority. The SAD between the reference blocks referred to by the initial MV candidate may be reduced by 1 / 4 of the SAD value.
[0179]
[0195] After integer sample search, fractional sample refinement may follow. To save the complexity of calculation, instead of additional search by SAD comparison, fractional sample refinement can be derived by using a parametric error surface equation. Fractional sample refinement may be conditionally activated based on the output of the integer sample search stage. If the integer sample search stage ends at the center while having the minimum SAD in either the first iterative search or the second iterative search, fractional sample refinement may be further applied.
[0180]
[0196] In parametric error surface-based sub-pixel offset estimation, the cost at the central position and the costs at four positions near the center can be used to fit a 2-D parabolic error surface equation based on Equation (37):
[0181]
Equation
[0182]
Equation
[0183]
[0197] Similar to VVC, it is possible to apply bilinear interpolation and sample padding. The resolution of the MV can be, for example, 1 / 16 luma samples. It is possible to interpolate samples at fractional positions using an 8-tap interpolation filter. In DMVR, since the search points can surround the initial fractional pel MV with integer sample offsets, samples at fractional positions need to be interpolated in the case of the DMVR search process. To reduce the computational complexity, it is possible to use a bilinear interpolation filter to generate fractional samples for the search process in DMVR. As another important effect, by using a bilinear filter with a 2-sample search range, DVMR does not access more reference samples compared to the normal motion compensation process. After a refined MV is obtained in the DMVR search process, a normal 8-tap interpolation filter can be applied to generate the final prediction. Samples that may not be required for the interpolation process based on the original MV to avoid accessing more reference samples compared to the normal MC process, but may be required for the interpolation process based on the refined MV, can be padded from the available samples.
[0184]
[0198] If the width and / or height of the CU is greater than 16 luma samples, the CU can further be divided into sub-blocks having a width and / or height equal to 16 luma samples. The maximum unit size of the DMVR search process can be limited to 16x16.
[0185]
[0199] In an embodiment, like in VVC, a merge with motion vector difference (MMVD) mode is used, in which case samples of a CU (e.g., the current CU) can be predicted using implicitly derived motion information. The MMVD mode is used for either the skip mode or the merge mode using a motion vector representation method. The MMVD merge flag can be signaled to specify whether the MMVD mode is used for a CU, for example, after signaling a skip flag or a merge flag.
[0186]
[0200] In some examples, MMVD reuses merge candidates. The candidates can be selected from among the merge candidates and are further extended by the motion vector representation method. MMVD provides a motion vector representation with simplified signaling. In some examples, the motion vector representation method includes a starting point, a magnitude of motion, and a direction of motion.
[0187]
[0201] In some examples (e.g., VVC), the MMVD technique can use a merge candidate list to select a candidate for the starting point. However, in one example, the only candidate (MRG_TYPE_DEFAULT_N), which is the default merge type, is considered for the extension of MMVD.
[0188]
[0202] In some examples, a basic candidate index is used to define the starting point. The basic candidate index indicates the best candidate among the candidates in the list shown in Table 1. For example, the list is a merge candidate list having motion vector predictors (MVPs). The basic candidate index can indicate the best candidate within the merge candidate list.
[0189] Table 1 - Example of basic candidate index (IDX)
[0190]
Table 1
[0203] Note that in one example, the number of basic candidates is equal to 1, in which case the basic candidate IDX is not signaled.
[0191]
[0204] In the MMVD mode, after a merge candidate (also called an MV base or an MV starting point) is selected, the merge candidate can be refined by additional information such as the signaled MVD information. The additional information can include an index (e.g., a distance index, e.g., mmvd_distance_idx[x0][y0]) used to specify the magnitude of the motion and an index (e.g., a direction index, e.g., mmvd_direction_idx[x0][y0]) used to indicate the direction of the motion. In the MMVD mode, one of the first two candidates in the merge list can be selected as the MV base. For example, a merge candidate flag (e.g., mmvd_cand_flag[x0][y0]) indicates one of the first two candidates in the merge list. The merge candidate flag can be signaled to indicate (e.g., specify) which of the first two candidates is selected. The additional information can indicate the MVD (or motion offset) with respect to the MV base. For example, the magnitude of the motion indicates the magnitude of the MVD, and the direction of the motion indicates the direction of the MVD.
[0192]
[0205] In one example, a merge candidate selected from the merge candidate list is used to provide a starting point or an MV starting point in the reference picture. The motion vector of the current block can be represented by the starting point and a motion offset (or MVD) including the magnitude and direction of the motion with respect to the starting point. On the encoder side, the selection of the merge candidate and the determination of the motion offset can be performed based on a search process (evaluation process) as shown in FIG. 19. On the decoder side, the selected merge candidate and the motion offset can be determined based on the signaling from the encoder side.
[0193]
[0206] FIG. 19 shows an example of a search process (1900) in the MMVD mode. FIG. 20 shows an example of search points in the MMVD mode. In some examples, a subset or the entire set of the search points in FIG. 20 is used in the search process (1900) of FIG. 19. For example, by executing the search process (1900) on the encoder side, merge candidate flags (e.g., mmvd_cand_flag[x0][y0]), distance indexes (e.g., mmvd_distance_idx[x0][y0]), direction indexes (e.g., mmvd_direction_idx[x0][y0]) and other additional information can be determined for the current block (1901) in the current picture (or current frame).
[0194]
[0207] The first motion vector (1911) and the second motion vector (1921) belonging to the first merge candidate are shown. The first motion vector (1911) and the second motion vector (1921) are the starting points of the MVs used in the search process (1900). The first merge candidate may be a merge candidate in the merge candidate list constructed for the current block (1901). The first and second motion vectors (1911) and (1921) can be associated with two reference pictures (1902) and (1903) in the reference picture lists L0 and L1, respectively. Referring to FIGS. 19-20, the first and second motion vectors (1911) and (1921) can point to two starting points (2011) and (2021) in the reference pictures (1902) and (1903), respectively, as shown in FIG. 20.
[0195]
[0208] Referring to FIG. 20, the two starting points (2011) and (2021) in FIG. 20 can be determined in the reference pictures (1902) and (1903). In one example, based on the starting points (2011) and (2021), a plurality of predetermined points extending in the vertical direction (+Y or -Y) or the horizontal direction (+X, -X) in the reference pictures (1902) and (1903) can be evaluated starting from the starting points (2011) and (2021). In one example, a pair of points (2014) and (2024) (e.g., as indicated by a shift of 1S in FIG. 19), or a pair of points (2015) and (2025) (e.g., as indicated by a shift of 2S in FIG. 19), i.e., a pair of points that are mirror images of each other with respect to each starting point (2011) or (2021), is used to determine a pair of motion vectors (e.g., MV(1913) and (1923) in FIG. 19) that can form a motion vector predictor candidate for the current block (1901). Also, a motion vector predictor candidate (e.g., MV(1913) and (1923) in FIG. 19) determined based on predetermined points around the starting point (2011) or (2021) can be evaluated.
[0196]
[0209] The distance index (e.g., mmvd_distance_idx[x0][y0]) specifies information about the magnitude of motion and can also indicate a predetermined offset (e.g., 1S or 2S in FIG. 19) from the starting point indicated by the merge candidate flag. Note that the predetermined offset is also referred to as the MMVD step in one example.
[0197]
[0210] Referring to FIG. 19, the offset (e.g., MVD(1912) or MVD(1922)) can be applied (e.g., added) to the horizontal or vertical component of the starting MV (e.g., MV(1911) or (1921)). An exemplary relationship between the distance index (IDX) and the predefined offset is specified in Table 2.
[0198] When full-pel MMVD (full-pel MMVD) is off, for example, if the full-pel MMVD flag (e.g., slice_fpel_mmvd_enabled_flag) is equal to 0, it is possible to assume that the range of the predefined offset of MMVD is from 1 / 4 luma sample to 32 luma samples. When full-pel MMVD is off, the predefined offset can have a non-integer value such as a fraction of a luma sample (e.g., 1 / 4 pixel or 1 / 2 pixel).
[0199] When full-pel MMVD is on, for example, if the full-pel MMVD flag (e.g., slice_fpel_mmvd_enabled_flag) is equal to 1, it is possible to assume that the range of the predefined offset of MMVD is from 1 luma sample to 128 luma samples. In one example, when full-pel MMVD is on, the predefined offset only has an integer value such as one or more luma samples.
[0200] Table 2 - Exemplary relationship between distance index and offset (e.g., a predetermined offset)
[0201]
Table 2
[0211] The direction index can represent the direction (or movement direction) of the MVD with respect to the starting point. In one example, the direction index represents one of the four directions shown in Table 3. The meaning of the MVD sign in Table 3 may vary depending on the information of the starting MV. In one example, when the starting MV is a single-predicted MV, or when the starting MV is a bi-predicted MV and both reference lists point to the same side of the current picture (for example, when the POCs of both reference pictures are greater than the POC of the current picture, or when the POCs of both reference pictures are less than the POC of the current picture), the MVD sign in Table 3 indicates the sign or sign of the MV offset (or MVD) added to the starting MV.
[0202]
[0212] When the starting MV is a bi-predicted MV and the two MVs point to different sides of the current picture (for example, when the POC of one reference picture is greater than the POC of the current picture and the POC of the other reference picture is less than the POC of the current picture), the MVD sign in Table 3 indicates the sign of the MV offset (or MVD) added to the list0 MV component of the starting MV, and the MVD sign of the list1 MV has the opposite value. Referring to Figure 19, the starting MVs (1911) and (1921) are bi-predicted MVs, and the two MVs (1911) and (1921) point to different sides of the current picture. The POC of the L1 reference picture (1903) is greater than the POC of the current picture, and the POC of the L0 reference picture (1902) is less than the POC of the current picture. The MVD sign (for example, the sign "+" with respect to the x-axis) indicated by the direction index (for example, 00) in Table 3 specifies the sign (for example, the sign "+" with respect to the x-axis) of the MVD (for example, MVD(1912)) added to the list0 MV component of the starting MV (for example, (1911)), and the MVD sign of MVD(1922) with respect to the list1 MV component of the starting MV (for example, (1921)) has the opposite value, for example, the sign "-" opposite to the sign "+" of MVD(1912).
[0203]
[0213] Referring to Table 3, direction index 00 indicates the positive direction on the x-axis, direction index 01 indicates the negative direction on the x-axis, direction index 10 indicates the positive direction on the y-axis, direction index 11 indicates the negative direction on the y-axis.
[0204] Exemplary relationship between the sign of the table 3 - MV offset and the direction index
[0205]
Table 3
[0214] The syntax element mmvd_merge_flag[x0][y0] can be used to represent the MMVD merge flag of the current CU. In one example, an MMVD merge flag equal to 1 (e.g., mmvd_merge_flag[x0][y0]) indicates that the MMVD mode is used to generate the inter-prediction parameters of the current CU. An MMVD merge flag equal to 0 (e.g., mmvd_merge_flag[x0][y0]) indicates that the MMVD mode is not used to generate the inter-prediction parameters. The array indices x0 and y0 can specify the position (x0, y0) of the top-left luma sample of the coding block (e.g., the current CB) with respect to the top-left luma sample of the picture (e.g., the current picture).
[0206]
[0215] If the MMVD merge flag (e.g., mmvd_merge_flag[x0][y0]) does not exist for the current CU, the MMVD merge flag (e.g., mmvd_merge_flag[x0][y0]) can be assumed to be equal to 0 for the current CU.
[0207]
[0216] In some examples, such as the VVC specification, a single context is used to signal the MMVD merge flag (e.g., mmvd_merge_flag). For example, a single context is used to code (e.g., encode and / or decode) the MMVD merge flag in context-adaptive binary arithmetic coding (CABAC).
[0208]
[0217] The syntax element mmvd_cand_flag[x0][y0] can represent a merge candidate flag. In one example, the merge candidate flag (e.g., mmvd_cand_flag[x0][y0]) specifies whether the first (0) or second (1) candidate in the merge candidate list is used with the MVD derived from the distance index (e.g., mmvd_distance_idx[x0][y0]) and the direction index (e.g., mmvd_direction_idx[x0][y0]). The array indices x0 and y0 can specify the position (x0, y0) of the top-left luma sample of the coding block (e.g., current CB) being considered relative to the top-left luma sample of the picture (e.g., current picture).
[0209]
[0218] If the merge candidate flag (e.g., mmvd_cand_flag[x0][y0]) does not exist, it is possible to assume that the merge candidate flag (e.g., mmvd_cand_flag[x0][y0]) is equal to 0.
[0210]
[0219] The syntax element mmvd_distance_idx[x0][y0] can represent a distance index. In one example, the distance index (e.g., mmvd_distance_idx[x0][y0]) specifies the index used to derive MmvdDistance[x0][y0] as specified in Table 4. The array indices x0 and y0 can specify the position (x0, y0) of the top-left luma sample of the coding block (e.g., current CB) with respect to the top-left luma sample of the picture (e.g., current picture).
[0211] Table 4 - Exemplary relationship between MmvdDistance[x0][y0] and mmvd_distance_idx[x0][y0]
[0212]
Table 4
[0220] The first column of Table 4 shows the distance index (e.g., mmvd_distance_idx[x0][y0]).
[0213] The second column of Table 4 shows the magnitude of motion (e.g., MmvdDistance[x0][y0]) when full - pel MMVD is off, e.g., when the full - pel MMVD flag (e.g., slice_fpel_mmvd_enabled_flag) is equal to 0.
[0214] The third column of Table 4 shows the magnitude of motion (e.g., MmvdDistance[x0][y0]) when full - pel MMVD is on, e.g., when the full - pel MMVD flag (e.g., slice_fpel_mmvd_enabled_flag) is equal to 1.
[0215]
[0221] In one example, the units of the second and third columns in Table 4 are 1 / 4 luma samples. Referring to the first column of Table 4, when the distance index (e.g., mmvd_distance_idx[x0][y0]) is 0, the motion magnitude (e.g., MmvdDistance[x0][y0]) is 1 if full pel MMVD is off (e.g., slice_fpel_mmvd_enabled_flag is 0). The motion magnitude (e.g., MmvdDistance[x0][y0]) is 1×1 / 4 luma samples or 1 / 4 luma samples.
[0216] When the distance index (e.g., mmvd_distance_idx[x0][y0]) is 0, the motion magnitude (e.g., MmvdDistance[x0][y0]) is 4 if full pel MMVD is on (e.g., slice_fpel_mmvd_enabled_flag is 1). The motion magnitude (e.g., MmvdDistance[x0][y0]) is 4×1 / 4 luma samples or 1 luma sample.
[0217]
[0222] In one example, the second column (1 / 4 luma sample unit) in Table 4 corresponds to the second row (luma sample unit) in Table 1, and the third column (1 / 4 luma sample unit) in Table 4 corresponds to the third row (luma sample unit) in Table 2.
[0218]
[0223] The syntax element mmvd_direction_idx[x0][y0] can represent a direction index. In one example, the direction index (e.g., mmvd_direction_idx[x0][y0]) indicates an index used to derive the motion direction (e.g., MmvdSign[x0][y0]) as shown in Table 5. The array indices x0 and y0 indicate the position (x0, y0) of the top-left luma sample of the coding block under consideration (e.g., the current CB) relative to the top-left luma sample of the picture (e.g., the current picture). The first column in Table 5 indicates the direction index (e.g., mmvd_distance_idx[x0][y0]). The second column of Table 5 indicates the first sign (e.g., MmvdSign[x0][y0][0]) of the first component of the MVD (e.g., MVD x or MmvdOffset[x0][y0][0]). The third column of Table 5 indicates the second sign (e.g., MmvdSign[x0][y0][1]) of the second component of the MVD (e.g., MVD y or MmvdOffset[x0][y0][1]).
[0219] Table 5 - Exemplary relationship between MmvdSign[x0][y0] and mmvd_direction_idx[x0][y0]
[0220]
Table 5
[0224] The first component (e.g., MmvdOffset[x0][y0][0]) and the second component (e.g., MmvdOffset[x0][y0][1]) of the MVD or the offset MmvdOffset[x0][y0] can be derived as follows:
[0221]
Number
[0225] In one example, the distance index (e.g., mmvd_distance_idx[x0][y0]) is 3 and the direction index (e.g., mmvd_distance_idx[x0][y0]) is 2. Based on Table 5 and the direction index (e.g., mmvd_direction_idx[x0][y0]) being 2, the first sign (e.g., MmvdSign[x0][y0][0]) of the first component of the MVD (e.g., MVD x orMmvdOffset[x0][y0][0]) is "0", and the second sign (e.g., MmvdSign[x0][y0][1]) of the second component of the MVD (e.g., MVD y orMmvdOffset[x0][y0][1]) is "+1". In this example, the MVD is along the positive vertical direction (+y) and has no horizontal component.
[0222]
[0226] When the full - pel MMVD flag (e.g., slice_fpel_mmvd_enabled_flag) is equal to 0 and the full - pel MMVD is off, based on Table 4 and the distance index (e.g., mmvd_distance_idx[x0][y0]) being 3, the magnitude of the motion indicated by MmvdDistance[x0][y0] is 8. Based on Equations 10 - 11, the first component of the MVD (e.g., MmvdOffset[x0][y0][0]) is (8 << 2)×0 = 0 and the second component of the MVD (e.g., MmvdOffset[x0][y0][1]) is (8 << 2)×(+1)=2 (luma samples) and this is the case.
[0223]
[0227] When the full - pel MMVD flag (e.g., slice_fpel_mmvd_enabled_flag) is equal to 1 and the full - pel MMVD is on, based on the fact that Table 4 and the distance index (e.g., mmvd_distance_idx[x0][y0]) is 3, the magnitude of the motion indicated by MmvdDistance[x0][y0] is 32. Based on Equation (40) and Equation (41), the first component of the MVD (e.g., MmvdOffset[x0][y0][0]) is (32 << 2)×0 = 0 and the second component of the MVD (e.g., MmvdOffset[x0][y0][1]) is (32 << 2)×(+1)=8 (luma samples) and becomes
[0224]
[0228] According to one aspect of the present disclosure, affine merge using motion vector difference (affine MMVD) can be used for video coding. Affine MMVD selects available affine merge candidates as base predictors from a sub - block - based merge list. Affine MMVD applies a motion vector offset to the motion vector value of each control point from the base predictor. In one example, if no affine merge candidates are available, affine MMVD will not be used. In some examples, the distance index and the offset direction index can be signaled thereafter.
[0225]
[0229] In some examples, a distance index is signaled to indicate which distance offset to use from an offset table as shown in Table 6: Table 6: Example of offset table
[0226]
Table 6
[0230] In some examples, the direction index can represent four directions as shown in Table 7, where only the x or y direction may have an MV difference and not both directions.
[0227] Table 7: Example of a direction table
[0228]
Table 7
[0231] In some examples, the inter prediction is a single - prediction, and to generate a result that includes the MV value of each control point, a signaled distance offset is applied in the offset direction of each control point predictor.
[0229]
[0232] In some examples, the inter prediction is a dual - prediction, and the signaled distance offset can be applied in the signaled offset direction with respect to the L0 motion vector of the control point predictor, and the offset applied to the L1 MV can be applied to a mirrored or scaled base as in the following specific examples.
[0230]
[0233] In a specific example, the inter prediction is a dual - prediction, and the signaled distance offset is applied in the signaled distance offset direction with respect to the L0 motion vector of the control point predictor. In the case of L1 CPMV, the offset is applied to a mirrored base, which means that the same amount of distance offset is applied in the opposite direction.
[0231]
[0234] In another specific example, an offset mirroring method based on the POC distance is used for dual-prediction. When the base candidate is dual-predicted, the offset applied to L0 is as signaled, and the offset for L1 depends on the temporal positions of the reference pictures in list L0 and list L1. For example, if both reference pictures are on the same side of the temporal current picture, the same distance offset and the same offset direction are applied to the CPMVs of both L0 and L1. In another example, if the two reference pictures are on different sides of the current picture, the CPMVs of L1 can have a distance offset applied in the opposite offset direction.
[0232]
[0235] In another specific example, an offset scaling method based on the POC distance is used for dual-prediction. When the base candidate is dual-predicted, the offset applied to L0 is as signaled, and the offset for L1 can be scaled based on the temporal distances of the reference pictures in list 0 and list 1.
[0233]
[0236] In some examples, the range of the distance offset values is extended. For example, it is possible to provide three sets of distance offset values, and one set of distance offset values can be adaptively selected based on the picture resolution. In one example, an offset table is selected based on the picture resolution. Table 8 shows an example of an extended distance offset table that includes three sets of distance offset values associated with different picture resolutions. The set of distance offset values can be selected based on the picture resolution.
[0234] Table 8: Example of an extended distance-offset table
[0235]
Table 8
[0237] Template matching (TM) technology can be used in video / image coding. To further improve the compression efficiency of the VVC standard, for example, TM can be used to refine the motion vector (MV). In one example, TM is used on the decoder side. In TM mode, a template (e.g., the current template) of a block (e.g., the current block) within the current picture is constructed, and the MV can be refined by determining the closest matching between the template of the block within the current picture and multiple possible templates (e.g., multiple possible reference templates) within the reference picture. In an embodiment, the template of a block within the current picture can include the reconstructed samples near the left side of the block and the reconstructed samples near the upper side of the block. TM can be used in video / image coding beyond VVC.
[0236]
[0238] Figure 21 shows an example of template matching (2100). TM can be used to derive the motion information of the current coding unit (CU) (e.g., the current block) (2101) by determining the closest match between a template (e.g., the current template) (2121) of the current CU (2101) within the current picture (2111) and a template of multiple possible templates (e.g., a reference template) (e.g., one of the multiple possible templates is template (2125)) within the reference picture (2110) (e.g., deriving the final motion information from the initial motion information such as the initial MV 2102). The template (2121) of the current CU (2101) can have any suitable shape and any suitable size.
[0237]
[0239] In an embodiment, the template (2121) of current CU (2101) includes a top template (2122) and a left template (2123). Each of the top template (2122) and the left template (2123) can have any suitable shape and any suitable size.
[0238]
[0240] The top template (2122) can include samples within one or more top-neighboring blocks of current CU (2101). In one example, the top template (2122) includes four sample rows within one or more top-neighboring blocks of current CU (2101).
[0239] The left template (2123) can include samples within one or more left-neighboring blocks of current CU (2101). In one example, the left template (2123) includes four sample columns within one or more left-neighboring blocks of current CU (2101).
[0240]
[0241] Each of a plurality of possible templates (e.g., template (2125)) within the reference picture (2111) corresponds to a template (2121) within the current picture (2110). In an embodiment, the initial MV (2102) points from the current CU (2101) to a reference block (2103) within the reference picture (2111). Each of a plurality of template candidates (e.g., template (2125)) within the reference picture (2111) and the template (2121) within the current picture (2110) can have the same shape and the same size. For example, the template (2125) of the reference block (2103) includes a top template (2126) of the reference picture (2111) and a left template (2127) of the reference picture (2111). The top template (2126) can include samples within one or more top-neighboring blocks of the reference block (2103). The left template (2127) can include samples within one or more left-neighboring blocks of the reference block (2103).
[0241]
[0242] The TM cost can be determined based on a pair of templates such as a template (e.g., the current template) (2121) and a template (e.g., the reference template) (2125). The TM cost can indicate the matching between the template (2121) and the template (2125). An optimized MV (or the final MV) can be determined based on the search around the initial MV (2102) of the current CU (2101) within the search range (2115). The search range (2115) can have any suitable shape and any suitable number of reference samples. In one example, the search range (2115) within the reference picture (2111) includes a range of [-L,L]-pel, where L is a positive integer such as 8 (e.g., 8 samples). For example, based on the search range (2115), a difference (e.g., [0,1]) is determined, and an intermediate MV is determined by the sum of the initial MV (2102) and the difference (e.g., [0,1]). Based on the intermediate MV, an intermediate reference block and the corresponding template within the reference picture (2111) can be determined. The TM cost can be determined based on the template (2121) and the intermediate template of the reference picture (2111). The TM cost can correspond to the difference (e.g., [0,0], [0,1], etc. corresponding to the initial MV (2102)) determined based on the search range (2115). In one example, the difference corresponding to the minimum TM cost is selected, and the optimized MV is the sum of the difference corresponding to the minimum TM cost and the initial MV (2102). As described above, TM can derive the final motion information (e.g., the optimized MV) from the initial motion information (e.g., the initial MV 2102).
[0242]
[0243] In the example of FIG. 21, within a search range such as [-8pel, +8pel], a better MV can be searched around the initial motion vector of the current CU.
[0243]
[0244] TM can be applied in affine modes such as the affine AMVP mode and the affine merge mode, and can be called affine TM. FIG. 22 shows an example of TM (2200) in the affine merge mode, for example. The template (2221) of the current block (for example, the current CU) (2201) may correspond to the template (for example, the template (2121) in FIG. 21) in the TM applied to the translational motion model. The reference template (2225) of the reference block in the reference picture can be a plurality of sub-block templates (for example, 4×4 sub-blocks), including those indicated by the control point MV (CPMV)-derived MVs of the neighboring sub-blocks (for example, A0-A3 and L0-L3 as shown in FIG. 22) at the block boundary.
[0244]
[0245] The search process for TM applied in the affine mode (for example, the affine merge mode) can start from CPMV0 while keeping other CPMVs (for example, (i) CPMV1 if the 4-parameter model is used, or (ii) CPMV1 and CPMV2 if the 6-parameter model is used) constant. The search can be performed in the horizontal and vertical directions. In one example, the search continues diagonally later only if the zero vector is not the best difference vector found from the horizontal and vertical searches. Affine TM can repeat the same search process for CPMV1. Affine TM can repeat the same search process for CPMV2 when the 6-parameter model is used. If the zero vector is not the best difference vector from the previous iteration and the search process has been repeated less than 3 times, the entire search process can resume from the refined CPMV0 based on the refined CPMVs.
[0245]
[0246] According to one aspect of the disclosure, it is possible to reduce signaling overhead by using a candidate reordering technique based on template matching. For example, a technique referred to as adaptive reordering of merge candidates with template matching (ARMC-TM) can be used.
[0246]
[0247] In some examples, using ARMC-TM, merge candidates are adaptively reordered using template matching (TM). ARMC-TM can be applied to regular merge mode, template matching (TM) merge mode, and affine merge mode (excluding SbTMVP candidates). In the case of TM merge mode, merge candidates are reordered before the refinement process.
[0247]
[0248] In some examples, after a merge candidate list is constructed using ARMC-TM, the merge candidates are divided into several subgroups. In one example, the subgroup size is set to 5 in regular merge mode. In another example, the subgroup size is set to 3 for affine merge mode. The merge candidates within each subgroup are reordered in ascending order according to a cost value based on template matching. For simplicity, in some examples, the merge candidates of the latest subgroup, rather than the first subgroup, are not reordered.
[0248]
[0249] The template matching cost of the merge candidates is measured by the sum of absolute differences (SAD) between the samples of the template of the current block and the reference samples corresponding to the template (referred to as the reference template in one example). The template includes a set of reconstructed samples near the current block. The reference samples of the template are arranged according to the motion information of the merge candidates.
[0249]
[0250] When the merge candidate uses bi-directional prediction, the reference samples of the template of the merge candidate are also generated by bi-directional prediction.
[0250]
[0251] FIG. 23 shows an illustration of the reference samples of the template of the current block for a bi-predicted merge candidate. In FIG. 23, the current picture (2310) includes the current block for coding. When the merge candidate is a bi-predicted merge candidate, the MV of the merge candidate may point to a first reference block in the reference picture (2320) and a second reference block in the second reference picture (2330). The template of the current block is indicated by (T), and the template includes a set of reconstructed samples adjacent to the current block. The first set of reference samples of the template is in the first reference picture (2320) adjacent to the first reference block, and the second set of reference samples of the template is in the second reference picture (2330) adjacent to the second reference block. In one example, the template matching cost of the bi-predicted merge candidate is calculated by adding the sum of absolute differences (SAD) between the samples of the template of the current block and the first set of reference samples of the template and the sum of the second absolute differences (SAD) between the samples of the template of the current block and the second set of reference samples of the template.
[0251]
[0252] In some examples, the merge candidate may be a sub-block based merge candidate. In one example, for a sub-block based merge candidate having a sub-block size equal to Wsub×Hsub, the top template may include several sub-templates having a size of Wsub×1, and the left template may include several sub-templates having a size of 1×Hsub. Wsub is the width of the sub-block and Hsub is the height of the sub-block.
[0252]
[0253] It is possible to assume that the exemplary derivation of the template for the current block having sub-block-based merge candidates and the reference sample of the template is as shown in FIG. 24. As shown in FIG. 24, the current block (2402) can be included in the current picture (2404). The current block (2402) can include sub-blocks A-G in the first row and the first column. The current block (2402) can include a template (2406) adjacent to the top side and the left side of the current block (2402). The collocated block (2408) for the current block (2402) is within the reference picture (2410). The collocated block (2408) can include sub-blocks A-G in the first row and the first column, which correspond to the sub-blocks A-G within the current block (2402). The sub-block motion information of sub-blocks A-G in the first row and the first column of the current block (2402) (for example, corresponding to the affine motion vectors) can be used to derive the reference sample of the sub-template (or sub-reference template) of the collocated block (2408).
[0253] For example, the motion information of sub-blocks A, E, F, and G of the current block (2402) can be applied to derive the reference sample of the sub-template arranged adjacent to the left side of sub-blocks A, E, F, and G of the collocated block (2408). The sub-template adjacent to the left side of sub-blocks A, E, F, G of the collocated block (2408) can form the left reference template of the collocated block (2408).
[0254] The motion information of sub - blocks A, B, C, and D of the current block (2402) can be applied to derive reference samples of the sub - template arranged adjacent to the upper side of sub - blocks A, B, C, and D of the equivalent - position block (2408). The sub - template adjacent to the top side of sub - blocks A, B, C, and D of the equivalent - position block (2408) can further form the top - side reference template of the equivalent - position block (2408).
[0255]
[0254] In some examples, it is possible to use ARMC based on the MV candidate type. For example, a merge candidate of a certain single - candidate type, such as TMVP or non - adjacent MVP (NA - MVP), is sorted based on the ARMC TM cost value. Then, the sorted candidates are added to the merge candidate list. For example, the TMVP candidate - type ARMC can add more TMVP candidates with more temporal positions and different inter - prediction directions to perform sorting and selection. Further, the NA - MVP candidate - type ARMC expands non - adjacent MVPs that are more spatially non - adjacent. The target reference picture of the TMVP candidate can be selected from any one of the reference pictures in the list according to the scaling factor. For example, the selected reference picture is the one whose scaling factor is closest to 1.
[0256]
[0255] According to one aspect of the disclosure, candidate sorting based on template matching can be performed for MMVD and affine MMVD.
[0257]
[0256] In some examples, the MMVD offset is extended to more positions for the MMVD mode and the affine MMVD mode.
[0258]
[0257] FIG. 25 is a diagram showing the directions in which refinement positions can be added with respect to the MMVD. In FIG. 25, additional refinement positions along the diagonal (diagonal direction angle) of k×π / 8 are added, where k is an integer. Position (2501) corresponds to the base candidate and can be the starting point, and positions (2511)-(2514) are in the directions of 0, π / 2, π, and 3π / 2, respectively. More directions may be added. For example, positions (2521)-(2524) are in the directions of π / 4, 3π / 4, 5π / 4, and 7π / 4, respectively; positions (2531)-(2538) are in the directions of π / 8, 3π / 8, 5π / 8, 7π / 8, 9π / 8, 11π / 8, 13π / 8, and 15π / 8, respectively. Therefore, the number of directions is increased from 4 to 16. Further, in one example, each direction may have six MMVD refinement positions. The total number of possible MMVD refinement positions is 16×6.
[0259]
[0258] According to one aspect of the present disclosure, the SAD cost between the current template (e.g., the top 1 row and the left 1 column with respect to the current block) and the reference template can be calculated for each refinement position. Based on the SAD cost of the refinement positions, all possible MMVD refinement positions (16×6) for each base candidate are sorted. Then, the upper part of the refinement positions, for example, the upper 1 / 8 of the refinement positions (e.g., 12), for example, those with the minimum template SAD cost, are retained as the positions available for MMVD index coding. The MMVD index is binarized by a rice code having a parameter equal to 2.
[0260]
[0259] In some examples, it is possible to increase the refinement positions for the affine MMVD, and it is possible to apply the candidate rearrangement based on template matching to the rearrangement of the affine MMVD. For example, the affine MMVD refinement positions are in the directions along the diagonal angles of k×π / 4, such as the 8 directions of 0, π / 4, π / 2, 3π / 4, π, 5π / 4, 3π / 2, and 7π / 4 respectively. Each direction may have 6 affine MMVD refinement positions. The total number of possible affine MMVD refinement positions is 8×6. In one example, it is possible to calculate the SAD cost between the current template (e.g., the top 1 row and the left 1 column for the current block) and the reference template for each refinement position. Based on the SAD cost of the refinement positions, all possible affine MMVD refinement positions (8×6) for each base candidate are rearranged. Then, the upper part of the refinement positions, for example, the upper 1 / 2 of the refinement positions (e.g., 24), for example, those with the minimum template SAD cost, are retained as the positions available for affine MMVD index coding as a result.
[0261]
[0260] In some examples, the MVD sign prediction technique is used. In one example, the possible combinations of MVD signs (e.g., various combinations of signs in the x and y directions) are sorted according to the template matching cost of the possible combinations of MVD signs, an index corresponding to the true combination of MVD signs is derived, and context coding is performed. According to the MVD sign prediction technique, the true combination of MVD signs has a high probability in the leading part of the sorted order. Therefore, an appropriate signaling technique can be used to signal the index with a low signaling cost.
[0262]
[0261] In one example, on the decoder side, the true MVD sign can be derived. For example, it is possible to analyze the magnitude of the MVD component, and the context-coded MVD sign prediction index is analyzed from the bitstream that conveys the video. Further, the MV candidates can be formed by creating combinations from among possible combinations of MVD signs, and the magnitude of the MVD component and the MV candidates can be added to the MV predictor list. It is possible to calculate the template matching cost for the MV candidates in the MV predictor list. The MV candidates in the MV predictor list can be sorted according to the template matching cost. Then, the context-coded MVD sign prediction index is used to select the combination of true MVD signs from the MV predictor list. The MVD sign prediction technique can be applied to various modes including MVD, such as the inter AMVP, affine AMVP, MMVD, and affine MMVD modes.
[0263]
[0262] In some examples, it is possible to use a technique called history-parameter-based affine model inheritance. Specifically, in some examples, a first history-parameter table (HPT) and a second HPT are set.
[0264]
[0263] FIG. 26 shows a diagram (2600) showing the first HPT and the second HPT in some examples.
[0265]
[0264] As shown in FIG. 26, the entry of the first HPT stores a set of affine parameters of an affine model such as a, b, c, d, and each affine parameter is represented by a 16-bit signed integer. The entries in the first HPT are classified by a reference list (e.g., reference picture list L0 or reference picture list L1) and a reference index. Five reference indexes are supported for each reference list in the first HPT. As an example, expressed in a mathematical formula, the category of the first HPT (denoted as HPTCat) is calculated as shown in Equation (42): HPTCat(RefList,RefIdx)=5×RefList+min(RefIdx,4) Eq.(42) Here, RefList represents the reference picture list (0 or 1), and RefIdx represents the reference index.
[0266]
[0265] For each category, it is possible to store a maximum of seven entries. As a result, there are a total of 70 entries in the first HPT. At the beginning of each CTU row, the number of entries in each category is initialized to 0. After decoding the CU coded affine using the reference list RefList cur and RefIdx cur , the affine parameters are used to update the entry of the category HPTCat(RefList cur ,RefIdx cur ) in the same way as the update of the HMVP table.
[0267]
[0266] In one example, a history-affine-parameter-based candidate (HAPC) is derived from one of seven neighboring 4×4 blocks shown as A0, A1, A2, B0, B1, B2, or B3 in FIG. 26, and a set of affine parameters stored in a corresponding entry within a first HPT. The MV of the neighboring 4×4 block is provided as a base MV. In the formulation, the MV of the current block at position (x, y) is calculated as in Equation (43):
[0268]
Number
[0269]
[0267] A second history parameter table (HPT) with base MV information is also added. The second HPT can include nine entries, and the entries can include a base MV, a reference index and four affine parameters for each reference list, and a base position. In one example, an additional merge HAPC can be generated from the base MV information and the corresponding affine model (e.g., affine parameters) stored in the entries of the second HPT.
[0270]
[0268] Further, in some examples, the pairwise affine merge candidates are generated by two affine merge candidates that are either history-derived or not history-derived. In one example, the pairwise affine merge candidates are generated by averaging the CPMVs of existing affine merge candidates in a candidate list.
[0271]
[0269] In some examples, in response to the introduction of a new HAPC, the size of the sub-block-based merge candidate list is increased from 5 to 15, all of which may be involved in the ARMC process.
[0272]
[0270] In the above description, the HPTs (e.g., the first HPT and the second HPT) are updated online. In addition to the HPT updated by one line, the HPT stored in the upper / upper-right CTU of the current CTU can, in some examples, be used by blocks within the current CTU. After coding / decoding the CTU, the HPT may be stored in a line buffer for use in the next CTU row.
[0273]
[0271] FIG. 27 is a diagram showing a history parameter table stored in a line buffer in some examples. In FIG. 27, a picture (2700) is partitioned into CTUs. FIG. 27 shows CTU row k and CTU row k + 1. The current CTU (2710) is within CTU row k + 1. For coding the current block (2711), the HPT (2701) stored in the CTU above the current CTU (2710) and the HPT (2702) stored in the upper-right CTU of the current CTU (2710) can, in some examples, be used by blocks within the current CTU. The HPTs (e.g., the first HPT and the second HPT) are updated online and stored in the line buffer of the current CTU (2710) after decoding the last coding block within the current CTU.
[0274]
[0272] According to one aspect of the present disclosure, non-adjacent spatial neighborhoods can be used for affine mode.
[0275]
[0273] In an affine mode with non-adjacent spatial neighbors (NA-AFF), it is possible to obtain non-adjacent spatial neighborhoods.
[0276]
[0274] FIGS. 28A-28B show patterns for obtaining non-adjacent spatial neighborhoods in some examples. Similar to existing non-adjacent regular merge candidates, the distance between the current CU and non-adjacent spatial neighbors in NA-AFF is also defined based on the width and height of the current CU.
[0277]
[0275] Using the motion information of non-adjacent spatial neighborhoods, additional inherited and constructed affine merge / AMVP candidates are generated. FIG. 28A shows generating additional inherited affine merge / AMVP candidates, and FIG. 28B shows generating additional constructed affine merge / AMVP candidates.
[0278]
[0276] Specifically, as shown in FIG. 28A, for inherited candidates, the same derivation process of inherited affine merge / AMVP candidates in VVC is maintained without change except that CPMVs are inherited from non-adjacent spatial neighborhoods. Non-adjacent spatial neighborhoods are checked based on their distances to the current block, i.e., from the nearest to the farthest. At a specific distance, only the first available neighborhoods (coded in affine mode) from each side (e.g., left and upper sides) of the current block are included for inherited candidate derivation. As indicated by the dashed arrows in FIG. 28A, the checking order of neighborhoods on the left and upper sides is from bottom to top and from right to left, respectively.
[0279]
[0277] For the first type of constructed candidate, as shown in FIG. 28B, the positions of one left and upper non-adjacent spatial neighborhood are first determined independently. Thereafter, the position of the top-left neighborhood can then be determined accordingly, which can surround a rectangular virtual block using the left and upper non-adjacent neighborhoods.
[0280]
[0278] Next, as shown in FIG. 29, the motion information of three non-adjacent neighborhoods is used to form CPMVs at the top-left (A), top-right (B), and bottom-left (C) of the virtual block, and then the motion information is projected onto the current CU to generate the corresponding constructed candidate.
[0281]
[0279] In some examples, for the second type of constructed candidate, the derivation process is similar to the construction scheme in history-based affine model inheritance (HAMI). However, instead of using a history-based look-up table, non-translational affine parameters are inherited from non-adjacent spatial neighborhoods. Specifically, the second type of affine construction candidate is 1) the translational affine parameters of adjacent neighborhood 4x4 blocks; and 2) the non-translational affine parameters inherited from non-adjacent spatial neighborhoods as defined in FIG. 28A; generated from a combination of.
[0282]
[0280] In some examples, the NA-AFF candidates are inserted into the existing affine merge candidate list and affine AMVP candidate list according to a specific order.
[0283]
[0281] In one example, in the affine merge mode, the order includes the following: 1. SbTMVP candidate (if available); 2. Inherited from adjacent neighborhood; 3. Inherited from non - adjacent neighborhood; 4. Constructed from adjacent neighborhood; 5. Second type of constructed affine candidate from non - adjacent neighborhood; 6. First type of constructed affine candidate from non - adjacent neighborhood; 7. Zero MV.
[0284]
[0282] In another example, in the affine AMVP mode, the order includes the following: 1. Inherited from adjacent neighborhood; 2. Constructed from adjacent neighborhood; 3. Translational MV from adjacent neighborhood; 4. Translational MV from temporal neighborhood; 5. Inherited from non - adjacent neighborhood; 6. First type of constructed affine candidate from non - adjacent neighborhood; 7. Zero MV.
[0285]
[0283] Due to including additional candidates generated by NA - AFF, the size of the affine merge candidate list is increased from 5 to 15. The subgroup size of ARMC for the affine merge mode is increased from 3 to 15.
[0286]
[0284] In some examples (e.g., ECM - 5.0 software), NA - AFF is implemented without adding constraints on memory usage.
[0287]
[0285] Some video coding standards, such as HEVC and AVC / H.264, use a fixed motion vector resolution of 1 / 4 luma samples. According to one aspect of the present disclosure, an optimal trade-off between the displacement vector rate and the prediction error rate can be selected to achieve overall rate-distortion optimality. Some video coding standards, such as VVC, allow selecting the motion vector resolution at the coding block level and thus allow a trade-off between fidelity and bitrate with respect to the signaling of motion parameters. In particular, VVC uses an adaptive motion vector resolution (AMVR) when the AMVR mode is enabled. The AMVR mode is signaled at the coding block level when at least one component of the MVD (e.g., the x component and / or the y component) is not equal to zero. The motion vector predictor is rounded to a given resolution, so that the resulting motion vector can be dropped onto a grid of the given resolution.
[0288] According to one aspect of the present disclosure, in some scenarios of affine-coded blocks, all CPVMs are the same for the reference picture, which makes the affine model for the reference picture equivalent to a translational motion model. In one example, in affine dual-prediction, the CPMVs for the first reference picture are the same, and the CPMVs for the second reference picture are different. In another example, the motion vector predictor for the current block is in the affine mode and has multiple control points with different CPMVs, but when the CPMVs of the motion vector predictor are combined with the MVD, the resulting MV of the control points of the current block may have the same MV value. In some related affine MVD coding techniques, all control points (two or three depending on the 4-parameter or 6-parameter affine type) are MVD-coded. In each MVD coding, the horizontal and vertical components are coded separately. The related affine MVD coding techniques may have a high signaling overhead compared to translational MVD coding.
[0289] Some aspects of the present disclosure provide techniques for combining affine motion and translational motion into an affine merge candidate called an affine-translational merge candidate. In some examples, the affine-translational merge candidate is a dual-prediction candidate that provides a prediction according to two reference pictures (e.g., the first reference picture from reference picture list L0 and the second reference picture from reference picture list L1), and one motion information of affine mono-prediction for one reference list (e.g., LX, where X is 0 or 1) and one motion information of translational mono-prediction for the other reference list (e.g., L(1 - X)) are combined into the affine-translational merge candidate. In some examples, in the case of an affine-translational merge candidate, the reference picture that provides a prediction according to the affine motion information is referred to as an affine reference picture, and the reference picture that provides a prediction according to the translational motion information is referred to as a translational reference picture. In some examples, for an affine - translational merge candidate, the reference list that provides a prediction according to the affine motion information is referred to as the affine reference list, and the reference list that provides a prediction according to the translational motion information is referred to as the translational reference list. In one example, for a reference picture using a translational motion vector (MV), it is possible to use an affine model in which the control point motion vectors (CPMVs) are set to the same MV value.
[0290]
[0288] According to one aspect of the present disclosure, regarding an affine - translational merge candidate, the derivation of the affine motion information portion can be from existing affine merge candidates such as existing single - prediction affine merge candidates, existing double - prediction affine merge candidates, or a combination of existing single - prediction affine merge candidates and existing double - prediction affine merge candidates. In one embodiment, the affine motion information of an affine - translational merge candidate is derived from one or more existing affine merge candidates by single - prediction using the available affine motion information of the reference list corresponding to the affine reference list of the affine - translational merge candidate. In another embodiment, the affine motion information of an affine - translational merge candidate is derived from existing affine merge candidates by double - prediction, and the affine motion information of the existing merge candidates for the reference list corresponding to the affine reference list of the affine - translational merge candidate is used in the derivation.
[0291]
[0289] According to one aspect of the present disclosure, regarding an affine - translational merge candidate, the derivation of the translational motion information portion can be from various sources and can also be from a combination of multiple sources.
[0292]
[0290] In one embodiment, the translational motion information of an affine - translational merge candidate is derived from existing translational merge candidates using single - prediction with the available translational motion information of the reference list corresponding to the translational reference list.
[0293]
[0291] In another embodiment, the translational motion information of the affine-translational merge candidate is derived from an existing translational merge candidate using bi-prediction, and the translational motion information of the existing translational merge candidate for the reference list corresponding to the translational reference list of the affine-translational merge candidate is used for the derivation.
[0294]
[0292] In another embodiment, the translational motion information of the affine-translational merge candidate is derived from the current translational merge candidates list based on a predefined rule that can be based on, for example, an index value, index parity, index position, etc. In one example, the derivation depends on the parity of the merge index. For example, in the case of an even merge index, the translational motion information of the candidate from the reference list L0 (or L1) is used for the derivation, and in the case of an odd merge index, the translational motion information of the candidate from the reference list L1 (or L0) can be used for the derivation. If there is no candidate from the reference list determined according to the predefined rule, it is possible to derive the translational motion information using another reference list.
[0295]
[0293] In another embodiment, the translational motion information of the affine-translational merge candidate is derived from an existing affine merge candidate. In one example, the motion vector of a predetermined control point, such as the first control point motion vector (CPMV) of the existing affine merge candidate, can be used as the translational motion information of the affine-translational merge candidate.
[0296]
[0294] In another embodiment, the translational motion information of the affine-translational merge candidate is derived from an existing affine merge candidate. In one example, the average of a plurality of control point motion vectors (CPMVs) can be used as the translational motion information of the affine-translational merge candidate.
[0297]
[0295] In another embodiment, the translational motion information of the affine-translational merge candidate is derived from an existing affine merge candidate. As an example, the MV of the central position of the current block (e.g., the position of (W / 2, H / 2), where W and H are the block width and block height in the luma samples of the current block) derived from the affine model of the affine merge candidate can be used as the translational motion information of the affine-translational merge candidate.
[0298]
[0296] According to one aspect of the present disclosure, the affine-translational merge candidate can be added to an existing sub-block merge candidate list (e.g., a sub-block merge list such as in VVC). In one example, the sub-block merge candidate list includes an affine merge candidate and an SbTMVP candidate. In another example, the sub-block merge candidate list includes an affine merge candidate but does not include an SbTMVP candidate. The affine transform merge candidate can be added to the sub-block based merge candidate list at various positions. In one example, the affine-translational merge candidate is added to the sub-block based merge candidate list after the inherited affine merge candidate. In one example, only the existing inherited affine merge candidate can be used as the source of the affine operation to derive the affine operation of the affine-translational merge candidate.
[0299]
[0297] In another example, the affine-translational merge candidate is added to the sub-block based merge candidate list after the constructed affine merge candidate. In one example, only the existing affine merge candidate before the affine-translational merge candidate can be used as the source of the affine operation to derive the affine operation of the affine-translational merge candidate.
[0300]
[0298] In another example, the affine-translation merge candidates are added to the sub-block-based merge candidate list after all the affine merge candidates except the zero MV candidates. In one example, in order to derive the affine operation of the affine-translation merge candidate, only the existing affine merge candidates before the affine-translation merge candidate can be used as the source of the affine operation.
[0301]
[0299] According to one aspect of the present disclosure, the affine-translation merge candidates are grouped into a new merge list called the affine-translation merge candidate list, and a flag indicated by affine_trans_merge_flag is signaled to specify whether the affine-translation merge candidate list is selected.
[0302]
[0300] In one example, when it is signaled that the current mode is the sub-block-based merge mode, the flag affine_trans_merge_flag is further signaled.
[0303] When the flag affine_trans_merge_flag is true, an index is signaled to indicate which affine-translation merge candidate is used from the affine-translation merge candidate list.
[0304] When the flag affine_trans_merge_flag is false, the regular sub-block-based merge candidate list is used, and candidates are appropriately selected from the regular sub-block-based merge candidate list.
[0305]
[0301] In one example, in response to the template matching flag in the high-level syntax such as SPS, PPS, and slice header, the affine-translation merge candidate list is sorted based on the template matching cost of the affine-translation merge candidates.
[0306]
[0302] In another example, an affine-translation merge candidate list is derived whether the template matching-based rearrangement of the affine-translation merge candidate list is enabled or disabled. In another example, the affine-translation merge candidate list is derived when the template matching-based rearrangement of the affine-translation merge candidate list is disabled.
[0307]
[0303] In some embodiments, there can be at most N affine-translation merge candidates added to a candidate list such as a regular sub-block-based merge candidate list, an affine-translation merge candidate list, etc. In one example, N is a predetermined value (e.g., in one example, N is set equal to 2). In another example, N is signaled by a high-level syntax such as sequence level, picture level, slice level, etc.
[0308]
[0304] In some embodiments, it is possible to derive affine-translation merge candidates using only the first M existing affine merge candidates. In one example, M is a predetermined value. In another example, M is signaled by a high-level syntax such as sequence level, picture level, slice level, etc.
[0309]
[0305] In some embodiments, it is possible to derive affine-translation merge candidates using only the first L existing translation merge candidates. In one example, L is a predetermined value. In another example, L is signaled by a high-level syntax such as sequence level, picture level, slice level, etc.
[0310]
[0306] In one embodiment, the affine-translation merge candidates are used only when bi-prediction is allowed in the current picture.
[0311]
[0307] In another embodiment, the affine-translation merge candidates are used only for pictures with all reference pictures that are temporally prior to the current picture, as in the low-delay configuration in VVC.
[0312]
[0308] In another embodiment, whether to apply the technique for processing the affine-translation merge candidates is controlled by a high-level syntax such as at the sequence level (e.g., by the SPS), picture level (e.g., by the picture header or PPS), slice level (e.g., by the slice header), tile / tile group level, etc.
[0313]
[0309] In some embodiments, when the local illumination compensation (LIC) tool is enabled but prohibited for bi-prediction, the affine-translation merge candidates can have an LIC flag associated with the affine-translation merge candidates set to false.
[0314]
[0310] According to one aspect of the present disclosure, when the bi-prediction with CU level weighting (BCW) tool using CU level weighting is enabled, the affine-translation merge candidates can be processed accordingly.
[0315]
[0311] In one embodiment, when the affine-translation merge candidates are included, the BCW index of the current block may be set to a default value indicating equal weighting with respect to reference lists L0 and L1. Thus, equal weighting is applied to the prediction based on the affine motion and the prediction based on the translational motion.
[0316]
[0312] In another embodiment, if only one of the source of the affine motion information and the source of the translational information of the affine-translational merge candidate exhibits a BCW index indicating unequal weighting, the corresponding unequal weighting is applied to the affine motion-based prediction and the translational motion-based prediction by using the BCW index for the affine-translational merge candidate.
[0317]
[0313] In another embodiment, if both the source of the affine motion information and the source of the translational information of the affine-translational merge candidate have a BCW index indicating unequal weighting, the BCW index of the source of the affine motion information is used for the affine-translational merge candidate in order to apply the corresponding unequal weighting to the affine motion-based prediction and the translational motion-based prediction.
[0318]
[0314] FIG. 30 shows a flowchart illustrating an overview of a process (3000) according to an embodiment of the present disclosure. The process (3000) can be used in a video encoder. In various embodiments, the process (3000) is executed by a processing circuit such as a processing circuit in terminal devices (310), (320), (330), (340), a processing circuit that executes the functions of video encoder (403), a processing circuit that executes the functions of video encoder (603), and a processing circuit that executes the functions of video encoder (703). In some embodiments, the process (3000) is executed by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (3000). The process starts at (S3001) and proceeds to (S3010).
[0319]
[0315] In (S3010), a first affine-translational merge candidate is selected from a candidate list for prediction of a current block in a current picture. The first affine-translational merge candidate provides affine motion information related to a first reference picture in a first reference list and translational motion information related to a second reference picture in a second reference list.
[0320]
[0316] In (S3020), the first index indicating the first affine-translational merge candidate in the candidate list is encoded in a bitstream that conveys a video including the current picture, the first reference picture, and the second reference picture.
[0321]
[0317] In some examples, the affine motion information and the translational motion information of the first affine-translational merge candidate are derived separately. The first affine-translational merge candidate is inserted into the candidate list.
[0322]
[0318] In one example, the affine motion information is derived according to the first affine motion information of the first affine merge candidate by uni-prediction. In another example, the affine motion information is derived according to the second affine motion information of the second affine merge candidate by bi-prediction, and the second affine motion information is associated with the first reference list.
[0323]
[0319] In one example, the translational motion information is derived according to the first translational motion information of the first translational merge candidate by uni-prediction. In another example, the translational motion information is derived according to the second translational motion information of the second translational merge candidate by bi-prediction, and the second translational motion information is associated with the second reference list.
[0324]
[0320] In one example, for a merge candidate having a merge index in the translational merge candidate list to derive the translational motion information, from the first reference list and the second reference list, according to a predetermined rule, based on the merge index, the first list is selected. The translational motion information is derived based on the first translational motion information of the merge candidate associated with the first list if the first translational motion information exists, and is derived based on the second translational motion information of the merge candidate associated with a list different from the first list if the first translational motion information does not exist.
[0325]
[0321] In some examples, the translational motion information can be derived from the affine merge candidates. In one example, the translational motion information is derived based on the control point motion vectors of the affine merge candidates. In other examples, the translational motion information is derived based on the average of multiple control point motion vectors of the affine merge candidates. Also, in other examples, the translational motion information is derived based on the motion vector at the center of the current block, and the motion vector is determined according to the affine model of the affine merge candidate.
[0326]
[0322] In one example, the first affine-translational merge candidate is inserted into the existing sub-block-based merge candidate list after the latest inherited affine merge candidate. In another example, the first affine-translational merge candidate is inserted into the existing sub-block-based merge candidate list after the latest constructed affine merge candidate. In another example, the first affine-translational merge candidate is inserted into the existing sub-block-based merge candidate list after the latest affine merge candidate and before the zero motion vector candidate.
[0327]
[0323] In some examples, the first affine-translational merge candidate is inserted into the affine-translational merge list separately from the sub-block-based merge candidate list. In one example, a flag indicating the selection of the affine-translational merge list is encoded in the bitstream.
[0328]
[0324] In some examples, the candidates in the affine-translational merge list are sorted according to the template matching cost of the candidates.
[0329]
[0325] In some examples, in response to the count of the affine-translational merge candidates in the candidate list being lower than a certain value, the first affine-translational merge candidate is inserted into the candidate list. The value can be determined in advance or can be signaled at a high level from the encoder side to the decoder side.
[0330]
[0326] In some examples, the affine motion information is a subset of the affine merge candidates, derived from a subset that is before other affine merge candidates in the affine merge candidate list, and the count of the affine merge candidates in the subset is equal to a value. The value can be determined in advance or can be signaled at a high level from the encoder side to the decoder side.
[0331]
[0327] In some examples, the translational motion information is a subset of the translational merge candidates, derived from a subset that is before other translational merge candidates in the translational merge candidate list, and the count of the translational merge candidates in the subset is equal to a value. The value can be determined in advance or can be signaled at a high level from the encoder side to the decoder side.
[0332]
[0328] In some examples, a high-level syntax indicating the acceptance of affine-translational merge candidates for prediction of the current block is encoded in the bitstream.
[0333]
[0329] Then, the process proceeds to (S3099) and ends.
[0334]
[0330] The process (3000) can be suitably adapted. The steps in the process (3000) can be modified and / or omitted. Additional steps can be added. Any suitable execution order can be used.
[0335]
[0331] Figure 30 shows a flowchart illustrating an overview of a process (3100) according to an embodiment of the present disclosure. The process (3100) can be used in a video decoder. In various embodiments, the process (3100) is executed by a processing circuit such as a processing circuit in a terminal device (310), (320), (330), (340), a processing circuit that executes the functions of a video decoder (410), and a processing circuit that executes the functions of a video decoder (510). In some embodiments, the process (3100) is executed by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (3100). The process starts at (S3101) and proceeds to (S3110).
[0336]
[0332] In (S3110), a first affine-translation merge candidate for prediction of a current block in a current picture is determined from a candidate list. In one example, a coded video bitstream including information associated with the current block in the current picture is received. The first affine-translation merge candidate provides affine motion information associated with a first reference picture in a first reference list and translational motion information associated with a second reference picture in a second reference list.
[0337]
[0333] In (S3120), a first prediction for samples in the current block is generated according to the affine motion information associated with the first reference picture.
[0338]
[0334] In (S3130), a second prediction for samples in the current block is generated according to the translational motion information associated with the second reference picture.
[0339]
[0335] In (S3140), samples of the current block are reconstructed according to a combination of the first prediction and the second prediction.
[0340]
[0336] In one example, the affine motion information and the translational motion information of the first affine-translational merge candidate are derived, and the first affine-translational merge candidate is inserted into the candidate list.
[0341]
[0337] In one example, the affine motion information is derived according to the first affine motion information of the first affine merge candidate using single-prediction. In another example, the affine motion information is derived according to the second affine motion information of the second affine merge candidate using double-prediction, and the second affine motion information is associated with the first reference list.
[0342]
[0338] In one example, the translational motion information is derived according to the first translational motion information of the first translational merge candidate using single-prediction. In another example, the translational motion information is derived according to the second translational motion information of the second translational merge candidate using double-prediction, and the second translational motion information is associated with the second reference list.
[0343]
[0339] In one example, to derive the translational motion information, for the merge candidate with a merge index in the translational merge candidate list, the first list is selected from the first reference list and the second reference list based on the merge index according to a pre-set rule. Depending on the presence of the first translational motion information, the translational motion information is derived based on the first translational motion information of the merge candidate associated with the first list; also, depending on the absence of the first translational motion information, the translational motion information is derived based on the second translational motion information of the merge candidate associated with another list different from the first list.
[0344]
[0340] In some examples, the translational motion information can be derived from the affine merge candidates. In one example, the translational motion information is derived based on the control point motion vectors of the affine merge candidates. In another example, the translational motion information is derived based on the average of multiple control point motion vectors of the affine merge candidates. In another example, the translational motion information is derived based on the motion vector at the center of the current block, and the motion vector is determined according to the affine model related to the affine merge candidate.
[0345]
[0341] In one example, the first affine-translational merge candidate is inserted into the existing sub-block-based merge candidate list after the latest inherited affine merge candidate. In another example, the first affine-translational merge candidate is inserted into the existing sub-block-based merge candidate list after the latest configured affine merge candidate. In another example, the first affine-translational merge candidate is inserted into the existing sub-block-based merge candidate list after the latest affine merge candidate and before the zero motion vector candidate.
[0346]
[0342] In one example, the first affine-translational merge candidate is inserted into the affine-translational merge list separately from the existing sub-block-based merge candidate list. In one example, in response to the decoded flag (decoded from the coded video bitstream) indicating the selection of the affine-translational merge list, the first affine-translational merge candidate is selected from the affine-translational merge list according to the decoded index. In response to the decoded flag indicating the selection of the sub-block-based merge candidate list, the merge candidate is selected from the sub-block-based merge candidate list according to the decoded index.
[0347]
[0343] In some examples, the candidates in the affine-translational merge list are reordered according to the template matching cost of the candidates.
[0348]
[0344] In some examples, in response to the count of affine-translation merge candidates in the candidate list being lower than a certain value, a first affine-translation merge candidate is inserted into the candidate list. That value can be determined in advance or can be signaled at a high level from the encoder side to the decoder side.
[0349]
[0345] In some examples, the affine motion information is a subset of the affine merge candidates, derived from a subset that is before other affine merge candidates in the affine merge candidate list, and the count of the affine merge candidates in the subset is equal to a certain value. That value can be determined in advance or can be signaled at a high level from the encoder side to the decoder side.
[0350]
[0346] In some examples, the translational motion information is a subset of the translational merge candidates, derived from a subset that is before other translational merge candidates in the translational merge candidate list, and the count of the translational merge candidates in the subset is equal to a certain value. That value can be determined in advance or can be signaled at a high level from the encoder side to the decoder side.
[0351]
[0347] In some examples, a determination is made that bi-prediction is allowed in the current picture, and then, in response to that determination, a first affine-translation merge candidate for predicting the current block is determined.
[0352]
[0348] In some examples, a determination is made that the reference picture of the current picture is temporally prior to the current picture, and in response to that determination, a first affine-translation merge candidate for predicting the current block is determined.
[0353]
[0349] In some examples, a high-level syntax indicating acceptance of a first affine-translation merge candidate for current block prediction is determined (e.g., decoded from a bitstream).
[0354]
[0350] In some examples, in response to local illumination compensation (LIC) being enabled but prohibited for bi-prediction, the LIC flag associated with the first affine-translation merge candidate is set to false.
[0355]
[0351] In one example, for combining the first prediction and the second prediction with equal weights, the bi-prediction (BCW) index by coding unit level weighting is set to a default value.
[0356]
[0352] In another example, only one of the affine motion information and the translational motion information has a bi-prediction (BCW) index by coding unit level weighting. The first prediction and the second prediction are combined according to the BCW index.
[0357]
[0353] In another example, the first BCW index is associated with the affine motion information, and the second BCW index is associated with the translational motion information. The first prediction and the second prediction are combined according to the first BCW index.
[0358]
[0354] The process then proceeds to (S3199) and ends.
[0359]
[0355] The process (3100) can be suitably adapted. The steps in the process (3100) can be modified and / or omitted. Additional steps can be added. Any suitable execution order can be used.
[0360]
[0356] The above-described technology can be implemented as computer software using computer-readable instructions and can be physically stored in one or more computer-readable media. For example, FIG. 32 shows a computer system (3200) suitable for implementing a particular embodiment of the disclosed subject matter.
[0361]
[0357] The computer software can be coded using any suitable machine code or computer language that can be the subject of assembly, compilation, linking, or similar mechanisms to create code that includes instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or instructions that pass through interpretation or microcode execution.
[0362]
[0358] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet-of-Things devices, etc.
[0363]
[0359] The components shown in FIG. 32 for the computer system (3200) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software for implementing embodiments of the present disclosure. Also, the configuration of the components should not be construed as having any dependency or requirement with respect to any one or combination of the components shown in the exemplary embodiment of the computer system (3200).
[0364]
[0360] A computer system (3200) can include a specific human interface input device. Such a human interface input device can respond to input by one or more human users via, for example, tactile input (e.g., keystrokes, swipes, movements of a data glove), auditory input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input (not shown). Also, the human interface device can be used to capture specific media such as audio (e.g., conversations, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic pictures) that are not necessarily directly related to conscious input by humans.
[0365]
[0361] The input human interface device can potentially include one or more of a keyboard (3201), a mouse (3202), a trackpad (3203), a touch screen (3210), a data glove (not shown), a joystick (3205), a microphone (3206), a scanner (3207), and a camera (3208) (although only one of each is depicted).
[0366]
[0362] The computer system (3200) may also include a specific human interface output device. Such a human interface output device can, for example, stimulate the senses of one or more human users through tactile output, sound, light, and smell / taste. Such a human interface output device can be a tactile output device (e.g., a touch screen (3210), a data glove (not shown), tactile feedback by a joystick (3205), although there may be a tactile feedback device that does not serve as an input device), an auditory output device (e.g., a speaker (3209), headphones (not shown)), a visual output device (e.g., a screen (3210) including a CRT screen, an LCD screen, a plasma screen, an OLED screen, each of which may or may not have a touch screen input function, each of which may or may not have a tactile feedback function, and some of them may be capable of outputting three-dimensional or higher-dimensional output by means such as two-dimensional visual output, stereoscopic output; virtual reality glasses (not shown), holographic display, and smoke tank (not shown)), and a printer (not shown).
[0367]
[0363] The computer system (3200) may also include an optical medium including a CD / DVD ROM / RW (3220) using a medium (3221) such as a CD / DVD, a thumb drive (3222), a removable hard drive or solid state drive (3223), a legacy magnetic medium (not shown) such as a tape and a floppy disk (not shown), and a human-accessible storage device and related media such as a specialized ROM / ASIC / PLD-based device such as a security dongle (not shown).
[0368] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not include a transmission medium, a carrier wave, or other transient signals.
[0369]
[0365] The computer system (3200) may also include an interface to one or more communication networks (3255). The network can be, for example, wireless, wired, or optical. The network can further be related to local, wide area, metropolitan, vehicle industry, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), wired or wireless wide area digital networks for TV (including cable TV, satellite TV, and over-the-air broadcast TV), vehicle industries including CANBus, etc. A particular network generally requires an external network interface adapter attached to a particular general-purpose data port or peripheral bus (3249) (e.g., the USB port of the computer system (3200)); others are generally integrated commonly into the core of the computer system (3200) by attaching to a system bus as described below (e.g., an Ethernet interface is integrated within a PC computer system, and a cellular network interface is integrated within a smartphone computer system). Using any of these networks, the computer system (3200) can communicate with other entities. Such communication can be one-way receive-only (e.g., broadcast TV), one-way transmit-only (e.g., CANbus for a particular CANbus device), or two-way, e.g., for other computer systems using local or wide area digital networks. Specific protocols and protocol stacks can be used with each of those networks and network interfaces as described above.
[0370]
[0366] The aforementioned human interface device, human accessible storage device, and network interface can be attached to the core (3240) of the computer system (3200).
[0371]
[0367] The core (3240) can include one or more central processing units (CPUs) (3241), a graphics processing unit (GPU) (3242), a special programmable processing device in the form of a field programmable gate array (FPGA) (3243), a hardware accelerator for specific tasks (3244), a graphics adapter (3250), etc. These devices can be connected via a system bus (3248) together with a read only memory (ROM) (3245), a random access memory (3246), and an internal mass storage device (e.g., an internal non-user accessible hard drive, SSD, etc.) (3247). In some computer systems, the system bus (3248) may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the system bus (3248) of the core or via a peripheral bus (3249). In one example, the screen (3210) can be connected to the graphics adapter (3250). The architecture of the peripheral bus includes PCI, USB, etc.
[0372]
[0368] The CPU (3241), GPU (3242), FPGA (3243), and accelerator (3244) can be combined to execute specific instructions capable of constituting the aforementioned computer code. The computer code can be stored in the ROM (3245) or RAM (3246). Temporary data can be stored in the RAM (3246), while persistent data can be stored, for example, in the internal mass storage (3247). Fast storage and retrieval for any memory device may be made possible by utilizing cache memory, which can be closely associated with one or more CPUs (3241), GPUs (3242), mass storage (3247), ROM (3245), RAM (3246), etc.
[0373]
[0369] A computer-readable medium can have computer code therein for performing various computer-implemented operations. The medium and the computer code can be considered to be specially designed and constructed for the purposes of this disclosure, or they can be considered to be of the kind well-known and available to those of ordinary skill in the field of computer software.
[0374] As an example, and not by way of limitation, a computer system having an architecture (3200), specifically a core (3240), can provide functionality by a processor (including a CPU, GPU, FPGA, accelerator, etc.) that executes software embodied on one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage as described above, as well as specific storage of the core (3240) of a non-transitory nature such as mass storage (3247) inside the core or ROM (3245). The software implementing various embodiments of the present disclosure can be stored on such a device and executed by the core (3240). The computer-readable media can include one or more memory devices or chips, depending on specific needs. The software includes defining a data structure stored in RAM (3246) and modifying such a data structure according to a process defined by the software, and causing a specific process or a specific part of a specific process described herein to be executed by the core (3240) and in particular by a processor (including a CPU, GPU, FPGA, etc.) therein. Further or alternatively, the computer system can provide functionality as a result of logic wired or otherwise incorporated within a circuit (e.g., an accelerator (3244)), which circuit can execute a specific process or a specific part of a specific process described herein instead of or in addition to software. References to software include logic and, if necessary, vice versa. References to computer-readable media can include a circuit (such as an integrated circuit (IC)) that stores software for execution, a circuit that embodies logic for execution, or, where appropriate, both. The present disclosure encompasses any suitable combination of hardware and software.
[0375] Appendix A: Acronyms JEM: joint exploration model VVC: versatile video coding BMS: benchmark set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplementary Enhancement Information VUI: Video Usability Information GOPs: Groups of Pictures TUs: Transform Units, PUs: Prediction Units CTUs: Coding Tree Units CTBs: Coding Tree Blocks PBs: Prediction Blocks HRD: Hypothetical Reference Decoder SNR: Signal Noise Ratio CPUs: Central Processing Units GPUs: Graphics Processing Units CRT: Cathode Ray Tube LCD: Liquid-Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile communications LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Areas SSD: solid-state drive IC: Integrated Circuit CU: Coding Unit
[0371] Although the present disclosure has described several exemplary embodiments, there are changes, substitutions, and various alternative equivalents that fall within the scope of the present disclosure. Therefore, although not explicitly illustrated or described herein, it will be understood that those skilled in the art will be able to devise many systems and methods that embody the principles of the present disclosure and are thus within its spirit and scope.
Claims
1. A video processing method in a decoder, comprising: receiving a coded video bitstream including a current block in a current picture; determining, from a candidate list, a first affine - translational merge candidate for prediction of the current block in the current picture, wherein the first affine - translational merge candidate provides affine motion information associated with a first reference picture in a first reference list and translational motion information associated with a second reference picture in a second reference list; generating a first prediction for samples in the current block according to the affine motion information associated with the first reference picture; generating a second prediction for samples in the current block according to the translational motion information associated with the second reference picture; and reconstructing samples of the current block according to a combination of the first prediction and the second prediction; A method comprising the above steps.
2. The method according to claim 1, further comprising: deriving the affine motion information and the translational motion information of the first affine - translational merge candidate; and inserting the first affine - translational merge candidate into the candidate list; A method comprising the above steps.
3. In the method according to claim 2, the step of deriving the affine motion information comprises: deriving the affine motion information according to the first affine motion information of the first affine merge candidate using single - prediction; and deriving the affine motion information according to the second affine motion information of the second affine merge candidate using dual - prediction, wherein the second affine motion information is associated with the first reference list; and A method comprising at least one of the above steps.
4. In the method according to claim 2, the step of deriving the translational motion information comprises: deriving the translational motion information according to the first translational motion information of the first translational merge candidate using single - prediction; and deriving the translational motion information according to the second translational motion information of the second translational merge candidate using dual - prediction, wherein the second translational motion information is associated with the second reference list; and A method comprising at least one of the following. **Claim 5** In the method according to claim 2, the step of deriving the translational motion information comprises: Regarding a merge candidate with a merge index in the translational merge candidate list; Selecting a first list from the first reference list and the second reference list based on the merge index according to a pre-set rule; Deriving the translational motion information based on the first translational motion information of the merge candidate associated with the first list, if the first translational motion information exists; Deriving the translational motion information based on the second translational motion information of the merge candidate associated with another list different from the first list, if the first translational motion information does not exist; A method comprising the above. **Claim 6** In the method according to claim 2, the step of deriving the translational motion information comprises: Deriving the translational motion information based on the control point motion vector of the affine merge candidate; Deriving the translational motion information based on the average of a plurality of control point motion vectors of the affine merge candidate; and Deriving the translational motion information based on the motion vector at the center of the current block, wherein the motion vector is determined according to the affine model from among the affine merge candidates; A method comprising at least one of the above. **Claim 7** In the method according to claim 2, the step of inserting the first affine-translational merge candidate into the candidate list comprises: Inserting the first affine-translational merge candidate into the existing sub-block-based merge candidate list after the latest inherited affine merge candidate; Inserting the first affine-translational merge candidate into the existing sub-block-based merge candidate list after the latest configured affine merge candidate; and Inserting the first affine-translational merge candidate into the existing sub-block-based merge candidate list after the latest affine merge candidate and before the zero motion vector candidate; A method further comprising at least one of the above. **Claim 8** In the method according to claim 2, further: Inserting the first affine-translation merge candidate into an affine-translation merge list separately from a sub-block-based merge candidate list; A method comprising.
9. In the method according to claim 8, the step of determining the first affine-translation merge candidate comprises: Selecting the first affine-translation merge candidate from the affine-translation merge list according to a decoded index from the coded video bitstream in response to a decoded flag from the coded video bitstream indicating a selection of the affine-translation merge list; A method comprising.
10. In the method according to claim 8, further: Sorting candidates in the affine-translation merge list according to candidate template matching costs; A method comprising.
11. In the method according to claim 2, the step of inserting the first affine-translation merge candidate into the candidate list further comprises: Inserting the first affine-translation merge candidate into the candidate list in response to the count of affine-translation merge candidates in the candidate list being lower than a predetermined value; A method comprising.
12. In the method according to claim 2, the step of deriving the affine motion information comprises: Deriving the affine motion information from a subset of affine merge candidates that is before other affine merge candidates in the affine merge candidate list, wherein the count of affine merge candidates in the subset is equal to a predetermined value; A method comprising.
13. In the method according to claim 2, the step of deriving the translational motion information comprises: Deriving the translational motion information from a subset of translational merge candidates that is before other translational merge candidates in the translational merge candidate list, wherein the count of translational merge candidates in the subset is equal to a predetermined value; A method comprising.
14. In the method according to claim 1, further: Determining that dual-prediction is allowed in the current picture; and Determining the first affine-translation merge candidate for prediction of the current block in response to the determination; A method comprising.
15. In the method according to claim 1, further: determining that a reference picture of the current picture is temporally prior to the current picture; and determining, in response to the determination, the first affine - translational merge candidate for prediction of the current block; A method comprising.
16. In the method according to claim 1, further: determining a high - level syntax indicating acceptance of the first affine - translational merge candidate for prediction of the current block; A method comprising.
17. In the method according to claim 1, further: setting the LIC flag associated with the first affine - translational merge candidate to false in response to local illumination compensation (LIC) being enabled; A method comprising.
18. In the method according to claim 1, the step of reconstructing samples of the current block further: setting a bi - prediction (BCW) index by coding unit level weighting to a default value to combine the first prediction and the second prediction with equal weight; A method comprising.
19. In the method according to claim 1, the step of reconstructing samples of the current block further: determining that one of the affine motion information and the translational motion information has a bi - prediction (BCW) index by coding unit level weighting; and combining the first prediction and the second prediction according to the BCW index; A method comprising.
20. In the method according to claim 1, the step of reconstructing samples of the current block further: determining that a first coding unit level weighting - based bi - prediction (BCW) index is associated with the affine motion information and a second BCW index is associated with the translational motion information; and combining the first prediction and the second prediction according to the first BCW index; A method comprising.
21. A computer program causing a computer to execute the method according to any one of claims 1 - 19.
22. An apparatus comprising a processor configured to execute the method according to any one of claims 1 - 19.
23. A video processing method in an encoder, comprising: generating a coded video bitstream including a current block in a current picture and transmitting the coded video bitstream to a decoder; having a first affine-translation merge candidate for prediction of the current block in the current picture determined from a candidate list, the first affine-translation merge candidate providing affine motion information associated with a first reference picture in a first reference list and translational motion information associated with a second reference picture in a second reference list; a method in which samples of the current block are reconstructed according to a combination of a first prediction generated for samples in the current block according to the affine motion information associated with the first reference picture and a second prediction generated for samples in the current block according to the translational motion information associated with the second reference picture.
Citation Information
Patent Citations
Video signal encoding / decoding method and device for said method
JP2022505874A
Method, apparatus, and computer program for video coding
JP2022521699A
Video or image coding deriving weight index information for bi-prediction - Patents.com
JP2022524432A
Image encoding / decoding method and device for performing weighted prediction, and bitstream transmission method
JP2022544844A
Image encoding / decoding method and device for performing weighted prediction, and method for transmitting bitstream
WO2021034123A1