Method and computer program for video encoding and decoding

By employing template matching to reorder displacement vector candidates for subblock-based temporal motion vector prediction, the proposed solution addresses the inefficiencies in existing video coding technologies, resulting in enhanced coding efficiency and reduced complexity.

JP2025517841APending Publication Date: 2025-06-12TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024515303
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-10
Filing Date
2022-11-11
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently predicting motion vectors, particularly in subblock-based temporal motion vector prediction, due to the complexity of directional prediction and the need for effective displacement vector reordering.

Method used

The proposed solution involves a processing circuit that receives prediction information for a current coding block and uses template matching to reorder displacement vector candidates, calculating cost values based on template comparisons to sort and select the most appropriate displacement vector for subblock-based temporal motion vector prediction.

Benefits of technology

This approach enhances the efficiency of motion vector prediction by optimizing displacement vector reordering through template matching, leading to improved coding efficiency and reduced computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025517841000001_ABST
    Figure 2025517841000001_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide a method and apparatus for video encoding / decoding. The apparatus is: receiving, from a coded video bitstream, prediction information for a current coding block in a current picture, the prediction information indicating that the current coding block is coded using a sub-block based temporal motion vector prediction (SbTMVP) mode; deriving a plurality of displacement vector (DV) candidates by applying a plurality of displacement vector (DV) offsets to a fixed DV predictor of the current coding block; comparing a template of the current coding block with each of the plurality of DV candidates, each template of the plurality of templates being placed at a position specified by a corresponding one of the plurality of DV candidates; calculating a cost value associated with each one of the plurality of DV candidates based on the comparison; and including a processing circuit for reordering DV offset indices of the plurality of DV candidates based on their calculated cost values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 346,283, filed May 26, 2022, entitled "SUBBLOCK BASED MOTION VECTOR PREDICTOR DISPLACEMENT VECTOR REORDERING USING TEMPLATE MATCHING", and claims the benefit of priority to U.S. Patent Application No. 17 / 985,127, filed November 10, 2022, entitled "SUBBLOCK BASED MOTION VECTOR PREDICTOR DISPLACEMENT VECTOR REORDERING USING TEMPLATE MATCHING". The disclosure of the prior application is hereby incorporated by reference in its entirety.

[0002] This application generally describes aspects related to video coding.

Background Art

[0003] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. The research of the presently named inventors, to the extent that it is not otherwise described in this background art and is not explicitly or implicitly admitted as prior art to the present disclosure, does not, in the context of that research, fall within the scope of what is normally considered prior art at the time of filing.

[0004] Uncompressed digital images and / or videos can include a series of pictures, each picture having spatial dimensions of, for example, 1920×1080 luminance samples and associated chrominance samples. The series of pictures can have a fixed or variable picture rate, for example 60 pictures per second or 60 Hz (also known informally as the frame rate). Uncompressed images and / or videos have specific bitrate requirements. For example, 1080p60 4:2:0 video at 8 bits per sample (1920×1080 luminance sample resolution at a frame rate of 60 Hz) requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires a storage area of over 600 GB.

[0005] One purpose of image and / or video coding and decoding can be the reduction of redundancy in the input image and / or video signal through compression. Compression can help reduce the aforementioned bandwidth and / or storage area requirements, in some cases by more than two orders of magnitude. The description here uses video encoding / decoding as an illustrative example, but the same techniques can be applied to image encoding / decoding in a similar manner without departing from the spirit of the present disclosure. Both reversible compression and irreversible compression, as well as combinations thereof, can be employed. Reversible compression refers to techniques that can reconstruct an exact copy of the original signal from the compressed signal. When using irreversible compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough to make the reconstructed signal useful for the intended application. In the case of video, irreversible compression is widely used. The amount of allowable distortion depends on the application. For example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio may reflect that higher allowable / tolerable distortion can result in a higher compression ratio.

[0006] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transform processing, quantization, and entropy coding.

[0007] Video codec technology can include techniques known as intra coding. In intra coding, sample values are represented without reference to samples from previously reconstructed reference pictures or other data. In some video codecs, a picture is spatially subdivided into blocks of samples. When all blocks of samples are coded in an intra mode, that picture can be an intra picture. Intra pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and thus can be used as the first picture in a coded video bitstream and video session or as a still image. Samples in an intra block can be subjected to a transform, and the transform coefficients can be quantized prior to entropy coding. Intra prediction can be a technique that minimizes sample values in a pre-transform domain. In some cases, the smaller the DC value and the AC coefficients after transformation, the fewer bits are required at a given quantization step size to represent the block after entropy coding.

[0008] For example, traditional intra coding used in MPEG-2 generation coding technology does not use intra prediction. However, some newer video compression technologies include techniques that attempt to perform prediction based on surrounding sample data and / or metadata obtained during the encoding and / or decoding of blocks of data. Such techniques are referred to hereinafter as "intra prediction" techniques. It should be noted that in at least some cases, intra prediction uses only reference data from the current picture being reconstructed and not that from reference pictures.

[0009] There can be many different forms of intra prediction. In a given video coding technology, when two or more of such technologies can be used, the particular technology in use can be coded as a particular intra prediction mode that uses that particular technology. In some cases, the intra prediction mode can have sub - modes and / or parameters, where the sub - modes and / or parameters can be coded individually or included in a mode codeword that defines the prediction mode being used. Which codeword to use for a given mode / sub - mode and / or parameter combination can affect the coding efficiency gain through intra prediction, and the same can be true for the entropy coding technology used to convert the codeword into the bitstream.

[0010] A particular mode of intra prediction was introduced in H.264, scrutinized in H.265, and further scrutinized in newer coding technologies such as the JEM (joint exploration model), VVC (versatile video coding), and benchmark set (BMS). A predictor block can be formed using the sample values of adjacent samples of an already available sample. The sample values of the adjacent samples are copied into the predictor block according to a direction. The reference to the direction in use can be coded in the bitstream or can itself be predicted.

[0011] Referring to FIG. 1A, shown at the lower right is a subset of nine predictor directions, out of 33 possible predictor directions (corresponding to the 33 angular modes of the 35 intra modes) defined in H.265. The point (101) where the arrows converge represents the predicted sample. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples at an angle of 45 degrees from horizontal, to the upper right. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples at an angle of 22.5 degrees from horizontal, to the lower left of sample (101).

[0012] Continuing to refer to FIG. 1A, shown at the upper left is a square block (104) of 4×4 samples (shown by the thick dashed line). The square block (104) contains 16 samples, each labeled with an "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample from the top in the Y dimension and the first sample from the left in the X dimension. Similarly, sample S44 is the fourth sample of block (104) in both the Y and X dimensions. Since the block size is 4×4 samples, S44 is at the lower right. Further, reference samples are shown following a similar numbering scheme. The reference samples are labeled with an R for the block (104), its Y position (e.g., row index), and its X position (column index). In both H.264 and H.265, the predicted samples are adjacent to the block being reconstructed, and thus, negative values need not be used.

[0013] Intra picture prediction can function by copying reference sample values from adjacent samples, as indicated by the predicted direction signaled. For example, assume that a coded video bitstream includes signaling indicating a prediction direction that coincides with arrow (102) for this block, i.e., samples are predicted at a 45-degree angle from the predicted sample to the upper right horizontally. In this case, samples S41, S32, S23, S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.

[0014] In some cases, especially when the direction does not divide evenly by 45 degrees, the values of multiple reference samples can be combined, for example, through interpolation, to calculate a reference sample.

[0015] As video coding technology has evolved, the number of possible directions has increased. In H.264 (2003), nine different directions can be represented. This has increased to 33 in H.265 (2013). Currently, JEM / VVC / BMS can support up to 65 directions. Experiments are conducted to identify the most likely directions, and certain techniques in entropy coding are used to represent those most likely directions with a small number of bits, accepting a specific penalty for less likely directions. Additionally, sometimes the direction itself can be predicted from adjacent directions used in adjacent already decoded blocks.

[0016] Figure 1B shows a schematic (110) indicating 65 intra prediction directions according to JEM, exemplifying the increase in the number of prediction directions over time.

[0017] The mapping of intra prediction direction bits representing directions within a coded video bitstream can vary for each video coding technology. Such mapping can range from, for example, a simple direct mapping to codewords, to complex adaptive schemes including the most probable mode and similar techniques. However, in most cases, there may be certain directions that are statistically less likely to occur within video content than certain other directions. Since the goal of video compression is redundancy reduction, those less likely directions are represented by more bits in well-functioning video coding technologies than the more likely directions.

[0018] Image and / or video coding and decoding can be performed using inter-picture prediction with motion compensation. Motion compensation can be an irreversible compression technique, and a block of sample data from a previously reconstructed picture or a part thereof (reference picture) may be spatially shifted in a direction indicated by a motion vector (hereinafter, MV) and then used for prediction of a newly reconstructed picture or a part of the picture. In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV can have two dimensions X and Y or three dimensions, and the third dimension is an indication of the reference picture in use (the latter may indirectly be the temporal dimension).

[0019] In some video compression techniques, the motion vector (MV) applicable to an area of sample data can be predicted from other MVs, for example, from an MV associated with another area of sample data that is spatially adjacent to the area being reconstructed and that precedes that MV in decoding order. Doing so can substantially reduce the amount of data required to code the MV, thereby removing redundancy and increasing compression. MV prediction can function effectively because, for example, when coding an input video signal (known as natural video) derived from a camera, there is a statistical likelihood that areas larger than the area to which a single MV is applicable will move in a similar direction, and thus, in some cases, prediction can be made using a similar motion vector derived from the MVs of adjacent areas. As a result, the MV found for a given area is similar or identical to the MV predicted from surrounding MVs, and this can then be represented in a smaller number of bits than would be used if the MV were coded directly after entropy coding. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., the MV) derived from the original signal (i.e., the sample stream). In other cases, MV prediction itself can be lossy, for example, due to rounding errors in calculating the predictor from several surrounding MVs.

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms proposed by H.265, the one described in relation to FIG. 2 is a technique called "spatial merge".

[0021] Referring to FIG. 2, the current block (201) contains samples that have been found to be predictable from the previous block of the same size that has been spatially shifted during the motion search process by the encoder. Instead of directly coding the MV, the MV can be derived using an MV associated with any one of five surrounding samples shown as A0, A1 and B0, B1, B2 (202 to 206 respectively) from the metadata associated with one or more reference pictures, for example from the latest reference picture (in decoding order). In H.265, the MV prediction can use predictors from the same reference picture as that used by adjacent blocks. Summary of the Invention

[0022] Aspects of the present disclosure provide methods and apparatus for video encoding and decoding. In some examples, the apparatus includes a processing circuit. The processing circuit is configured to receive, from a coded video bitstream, prediction information for a current coding block in a current picture, the prediction information indicating that the current coding block is coded using a subblock-based temporal motion vector prediction (SbTMVP) mode. The processing circuit is further configured to obtain a plurality of DV offsets from the coded video bitstream, where each DV offset corresponds to a displacement vector candidate, and to derive a plurality of displacement vector (DV) candidates by applying the plurality of DV offsets to a fixed DV predictor of the current coding block. The processing circuit is further configured to compare a template of the current coding block with each of a plurality of templates, where each template of the plurality of templates is located at a position specified by a corresponding one of the plurality of DV candidates. The processing circuit is further configured to calculate a cost value associated with each one of the plurality of DV candidates based on the comparison. The processing circuit is further configured to sort the DV offset indexes of the plurality of DV candidates based on their calculated cost values, and to predict the current coding block in the SbTMVP mode based at least on a DV offset index selected from the sorted DV offset indexes.

[0023] In some aspects, the processing circuit is further configured to receive an index signaled in the coded video bitstream, where the index indicates which DV offset candidate is selected from among the sorted DV offset indexes for performing SbTMVP.

[0024] In some embodiments, after the above rearrangement, the processing circuit is further configured to select, by default, the DV offset candidate having the lowest calculated template matching cost for executing SbTMVP.

[0025] In some embodiments, the cost value is calculated by performing sum of absolute differences (SAD), sum of absolute transformed differences (SATD), sum of squared error (SSE), sub-sampled SAD, or mean-removed SAD.

[0026] In some embodiments, the plurality of DV candidates include merge with motion vector difference (MMVD) candidates.

[0027] In some embodiments, the above comparison, the above calculation, and the above rearrangement are performed only on a subset of the MMVD candidates, and the relative order of one or more other MMVD candidates among the MMVD candidates is maintained without being changed.

[0028] In some embodiments, the above comparison, the above calculation, and the above rearrangement are performed on all of the MMVD candidates, and after the above rearrangement, only N of the MMVD candidates having the lowest cost are used, where the number N is less than or equal to the total number of MMVD candidates.

[0029] In some embodiments, to indicate which MMVD candidates are used, indices in the range [0, N−1] are signaled in the bitstream, where the number N is predefined or signaled in the high-level syntax.

[0030] In some embodiments, the DV offset indices are rearranged in descending or ascending order of their calculated cost values.

[0031] A further aspect of the present disclosure provides methods and apparatuses for video encoding and decoding. In some examples, the apparatus includes a processing circuit. The processing circuit is configured to receive prediction information for a current coding block in a current picture from a coded video bitstream, the prediction information indicating that the current coding block is coded using a sub-block based temporal motion vector prediction (SbTMVP) mode. The processing circuit is further configured to compare a template of the current coding block with each of a plurality of templates, each template of the plurality of templates being located at a position specified by a corresponding one of a plurality of displacement vector (DV) predictor candidates. The processing circuit is further configured to calculate a cost value associated with each one of the plurality of DV predictor candidates based on the comparison. The processing circuit is further configured to sort a list of the plurality of DV predictor candidates based on their calculated cost values and predict / reconstruct the current coding block in the SbTMVP mode based at least on a DV predictor selected from the sorted list of the plurality of DV predictor candidates.

[0032] In some aspects, the processing circuit is further configured to receive an index signaled in the coded video bitstream, the index indicating which DV predictor candidate is selected from the sorted list of the plurality of DV predictor candidates for performing SbTMVP.

[0033] In some aspects, after the sorting, the processing circuit is further configured to default to select a DV predictor candidate having the lowest calculated template matching cost for performing SbTMVP.

[0034] In some aspects, the list of multiple DV predictor candidates is constructed from spatial neighboring coding units (CUs) or from history-based motion vector prediction (HMVP) candidates.

[0035] In some aspects, after sorting, only the first N of the multiple DV predictor candidates on the list are signaled.

[0036] In some aspects, the multiple DV predictor candidates within the list are sorted in descending or ascending order of their calculated cost values.

[0037] Further aspects of the present disclosure provide methods and apparatuses for video encoding and decoding. In some examples, the apparatus includes a processing circuit. The processing circuit is configured to receive prediction information for a current coding block in a current picture from a coded video bitstream, the prediction information indicating that the current coding block is coded using a sub-block based temporal motion vector prediction (SbTMVP) mode. The processing circuit is further configured to generate an SbTMVP candidate list that includes a plurality of displacement vector (DV) candidates for the current coding block, the SbTMVP candidate list including at least one DV predictor candidate derived without applying any DV offset and at least one other DV predictor candidate derived by applying a DV offset to a base DV predictor. The processing circuit is further configured to compare a template of the current coding block with each of a plurality of templates, each template of the plurality of templates being positioned at a location specified by a corresponding one of the plurality of DV candidates within the SbTMVP candidate list. The processing circuit is further configured to calculate a cost value associated with each one of the plurality of DV candidates based on the comparison. The processing circuit is further configured to sort the plurality of DV candidates of the SbTMVP candidate list based on their calculated cost values.

[0038] In some aspects, the processing circuit is further configured to receive an index signaled in a coded video bitstream, the index indicating which DV candidate is selected from a reordered SbTMVP candidate list for performing SbTMVP.

[0039] In some aspects, the selected DV candidate is either an SbTMVP DV candidate derived without applying any DV offset, or an SbTMVP MMVD (Merge with Motion Vector Difference) candidate derived by applying respective DV offsets to respective base DV predictors.

[0040] In some aspects, when the selected DV candidate is an SbTMVP MMVD candidate, the signaling of each base DV predictor and the signaling of the index of each DV offset are performed separately.

[0041] In some aspects, the SbTMVP candidate list is constructed independently of the affine merge candidate list when template matching-based reordering of candidates is enabled for the current frame, and additional syntax is signaled at the coding block level to indicate whether to construct the SbTMVP candidate list or the affine merge candidate list when the sub-block merge mode is signaled.

[0042] Aspects of the present disclosure also provide a non-transitory computer-readable storage medium storing a program executable by at least one processor to perform any one of the above methods.

Brief Description of the Drawings

[0043] Further features, properties and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.

[0044]

Figure 1A

[0045]

Figure 1B

[0046]

Figure 2

[0047]

Figure 3

[0048]

Figure 4

[0049]

Figure 5

[0050]

Figure 6

[0051]

Figure 7

[0052]

Figure 8

[0053]

Figure 9

[0054]

Figure 10

[0055]

Figure 11

[0056]

Figure 12

[0057]

Figure 13

[0058]

Figure 14

[0059]

Figure 15

[0060]

Figure 16

[0061]

Figure 17

[0062]

Figure 18

[0063]

Figure 19

[0064]

Figure 20

[0065]

Figure 21

[0066]

Figure 22

DETAILED DESCRIPTION OF THE INVENTION

[0067] Figure 3 shows an exemplary block diagram of a communication system (300). The communication system (300) includes a plurality of terminal devices that can communicate with each other, for example, via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) perform one-way transmission of data. For example, the terminal device (310) may code video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The coded video data can be transmitted in the form of one or more coded video bitstreams. The terminal device (320) can receive the coded video data from the network (350), decode the coded video data to restore the video pictures, and display the video pictures according to the restored video data. One-way data transmission can be common in media delivery applications and the like.

[0068] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform two-way transmission of coded video data, for example, during a video conference. For two-way data transmission, in one example, each of the terminal devices (330) and (340) can code video data (e.g., a stream of video pictures captured by that terminal device) for transmission to the other of the terminal devices (330) and (340) via the network (350). Each of the terminal devices (330) and (340) can also receive the coded video data transmitted by the other of the terminal devices (330) and (340), decode the coded video data to restore the video pictures, and display the video pictures on an accessible display device according to the restored video data.

[0069] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) may each be exemplified as a server, a personal computer, and a smartphone, respectively, but the principles of the present disclosure cannot be so limited. Aspects of the present disclosure find application to laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network (350) represents any number of networks that transmit coded video data among the terminal devices (310), (320), (330), and (340), including, for example, wireline (wired) and / or wireless communication networks. The communication network (350) may exchange data in circuit-switch and / or packet-switch channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of the network (350) may not be important to the operation of the present disclosure, unless otherwise described herein below.

[0070] FIG. 4 shows a video encoder and a video decoder in a streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter may be similarly applicable to other video-enabled applications, including, for example, video conferencing, storage of compressed video in digital media including digital TV and streaming services, CD, DVD, memory stick, etc.

[0071] A streaming system may include a video source (401) that creates a stream of, for example, uncompressed video pictures (402), and a video capture subsystem (413) that may include, for example, a digital camera. In one example, the stream of video pictures (402) includes samples taken by a digital camera. The stream of video pictures (402) is shown as a thick line to emphasize the high data volume when compared to the encoded video data (404) (or coded video bitstream), and can be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter as described in more detail below. The encoded video data (404) (or encoded video bitstream) is shown as a thin line to emphasize the low data volume when compared to the stream of video pictures (402), and can be stored in a streaming server (405) for future use. One or more streaming client subsystems, such as client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include, for example, a video decoder (410) within an electronic device (430). The video decoder (410) decodes an input copy (407) of the encoded video data and creates an output stream of video pictures (411) that can be rendered on a display (412) (such as a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (such as video bitstreams) can be encoded according to a particular video coding / compression standard. Examples of these standards include ITU-T Recommendation H.265.In one example, the video coding standard under development is informally known as VVC (Versatile Video Coding). The disclosed subject matter may be used in the context of VVC.

[0072] Note that electronic devices (420) and (430) can include other components (not shown). For example, electronic device (420) can include a video decoder (not shown), and electronic device (430) can also include a video encoder (not shown).

[0073] FIG. 5 shows an exemplary block diagram of a video decoder (510). The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). Instead of the video decoder (410) in the example of FIG. 4, the video decoder (510) can be used.

[0074] The receiver (531) may receive one or more coded video sequences to be decoded by the video decoder (510). In one aspect, one coded video sequence is received at a time, in which case the decoding of each coded video sequence is independent of the decoding of other coded video sequences. The coded video sequence may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (531) may receive the encoded video data together with other data, such as coded audio data and / or auxiliary data streams, and these data may be transferred to their respective using entities (not shown). The receiver (531) may separate the coded video sequence from other data. To eliminate network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter, “parser (520)”). In certain applications, the buffer memory (515) is part of the video decoder (510). Otherwise, the buffer memory (515) may be external to the video decoder (510) (not shown). Still otherwise, there may be a buffer memory (not shown) external to the video decoder (510) to eliminate network jitter, in addition to another buffer memory (515) internal to the video decoder (510) to handle, for example, playback timing. When the receiver (531) is receiving data from a store / forward device with sufficient bandwidth and controllability or from an isosynchronous network, the buffer memory (515) may not be necessary or may be made small. For use in a best-effort packet network such as the Internet, a buffer memory (515) may be required, and its size may be relatively large, advantageously an adaptive size, and may be implemented at least in part in an operating system or similar element (not shown) external to the video decoder (510).

[0075] Video decoder (510) may include a parser (520) to reconstruct symbols (521) from a coded video sequence. The categories of these symbols include information used to manage the operation of the video decoder (510) and potentially information for controlling a rendering device (512) (such as a display screen) that is not an integral part of the electronic device (530) but can be coupled to the electronic device (530) as shown in FIG. 5. The control information for the rendering device may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (520) may syntax analyze / entropy decode the received, coded video sequence. The coding of the coded video sequence can follow video coding techniques or standards and can follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) can extract a set of subgroup parameters for at least one subgroup of a subgroup of pixels in the video decoder based on at least one parameter corresponding to the group. Subgroups can include Groups of Pictures (GOPs), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), etc. The parser (520) may also extract from coded video sequence information such as transform coefficients, quantization parameter values, motion vectors, etc.

[0076] The parser (520) may perform an entropy decoding / syntax analysis operation on the video sequence received from the buffer memory (515) to generate symbols (521).

[0077] The reconstruction of the symbol (521) can involve multiple different units depending on the type of the coded video picture or a portion thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. How each unit is involved can be controlled by subgroup control information parsed by the parser (520) from the coded video sequence. The flow of such subgroup control information between the parser (520) and the multiple units described below is not illustrated for clarity.

[0078] In addition to the function blocks already described, the video decoder (510) can be conceptually subdivided into multiple functional units as described below. In a practical implementation operating under commercial constraints, many of these units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.

[0079] The first unit is the scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives, as symbols (521) from the parser (520), not only control information including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc., but also quantized transform coefficients. The scaler / inverse transform unit (551) can output a block including sample values that can be input to the aggregator (555).

[0080] In some cases, the output samples of the scaler / inverse transform unit (551) may be related to intra-coded blocks. An intra-coded block is a block that does not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed part of the current picture. Such prediction information can be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) uses the surrounding already reconstructed information fetched from the current picture buffer (558) to generate blocks of the same size and shape as the block being reconstructed. The current picture buffer (558) buffers, for example, the partially reconstructed current picture and / or the fully reconstructed current picture. The aggregator (555) may, in some cases, add, for each sample, the prediction information generated by the intra prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).

[0081] In other cases, the output samples of the scaler / inverse transform unit (551) may be related to blocks that are inter-coded and potentially motion-compensated. In such cases, the motion-compensation prediction unit (553) can access the reference picture memory (557) to fetch the samples to be used for prediction. According to the symbols (521) related to the block, after motion-compensating the fetched samples, these samples can be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (in this case, called the residual samples or residual signal) to generate output sample information. The address in the reference picture memory (557) from which the motion-compensation prediction unit (553) fetches the prediction samples can be controlled by the motion vectors available to the motion-compensation prediction unit (553) in the form of symbols (521) that can have, for example, X, Y and reference picture components. Motion compensation can also include interpolation of sample values fetched from the reference picture memory (557) when accurate sub-sample motion vectors are used, motion vector prediction mechanisms, etc.

[0082] The output samples of the aggregator (555) can be subject to various loop-filtering techniques within the loop filter unit (556). The video compression technique is controlled by the parameters included in the coded video sequence (also called the coded video bitstream) and made available to the loop filter unit (556) as symbols (521) from the parser (520), and can include in-loop filtering techniques. Video compression can also respond to meta-information obtained during the decoding of the previous part of the coded picture or coded video sequence (in decoding order), and can respond to previously reconstructed and loop-filtered sample values.

[0083] The output of the loop filter unit (556) can be output to the rendering device (512) and can be stored in the reference picture memory (557) for use in future inter-picture prediction, and can be a sample stream.

[0084] Once a particular coded picture is completely reconstructed, it can be used as a reference picture for future prediction. For example, when the coded picture corresponding to the current picture is completely reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a fresh current picture buffer can be reallocated before starting the reconstruction of subsequent coded pictures.

[0085] The video decoder (510) may perform a decoding operation according to a predetermined video compression technique or a standard such as ITU-T Rec. H.265. The coded video sequence may conform to the syntax specified by the video compression technique or standard being used in the sense that the coded video sequence adheres to both the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. Specifically, the profile can select a specific tool as the only tool available for use under that profile. Also, it may be necessary for compliance that the complexity of the coded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level restricts the maximum picture size, maximum frame rate, maximum reconstruction sample rate (measured, for example, in megasamples per second), maximum reference pixel size, etc. The limits set by the level may, in some cases, be further restricted through the HRD specifications and metadata for buffer management of a Hypothetical Reference Decoder (HRD) signaled in the coded video sequence.

[0086] In one aspect, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0087] FIG. 6 shows an exemplary block diagram of a video encoder (603). The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). Instead of the video encoder (403) in the example of FIG. 4, the video encoder (603) can be used.

[0088] The video encoder (603) can receive video samples from a video source (601) (not part of the electronic device (620) in the example of FIG. 6) that can capture a video image to be coded by the video encoder (603). In another example, the video source (601) is part of the electronic device (620).

[0089] The video source (601) can provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream that can have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits,...), any color space (e.g., BT.601 YCrCb, RGB,...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media supply system, the video source (601) can be a storage device that stores pre-prepared video. In a video conferencing system, the video source (601) can be a camera that captures local image information as a video sequence. The video data can be provided as a plurality of individual pictures that convey motion when viewed in sequence. The picture itself can be organized as a spatial array of pixels, in which case each pixel can include one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0090] According to one aspect, a video encoder (603) can code and compress pictures of a source video sequence in real time or under any other required time constraints to obtain a coded video sequence (643). Implementing an appropriate coding speed is one function of the controller (650). In some aspects, the controller (650) controls and is functionally coupled to other functional units as described below. This coupling is not shown for clarity. Parameters set by the controller (650) can include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization techniques,...), picture size, layout of group of pictures (GOP), maximum motion vector search range, etc. The controller (650) can be configured to have other appropriate functions related to the video encoder (603) optimized for a particular system design.

[0091] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As a gross oversimplification, in one example, the coding loop can include a source coder (630) (which is responsible for creating symbols such as a symbol stream based on an input picture and reference pictures to be coded) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in the same way as a (remote) decoder also does. The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream results in bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory (634) is also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the same sample values as the reference picture samples that the decoder "sees" when using the prediction during decoding. This basic principle of the synchronicity of the reference picture (and the resulting drift if the synchronicity cannot be maintained, for example due to channel errors) is also used similarly in several related technologies.

[0092] The operation of the "local" decoder (633) can be the same as that of a "remote" decoder such as the video decoder (510), which has already been described above in connection with FIG. 5. However, referring briefly to FIG. 5 as well, since the symbols are available and the encoding / decoding of the symbols into the coded video sequence by the entropy encoder (645) and the parser (520) can be reversible, the entropy decoding part of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully implemented in the local decoder (633).

[0093] In one aspect, any decoder technology other than the parsing / entropy decoding present in the decoder exists in the corresponding encoder in the same or substantially the same functional form. Thus, the disclosed subject matter focuses on the operation of the decoder. Since the description of the encoder technology is the opposite of the decoder technology described comprehensively, it may be omitted. In certain areas, more detailed descriptions are provided below.

[0094] In some examples, during operation, the source coder (630) may perform motion-compensated predictive coding, which predictively codes an input picture in relation to one or more previously coded pictures from a video sequence designated as a "reference picture". In this way, the coding engine (632) codes the difference between a pixel block of the input picture and a pixel block of a reference picture that can be selected as a prediction reference for the input picture.

[0095] The local video decoder (633) may decode the coded video data of a picture that can be designated as a reference picture based on the symbols created by the source coder (630). The operation of the coding engine (632) may advantageously be an irreversible process. When the coded video data can be decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (633) may replicate the decoding process that can be performed by the video decoder for the reference picture and store the reconstructed reference picture in the reference picture memory (634). In this way, the video encoder (603) may locally store a copy of the reconstructed reference picture with common content as the reconstructed reference picture obtained by the remote video decoder (without transmission errors).

[0096] Predictor (635) may perform predictive search for the coding engine (632). That is, for a new picture to be coded, Predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc., which may function as appropriate predictive references for the new picture. Predictor (635) may operate on a sample block-by-pixel block basis to find an appropriate predictive reference. In some cases, the input picture may have predictive references drawn from a plurality of reference pictures stored in the reference picture memory (634) as determined by the search results obtained by Predictor (635).

[0097] Controller (650) may manage the coding operations of source coder (630), including setting parameters and subgroup parameters used, for example, to encode video data.

[0098] Outputs of all of the aforementioned functional units may be subject to entropy coding in entropy coder (645). Entropy coder (645) converts symbols generated by various functional units into a coded video sequence by applying reversible compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0099] The transmitter (640) may buffer the coded video sequence created by the entropy coder (645) to prepare for transmission via the communication channel (660), which may be a hardware / software link to a storage device storing the encoded video data. The transmitter (640) may merge the coded video data from the video encoder (603) with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (source not shown).

[0100] The controller (650) may manage the operation of the video encoder (603). During coding, the controller (650) may assign a specific coded picture type to each coded picture, which may affect the coding applicable to each picture. For example, a picture may often be assigned as one of the following picture types:

[0101] An intra picture (I picture) may be coded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example independent decoder refresh (IDR: Independent Decoder Refresh) pictures. Those skilled in the art are aware of these variations of I pictures and their respective uses and characteristics.

[0102] A predicted picture (P picture) may be coded and decoded using intra prediction or inter prediction using at most one motion vector and a reference index to predict the sample values of each block.

[0103] A bi-directional predicted picture (B picture) can be coded and decoded using intra prediction or inter prediction, using up to two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple-predictive pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0104] A source picture is typically subdivided spatially into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8 or 16×16 samples each) and can be coded block by block. The blocks can be coded predictively in relation to other (already coded) blocks, as determined by the coding assignment applied to each picture of the block. For example, blocks of an I picture may be coded non-predictively, or they may be coded predictively in relation to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be coded predictively via spatial prediction or via temporal prediction in relation to one previously coded reference picture. Blocks of a B picture can be coded predictively via spatial prediction or via temporal prediction in relation to one or two previously coded reference pictures.

[0105] A video encoder (603) can perform coding operations according to a given video coding technology or standard such as ITU-T Rec. H.265. In that operation, the video encoder (603) can perform various compression operations, including predictive coding operations that utilize the temporal and spatial redundancy in the input video sequence. The coded video data can thus conform to a syntax specified by the video coding technology or standard being used.

[0106] In one aspect, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include such data as part of the coded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, and the like.

[0107] Video may be captured as a plurality of source pictures (video pictures) in a temporal sequence. Intra picture prediction (often abbreviated as intra prediction) uses the spatial correlation in a given picture, and inter picture prediction uses the (temporal or other) correlation between pictures. In one example, a particular picture during encoding / decoding is called the current picture and is partitioned into blocks. When a block within the current picture is similar to a reference block within a reference picture that has been previously coded and is still buffered within the video, the block within the current picture can be coded by a vector called a motion vector. The motion vector points to the reference block within the reference picture and can have a third dimension that identifies the reference picture in cases where multiple reference pictures are being used.

[0108] In some aspects, dual prediction techniques can be used in inter picture prediction. According to the dual prediction technique, two reference pictures, such as a first reference picture and a second reference picture, both of which precede the current picture within the video in decoding order (although in display order they can be past and future respectively), are used. A block within the current picture can be coded by a first motion vector that points to a first reference block within the first reference picture and a second motion vector that points to a second reference block within the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.

[0109] Furthermore, in order to improve coding efficiency, a merge mode technique can be used in inter-picture prediction.

[0110] According to some aspects of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and the CTUs within a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luma CTB and two chroma CTBs. Each CTU can be recursively quadtree split into one or more coding units (CUs). For example, a 64×64 pixel CTU can be split into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine the prediction type of the CU, such as an inter-prediction type or an intra-prediction type. The CU is split into one or more prediction units (PUs) depending on its temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one aspect, the prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block may include a matrix of values (e.g., luma values) for pixels, such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0111] FIG. 7 shows an exemplary diagram of a video encoder (703). The video encoder (703) receives a processing block (e.g., a prediction block) of sample values within a current video picture in a sequence of video pictures, and is configured to encode the processing block into a coded picture that is part of a coded video sequence. In one example, the video encoder (703) is used in place of the video encoder (403) of the example of FIG. 4.

[0112] In an example of HEVC, the video encoder (703) receives a matrix of sample values of a processing block such as a prediction block of 8×8 samples. The video encoder (703) determines whether the processing block is best coded using an intra mode, an inter mode, or a bi-prediction mode, for example, using rate-distortion optimization. When the processing block is to be coded in the intra mode, the video encoder (703) may use intra prediction techniques to encode the processing block into a coded picture, and when the processing block is to be coded in the inter mode or the bi-prediction mode, the video encoder (703) may use inter prediction techniques or bi-prediction techniques, respectively, to encode the processing block into a coded picture. In certain video coding techniques, the merge mode can be an inter-picture prediction sub-mode when the motion vector is derived from one or more motion vector predictors without the benefit of the coded motion vector components outside the predictor. In certain other video coding techniques, there may be motion vector components applicable to the target block. In one example, the video encoder (703) includes other components such as a mode decision module (not shown) to determine the mode of the processing block.

[0113] In the example of FIG. 7, video encoder (703) includes inter-encoder (730), intra-encoder (722), residual calculator (723), switch (726), residual encoder (724), general controller (721), and entropy encoder (725) that are coupled together as shown in FIG. 7.

[0114] Inter-encoder (730) receives samples of a current block (e.g., a processing block), compares the block to one or more reference blocks (e.g., blocks in a previous picture and a subsequent picture) in a reference picture, generates inter-prediction information (e.g., a description of redundant information according to an inter-coding technique, a motion vector, merge mode information), and is configured to calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on the encoded video information.

[0115] Intra-encoder (722) receives samples of a current block (e.g., a processing block), optionally compares the block to already-coded blocks in the same picture, generates quantized coefficients after transformation, and is optionally also configured to generate intra-prediction information (e.g., intra-prediction direction information according to one or more intra-coding techniques). In one example, intra-encoder (722) also calculates an intra-prediction result (e.g., a predicted block) based on the intra-prediction information and reference blocks in the same picture.

[0116] The general controller (721) is configured to determine general control data and control other components of the video encoder (703) based on the general control data. In one example, the general controller (721) determines the mode of a block and provides a control signal to the switch (726) based on the mode. For example, when the mode is the intra mode, the general controller (721) controls the switch (726) to select the intra mode result for use by the residual calculator (723), controls the entropy encoder (725) to select the intra prediction information and include the intra prediction information in the bitstream, and when the mode is the inter mode, the general controller (721) controls the switch (726) to select the inter prediction result for use by the residual calculator (723), controls the entropy encoder (725) to select the inter prediction information and include the inter prediction information in the bitstream.

[0117] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result, selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) operates based on the residual data and is configured to encode the residual data to generate transform coefficients. In one example, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain to generate transform coefficients. The transform coefficients are then subject to quantization processing to obtain quantized transform coefficients. In various aspects, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform inverse transformation to generate decoded residual data. The decoded residual data can be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and the intra prediction information. The decoded block is appropriately processed to generate a decoded picture, and the decoded picture is buffered in a memory circuit (not shown) and can be used as a reference picture in some examples.

[0118] The entropy encoder (725) is configured to format the bitstream to include the encoded block. The entropy encoder (725) is configured to include various information in the bitstream according to an appropriate standard such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other appropriate information in the bitstream. Note that there is no residual information when coding a block in either the merge submode of the inter mode or the bi-prediction mode according to the disclosed subject matter.

[0119] FIG. 8 shows an exemplary diagram of a video decoder (810). The video decoder (810) is configured to receive a coded picture that is part of a coded video sequence and decode the coded picture to generate a reconstructed picture. In one example, the video decoder (810) is used instead of the video decoder (410) of the example of FIG. 4.

[0120] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) that are coupled together as shown in FIG. 8.

[0121] The entropy decoder (871) can be configured to reconstruct from the coded picture specific symbols that represent syntax elements from which the coded picture is composed. Such symbols can include, for example, the mode in which a block is coded (e.g., intra mode, inter mode, bi-prediction mode, two latter merge sub-modes, or another sub-mode, etc.), and prediction information (e.g., intra prediction information or inter prediction information, etc.) that can identify specific samples or metadata used for prediction by the intra decoder (872) or the inter decoder (880), respectively. The symbols can also include, for example, residual information in the form of quantized transform coefficients. In one example, when the prediction mode is inter mode or bi-prediction mode, the inter prediction information is provided to the inter decoder (880), and when the prediction mode is intra prediction mode, the intra prediction information is provided to the intra decoder (872). The residual information can be subject to inverse quantization and is provided to the residual decoder (873).

[0122] The inter decoder (880) is configured to receive the inter prediction information and generate an inter prediction result based on the inter prediction information.

[0123] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0124] The residual decoder (873) may be configured to perform inverse quantization to extract de-quantized transform coefficients, and process the de-quantized transform coefficients to convert residual information from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (including quantizer parameters (QP)), which may be provided by the entropy decoder (871) (since this may be only a small amount of control information, the data path is not shown).

[0125] The reconstruction module (874) is configured to combine, in the spatial domain, the residual information as the output by the residual decoder (873) and the prediction result (optionally as the output by the inter or intra prediction module) to form a reconstruction block, which may be part of a reconstruction picture that may be part of the reconstructed video. Note that other appropriate operations such as a deblocking operation can be performed to improve visual quality.

[0126] Note that the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) can be implemented using any appropriate technology. In one aspect, the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) can be implemented using one or more integrated circuits. In another aspect, the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) can be implemented using one or more processors that execute software instructions.

[0127] In VVC, various inter prediction modes can be used. In the case of an inter prediction CU, the motion parameters can include an MV, one or more reference picture indexes, a reference picture list utilization index, and additional information on specific coding features used for inter prediction sample generation. The motion parameters can be signaled explicitly or implicitly. When a CU is coded in skip mode, the CU may be associated with a prediction unit (PU) and may not have significant residual coefficients, a coded motion vector delta or MV difference (e.g., MVD), or a reference picture index.

[0128] The motion parameters for the current CU can be obtained from adjacent CUs, including spatial and / or temporal candidates and optionally additional information introduced in VVC, to specify the merge mode. The merge mode can be applied not only to skip mode but also to inter predicted CUs. In one example, an alternative to the merge mode is the explicit transmission of motion parameters, in which case the corresponding reference picture indexes of the MV, each reference picture list, and the reference picture list usage flag, and other information are signaled explicitly for each CU.

[0129] In a mode such as VVC, the VVC Test Model (VTM) reference software includes one or more improved inter prediction coding tools, including extended merge prediction, merge with motion vector difference (MMVD) mode, adaptive motion vector prediction (AMVP) mode with symmetric MVD signaling, affine motion compensation prediction, subblock-based temporal motion vector prediction (SbTMVP), adaptive motion vector resolution (AMVR), motion field storage (1 / 16 luma sample MV storage and 8×8 motion field compression), bi-prediction with CU-level weights (BCW), bi-directional optical flow (BDOF), prediction refinement using optical flow (PROF), decoder side motion vector refinement (DMVR), combined inter and intra prediction (CIIP), geometric partitioning mode (GPM), etc. Inter prediction and related methods are described in detail below.

[0130] In some examples, extended merge prediction can be used. In an example such as VTM4, the merge candidate list includes the following five types of candidates in order: namely, a spatial motion vector predictor (MVP) from a spatially adjacent CU, a temporal MVP from a co-located CU, a history-based MVP from a first-in first-out (FIFO) table, a pairwise average MVP, and a zero MV.

[0131] The size of the merge candidate list can be signaled in the slice header. In one example, the maximum allowable size of the merge candidate list is 6 in VTM4. For each CU coded in merge mode, the index of the best merge candidate (e.g., the merge index) can be coded using truncation-unary binary (TU). The first bin of the merge index can be coded in a context (e.g., using context-adaptive binary arithmetic coding (CABAC)), and bypass coding can be used for the other bins.

[0132] Some examples of the generation process for each category of merge candidates are provided below. In one aspect, the spatial candidates are derived as follows. The derivation of spatial merge candidates in VVC can be the same as that in HEVC. In one example, up to four merge candidates are selected from the candidates located at the positions shown in FIG. 9. FIG. 9 shows the positions of the spatial merge candidates according to an aspect of the present invention. Referring to FIG. 9, the order of derivation is B1, A1, B0, A0, and then B2. The position B2 is considered only when any of the CUs at positions A0, B0, B1, and A1 are not available (e.g., because the CU belongs to another slice or another tile) or when it is intra-coded. After the candidate at position A1 is added, the addition of the remaining candidates is subject to a redundancy check that ensures candidates with the same motion information are excluded from the candidate list, thereby improving coding efficiency.

[0133] To reduce the computational complexity, not necessarily all possible candidate pairs are considered in the aforementioned redundancy check. Instead, only the pairs linked by the arrows in FIG. 10 are considered, and a candidate is added to the candidate list only if the corresponding candidate used in the redundancy check does not have the same motion information. FIG. 10 shows the candidate pairs considered for the redundancy check of the spatial merge candidates according to an aspect of the present disclosure. Referring to FIG. 10, the pairs linked by each arrow include A1 and B1, A1 and A0, A1 and B2, B1 and B0, and B1 and B2. Therefore, the candidates at positions B1, A0 and / or B2 can be compared with the candidates at position A1, and the candidates at positions B0 and / or B2 can be compared with the candidates at position B1.

[0134] In one aspect, the time candidates are derived as follows. In one example, only one time merge candidate is added to the candidate list. FIG. 11 shows an exemplary motion vector scaling for time merge candidates. To derive the time merge candidates for the current CU (1111) in the current picture (1101), the scaled MV (1121) (e.g., shown by the dotted line in FIG. 11) can be derived based on the collocated CU (1112) belonging to the collocated reference picture (1104). The reference picture list used to derive the collocated CU (1112) can be explicitly signaled within the slice header. As shown by the dotted line in FIG. 11, the scaled MV (1121) of the time merge candidate can be obtained. The scaled MV (1121) can be scaled from the MV of the collocated CU (1112) using the picture order count (POC) distances tb and td. The POC distance tb can be defined as the POC difference between the current reference picture (1102) of the current picture (1101) and the current picture (1101). The POC distance td can be defined as the POC difference between the collocated reference picture (1104) of the collocated picture (1103) and the collocated picture (1103). The reference picture index of the time merge candidate can be set to zero.

[0135] FIG. 12 shows exemplary candidate positions (e.g., C0 and C1) for the time merge candidates of the current CU. The position of the time merge candidate can be selected between the candidate positions C0 and C1. The candidate position C0 is located at the lower right corner of the collocated CU (1210) of the current CU. The candidate position C1 is located at the center of the collocated CU (1210) of the current CU. If the CU at the candidate position C0 is not available, is intra-coded, or is outside the current row of the CTU, the candidate position C1 is used to derive the time merge candidate. Otherwise, for example, if the CU at the candidate position C0 is available, is intra-coded, is in the current row of the CTU, the candidate position C0 is used to derive the time merge candidate.

[0136] In some systems, merge with motion vector difference (MMVD) is used for the skip mode or merge mode according to the motion vector representation method. MMVD reuses merge candidates in VVC. FIG. 13 shows an exemplary MMVD search process for the current block (1308) of the current frame (1304) using reference picture list 0 (L0) (1302) and reference picture list 1 (L1) (1306). In some aspects, the search is performed over the MMVD search points in L0 (1302) and L1 (1306), where exemplary search points are shown in FIG. 14. A candidate can be selected from among the merge candidates and further extended by the motion vector representation. MMVD provides a motion vector representation with simplified signaling. The representation method includes a starting point, a motion magnitude, and a motion direction.

[0137] The MMVD technique uses the merge candidate list in VVC, but only candidates of the default merge type (MRG_TYPE_DEFAULT_N) are considered for the MMVD extension. The base candidate index defines the starting point. The base candidate index indicates the best candidate among the candidates in the list as follows, where MVP represents a motion vector predictor.

Table 1

[0138] When the number of base candidates is equal to 1, the base candidate IDX is not signaled. The distance index is the motion magnitude information. The distance index indicates a predefined distance from the starting point information. The predefined distances are as follows:

Table 2

Table 3

[0139] The direction index represents the direction of the MVD with respect to the starting point. The direction index can represent one of four directions. The MMVD flag is signaled immediately after the skip flag and the merge flag are sent. If the skip flag and the merge flag are true, the MMVD flag is parsed. If the MMVD flag is equal to 1, the MMVD syntax is parsed; otherwise, the AFFINE flag is parsed. If the AFFINE flag is equal to 1, i.e., in AFFINE mode, but otherwise, the skip / merge index is parsed for the skip / merge mode of the VTM.

[0140] In some examples, template matching-based candidate reordering on MMVD and affine MMVD can be used. In some systems, the MMVD offset (offset in MMVD mode) is extended for both MMVD mode and affine MMVD mode. Additional refinement positions along the diagonal of k×π / 8 are added as shown in the exemplary search points (1500) of FIG. 15, and thus the number of directions increases from 4 to 16. Second, all possible MMVD refinement positions (16×6) for each base candidate are reordered based on the sum of absolute difference (SAD) cost between the template (one row above and one column to the left of the current block) and its reference for each refinement position. The SAD between two blocks is the sum over all pixels of the absolute value of the difference between the pixel values of the two blocks. Finally, the top 1 / 8 of the refinement positions with the smallest template SAD cost are retained as available positions for MMVD index coding. The MMVD index is binarized by a Rice code with a parameter equal to 2.

[0141] In one example, the (1,1) offset applied to the (1,1) base MV results in a (2,2) MV. Thus, the offset is an MVD applied on top of (i.e., in addition to) the MVP.

[0142] In another way, on top of the MMVD extension as described above, the affine MMVD rearrangement is also extended, with additional refinement positions along the k×π / 4 diagonal added. After the rearrangement, half of the refinement positions with the minimum template SAD cost are retained.

[0143] To improve coding efficiency and reduce the transmission overhead of motion vector information, sub-block level motion vector refinement is applied to extend the CU level temporal motion vector prediction (TMVP). The sub-block based TMVP (SbTMVP) enables inheriting motion information at the sub-block level from the collocated reference picture. Each sub-block of a large-sized CU can have its own motion information without explicitly transmitting the block partition structure or motion information. SbTMVP obtains the motion information of each sub-block in three steps. The first step is the derivation of the displacement vector (DV) of the current CU. Next, the availability of SbTMVP candidates is checked and the central motion is derived. Finally, the sub-block motion information can be derived from the corresponding sub-block based on the DV. Different from the TMVP candidate derivation that always derives the temporal motion vector from the collocated block within the reference frame, SbTMVP applies the DV derived from the MV of the left adjacent CU of the current CU to find the corresponding sub-block within the collocated picture for each sub-block of the current CU. If the corresponding sub-block is not inter-coded, the motion information of the current sub-block is set to the central motion.

[0144] MV is a common term for vectors used for motion compensation in various modes, while DV is specifically defined for the SbTMVP mode. In SbTMVP, a CU is split into multiple sub-blocks, the positions within the reference picture are identified, and the corresponding sub-blocks are used for SbTMVP prediction. In this case, the vector pointing to the reference block position within the reference picture of SbTMVP is the DV. Generally, depending on the prediction mode, the MV can be used for a block or sub-block. In the SbTMVP mode, the DV points to the whole block for SbTMVP prediction, while the MV is used for sub-block motion compensation after SbTMVP prediction.

[0145] VVC supports the SbTMVP method. Similar to TMVP in HEVC, SbTMVP uses the motion field in the collocated picture to improve motion vector prediction and the merge mode for the CU in the current picture. The same collocated picture used by TMVP is used by SbTMVP. SbTMVP differs from TMVP in the following two main aspects. First, TMVP predicts motion at the CU level, while SbTMVP predicts motion at the sub-CU level. Second, TMVP fetches the temporal motion vector from the collocated block in the collocated picture (the collocated block is the block at the lower right or center relative to the current CU), whereas SbTMVP applies a motion shift before fetching the temporal motion information from the collocated picture, where the motion shift is obtained from the motion vector of one of the spatial neighboring blocks of the current CU.

[0146] The SbTMVP process is illustrated in FIGS. 16 and 17. SbTMVP predicts the motion vectors of sub-CUs within the current CU (1600) in the current picture (1700) in two steps. In the first step, the spatial neighborhood A1 (1602) is examined. If A1 (1602) has a motion vector (1704) that uses the collocated picture (1702) as its reference picture, this motion vector (1704) is selected to be the motion shift (or displacement vector) to be applied. If no such motion is identified, the motion shift is set to (0,0).

[0147] In the second step, the motion shift identified in step 1 is applied (i.e., added to the coordinates of the current block), and as shown in FIG. 17, sub-CU level motion information (motion vectors and reference indices) is obtained from the collocated picture (1702). In the example of FIG. 17, it is assumed that the motion shift is set for the motion of block A1. Next, for each sub-CU, the motion information of the sub-CU is derived using the motion information of its corresponding block in the collocated picture (1702) (the smallest motion grid covering the central sample). After the motion information of the collocated sub-CU is identified, it is converted into the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current CU (1600).

[0148] In VVC, a combined sub-block-based merge list that includes both SbTMVP candidates and affine merge candidates is used for signaling the sub-block-based merge mode. The SbTMVP mode is enabled / disabled by a sequence parameter set (SPS) flag. When the SbTMVP mode is enabled, the SbTMVP predictor is added as the first entry in the list of sub-block-based merge candidates, followed by the affine merge candidates. The size of the sub-block-based merge list is signaled in the SPS, and the maximum allowable size of the sub-block-based merge list is 5 in VVC.

[0149] In VVC, the sub-CU size used for SbTMVP is fixed to be 8×8, and the SbTMVP mode is applicable only to CUs having both a width and a height of 8 or more, as is done in the affine merge mode. The sub-block size may be configurable to other sizes such as 4×4 in an enhanced compression model (ECM) software model used for exploration beyond VVC.

[0150] In VVC and ECM, sub-block-based TMVP is made possible by using a DV derived only from the MVs of adjacent CUs of the current CU. However, SbTMVP with the derived DV may not be the best match.

[0151] Aspects of the present disclosure may be used separately or combined in any order. Further, aspects of the present disclosure may be implemented by a processing circuit (e.g., one or more processors including a processing circuit such as an integrated circuit). In one example, the one or more processors execute a program stored in a non-transitory computer-readable medium.

[0152] As used herein, a template refers to a predefined adjacent reconstruction area of the current block. In one example, the template may include the top N rows of the adjacent reconstruction samples above, and / or the left M columns of the adjacent reconstruction samples to the left. Exemplary values of M and N include, but are not limited to, 1, 2, 3, 4,....

[0153] In some aspects, there may be a single DV predictor, but the DV offset is further signaled as a correction (such as MMVD) to the DV predictor to derive the final DV used in SbTMVP. The offset may be signaled, for example, by signaling an index from a predefined set of sets of indexes associated with a set of offsets. Alternatively, a set of DV offsets may be signaled in the coded bitstream. In some other aspects, there may be multiple DV predictors (i.e., multiple base predictors) for deriving the final DV used in SbTMVP. In these cases, in addition to the selected DV predictor, it may or may not be possible to signal the DV offset.

[0154] When signaling the DV offset for sub-block based TMVP (SbTMVP), some current aspects sort the candidate DV offsets based on their associated cost values using template matching. The template of the current coding block is compared with the templates of multiple blocks located at different candidate positions specified by the candidate DV offsets and the DV predictor (derived from the motion information of adjacent blocks such as any spatial neighborhood or any temporal neighborhood with scaling), and the cost value C is calculated for each candidate DV offset.

[0155] The cost value C represents the difference between the sample value of the template of the current coding block and the sample values of each candidate position specified by the candidate DV offset and the DV predictor. That is, the cost value C represents the similarity between the sample value of the template of the current coding block and the sample value of the corresponding candidate position. For example, a lower value of C indicates a higher similarity between the sample value of the template of the current block and the sample value of the corresponding candidate position, and a higher value of C may indicate a lower similarity.

[0156] Therefore, the cost value C calculated for each candidate DV offset represents the similarity between the sample value of the template of the current coding block and the sample value of the candidate position represented by the combination of the DV predictor and each candidate DV offset. Based on the cost value, the DV offset candidates are sorted, for example, based on ascending or descending order.

[0157] This example is shown in FIG. 18, in which the template of the current block (1600) is Tc, and after template matching, T0 and T1 are templates associated with two exemplary candidate DVs, and the two exemplary candidate DVs are derived by different DV offset values applied to the DV predictor. The cost value is calculated for each DV offset candidate, and then, based on the cost, the DV offset candidate index (e.g., MMVD index) is sorted, and the selected index among the sorted indexes is further signaled by the encoder to the decoder to indicate the selected DV offset. The selected DV offset is then used in combination with the DV predictor to predict / reconstruct the current coding block in the SbTMVP mode.

[0158] In some aspects, the selected index may be selected by default, without additional signaling, as the index associated with the minimum cost calculated above. In some other aspects, the selected index may be associated with the minimum prediction cost.

[0159] The shape of the template used in various aspects is not limited to the shape shown in FIG. 18. For example, in various aspects, the template may include only one or more adjacent cells above the block, only one or more adjacent cells to the left of the block, any combination of one or more adjacent cells above the block and one or more adjacent cells to the left of the block, and the like.

[0160] In various alternative aspects, the cost value C can be calculated by the sum of absolute differences (SAD) between the sample value of the template of the current block and the sample value of the template of the candidate block, the sum of absolute transformed differences (SATD), the sum of squared error (SSE), the subsampled SAD, the mean-removed SAD, etc. Various calculation methods of the cost value C represent the similarity between the sample value of the template of the current block and the sample value of the corresponding candidate position. For example, when the sample of the template of the current block and the sample of the template at one of the candidate positions are very similar, the cost value C may be very low so as to represent a high similarity between the template of the current block and the template at that candidate position.

[0161] In one aspect, only some of the candidates on the list (which may include MMVD or other candidates) are measured by template matching, and their associated indexes are sorted. For the remaining candidates on the list, the relative order remains unchanged. For example, assuming there are 8 directions including 2 horizontal directions, 2 vertical directions, and 4 45-degree directions, the probability of using one of the 4 45-degree directions may be relatively low compared to the other directions. In this case, template matching, cost calculation, and index sorting may be performed only for the 2 horizontal directions and 2 vertical directions at the top of the candidate list, and the order of the indexes of the 4 45-degree directions at the bottom of the candidate list is maintained without change. This reduces the computational complexity.

[0162] In one aspect, the template matching costs of all candidates (which may include MMVD or other candidates) on the list are measured, and all candidates are sorted. After sorting, indexes in the range of [0, N-1] are assigned only to the top N candidates (where N is less than or equal to the total number of candidates) having the lowest costs, and the indexes are signaled in the bitstream to indicate which candidates are used. The number N can be predefined or signaled in a high-level syntax such as a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), a slice header (SH), etc.

[0163] In SbTMVP, multiple DV predictors are available, and when the selection of the DV predictors is signaled, some current embodiments sort the DV predictors based on their associated cost values derived by template matching. Based on the cost values, the candidate DV predictors are sorted in ascending or descending order, and the index of the selected DV predictor within the sorted index is signaled to indicate which DV predictor is applied to derive the DV used in SbTMVP. The selected DV predictor may be the first or last predictor on the list, or another DV predictor on the list with the lowest prediction cost.

[0164] In some systems, for SbTMVP, a list of derived DVs (i.e., when the DV predictors are used directly without further signaling of DV offsets) can be constructed from a method such as spatially adjacent CUs or from history-based motion vector prediction (HMVP) candidates. Across all candidate DVs in the list, the cost value C between these and the current block is calculated by template matching. The list is sorted in descending or ascending order of the cost value C. An index is signaled in the bitstream to indicate which of the derived DVs in the list is used. Template matching-based sorting is beneficial in that the most useful DV candidates have shorter codewords and an efficient context model for entropy coding.

[0165] In one embodiment, after sorting, an index for signaling may be assigned only to the first N candidate DVs on the list, such that the signaling cost can be reduced according to the value of C.

[0166] In another embodiment, after sorting the candidate DVs, the first candidate DV having the minimum template matching cost may be used by default, such that signaling is not required to indicate the index of the candidate DVs.

[0167] In another aspect, any of the candidate DVs on the rearrangement list may be signaled by an index within the bitstream. For example, the candidate DV having the lowest prediction cost may be signaled by an index within the bitstream.

[0168] Some aspects may combine a plurality of DV predictor candidates and / or a DV predictor having a plurality of DV offsets into a single DV candidate list. Then, all the DV candidates may be sorted based on the template matching cost. The index of the candidate selected from the sorted list may be signaled in the bitstream.

[0169] In one aspect, all possible DV candidates generated by adding an offset over all DV predictors are sorted in ascending or descending order by using the template matching cost to construct a template matching-based candidate list. An index is signaled in the bitstream to indicate which candidate in the above list is selected. Thus, base DV predictor signaling is not required.

[0170] In one aspect, all possible DV predictor candidates and all possible candidates generated by adding an offset to the base predictor are combined and sorted in ascending order based on the template matching cost to form a sorted candidate list. An index is signaled to indicate which candidate in the list is selected. This indicated candidate may be an SbTMVP candidate having one of a plurality of base DV candidates, or an SbTMVP MMVD candidate having an offset on one of a plurality of base DV candidates.

[0171] In one example, this template matching based rearrangement SbTMVP candidate list may be constructed as an independent candidate list, and when the template matching based rearrangement of the SbTMVP candidate list is enabled for the current frame, the SbTMVP candidates are not included in the affine merge candidate list. When the sub-block merge mode is signaled, additional syntax may be signaled at the coding block level to indicate whether to construct this SbTMVP list or the affine merge list.

[0172] In one aspect, the candidate rearrangement based on template matching may be applied when the base DV predictor is determined, and all candidates generated by adding a DV offset on top of the DV predictor may be rearranged based on their template matching costs, such that the signaling of the base DV predictor and the signaling of the candidate indices of the rearranged list may be done separately.

[0173] Figures 19 to 21 show exemplary flowcharts outlining the encoding / decoding processes (1900, 2000, 2100) of SbTMVP according to various aspects of the present disclosure. In various aspects, each one of the processes (1900, 2000, 2100) may be executed by a processing circuit in the terminal devices (310), (320), (330) and (340), a processing circuit that executes the functions of a video encoder (e.g., (403), (603), (703)), a processing circuit that executes the functions of a video decoder (e.g., (410), (510), (810)), etc. In some aspects, each one of the processes (1900, 2000, 2100) may be implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes one of the processes (1900, 2000, 2100). Each one of the processes (1900, 2000, 2100) can be appropriately adapted to various scenarios, and the steps in each one of the processes (1900, 2000, 2100) can be adjusted accordingly. One or more of the steps in each one of the processes (1900, 2000, 2100) can be adapted, omitted, repeated, and / or combined. Each one of the processes (1900, 2000, 2100) can be implemented using any suitable order. Additional steps can be added.

[0174] First, referring to FIG. 19, the process (1900) starts at (S1901) and proceeds to (S1910).

[0175] (S1905), the process (1900) receives prediction information for the current coding block in the current picture from the coded video bitstream, the prediction information indicating that the current coding block is coded using the sub-block based temporal motion vector prediction (SbTMVP) mode.

[0176] In (S1910), process (1900) derives a plurality of DV candidates by applying a plurality of displacement vector (DV) offset candidates to the current coding block's fixed DV predictor.

[0177] In (S1920), process (1900) compares the template of the current coding block with each of a plurality of templates, where each of the plurality of templates is placed at a position specified by a corresponding one of the plurality of DV candidates.

[0178] In (S1930), process (1900) calculates a cost value associated with each one of the plurality of DV offset candidates based on the comparison.

[0179] In (S1940), process (1900) reorders the DV offset indices of the plurality of DV offset candidates based on their calculated cost values.

[0180] In one example, process (1900) may further receive an index signaled within the coded video bitstream, where the index indicates which DV offset candidate is selected for performing SbTMVP.

[0181] In one example, after reordering, process (1900) may further select the DV offset candidate having the lowest calculated template matching cost by default for performing SbTMVP. Thus, signaling is not required for the selected DV offset.

[0182] In one example, the cost value is calculated by performing sum of absolute differences (SAD), sum of absolute transformed differences (SATD), sum of squared error (SSE), sub-sampled SAD, or mean removed SAD.

[0183] In one example, the plurality of DV candidates includes merge with motion vector difference (MMVD) candidates based on motion vector differences.

[0184] In one example, the comparison, calculation, and sorting are performed only on a subset of the MMVD candidates, where the relative order of one or more of the other MMVD candidates among the MMVD candidates is maintained without being changed.

[0185] In one example, the comparison, calculation, and sorting are performed on all of the MMVD candidates, where after sorting, only the N MMVD candidates having the lowest cost among the MMVD candidates are used, where the number N is less than or equal to the total number of MMVD candidates.

[0186] In one example, to indicate which MMVD candidates are used, indexes in the range of [0, N−1] are signaled in the bitstream, where the number N is predefined or signaled in the high-level syntax.

[0187] In one example, the DV offset indexes are sorted in descending or ascending order of their calculated cost values.

[0188] The process (1900) then proceeds to (S1999) and ends.

[0189] Next, referring to FIG. 20, the process (2000) starts at (S2001) and proceeds to (S2010).

[0190] In (S2005), the method (2000) receives prediction information of a current coding block in a current picture from a coded video bitstream, where the prediction information indicates that the current coding block is coded using a sub-block based temporal motion vector prediction (SbTMVP) mode.

[0191] In (S2010), process (2000) compares the template of the current coding block with each of a plurality of templates, and each of the plurality of templates is placed at a position specified by a corresponding one of a plurality of displacement vector (DV) predictor candidates.

[0192] In (S2020), process (2000) calculates a cost value associated with each one of a plurality of DV predictor candidates based on the comparison.

[0193] In (S2030), process (2000) sorts a list of a plurality of DV predictor candidates based on their calculated cost values.

[0194] In one example, process (2000) further receives an index signaled in the coded video bitstream, where the index indicates which DV predictor candidate is selected from the sorted list of a plurality of DV predictor candidates for performing SbTMVP.

[0195] In one example, after the sorting, process (2000) further selects the DV predictor candidate with the lowest template matching cost calculated by default for performing SbTMVP.

[0196] In one example, the list of a plurality of DV predictor candidates is constructed from spatial neighboring coding units (CUs) or from history-based motion vector prediction (HMVP) candidates.

[0197] In one example, after the sorting, only the first N of the plurality of DV predictor candidates on the list are signaled.

[0198] In one example, the plurality of DV predictor candidates in the list are sorted in descending or ascending order of their calculated cost values.

[0199] Process (2000) then proceeds to (S2099) and ends.

[0200] Next, referring to FIG. 21, the process (2100) starts at (S2101) and proceeds to (S2110).

[0201] In (S2105), the process (2100) receives prediction information of the current coding block in the current picture from the coded video bitstream, where the prediction information indicates that the current coding block is coded using a sub-block based temporal motion vector prediction (SbTMVP) mode.

[0202] In (S2110), the process (2100) generates an SbTMVP candidate list including a plurality of displacement vector (DV) candidates of the current coding block, where the SbTMVP candidate list includes at least one DV predictor candidate derived without applying any DV offset and at least one other DV predictor candidate derived by applying a DV offset to a base DV predictor.

[0203] In (S2120), the process (2100) compares a template of the current coding block with each of a plurality of templates, and each template of the plurality of templates is placed at a position specified by a corresponding one of the plurality of DV candidates in the SbTMVP candidate list.

[0204] In (S2130), the process (2100) calculates a cost value associated with each one of the plurality of DV candidates based on the comparison.

[0205] In (S2140), the process (2100) sorts the plurality of DV candidates in the SbTMVP candidate list based on their calculated cost values.

[0206] In one example, process (2100) further receives an index signaled in a coded video bitstream, where the index indicates which DV candidate is selected from a reordered SbTMVP candidate list for performing SbTMVP.

[0207] In one example, the selected DV candidate is either an SbTMVP DV candidate derived without applying any DV offset, or an SbTMVP MMVD (Merge with Motion Vector Difference) candidate derived by applying respective DV offsets to respective base DV predictors.

[0208] In one example, when the selected DV candidate is an SbTMVP MMVD candidate, the signaling of respective base DV predictors and the indexing of respective DV offsets are performed separately.

[0209] In one example, the SbTMVP candidate list is constructed independently of the affine merge candidate list when candidate template matching-based reordering is enabled for the current frame, where additional syntax is signaled at the coding block level to indicate whether to construct the SbTMVP candidate list or the affine merge candidate list when the sub-block merge mode is signaled.

[0210] Process (2100) then proceeds to (S2199) and ends.

[0211] Aspects in the present disclosure may be used separately or combined in any order. Further, each of the method (or aspect), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.

[0212] The above-described technology can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, FIG. 22 shows a computer system (2200) suitable for implementing certain aspects of the disclosed subject matter.

[0213] The computer software can be coded using any suitable machine code or computer language that can be the subject of assembly, compilation, linking, or similar mechanisms, and can create code that includes instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or through interpretation, microcode execution, etc.

[0214] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, game devices, Internet of Things devices, etc.

[0215] The components shown in FIG. 22 for the computer system (2200) are illustrative in nature and are not intended to imply any limitations regarding the scope of use or functionality of the computer software implementing the aspects of the present disclosure. Also, the configuration of the components should not be construed as having any dependencies or requirements regarding any one or combination of the components shown in the exemplary aspects of the computer system (2200).

[0216] The computer system (2200) may include a specific human interface input device. Such a human interface input device may respond to input by one or more human users through, for example, tactile input (keystrokes, swipes, movement of a data glove, etc.), audio input (voice, clapping, etc.), visual input (gestures, etc.), and olfactory input (not shown). Also, the human interface input device can be used to capture a specific medium, such as audio (voice, music, ambient sound, etc.), images (scanned images, photographic images obtained from a still image camera, etc.), video (2D video, 3D video including stereoscopic video, etc.), which is not necessarily directly related to conscious human input.

[0217] The human interface input device may include one or more of a keyboard (2201), a mouse (2202), a trackpad (2203), a touch screen (2210), a data glove (not shown), a joystick (2205), a microphone (2206), a scanner (2207), and a camera (2208) (only one of each is shown).

[0218] The computer system (2200) may also include certain human interface output devices. Such human interface output devices may stimulate the senses of one or more human users, for example, through tactile output, acoustics, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by a touch screen (2210), a data glove (not shown), or a joystick (2205), although there may be a tactile feedback device that does not function as an input device), audio output devices (such as speakers (2209), headphones (not shown), etc.), visual output devices (each, regardless of the presence or absence of a touch screen input function and also regardless of the presence or absence of a tactile feedback function, some of which can output two-dimensional visual output or output beyond three dimensions through means such as stereoscopic image output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), including screens (2210) such as CRT screens, LCD screens, plasma screens, and OLED screens), and printers (not shown).

[0219] The computer system (2200) may also include human-accessible storage devices and their associated media, such as optical media or similar media (2221) including CD / DVD ROM / RW (2220) with CD / DVDs, thumb drives (2222), removable hard drives or solid state drives (2223), legacy magnetic media such as tapes and floppy disks (not shown), and special ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0220] One of ordinary skill in the art should also understand that when used in connection with the presently disclosed subject matter, the term "computer-readable medium" does not include a transmission medium, a carrier wave, or other transient signals.

[0221] The computer system (2200) can also include an interface (2254) to one or more communication networks (2255). The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet (registered trademark), wireless LAN, cellular networks including GSM (registered trademark), 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, and vehicle and industrial networks including CANBus. A particular network generally requires attachment to a specific general-purpose data port or peripheral bus (2249) and an external network interface adapter (such as a USB port of the computer system (2200)), while others are generally integrated into the core of the computer system (2200) by attachment to the system bus described later (such as an Ethernet (registered trademark) interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (2200) can communicate with other entities. Such communication can be, for example, one-way receive-only (such as broadcast TV), one-way transmit-only (such as from a specific CANbus to a specific CANbus device), or two-way using a local or wide area digital network. As described above, specific protocols and protocol stacks can be used in each of these networks and network interfaces.

[0222] The aforementioned human interface device, human-accessible storage device, and network interface can be attached to the core (2240) of the computer system (2200).

[0223] The core (2240) can include one or more central processing units (CPUs) (2241), a graphics processing unit (GPU) (2242), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (2243), a hardware accelerator for specific tasks (2244), a graphics adapter (2250), etc. These devices can be connected through a system bus (2248) together with a read-only memory (ROM) (2245), a random access memory (RAM) (2246), an internal mass storage such as an internal non-user-accessible hard drive, SSD, etc. (2247). In some computer systems, the system bus (2248) is accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the system bus (2248) of the core or via a peripheral bus (2249). In one example, a screen (2210) can be connected to a graphics adapter (2250). The architecture of the peripheral bus includes PCI, USB, etc.

[0224] The CPU (2241), GPU (2242), FPGA (2243), and accelerator (2244) can execute specific instructions that can be combined to form the above-mentioned computer code. That computer code can be stored in the ROM (2245) or RAM (2246). Also, temporary data can be stored in the RAM (2246), while permanent data can be stored, for example, in the internal mass storage (2247). Through the use of cache memory that can be closely associated with one or more CPUs (2241), GPUs (2242), mass storage (2247), ROM (2245), RAM (2246), etc., fast storage and retrieval for any of the memory devices can be enabled.

[0225] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be specially designed and constructed for the purposes of this disclosure, or they can be of the kind well known and available to those having skill in the computer software arts.

[0226] By way of example and not limitation, a computer system having an architecture (2200) and in particular a core (2240) can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied on one or more tangible computer-readable media. Such computer-readable media can be associated with user-accessible mass storage as introduced above, as well as specific storage of the core (2240) of a non-transitory nature such as internal core mass storage (2247) or ROM (2245). Software implementing various aspects of the present disclosure can be stored on such devices and executed by the core (2240). The computer-readable media can include one or more memory devices or chips according to specific needs. The software can cause the core (2240) and in particular the processors therein (including a CPU, GPU, FPGA, etc.) to define data structures stored in RAM (2246) and modify such data structures according to processes defined by the software, thereby executing specific processes or specific parts of specific processes described herein. Additionally or alternatively, the computer system can provide functionality as a result of being embodied in a circuit (e.g., an accelerator (2244)) in a logical hardwire or other manner, and this circuit can operate instead of or together with the software to execute specific processes or specific parts of specific processes described herein. References to software include logic and, if necessary, vice versa. References to computer-readable media can include circuits (such as integrated circuits (ICs), etc.) that store software for execution, circuits that embody logic for execution, or both as appropriate. The present disclosure encompasses any suitable combination of hardware and software.

[0227] Some acronyms used herein include the following.

[0228] JEM: joint exploration model

[0229] VVC: versatile video coding

[0230] BMS: benchmark set

[0231] MV: Motion Vector

[0232] HEVC: High Efficiency Video Coding

[0233] SEI: Supplementary Enhancement Information

[0234] VUI: Video Usability Information

[0235] GOPs: Groups of Pictures

[0236] TUs: Transform Units

[0237] PUs: Prediction Units

[0238] CTUs: Coding Tree Units

[0239] CTBs: Coding Tree Blocks

[0240] PBs: Prediction Blocks

[0241] HRD: Hypothetical Reference Decoder

[0242] SNR: Signal Noise Ratio

[0243] CPUs: Central Processing Units

[0244] GPUs: Graphics Processing Units

[0245] CRT: Cathode Ray Tube

[0246] LCD: Liquid-Crystal Display

[0247] OLED: Organic Light-Emitting Diode

[0248] CD: Compact Disc

[0249] DVD: Digital Video Disc

[0250] ROM: Read-Only Memory

[0251] RAM: Random Access Memory

[0252] ASIC: Application-Specific Integrated Circuit

[0253] PLD: Programmable Logic Device

[0254] LAN: Local Area Network

[0255] GSM: Global System for Mobile communications (a pan-European digital mobile phone system)

[0256] LTE: Long-Term Evolution

[0257] CANBus: Controller Area Network Bus

[0258] USB: Universal Serial Bus

[0259] PCI: Peripheral Component Interconnect

[0260] FPGA: Field Programmable Gate Array

[0261] SSD: solid-state drive

[0262] IC: Integrated Circuit

[0263] CU: Coding Unit

[0264] R-D: Rate-Distortion

[0265] HDR: High dynamic range

[0266] SDR: Standard dynamic range

[0267] JVET: Joint Video Exploration Team

[0268] AMVR: Adaptive Motion Vector Resolution

[0269] POC: Picture Order Count

[0270] SbTMVP: Subblock-based Temporal Motion Vector Predictor

[0271] DV: displacement vector

[0272] Although the present disclosure describes some exemplary aspects, there are changes, substitutions, and various alternative equivalents within the scope of the present disclosure. Accordingly, it will be understood by those skilled in the art that although not explicitly shown or described herein, various systems and methods that embody the principles of the present disclosure and thus fall within the spirit and scope of the present disclosure can be devised.

Claims

1. A method performed by a processing circuit for video decoding, comprising: receiving, from a coded video bitstream, prediction information for a current coding block in a current picture, the prediction information indicating that the current coding block is coded using a sub-block based temporal motion vector prediction (SbTMVP) mode; obtaining, from the coded video bitstream, a plurality of displacement vector (DV) offsets, each DV offset corresponding to a displacement vector candidate; deriving a plurality of DV candidates by applying the plurality of DV offsets to a fixed DV predictor of the current coding block; comparing a template of the current coding block with each of a plurality of templates, each template of the plurality of templates being located at a position specified by a corresponding one of the plurality of DV candidates; calculating, based on the comparison, a cost value associated with each one of the plurality of DV candidates; sorting DV offset indexes of the plurality of DV candidates based on their calculated cost values; predicting the current coding block in the SbTMVP mode based at least on a DV offset index selected from the sorted DV offset indexes; A method comprising the above steps.

2. The method of claim 1, further comprising receiving an index signaled in the coded video bitstream, the index indicating which DV offset candidate is selected from among the sorted DV offset indexes for performing SbTMVP. The method according to claim 1.

3. After the sorting, the method further comprises, by default, selecting a DV offset candidate having the lowest calculated template matching cost for performing SbTMVP. The method according to claim 1.

4. The cost value is calculated by performing sum of absolute differences (SAD), sum of absolute transform differences (SATD), sum of squared error differences (SSE), sub-sampled SAD, or mean removed SAD. The method according to claim 1.

5. The plurality of DV candidates includes merge candidates by motion vector difference (MMVD). The method according to claim 1.

6. The comparing step, the calculating step, and the sorting step are performed only on a subset of the MMVD candidates, and the relative order of one or more other MMVD candidates among the MMVD candidates is maintained without being changed. The method according to claim 5.

7. The comparing step, the calculating step, and the sorting step are performed on all of the MMVD candidates, and after the sorting, only N of the MMVD candidates having the lowest cost are used, where the number N is less than or equal to the total number of the MMVD candidates. The method according to claim 5.

8. To indicate which MMVD candidates are used, indexes in the range [0, N−1] are signaled in a bitstream, where the number N is predefined or signaled in a high-level syntax. The method according to claim 7.

9. The DV offset index is sorted in descending or ascending order of the cost values calculated for the corresponding DV offset candidates. The method according to claim 1.

10. A method performed by a processing circuit for video decoding, receiving, from a coded video bitstream, prediction information of a current coding block in a current picture, the prediction information indicating that the current coding block is coded using a sub-block-based temporal motion vector prediction (SbTMVP) mode; comparing a template of the current coding block with each of a plurality of templates, each of the plurality of templates being located at a position specified by a corresponding one of a plurality of displacement vector (DV) predictor candidates; calculating, based on the comparison, a cost value associated with each one of the plurality of DV predictor candidates; sorting a list of the plurality of DV predictor candidates based on their calculated cost values; predicting the current coding block in the SbTMVP mode based at least on a DV predictor selected from the sorted list of the plurality of DV predictor candidates; comprising a method.

11. The list of the plurality of DV predictor candidates is constructed from a spatial adjacent coding unit (CU) or from history-based motion vector prediction (HMVP) candidates. The method according to claim 10.

12. After the rearrangement, only the first N of the plurality of DV predictor candidates in the list are signaled. The method according to claim 10.

13. A method executed by a processing circuit for video decoding, comprising: receiving, from a coded video bitstream, prediction information for a current coding block in a current picture, the prediction information indicating that the current coding block is coded using a sub-block-based temporal motion vector prediction (SbTMVP) mode; generating an SbTMVP candidate list including a plurality of displacement vector (DV) candidates for the current coding block, the SbTMVP candidate list including at least one DV predictor candidate derived without applying any DV offset and at least one other DV predictor candidate derived by applying a DV offset to a base DV predictor; comparing a template of the current coding block with each of a plurality of templates, each template of the plurality of templates being located at a position specified by a corresponding one of the plurality of DV candidates in the SbTMVP candidate list; calculating a cost value associated with each one of the plurality of DV candidates based on the comparison; rearranging the plurality of DV candidates in the SbTMVP candidate list based on their calculated cost values; predicting the current coding block in the SbTMVP mode based on a DV predictor candidate selected from the rearranged DV candidates in the SbTMVP candidate list; A method comprising.

14. Further comprising receiving an index signaled in the coded video bitstream, the index indicating which DV candidate is selected from the rearranged DV candidates of the SbTMVP candidate list for performing SbTMVP. The method according to claim 13.

15. The selected DV candidate is either the SbTMVP DV candidate derived without applying any DV offset, or the SbTMVP MMVD candidate derived by applying each DV offset to its respective base DV predictor. The method according to claim 14.

16. When the selected DV candidate is the SbTMVP MMVD candidate, the signaling of each of the base DV predictors and the signaling of the index of each of the DV offsets are performed separately. The method according to claim 15.

17. The SbTMVP candidate list is constructed independently of the affine merge candidate list when template matching-based reordering of candidates is enabled for the current frame, and additional syntax is signaled at the coding block level to indicate whether to construct the SbTMVP candidate list or the affine merge candidate list when the sub-block merge mode is signaled. The method according to claim 13.

18. A method executed by a processing circuit for video encoding, comprising: adding prediction information of a current coding block in a current picture to a coded video bitstream, wherein the prediction information indicates that the current coding block is coded using a sub-block based temporal motion vector prediction (SbTMVP) mode, acquiring a plurality of displacement vector (DV) offsets from the coded video bitstream, deriving a plurality of DV candidates by applying the plurality of DV offsets to a fixed DV predictor of the current coding block, comparing a template of the current coding block with each of a plurality of templates, each of the plurality of templates being located at a position specified by a corresponding one of the plurality of DV candidates, calculating a cost value associated with each one of the plurality of DV candidates based on the comparison, reordering the DV offset indices of the plurality of DV candidates based on their calculated cost values. Based at least on the DV offset index selected from the sorted DV offset indexes, the current coding block in the SbTMVP mode is predicted. Method. Claim 19 A method executed by a processing circuit for video coding, comprising: adding prediction information of a current coding block in a current picture to a coded video bitstream; the prediction information indicating that the current coding block is coded using a sub-block based temporal motion vector prediction (SbTMVP) mode; a template of the current coding block is compared with each of a plurality of templates, each of the plurality of templates being located at a position specified by a corresponding one of a plurality of displacement vector (DV) predictor candidates; based on the comparison, a cost value associated with each one of the plurality of DV predictor candidates is calculated; a list of the plurality of DV predictor candidates is sorted based on their calculated cost values; based at least on the DV predictor selected from the sorted list of the plurality of DV predictor candidates, the current coding block in the SbTMVP mode is predicted. Method. Claim 20 A method executed by a processing circuit for video coding, comprising: adding prediction information of a current coding block in a current picture to a coded video bitstream; the prediction information indicating that the current coding block is coded using a sub-block based temporal motion vector prediction (SbTMVP) mode; an SbTMVP candidate list including a plurality of displacement vector (DV) candidates of the current coding block is generated, the SbTMVP candidate list including at least one DV predictor candidate derived without applying any DV offset and at least one other DV predictor candidate derived by applying a DV offset to a base DV predictor. The template of the current coding block is compared with each of a plurality of templates, and each template of the plurality of templates is placed at a position specified by a corresponding one of the plurality of DV candidates in the SbTMVP candidate list. Based on the comparison, a cost value associated with each one of the plurality of DV candidates is calculated. The plurality of DV candidates in the SbTMVP candidate list are sorted based on their calculated cost values. Based on a DV predictor candidate selected from the sorted DV candidates in the SbTMVP candidate list, the current coding block in the SbTMVP mode is predicted. Method. Claim 21 A computer program which, when executed by a processing circuit, causes the processing circuit to execute the method according to any one of claims 1 to 20.

Citation Information

Patent Citations

  • Device and method for processing video signal by using inter prediction

    US20220078408A1