Sub-block level temporal motion vector prediction using multiple displacement vector predictors and offsets

The SbTMVP mode in video decoding apparatuses addresses the challenge of predicting motion vectors for sub-blocks by determining displacement vectors through predictors and offsets, resulting in improved compression efficiency and reduced data requirements.

JP2025517843APending Publication Date: 2025-06-12TENCENT AMERICA LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024522498
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-09
Filing Date
2022-11-11
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently predicting motion vectors for sub-blocks within a picture, leading to increased data requirements and reduced compression efficiency.

Method used

The proposed solution involves an apparatus for video decoding that uses a sub-block based temporal motion vector prediction (SbTMVP) mode. This mode determines a displacement vector (DV) for a current block by summing a DV predictor and a DV offset, allowing for the reconstruction of sub-blocks based on motion information from corresponding sub-blocks in a collocated reference picture.

Benefits of technology

This approach enhances compression efficiency by accurately predicting motion vectors at the sub-block level, thereby reducing the data required to encode motion information and improving overall video coding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025517843000001_ABST
    Figure 2025517843000001_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide a method and an apparatus including a processing circuit that receives a coded video bitstream including a current picture including a current block. The processing circuit determines that a current block including a plurality of sub-blocks is coded in a sub-block based temporal motion vector prediction (SbTMVP) mode based on a syntax element in the coded video bitstream. The processing circuit determines a plurality of displacement vector (DV) predictor (DVP) candidates, and receives a base index indicating a DVP among the plurality of DVP candidates and a DV offset of the current block. The processing circuit determines a DV based on the DVP and the DV offset. The DV indicates a block at the same position as the current block within the same position reference picture. The processing circuit reconstructs sub-blocks within the plurality of sub-blocks based on motion information of corresponding sub-blocks within the co-located block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally describes embodiments related to video coding.

Background Art

[0002] Cross - reference to related applications of the present invention This application claims the benefit of priority of U.S. Patent Application No. 17 / 984,107, filed on November 9, 2022, "SUBBLOCK LEVEL TEMPORAL MOTION VECTOR PREDICTION WITH MULTIPLE DISPLACEMENT VECTOR PREDICTORS AND AN OFFSET", which claims the benefit of priority of U.S. Provisional Application No. 63 / 345,802, filed on May 25, 2022, "SUBBLOCK BASED MOTION VECTOR PREDICTOR WITH MULTIPLE BASE INDEX AND MV OFFSET". The disclosure of the prior application is hereby incorporated by reference in its entirety.

[0003] The description of the background art provided herein is for the purpose of generally presenting the context of the present disclosure. The research of the present inventors is not recognized as prior art to the present disclosure, either explicitly or implicitly, in the same way as aspects of the description that are not recognized as prior art at the time of filing, to the extent described in this background art section.

[0004] Uncompressed digital images and / or videos can include a series of pictures, and each picture can have, for example, a spatial dimension of 1920×1080 luminance samples and associated chrominance samples. The series of pictures can have, for example, a fixed or variable picture rate of 60 pictures per second, i.e., 60 Hz (informally also known as the frame rate). Uncompressed images and / or videos have specific bitrate requirements. For example, an 8-bit-per-sample 1080p60 4:2:0 video (1920×1080 luminance sample resolution at a 60 Hz frame rate) requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires more than 600 gigabytes of storage space.

[0005] One purpose of image and / or video coding and decoding can be to reduce redundancy in the input image and / or video signal through compression. Compression can, in some cases, help reduce the aforementioned bandwidth and / or storage space requirements by more than two orders of magnitude. The description in this specification uses video encoding / decoding as an illustrative example, but the same techniques can be applied to image encoding / decoding in a similar manner without departing from the spirit of the present disclosure. Both reversible compression and irreversible compression, as well as combinations thereof, can be used. Reversible compression (lossless compression) refers to a technique where an exact copy of the original signal can be reconstructed from the compressed original signal. When using irreversible compression (lossy compression), the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough to use the reconstructed signal for its intended purpose. In the case of video, irreversible compression is widely used. The amount of acceptable distortion depends on the application. For example, a user of a consumer streaming application may tolerate higher distortion than a user of a television distribution application. The achievable compression rate can reflect the following: the higher the acceptable / tolerable distortion, the higher the compression rate that can be obtained.

[0006] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transform processing, quantization, and entropy coding.

[0007] Video codec technology can include techniques known as intra coding. In intra coding, sample values are represented without reference to samples from previously reconstructed reference pictures or other data. In some video codecs, a picture is spatially subdivided into blocks of samples. When all blocks of samples are coded in an intra mode, that picture can be an intra picture. Intra pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and thus can be used as the first picture in a coded video bitstream and video session or as a still image. Samples of an intra block can be subjected to a transform, and the transform coefficients can be quantized prior to entropy coding. Intra prediction can be a technique that minimizes sample values in a pre-transform region. In some cases, the smaller the post-transform DC value and the smaller the AC coefficients, the fewer bits are required at a given quantization step size to represent the block after entropy coding.

[0008] For example, conventional intra coding used in MPEG-2 generation coding technology does not use intra prediction. However, some newer video compression technologies include techniques that attempt to perform prediction based on, for example, surrounding sample data and / or metadata obtained during coding and / or decoding of a block of data. Such techniques are hereinafter referred to as "intra prediction" techniques. It should be noted that in at least some cases, intra prediction uses only reference data from the currently reconstructed picture during reconstruction rather than from a reference picture.

[0009] Intra prediction can have many different forms. When two or more of such techniques can be used in a given video coding technique, the particular technique in use can be coded as a particular intra prediction mode that uses that particular technique. In some cases, the intra prediction mode can have sub - modes and / or parameters, and the sub - modes and / or parameters can be coded individually or can be included in the mode - code word that defines the prediction mode in use. Which code word should be used for a given combination of mode, sub - mode, and / or parameter can affect the coding efficiency gain through intra prediction, and the same can be true for the entropy coding technique used to convert the code word into the bitstream.

[0010] One mode of intra prediction was introduced with H.264, improved in H.265, and further improved in more recent coding techniques such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). The predictor block can be formed using the adjacent sample values of already available samples. The sample values of the adjacent samples are copied into the predictor block according to a direction. The reference to the direction in use may be coded within the bitstream or the direction itself may be predicted.

[0011] Referring to FIG. 1A, in the lower right, a subset of 9 predictor directions known from 33 possible predictor directions (corresponding to 33 of the 35 intra - modes defined in H.265) is shown. The point (101) where the arrows converge represents the predicted sample. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples in the upper right at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples in the lower left of sample (101) at an angle of 22.5 degrees from the horizontal.

[0012] Referring further to FIG. 1A, in the upper left, a square block (104) of 4×4 samples (shown by the thick dashed line) is shown. The square block (104) contains 16 samples, each labeled with an "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample from the top in the Y dimension and the first sample from the left in the X dimension. Similarly, sample S44 is the fourth sample in block (104) in both the Y and X dimensions. Since the block is 4×4 samples in size, S44 is in the lower right. Further, reference samples following a similar numbering scheme are shown. The reference samples are labeled with an "R", its Y position (e.g., row index) relative to block (104), and its X position (column index). In both H.264 and H.265, the predicted samples are adjacent to the block being reconstructed, and thus negative values need not be used.

[0013] Intra-picture prediction can function by copying the reference sample value from adjacent samples indicated by the signaled prediction direction. For example, assume that the coded video bitstream includes signaling indicating a prediction direction that matches arrow (102) for this block, i.e., the samples are predicted from the sample in the upper right at a 45-degree angle from the horizontal. In that case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Sample S44 is then predicted from reference sample R08.

[0014] In some cases, especially when the direction is not evenly divisible by 45 degrees, the values of multiple reference samples can be combined, for example, through interpolation, to calculate the reference sample.

[0015] As video coding technology develops, the number of possible directions is increasing. In H.264 (2003), nine different directions can be represented. This increased to 33 in H.265 (2013). Currently, JEM / VVC / BMS can support up to 65 directions. Experiments are conducted to identify the most likely directions, and a certain technique in entropy coding is used to represent the likely directions with fewer bits while accepting a certain penalty for the less likely directions. Furthermore, the directions themselves may be predicted from adjacent directions used in adjacent, already decoded blocks.

[0016] FIG. 1B shows a schematic diagram (110) showing 65 intra prediction directions by JEM to show that the number of prediction directions increases over time.

[0017] The mapping of the intra prediction direction bits representing the directions in the coded video bitstream can vary for each video coding technology. Such mapping can range from a simple direct mapping to complex adaptive schemes including codewords, maximum likelihood modes, and similar techniques. However, in most cases, there may be specific directions in the video content that are statistically less likely to occur than other directions. Since the goal of video compression is redundancy reduction, in well-functioning video coding technologies, these less likely directions are represented by more bits than the likely directions.

[0018] Coding and decoding of images and / or videos can be performed using inter-picture prediction with motion compensation. Motion compensation can be a lossy compression technique, and a block of sample data from a previously reconstructed picture or a part thereof (reference picture) is spatially shifted in the direction indicated by a motion vector (hereinafter, MV) and then used for prediction of a newly reconstructed picture or picture part. In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV can have two dimensions X and Y, or can have three dimensions, where the third dimension can be an indication of the reference picture in use (the latter can be, indirectly, the temporal dimension).

[0019] In some video compression techniques, the motion vectors (MVs) applicable to an area of sample data can be predicted from other MVs. For example, an MV for another area of sample data that is spatially adjacent to an area being reconstructed can be predicted from an MV that precedes it in decoding order. Doing so can substantially reduce the amount of data required to code the MVs, thereby removing redundancy and increasing the compression ratio. MV prediction can work effectively, for example, when coding an input video signal obtained from a camera (known as natural video), because areas larger than the area to which a single MV is applicable move in a similar direction, and thus, in some cases, there is a statistical likelihood that similar motion vectors derived from the MVs of adjacent areas can be used for prediction. As a result, the MVs found for an area are similar to or the same as the predicted MVs from surrounding MVs and can be represented with fewer bits than when directly coding the MVs after entropy coding. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., an MV) derived from the original signal (i.e., the sample stream). In other cases, the MV prediction itself can be lossy, for example, due to rounding errors when calculating a predictor from several surrounding MVs.

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms provided by H.265, the one described with reference to FIG. 2 is a technique hereinafter referred to as "spatial merge".

[0021] Referring to FIG. 2, the current block (201) comprises samples discovered by the encoder during the motion search process that it is predictable from the previous block of the same size that has been spatially shifted. Instead of directly coding the MV, the MV can be derived from metadata related to one or more reference pictures, for example, from the most recent reference picture (in decoding order), using an MV associated with any one of five surrounding samples shown as A0, A1, and B0, B1, B2 (202-206 respectively). In H.265, MV prediction can use predictors of the same reference picture that adjacent blocks are using. SUMMARY OF THE INVENTION

[0022] Aspects of the present disclosure provide methods and apparatus for video encoding / decoding. In some examples, an apparatus for video decoding includes processing circuitry that receives a base index and displacement vector (DV) offset information of a current block in a current picture from a coded video bitstream. The current block includes a plurality of sub-blocks that are reconstructed using a sub-block based temporal motion vector prediction (SbTMVP) mode. The base index indicates a DV predictor (DVP) among a plurality of DVP candidates of the current block. The processing circuitry determines a DV of the current block based on the DVP of the current block and the DV offset of the current block indicated by the DV offset information. The DV indicates a block at the same position within a collocated reference picture. The collocated block is collocated with the current block. The processing circuitry determines motion information of sub-blocks in the plurality of sub-blocks based on motion information of corresponding sub-blocks in the collocated block, and reconstructs sub-blocks in the plurality of sub-blocks based on the motion information of sub-blocks in the plurality of sub-blocks.

[0023] In one example, the processing circuitry determines the DV as a vector sum of the DVP and the DV offset.

[0024] In one example, the processing circuit determines a plurality of DVP candidates based on the motion vectors (MVs) of the spatial neighbors of the current block. Each reference picture of each of the spatial neighbors used to determine one of the plurality of DVP candidates is a collocated reference picture.

[0025] In one example, if the number of MVs of the spatial neighbors is less than a threshold, the processing circuit inserts zero MVs into the plurality of DVP candidates.

[0026] In one embodiment, the processing circuit determines a plurality of DVP candidates based on candidates in a merge candidate list. The candidates include at least one of (a) a spatial motion vector predictor (MVP) candidate, (b) a history-based MVP (HMVP) candidate, (c) a pairwise average candidate, or (d) a zero motion vector (MV). The candidates do not include a temporal MVP (TMVP) candidate. Each reference picture of each of the candidates used to determine one of the plurality of DVP candidates is a collocated reference picture.

[0027] In one embodiment, the processing circuit constructs a sub-block-based merge list for the current block. The sub-block-based merge candidates in the sub-block-based merge list include a plurality of SbTMVP merge candidates and at least one affine merge candidate. Each of the plurality of SbTMVP merge candidates corresponds to one of the respective plurality of DVP candidates.

[0028] In one example, the base index indicates which SbTMVP merge candidate within a subset of the sub-block-based merge candidates that are the plurality of SbTMVP merge candidates should be selected. The selected SbTMVP merge candidate corresponds to the DVP.

[0029] In one example, a first number K of the plurality of SbTMVP merge candidates 0 and a second number K of at least one affine merge candidate 1is signaled during high-level syntax. K 0 , K 1 is a positive integer. The base index indicates which sub-block-based merge candidate within a subset of sub-block-based merge candidates that includes (i) a plurality of SbTMVP merge candidates and (ii) at least one affine merge candidate should be selected. The selected sub-block-based merge candidate corresponds to a DVP and one of the plurality of SbTMVP merge candidates.

[0030] In one example, the processing circuit constructs an SbTMVP merge candidate list for the current block that includes a plurality of SbTMVP merge candidates. Each of the plurality of SbTMVP merge candidates corresponds to one of each of the plurality of DVP candidates. The SbTMVP merge candidate list does not include affine merge candidates.

[0031] In one example, a merge mode or a merge motion vector difference (MMVD) mode is applied to the current block. The processing circuit receives a flag from the coded video bitstream. The flag indicates that the SbTMVP mode is applied to the current block.

[0032] In one example, a merge mode or a merge motion vector difference (MMVD) mode is applied to the current block. The processing circuit receives a first flag from the coded video bitstream. The first flag indicates that a sub-block-based merge mode is applied to the current block. The processing circuit receives a second flag from the coded video bitstream, and the second flag indicates that the sub-block-based merge mode is the SbTMVP mode.

[0033] In one example, the processing circuit constructs a sub-block based merge list for the current block. The sub-block based merge list includes one SbTMVP merge candidate and at least one affine merge candidate. One SbTMVP merge candidate corresponds to DVP. The processing circuit receives an index that points to one SbTMVP merge candidate in the sub-block based merge list. The index indicates that the SbTMVP mode is applied to the current block.

[0034] In one embodiment, the processing circuit receives a coded video bitstream including the current picture. The current picture includes the current block. The current block includes a plurality of sub-blocks. The processing circuit determines that the current block including the plurality of sub-blocks is coded in the SbTMVP mode based on syntax elements in the coded video bitstream. The processing circuit determines a plurality of DVP candidates for the current block. Each DVP candidate is derived from one or more motion vectors (MVs). The processing circuit receives a base index and a DV offset for the current block. The base index indicates the DVP among the plurality of DVP candidates for the current block. The processing circuit determines the DV for the current block based on the DVP indicated among the plurality of DVP candidates for the current block and the DV offset of the current block. The DV indicates a block at the same position as the current block in the same position reference picture. The processing circuit reconstructs the sub-blocks within the plurality of sub-blocks based on the motion information of the corresponding sub-blocks within the same position block (the block at the same position as the current block within the same position reference picture).

[0035] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to perform a method for video decoding.

Brief Description of the Drawings

[0036] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.

Figure 1A

Figure 1B

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16A

Figure 16B

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23A

Figure 23B

Figure 24

DETAILED DESCRIPTION OF THE INVENTION

[0037] FIG. 3 shows an exemplary block diagram of a communication system (300). The communication system (300) includes a plurality of terminal devices that can communicate with each other, for example, via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) perform unidirectional transmission of data. For example, the terminal device (310) can code video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The coded video data can be transmitted in the form of one or more coded video bitstreams. The terminal device (320) can receive the coded video data from the network (350), decode the coded video data to restore the video pictures, and display the video pictures according to the restored video data. Unidirectional data transmission can be common in media serving applications and the like.

[0038] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform two-way transmission of coded video data, for example, during a video conference. In the case of two-way data transmission, in one embodiment, each of the terminal devices (330) and (340) can code video data (e.g., a stream of video pictures captured by that terminal device) for transmission to the other of the terminal devices (330) and (340) via the network (350). Also, each of the terminal devices (330) and (340) can receive the coded video data transmitted by the other of the terminal devices (330) and (340), can decode the coded video data to restore the video pictures, and can display the video pictures on a display device accessible according to the restored video data.

[0039] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) are shown as a server, a personal computer, and a smartphone, respectively, but the principles of the present disclosure need not be so limited. Embodiments of the present disclosure find application in laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (350) represents any number of networks that transfer coded video data between the terminal devices (310), (320), (330), and (340), including, for example, wired and / or wireless communication networks. The communication network (350) can exchange data over circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this description, the architecture and topology of the network (350) may not be important for the operation of the present disclosure, unless otherwise described herein below.

[0040] FIG. 4 shows a video encoder and a video decoder in a streaming environment as an example of an application for the disclosed subject matter. The disclosed subject matter is equally applicable to other video-enabled applications including, for example, storing compressed video on digital media including videoconferencing, digital TV, streaming services, CD-Rs, DVD-Rs, memory sticks, etc.

[0041] A streaming system can include a capture subsystem (413) that creates, for example, a stream (402) of uncompressed video pictures, including a video source (401) such as a digital camera. In one embodiment, the stream (402) of video pictures includes samples taken by a digital camera. The stream (402) of video pictures, shown in bold lines to emphasize the high data volume compared to the coded video data (404) (or coded video bitstream), can be processed by an electronic device (420) including a video encoder (403) coupled to the video source (401). The video encoder (403) can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The coded video data (404) (or coded video bitstream), shown as thin lines to emphasize the lower data volume compared to the stream (402) of video pictures, can be stored in a streaming server (405) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to read copies (407) and (409) of the coded video data (404). The client subsystem (406) can include a video decoder (410) within, for example, an electronic device (430). The video decoder (410) decodes an input copy (407) of the coded video data and creates an output stream (411) of video pictures that can be rendered on a display (412) (e.g., display screen) or other rendering device (not shown). In some streaming systems, the coded video data (404), (407), and (409) (e.g., video bitstream) can be coded according to some video coding / compression standards. Examples of these standards include ITU-T Recommendation H.265.In one example, the video coding standard under development is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.

[0042] Note that electronic devices (420) and (430) can include other components (not shown). For example, electronic device (420) can include a video decoder (not shown), and electronic device (430) can also include a video encoder (not shown).

[0043] FIG. 5 shows an exemplary block diagram of a video decoder (510). The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). The video decoder (510) can be used in place of the video decoder (410) in the example of FIG. 4.

[0044] The receiver (531) can receive one or more coded video sequences to be decoded by the video decoder (510). In one embodiment, one coded video sequence is received at a time, and the decoding of each coded video sequence is independent of the decoding of other coded video sequences. The coded video sequence may be received from a channel (501), and the channel (510) may be a hardware / software link to a storage device that stores the coded video data. The receiver (531) can receive the coded video data together with other data, such as coded audio data and / or auxiliary data streams, which can be transferred to their respective using entities (not shown). The receiver (531) can separate the coded video sequence from other data. To handle network jitter, a buffer memory (515) can be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter, the "parser (520)"). For certain applications, the buffer memory (515) is part of the video decoder (510). In other embodiments, it may be external to the video decoder (510) (not shown). In still other embodiments, for example, to handle network jitter, a buffer memory (not shown) may exist external to the video decoder (510), and further, for example, to handle playback timing, another buffer memory (515) may exist inside the video decoder (510). When the receiver (531) is receiving data from a storage / transfer device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory (515) may not be necessary or may be small.A buffer memory (515) may be required for use on a best-effort packet network such as the Internet, can be made relatively large, advantageously can be of an adaptive size, and can be implemented at least partially within an operating system external to the video decoder (510) or a similar element (not shown).

[0045] The video decoder (510) can include a parser (520) that reconstructs symbols (521) from a coded video sequence. The categories of these symbols, as shown in FIG. 5, include information used to manage the operation of the video decoder (510) and potentially information for controlling a rendering device (512) (e.g., a display screen), which is not an integral part of the electronic device (530) but can be coupled to the electronic device (530). The control information for the rendering device can be in the form of supplementary enhancement information (SEI) messages or video usability information (VUI) parameter set fragments (not shown). The parser (520) can parse / entropy decode the received coded video sequence. The coding of the coded video sequence can follow video coding techniques or standards and can follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) can extract a set of subgroup parameters regarding at least one subgroup of pixels within the video decoder from the coded video sequence based on at least one parameter corresponding to a group. The subgroups can include picture groups (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (520) can also extract information such as transform coefficients, quantization parameter values, motion vectors, etc. from the coded video sequence.

[0046] The parser (520) can perform an entropy decoding / syntax analysis operation on the video sequence received from the buffer memory (515) to create symbols (521).

[0047] The reconstruction of the symbol (521) can include multiple different units depending on the type of the coded video picture or a part thereof (inter-picture and intra-picture, inter-block and intra-block, etc.) and other factors. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the video sequence coded by the parser (520). Such a flow of subgroup control information between the parser (520) and the following multiple units is not shown for clarity.

[0048] In addition to the function blocks already described, the video decoder (510) can be conceptually subdivided into several functional units as described below. In actual embodiments operating under commercial constraints, many of these units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.

[0049] The first unit is the scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives, as symbols (521), the quantized transform coefficients and control information including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc. from the parser (520). The scaler / inverse transform unit (551) can output a block containing sample values, and this block can be input to the aggregator (555).

[0050] In some cases, the output samples of the scaler / inverse transform unit (551) can be related to the intra-coded blocks. An intra-coded block is a block that does not use the prediction information from the previously reconstructed picture but can use the prediction information from the previously reconstructed part of the current picture. Such prediction information can be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) uses the surrounding already reconstructed information fetched from the current picture buffer (558) to generate a block of the same size and shape as the block being reconstructed. The current picture buffer (558) buffers, for example, the partially reconstructed current picture and / or the fully reconstructed current picture. The aggregator (555) adds, in some cases, for each sample, the prediction information generated by the intra prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).

[0051] In other cases, the output samples of the scaler / inverse transform unit (551) may be related to inter-coded and potentially motion-compensated blocks. In such cases, the motion compensation prediction unit (553) can access the reference picture memory (557) to fetch the samples used for prediction. After motion-compensating the samples fetched according to the symbols (521) related to the block, these samples can be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (in this case, called the residual samples or residual signal) to generate output sample information. The address in the reference picture memory (557) where the motion compensation prediction unit (553) fetches the prediction samples can be controlled by, for example, the motion vectors available to the motion compensation prediction unit (553) in the form of symbols (521) having X, Y, and reference picture components. Motion compensation can also include interpolation of the sample values fetched from the reference picture memory (557) when exact sub-sample motion vectors are used, a motion vector prediction mechanism, etc.

[0052] The output samples of the aggregator (555) can undergo various loop filtering techniques in the loop filter unit (556). The video compression technology can be controlled by the parameters included in the coded video sequence (also called the coded video bitstream), and can include in-loop filter techniques made available to the loop filter unit (556) as symbols (521) from the parser (520). Video compression can also respond to the meta information obtained during the decoding of the previous part (in decoding order) of the coded picture or coded video sequence, and can also respond to the previously reconstructed and loop-filtered sample values.

[0053] The output of the loop filter unit (556) can be a sample stream that is output to the rendering device (512) and stored in the reference picture memory (557) for use in future inter-picture prediction.

[0054] Once a coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, when the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and the new current picture buffer can be reallocated before starting the reconstruction of subsequent coded pictures.

[0055] The video decoder 510 can perform decoding operations according to a predetermined video compression technology or standard (e.g., ITU-T Rec. H.265). The coded video sequence can conform to the syntax specified by the video compression technology or standard being used in the sense that the coded video sequence conforms to both the syntax of the video compression technology or standard and the profile documented in the video compression technology or standard. Specifically, a profile can select some tools from all the tools available in the video compression technology or standard as the only tools available for use under that profile. Also, what is required for compliance can be that the complexity of the coded video sequence is within the bounds as defined by the level of the video compression technology or standard. In some cases, the level can limit the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limitations set by the level can, in some cases, be further restricted through the hypothetical reference decoder (HRD) specification and the metadata for HRD buffer management signaled in the coded video sequence.

[0056] In one embodiment, the receiver (531) can receive additional (redundant) data along with the coded video. The additional data may be included as part of the coded video sequence. The additional data can be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data can be in the form of, for example, a temporal, spatial, or signal-to-noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, forward error correction codes, and the like.

[0057] FIG. 6 shows an exemplary block diagram of a video encoder (603). The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used in place of the video encoder (403) in the example of FIG. 4.

[0058] The video encoder (603) can receive video samples from a video source (601) (not part of the electronic device (620) in the example of FIG. 6) that can capture the video images to be coded by the video encoder (603). In another example, the video source (601) is part of the electronic device (620).

[0059] The video source (601) can provide a source video sequence to be coded by a video encoder (603) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, ...), any color space (e.g., BT.601 YCrCb, RGB, ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media supply system, the video source (601) may be a storage device that stores previously prepared video. In a videoconference system, the video source (601) may be a camera that captures local image information as a video sequence. The video data can be provided as a plurality of individual pictures that give motion when viewed sequentially. Each picture itself can be organized as a spatial array of pixels, and each pixel can comprise one or more samples depending on the sampling structure, color space, etc. in use. A person skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0060] According to one embodiment, the video encoder (603) can code and compress pictures of the source video sequence into a coded video sequence (643) in real time or under any other time constraints as required. Enforcing an appropriate coding speed is one function of the controller (650). In some embodiments, the controller (650) controls other functional units and is functionally coupled to other functional units as described below. This coupling is not shown for clarity. The parameters set by the controller (650) can include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization techniques, ...), picture size, picture group (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other suitable functions related to the video encoder (603) optimized for a particular system design.

[0061] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As an overly simplified explanation, in one example, the coding loop may include a source coder (630) (which is responsible for creating symbols such as a symbol stream based on, for example, an input picture to be coded and one or more reference pictures) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in the same way as a (remote) decoder does. The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream results in a bit-exact result regardless of the position of the decoder (local or remote), the content in the reference picture memory (634) is also bit-exact between the local encoder and the remote encoder. In other words, the prediction unit of the encoder "sees" the same sample values as the sample values that the decoder will "see" when using prediction during decoding as the reference picture samples. This basic principle of the synchronization of the reference pictures (and the resulting drift if the synchronization cannot be maintained, for example, due to channel errors) is also used in some related technologies.

[0062] The operation of the "local" decoder (633) may be the same as that of a "remote" decoder such as the video decoder (510) already described in detail in relation to FIG. 5. However, referring briefly to FIG. 5 as well, since the symbols are available and the coding / decoding of the symbols into the coded video sequence by the entropy encoder (645) and the parser (520) can be lossless, the entropy decoding part of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully implemented in the local decoder (633).

[0063] In one embodiment, decoder techniques other than syntax analysis / entropy decoding present in the decoder exist in the corresponding encoder in the same or substantially the same functional form. Accordingly, the disclosed subject matter focuses on decoder operations. The description of encoder techniques may be omitted since it is the reverse of the decoder techniques described comprehensively. In some areas, more detailed descriptions are provided below.

[0064] During operation, in some examples, the source coder (630) can perform motion-compensated predictive coding. This is to predictively code an input picture by referring to one or more previously coded pictures from a video sequence designated as a "reference picture". In this way, the coding engine (632) codes the difference between a pixel block of the input picture and a pixel block of a reference picture that can be selected as a prediction reference for the input picture.

[0065] The local video decoder (633) can decode the coded video data of a picture that can be designated as a reference picture based on the symbols created by the source coder (630). The operation of the coding engine (632) may advantageously be a lossy process. When the coded video data can be decoded in a video decoder (not shown in FIG. 6), the reconstructed video sequence can generally be a replica of the source video sequence with some errors. The local video decoder (633) can replicate the decoding process that can be performed on the reference picture by the video decoder and store the reconstructed reference picture in the reference picture memory (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference picture having the same content as the reconstructed reference picture obtained by the remote video decoder (in the absence of transmission errors).

[0066] Predictor (635) can perform predictive search for the coding engine (632). That is, for a new picture to be coded, Predictor (635) can find sample data (as candidate reference pixel blocks), or specific metadata such as reference picture motion vectors, block shapes, etc. that can function as appropriate predictive references for the new picture, and search the reference picture memory (634). Predictor (635) can operate based on block-by-sample and block-by-pixel to find appropriate predictive references. In some cases, the input picture may have predictive references drawn from a plurality of reference pictures stored in the reference picture memory (634) as determined by the search results obtained by Predictor (635).

[0067] Controller (650) can manage the coding operation of source coder (630), including setting parameters and subgroup parameters used for encoding video data, for example.

[0068] The outputs of all the above-described functional units are entropy-coded by entropy coder 645. Entropy coder (645) converts the symbols generated by various functional units into a coded video sequence by applying reversible (lossless) compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0069] The transmitter (640) may buffer the coded video sequence created by the entropy coder (645) for transmission via the communication channel (660), which may be a hardware / software link to a storage device that stores the coded video data. The transmitter (640) may merge the coded video data from the video encoder (603) with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (source not shown).

[0070] The controller (650) can manage the operation of the video encoder (603). During coding, the controller (650) may assign to each coded picture a specific coded picture type that can affect the coding technique applicable to that picture. For example, a picture may often be assigned as one of the following picture types.

[0071] An intra picture (I picture) can be coded and decoded without using other pictures in the sequence as a source of prediction. Some video codecs allow for different types of intra pictures, including for example, an instantaneous decoder refresh (「IDR」) picture. Those skilled in the art are aware of these variations of I pictures, as well as their respective uses and characteristics.

[0072] A predicted picture (P picture) can be coded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.

[0073] A bi-directional predicted picture (B picture) can be coded and decoded using intra prediction or inter prediction that uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0074] A source picture can generally be spatially re-divided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and coded block by block. The blocks can be predictedly coded by referring to other (already coded) blocks as determined by the coding assignment applied to each respective picture of the block. For example, blocks of an I picture may be coded non-predictively or may be predictedly coded by referring to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be predictedly coded via spatial prediction or via temporal prediction by referring to one previously coded reference picture. Blocks of a B picture can be predictedly coded via spatial prediction or via temporal prediction by referring to one or two previously coded reference pictures.

[0075] A video encoder (603) can perform coding operations according to a predetermined video coding technology or standard such as ITU-T Rec. H.265. In its operation, the video encoder (603) can perform various compression operations, including predictive coding operations that utilize temporal redundancy and spatial redundancy in the input video sequence. Thus, the coded video data can conform to the syntax specified by the video coding technology or standard being used.

[0076] In one embodiment, the transmitter (640) can transmit additional data along with the encoded video. The source coder (630) can include such data as part of the coded video sequence. The additional data can include, for example, temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, and the like.

[0077] Video can be captured as a plurality of source pictures (video pictures) in a time series. Intra-picture prediction (often abbreviated as intra prediction) utilizes the spatial correlation within a given picture, and inter-picture prediction utilizes the (temporal or other) correlation between pictures. In one example, the picture being encoded / decoded, called the current picture, is divided into blocks. When a block in the current picture is similar to a reference block in a reference picture that has been previously coded and is still buffered in the video, the block in the current picture can be coded by a vector called a motion vector. The motion vector points to the reference block within the reference picture and can have a third dimension that identifies the reference picture when multiple reference pictures are being used.

[0078] In some embodiments, dual prediction techniques can be used in inter-picture prediction. According to the dual prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture that are both earlier in decoding order than the current picture in the video (although they can be in the past and future respectively in display order). A block within the current picture can be coded by a first motion vector that points to a first reference block within the first reference picture and a second motion vector that points to a second reference block within the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.

[0079] In addition, the coding efficiency can be improved by using the merge mode technology for inter-picture prediction.

[0080] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are divided into coding tree units (CTUs) for compression, and CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luma CTB and two chroma CTBs. Each CTU can be recursively quad-tree divided into one or more coding units (CUs). For example, a 64×64 pixel CTU can be divided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine the prediction type of the CU, such as an inter prediction type or an intra prediction type. The CU is divided into one or more prediction units (PUs) according to the temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation during coding (encoding / decoding) is performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luma values) such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels.

[0081] FIG. 7 shows an exemplary diagram of a video encoder (703). The video encoder (703) receives a processing block (e.g., a prediction block) of sample values within the current video picture of a sequence of video pictures and is configured to encode the processing block into a coded picture that is part of a coded video sequence. In one example, the video encoder (703) is used instead of the video encoder (403) in the example of FIG. 4.

[0082] In an example of HEVC, the video encoder (703) receives a matrix of sample values of a processing block such as an 8×8 sample prediction block. The video encoder (703) determines whether the processing block is best coded using an intra mode, an inter mode, or a bi-prediction mode, for example using rate-distortion optimization. If the processing block is coded in the intra mode, the video encoder (703) may encode the processing block into a coded picture using intra prediction techniques, and if the processing block is coded in the inter mode or the bi-prediction mode, the video encoder (703) may encode the processing block into a coded picture using inter prediction or bi-prediction techniques, respectively. In some video coding techniques, the merge mode may be an inter-picture prediction sub-mode, and the motion vector is derived from one or more motion vector predictors without benefit of the coded motion vector components outside the predictor. In some other video coding techniques, there may be motion vector components applicable to the target block. In one example, the video encoder (703) includes other components such as a mode decision module (not shown) that determines the mode of the processing block.

[0083] In the example of FIG. 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general-purpose controller (721), and an entropy encoder (725) that are coupled to each other as shown in FIG. 7.

[0084] The inter-encoder (730) receives samples of the current block (e.g., a processing block), compares the block with one or more reference blocks (e.g., blocks in the previous and subsequent pictures) in the reference picture, generates inter-prediction information (e.g., a description of redundant information by an inter-coding technique, a motion vector, merge mode information), and is configured to calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture that is decoded based on the encoded video information.

[0085] The intra-encoder (722) receives samples of the current block (e.g., a processing block), optionally compares the block with blocks that have already been coded within the same picture, generates quantized coefficients after transformation, and optionally also generates intra-prediction information (e.g., intra-prediction direction information by one or more intra-coding techniques). In one example, the intra-encoder (722) also calculates an intra-prediction result (e.g., a predicted block) based on the intra-prediction information and reference blocks within the same picture.

[0086] The general-purpose controller (721) is configured to determine general-purpose control data and control other components of the video encoder (703) based on the general-purpose control data. In one example, the general-purpose controller (721) determines the mode of a block and provides a control signal to the switch (726) based on that mode. For example, when the mode is the intra mode, the general-purpose controller (721) controls the switch (726) to select the intra mode result used by the residual calculator (723), and controls the entropy encoder (725) to select the intra prediction information and include it in the bitstream. When the mode is the inter mode, the general-purpose controller (721) controls the switch (726) to select the inter prediction result used by the residual calculator (723), and controls the entropy encoder (725) to select the inter prediction information and include it in the bitstream.

[0087] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) is configured to operate based on the residual data to encode the residual data and generate transformation coefficients. In one example, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain and generate transformation coefficients. The transformation coefficients then undergo quantization processing to obtain quantized transformation coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform inverse transformation and generate decoded residual data. The decoded residual data can be suitably utilized by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and intra prediction information. The decoded blocks are appropriately processed to generate a decoded picture, and the decoded picture is buffered in a memory circuit (not shown) and can be used as a reference picture in some examples.

[0088] The entropy encoder (725) is configured to format the bitstream to include the coded blocks. The entropy encoder (725) is configured to include various information in the bitstream according to a suitable standard such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other suitable information in the bitstream. Note that according to the disclosed subject matter, there is no residual information when coding a block in either the merge submode of the inter mode or the bi-prediction mode.

[0089] FIG. 8 shows an exemplary diagram of a video decoder (810). The video decoder (810) is configured to receive a coded picture that is part of a coded video sequence and decode the coded picture to generate a reconstructed picture. In one example, the video decoder (810) is used in place of the video decoder (410) in the example of FIG. 4.

[0090] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) that are coupled to each other as shown in FIG. 8.

[0091] The entropy decoder (871) may be configured to reconstruct from the coded picture specific symbols that represent syntax elements that make up the coded picture. Such symbols can include, for example, the mode in which a block is coded (e.g., intra mode, inter mode, bi-prediction mode, the latter two in the merge sub-mode, or another sub-mode, etc.) and prediction information (e.g., intra prediction information or inter prediction information, etc.) that can respectively identify specific samples or metadata used for prediction by the intra decoder (872) or the inter decoder (880). The symbols can also include residual information in the form of, for example, quantized transform coefficients. In one example, when the prediction mode is an inter prediction mode or a bi-prediction mode, the inter prediction information is provided to the inter decoder (880), and when the prediction type is an intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information can undergo inverse quantization and is provided to the residual decoder (873).

[0092] The inter decoder (880) is configured to receive inter prediction information and generate an inter prediction result based on the inter prediction information.

[0093] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0094] The residual decoder (873) is configured to perform inverse quantization to extract inverse quantized transform coefficients, and process the inverse quantized transform coefficients to convert the residual information from the frequency domain to the spatial domain. The residual decoder (873) may require certain control information (since it includes quantization parameter (QP)), and that information may be provided by the entropy decoder (871) (since it may be only low volume control information, the data path is not shown).

[0095] The reconstruction module (874) is configured to combine, in the spatial domain, the residual information output by the residual decoder (873) and the prediction result (which may be output by an inter prediction module or an intra prediction module in some cases) to form a reconstructed block that may be part of a reconstructed picture, and the reconstructed picture may be part of a reconstructed video. Note that other appropriate operations such as a deblocking operation may be performed to improve the visual quality.

[0096] Note that the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) may be implemented using any suitable technology. In one embodiment, the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (603), and the video decoders (410), (510), and (810) may be implemented using one or more processors that execute software instructions.

[0097] In VVC, various inter prediction modes can be used. For an inter-predicted CU, the motion parameters can include an MV, one or more reference picture indexes, a reference picture list use index, and additional information about some coding features to be used for generating the inter-predicted samples. The motion parameters can be signaled explicitly or implicitly. When a CU is coded in skip mode, the CU may be associated with a PU and may not have significant residual coefficients, a coded motion vector delta or MV difference (e.g., MVD), or a reference picture index. The merge mode can be specified, where the motion parameters of the current CU are obtained from one or more adjacent CUs, including spatial candidates and / or temporal candidates, and optionally additional information introduced in VVC. The merge mode can be applied not only to skip mode but also to inter-predicted CUs. In one example, an alternative to the merge mode is the explicit transmission of motion parameters, where the MV(s), the corresponding reference picture index for each reference picture list, and the reference picture list use flag, and other information are signaled explicitly for each CU.

[0098] In one embodiment such as VVC, the VVC Test Model (VTM) reference software includes one or more refined inter prediction coding tools including extended merge prediction, merge motion vector difference (MMVD) mode, adaptive motion vector prediction (AMVP) mode using symmetric MVD signaling, affine motion compensation prediction, sub-block based temporal motion vector prediction (SbTMVP), adaptive motion vector resolution (AMVR), motion field storage (1 / 16 luma sample MV storage and 8×8 motion field compression), bi-prediction using CU level weights (BCW), bidirectional optical flow (BDOF), prediction refinement using optical flow (PROF), decoder side motion vector refinement (DMVR), combined inter and intra prediction (CIIP), geometric partitioning mode (GPM), etc. Inter prediction and related methods will be described in detail below.

[0099] Extended merge prediction can be used in some examples. In one example such as VTM4, the merge candidate list is constructed by sequentially including the following five types of candidates: namely, spatial motion vector predictors (MVPs) from spatially adjacent CUs, temporal MVPs from collocated CUs, history-based MVPs (HMVP) from a first-in first-out (FIFO) table, pairwise average MVPs, and zero MV.

[0100] The size of the merge candidate list can be signaled in the slice header. In one example, the maximum allowable size of the merge candidate list is 6 in VTM4. For each CU coded in merge mode, the index of the best merge candidate (e.g., merge index) can be encoded using truncated unary binarization (TU). The first bin of the merge index can be coded using context (e.g., context adaptive binary arithmetic coding (CABAC)), and bypass coding can be used for the other bins.

[0101] Some examples of the generation process for each category of merge candidates are given below. In one embodiment, the spatial candidates are derived as follows. The derivation of spatial merge candidates in VVC can be the same as that in HEVC. As an example, among the candidates at the positions shown in FIG. 9, up to four merge candidates are selected. FIG. 9 shows the positions of spatial merge candidates according to an embodiment of the present disclosure. Referring to FIG. 9, the order of derivation is B1, A1, B0, A0, and B2. The position B2 is considered only when any of the CUs at positions A0, B0, B1, and A1 are not available (for example, because the CU belongs to another slice or another tile) or is intra-coded. After the candidate at position A1 is added, the addition of the remaining candidates is subject to a redundancy check. This ensures that candidates having the same motion information are excluded from the candidate list, thereby improving the coding efficiency.

[0102] To reduce the computational complexity, not all possible candidate pairs are considered in the above redundancy check. Instead, only the pairs linked by the arrows in FIG. 10 are considered, and if the corresponding candidates used for the redundancy check do not have the same motion information, the candidate is simply added to the candidate list. FIG. 10 shows the candidate pairs considered for the redundancy check of spatial merge candidates according to an embodiment of the present disclosure. Referring to FIG. 10, the pairs connected by arrows are A1 and B1, A1 and A0, A1 and B2, B1 and B0, B1 and B2. Thereby, the candidates at positions B1, A0, and / or B2 can be compared with the candidate at position A1, and the candidates at positions B0 and / or B2 can be compared with the candidate at position B1.

[0103] In one embodiment, the temporal candidates are derived as follows. In one example, only one temporal merge candidate is added to the candidate list. FIG. 11 shows an exemplary motion vector scaling for temporal merge candidates. To derive the temporal merge candidate for the current CU (1111) in the current picture (1101), a scaled MV (1121) (e.g., indicated by the dotted line in FIG. 11) can be derived based on the co-located CU (1112) belonging to the same position reference picture (1104). In one example, the same position reference picture (also called the same position picture) is a specific reference picture used, for example, for temporal motion vector prediction. The same position reference picture used for temporal motion vector prediction can be indicated by a reference index in the syntax such as high level syntax (e.g., picture header, slice header).

[0104] The reference picture list used to derive the co-located CU (1112) can be explicitly signaled in the slice header. The scaled MV (1121) for the temporal merge candidate can be obtained as indicated by the dotted line in FIG. 11. The scaled MV (1121) can be scaled from the MV of the co-located CU (1112) using the picture order count (POC) distances tb and td. The POC distance tb can be defined as the POC difference between the current reference picture (1102) of the current picture (1101) and the current picture (1101). The POC distance td can be defined as the POC difference between the co-located reference picture (1104) of the co-located picture (1103) and the co-located picture (1103). The reference picture index of the temporal merge candidate can be set to zero.

[0105] FIG. 12 shows exemplary candidate positions (e.g., C0 and C1) for the current CU's temporal merge candidates. The position of the temporal merge candidate can be selected from the candidate positions C0 and C1. Candidate position C0 is located at the lower right corner of the co-located CU (1210) at the same position as the current CU. Candidate position C1 is located at the center of the CU (1210) at the same position as the current CU. If the CU at candidate position C0 is not available, is intra-coded, or is outside the current row of the CTU, candidate position C1 is used to derive the temporal merge candidate. Otherwise, for example, if the CU at candidate position C0 is available, is intra-coded, is in the current row of the CTU, candidate position C0 is used to derive the temporal merge candidate.

[0106] For the merge mode with the skip mode or the motion vector representation method, a merge with motion vector difference (MMVD) mode can be used. The merge candidates as used in VVC can be reused in the MMVD mode. The candidate can be selected from among the merge candidates as the starting point (e.g., the MV predictor (MVP)) and can be further extended by the MMVD mode. The MMVD mode can provide a new motion vector representation using simplified signaling. The motion vector representation method includes a starting point and an MV difference (MVD). In one example, the MVD is indicated by the magnitude of the MVD (or the magnitude of the motion) and the direction of the MVD (e.g., the direction of the motion).

[0107] The MMVD mode can use a merge candidate list as used in VVC. In one embodiment, only candidates that are of the default merge type (e.g., MRG_TYPE_DEFAULT_N) are considered for the MMVD mode. The starting point can be indicated or defined by a base candidate index (IDX). The base candidate index can indicate a candidate (e.g., the best candidate) among the candidates (e.g., base candidates) in the merge candidate list. Table 1 shows an exemplary relationship between the base candidate index and the corresponding starting point. A base candidate index of 0, 1, 2, or 3 indicates that the corresponding starting point is the 1st MVP, 2nd MVP, 3rd MVP, or 4th MVP. In one example, when the number of (one or more) base candidates is equal to 1, the base candidate IDX is not signaled.

[0108]

Table 1

[0109]

Table 2

[0110]

Table 3

[0111] FIG. 13 and FIG. 14 show an example of the search process in the MMVD mode. By executing the search process, an index including a base candidate index, a direction index, and / or a distance index can be determined for the current block (1300) in the current picture (or called the current frame) (1301).

[0112] A first motion vector (MV) (1311) belonging to the first merge candidate and a second MV (1321) are shown. The first merge candidate can be a merge candidate in the merge candidate list configured for the current block (1300). The first MV (1311) and the second MV (1321) can be associated with two reference pictures (1302) and (1303) in the reference picture lists L0 and L1, respectively. Therefore, the two starting points (1411) and (1421) in FIGS. 13 to 14 can be determined in the reference pictures (1302) and (1303), respectively.

[0113] In one example, based on start points (1411) and (1421), a plurality of predefined points (e.g., 1-12 shown in FIG. 14) extending from the start points (1411) and (1421) in the vertical direction (represented by +Y or -Y) or the horizontal direction (represented by +X and -X) within reference pictures (1302) and (1303) can be evaluated. In one example, pairs of points such as the pair of points (1414) and (1424), or the pair of points (1415) and (1425), which are mirroring each other with respect to their respective start points (1411) or (1421), can be used to determine pairs of MVs (1314) and (1324) or pairs of MVs (1315) and (1325) that form MV predictor (MVP) candidates for the current block (1300). The MVP candidates determined based on predetermined points surrounding the start points (1411) and / or (1421) can be evaluated. Referring to FIG. 13, the MVD (1312) between the first MV (1311) and MV (1314) has a magnitude of 1S. The MVD (1322) between the second MV (1321) and MV (1324) has a magnitude of 1S. Similarly, the MVD between the first MV (1311) and MV (1315) has a magnitude of 2S. The MVD between the second MV (1321) and MV (1325) has a magnitude of 2S.

[0114] In addition to the first merge candidate, other available or valid merge candidates within the merge candidate list of the current block (1300) can be evaluated in the same way. In one example, for a single prediction merge candidate, only one prediction direction associated with one of the two reference picture lists is evaluated.

[0115] In one example, based on the evaluation, the best MVP candidate can be determined. Therefore, the optimal merge candidate corresponding to the optimal MVP candidate can be selected from the merge list, and the movement direction and movement distance can also be determined. For example, based on the selected merge candidate and Table 1, the base candidate index can be determined. Based on the selected MVP corresponding to a predetermined point (1415) (or (1425)), the direction and distance (e.g., 2S) of point (1415) relative to the starting point (1411) can be determined. According to Table 2 and Table 3, the direction index and distance index can be determined as appropriate.

[0116] As described above, two indexes such as a distance index and a direction index can be used to indicate MVD in the MMVD mode. Alternatively, a single index can be used to indicate MVD in the MMVD mode, for example, using a table that pairs the single index with MVD.

[0117] Template matching (TM)-based candidate reordering can be used in some prediction modes such as the MMVD mode and the affine MMVD mode. In one embodiment, the MMVD offset is extended for the MMVD mode and the affine MMVD mode. FIG. 15 shows additional refinement positions along a plurality of diagonals such as k×π / 8 diagonals (where k is an integer from 0 to 15). The additional refinement positions along the plurality of diagonals can increase the number of directions from, for example, four directions (e.g., +X, -X, +Y, and -Y) to 16 directions (e.g., k = 0, 1, 2,..., 15). In one example, each of the 16 directions is represented by the angle between the +X direction and the direction indicated by the center point (1500) and one of points 1-16. For example, point 1 indicates the +X direction with an angle of 0 (i.e., k = 0), point 2 indicates the direction along an angle of 1×π / 8 (i.e., k = 1), and so on.

[0118] TM can be executed in MMVD mode. In one example, for each MMVD refinement position, a TM cost can be determined based on the current template of the current block and one or more reference templates. The TM cost can be determined using any method such as the sum of absolute differences (SAD) (e.g., SAD cost), the sum of absolute transform differences (SATD), the sum of squared errors (SSE), the mean removed SAD / SATD / SSE, variance, partial SAD, partial SSE, partial SATD, or the like.

[0119] The current template of the current block can include any suitable samples, such as one row of samples above the current block and / or one column of samples to the left of the current block. Based on the TM cost (e.g., SAD cost) between the current template and the corresponding reference template for a refinement position, e.g., an MMVD refinement position, all possible MMVD refinement positions (e.g., 16×6 representing 16 directions and 6 sizes) for each base candidate (e.g., MVP) can be reordered. In one example, the top MMVD refinement position having the minimum TM cost (e.g., minimum SAD cost) is retained as the available MMVD refinement position for MMVD index coding. For example, a subset (e.g., 8) of the MMVD refinement positions having the minimum TM cost is used for MMVD index coding. For example, the MMVD index indicates which of the subset of MMVD refinement positions having the minimum TM cost is selected to code the current block. In one example, an MMVD index of 0 indicates that the MVD (e.g., MMVD refinement position) corresponding to the minimum TM cost is used to code the current block. The MMVD index can be binarized, for example, by a Rice code with a parameter of 2.

[0120] In one embodiment, in addition to the above-described MMVD offset extension, as shown in FIG. 15, the affine MMVD rearrangement is extended, and additional refinement positions along k×π / 4 diagonals are added. After the rearrangement, the top 1 / 2 of the refinement positions having the minimum TM cost (e.g., SAD cost) are retained for coding the current block.

[0121] In some embodiments, block-based affine transform motion compensation prediction is applied. In FIG. 16A, the affine motion field of a block is described by two control point motion vectors (CPMVs), CPMV0 and CPMV1, of two control points (CPs), CP0 and CP1, when a 4-parameter affine model is used. In FIG. 16(b), when a 6-parameter affine model is used, the affine motion field of the block is described by three CPMVs, CPMV0, CPMV1, and CPMV3, of CP0, CP1, and CP2.

[0122] For the 4-parameter affine motion model, the motion vector at the sample position (x, y) within the block is derived as follows.

[0123]

Equation

[0124]

Equation

[0125] To simplify motion compensation prediction, in some embodiments, sub-block based affine transform prediction is applied. For example, in FIG. 17, a 4-parameter affine motion model is used, and two CPMVs, (External 1) JPEG2025517843000007.jpg12119 and (External 2) JPEG2025517843000008.jpg14119 are determined. To derive the motion vector of each 4×4 (sample) luma sub-block (1702) divided from the current block (1710), the motion vector (1701) of the central sample of each sub-block (1702) is calculated according to Equation 3 and rounded to 1 / 16 fractional precision. Then, a motion compensation interpolation filter is applied to generate a prediction for each sub-block (1702) using the derived motion vector (1701). The sub-block size of the chroma component is set to 4×4. The MV of a 4×4 chroma sub-block is calculated as the average of the MVs of four corresponding 4×4 luma sub-blocks.

[0126] Similar to translational motion inter prediction, in some embodiments, two affine motion inter prediction modes including an affine merge mode and an affine AMVP mode are adopted. The affine skip mode is similar to the affine merge mode, but differs in that in the affine merge mode, a residual can be added to the predicted sample, while in the affine skip mode, no residual is added to the predicted sample. In the affine skip mode, no significant residual coefficients are used.

[0127] In some embodiments, the affine merge mode can be applied to a CU where both the width and height are 8 or more. The affine merge candidates for the current CU can be generated based on the motion information of spatially adjacent CUs. There can be up to 5 affine merge candidates, and an index indicating which one should be used for the current CU is signaled. For example, the following three types of affine merge candidates are used to create an affine merge candidate list.

[0128] Inherited affine merge candidates extrapolated from the CPMV of the adjacent CU Composed affine merge candidates derived using the translational MV of the adjacent CU. Zero MV.

[0129] In some embodiments, there can be at most two inherited affine candidates derived from the affine motion models of adjacent blocks, i.e., one from the left adjacent CU and one from the upper adjacent CU. The candidate blocks can be placed, for example, at the positions shown in FIG. 9. For the left predictor, the scan order is A0 > A1, and for the upper predictor, the scan order is B0 > B1 > B2. Only the first inherited candidate from each side is selected. Between the two inherited candidates, no pruning check is performed.

[0130] When an adjacent affine CU is identified, the CPMV of the identified adjacent affine CU is used to derive the CPMV candidates in the affine merge list of the current CU. As shown in FIG. 18, the adjacent lower left block A of the current CU (1810) is coded in affine mode. The motion vectors of the upper left corner, upper right corner, and lower left corner of the CU (1820) containing block A (Outer 3) JPEG2025517843000009.jpg8119,[[]] (Outer 4) JPEG2025517843000010.jpg10119 and (Outer 5) JPEG2025517843000011.jpg9119 are obtained. When block A is encoded with a 4-parameter affine model, the two CPMVs of the current CU (1810), (Outer 6) JPEG2025517843000012.jpg9119,[[]] (Outer 7) JPEG2025517843000013.jpg9119 are,[[]] (Outer 8) JPEG2025517843000014.jpg9119 and (Outer 9) It is calculated according to JPEG2025517843000015.jpg8119. When block A is coded with a 6-parameter affine model, three CPMVs (not shown) of the current CU are (Outer 10) JPEG2025517843000016.jpg9119, (Outer 11) JPEG2025517843000017.jpg8119 and (Outer 12) JPEG2025517843000018.jpg9119 are calculated according to.

[0131] The constructed affine candidates are formed by combining the adjacent translational motion information of each control point. The motion information of the control points is derived from the specific spatial neighbors and temporal neighbors shown in Figure 19. CPMVk (k = 1, 2, 3, 4) represents the k-th control point. For CPMV1, the B2 > B3 > A2 blocks are checked in order, and the MV of the first available block is used. For CPMV2, the B1 > B0 blocks are checked, and for CPMV3, the A1 > A0 blocks are checked. The TMVP in block T is used as CPMV4 if available.

[0132] After obtaining the MVs of the four control points, affine merge candidates are constructed based on their motion information. The following combinations of control point MVs are used to construct in order: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}.

[0133] The combination of three CPMVs constitutes a 6-parameter affine merge candidate, and the combination of two CPMVs constitutes a 4-parameter affine merge candidate. To avoid the motion scaling process, when the reference indices of the control points are different, the related combinations of control point MVs are discarded.

[0134] After the inherited affinity merge candidates and the configured affinity merge candidates are checked, if the list is not yet full, zero MVs are inserted at the end of the merge candidate list.

[0135] To improve coding efficiency and reduce the transmission overhead of MVs, sub-block level MV refinement can be applied to extend the CU level temporal motion vector prediction (TMVP). In one example, the sub-block based TMVP (SbTMVP) mode enables inheriting motion information at the sub-block level from a reference picture at the same position. As described above, the reference picture at the same position can be indicated by a reference index in the syntax such as a high-level syntax (e.g., picture header, slice header).

[0136] Each sub-block of the current CU in the current picture (e.g., the current CU having a large size) can have its respective motion information without explicitly transmitting the block partition structure or the respective motion information. In the SbTMVP mode, the motion information of each sub-block can be obtained as follows in, for example, three steps. In the first step, the displacement vector (DV) of the current CU can be derived. The DV can indicate a block in the reference picture at the same position. For example, the DV points to a block in the reference picture collocated with the current block in the current picture. Thus, the block indicated by the DV is considered to be at the same position as the current block and is called the collocated block of the current block. In the second step, the availability of SbTMVP candidates can be checked and the central motion (e.g., the central motion of the current CU) can be derived. In the third step, the sub-block motion information can be derived from the corresponding sub-block in the collocated block using the DV. The three steps can be combined into one or two steps, and / or the order of the three steps can be adjusted.

[0137] Unlike TMVP candidate derivation that derives temporal MVs from the same-position blocks in a reference frame or reference picture, in the SbTMVP mode, a DV (e.g., a DV derived from the MV of the left adjacent CU of the current CU) is applied to locate the corresponding sub-blocks in the same-position reference picture for each sub-block in the current CU in the current picture.

[0138] If the corresponding sub-blocks are not inter-coded, the motion information of the current sub-block can be set to the center motion of the same-position blocks.

[0139] The SbTMVP mode can be supported by various video coding standards including, for example, VVC. Similar to the TMVP mode, in HEVC, for example, in the SbTMVP mode, the motion field (also called the motion information field or MV field) in the reference picture at the same position can be used to improve the MV prediction and merge mode for the CUs in the current picture. In one example, the same same-position reference picture used by the TMVP mode is used in the SbTMVP mode. In one example, the SbTMVP mode differs from the TMVP mode in the following aspects: (i) the TMVP mode predicts motion information at the CU level, while the SbTMVP mode predicts motion information at the sub-CU level; (ii) the TMVP mode fetches temporal MVs from the same-position blocks in the same-position reference picture (e.g., the same-position block is the bottom-right or central block with respect to the current CU), while the SbTMVP mode can apply a motion shift before fetching temporal motion information from the same-position reference picture. In one example, the motion shift used in the SbTMVP mode is obtained from the MV of one of the spatial adjacent blocks of the current CU.

[0140] Figures 20-21 illustrate an exemplary SbTMVP process used in SbTMVP mode. The SbTMVP process can predict the motion vector (MV) of a sub-CU (e.g., sub-block) within a current CU (e.g., current block) (2001) in a current picture (2111) in, for example, two steps. In the first step, the spatial neighbors (e.g., A1) of the current block (2001) in FIGS. 20-21 are examined. If a spatial neighbor (e.g., A1) has an MV (2121) that uses the same-position reference picture (2112) as the reference picture of the spatial neighbor (e.g., A1), the MV (2121) can be selected to be the motion shift (or DV) to be applied to the current block (2001). If no such MV (e.g., an MV that uses the same-position reference picture (2112) as the reference picture) is identified, the motion shift or DV can be set to a zero MV (e.g., (0,0)). In some examples, the MV(s) of additional spatial neighbors such as A0, B0, B1, etc. are checked if such an MV is not identified for spatial neighbor A1.

[0141] In a second step, apply the motion shift or DV (2121) identified in the first step to the current block (2001) (e.g., by adding DV (2121) to the coordinates of the current block) to obtain motion information at the sub-CU level (e.g., including MV and reference index) from a reference picture (2112) at the same position. In the example shown in FIG. 21, the motion shift or DV (2121) is set to be the MV of the spatial neighbor A1 (e.g., block A1) of the current block (2001). For each sub-CU or sub-block (2131) within the current block (2001), the motion information of the corresponding co-located block (2101) within the co-located reference picture (2112) (e.g., the motion information of the smallest motion grid covering the center sample of the co-located block (2101)) can be used to derive the motion information of the sub-CU or sub-block (2131). After the motion information of the co-located sub-CU (2132) within the co-located block (2101) is identified, the motion information of the co-located sub-CU (2132) can be converted to the motion information (e.g., MV and one or more associated reference indices) of the current sub-CU (2131) using a scaling method such as a method similar to the TMVP process used in HEVC, where temporal motion scaling is applied to align the reference picture of the temporal MV with the reference picture of the current CU.

[0142] The motion field of the current block (2001) derived based on DV (2121) can include the motion information of each sub-block (2131) within the current block (2001), such as MV(s) and one or more associated reference indices. The motion field of the current block (2001) is also referred to as an SbTMVP candidate and corresponds to DV (2121).

[0143] FIG. 21 shows an example of the motion field of the current block 2001 or an SbTMVP candidate (also referred to as an SbTMVP merge candidate). The motion information of the bi-predicted sub-block (2131(1)) includes a first MV, a first index indicating the first reference picture in the reference picture list 0 (L0), a second MV, and a second index indicating the second reference picture in the reference picture list 1 (L1). In one example, the motion information of the uni-predicted sub-block (2131(2)) includes an MV and an index indicating a reference picture in L0 or L1.

[0144] In one example, DV (2121) is applied to the center position of the current block (2001) to locate the displaced center position within the same-position reference picture (2112). If the block including the displaced center position is not inter-coded, the SbTMVP candidate is considered unavailable. Otherwise, if the block including the displaced center position (e.g., the same-position block (2101)) is inter-coded, the motion information of the center position of the current block (2001), called the center motion of the current block (2001), can be derived from the motion information of the block including the displaced center position within the same-position reference picture (2112). In one example, a scaling process can be used to derive the central motion of the current block (2001) from the motion information of the block including the displaced center position within the same-positioned reference picture (2112). When the SbTMVP candidate is available, DV (2121) can be applied to find the corresponding sub-block (2132) within the same-position reference picture (2112) for each sub-block (2131) of the current block (2001). The motion information of the corresponding sub-block (2132) can be used to derive the motion information of the sub-block (2131) within the current block (2001) in the same way as used to derive the center motion of the current block (2001). In one example, if the corresponding sub-block (2132) is not inter-coded, the motion information of the current sub-block (2131) is set to be the center motion of the current block (2001).

[0145] In some examples, such as VVC, a combined sub-block-based merge list that includes SbTMVP candidates and (one or more) affine merge candidates is used in the signaling of the sub-block-based merge mode. The SbTMVP mode is enabled or disabled by a sequence parameter set (SPS) flag. When the SbTMVP mode is enabled, the SbTMVP candidate (or SbTMVP predictor) is added as the first entry of the sub-block-based merge list that includes sub-block-based merge candidates, and then the affine merge candidates can follow. The size of the sub-block-based merge list can be signaled in the SPS. In one example, the maximum allowable size of the sub-block-based merge list is 5 in VVC. In one example, multiple SbTMVP candidates are included in the sub-block-based merge list.

[0146] In some examples, such as VVC, the sub-CU size used in the SbTMVP mode is fixed to 8×8 as used in the affine merge mode. In one example, the SbTMVP mode is applicable only to CUs where both the width and height are 8 or more. The sub-block size (e.g., 8×8) can be set to other sizes, such as 4×4, in the ECM software model used for exploration beyond VVC. In one example, multiple same-position reference pictures, such as two same-position frames, are utilized to provide temporal motion information for SbTMVP and / or TMVP in the AMVP mode.

[0147] In the SbTMVP mode, for example, in order to generate a better match, an offset vector such as a displacement vector offset (DVO) can be used. The DVO can be an offset with respect to the DV used in the SbTMVP mode (for example, the initial DV which is the MV of an adjacent block of the current CU). In one example, the DVO is added to the DV to determine the updated DV'. In some examples, the DVO is called a motion vector offset (MVO). In one example, the updated DV' is the vector sum of the DV and the DVO. By using the DVO, the position of the same-position CU corresponding to the current CU can be adjusted, and thereby the MV field within the same-position CU can be adjusted. When the DVO is not 0, the updated DV' is used as the displacement vector to indicate the position of the CU at the same position for deriving the SbTMVP candidate (or SbTMVP merge candidate). Referring to FIG. 21, the updated DV' can be used as the DV (2121) to obtain the SbTMVP candidate.

[0148] This application describes the derivation of SbTMVP candidates (or SbTMVP merge candidates) in the SbTMVP mode using a plurality of displacement motion vectors (for example, DV predictors) having a DVO (or motion vector offset). An index is signaled to indicate which displacement motion vector is used as the initial displacement motion vector for deriving the SbTMVP merge candidate. In one example, the DVO is determined using the MMVD mode.

[0149] According to embodiments of the present disclosure, a plurality of displacement vector predictors (DVPs) and a DVO can be used to determine an updated displacement vector (or updated DV') of a current CU (e.g., a current block) to be coded in the SbTMVP mode. The plurality of DVPs may also be referred to as DVP candidates. For a current block, index information indicating a base index and DV offset information (or DVO information) indicating a DVO are signaled and can be parsed to determine the updated DV'. The base index can indicate which DVP is selected from the plurality of DVPs. The DVO can be used as an offset to the selected DVP. The position of the same-position CU (e.g., the same-position block) of the current CU (e.g., the current block) can be determined by using the selected DVP indicated by the base index and the DVO indicated by the DVO information. For example, the updated DV' is determined as the vector sum of the selected DVP and the DVO. The updated DV' indicates the position of the block at the same position and can thus be used as a displacement vector to determine an SbTMVP merge candidate.

[0150] The DVO can be indicated, for example, by signaling an index indicating the DVO from DVO candidates, where the DVO information includes the index. A predetermined DVO list can include DVO candidates. One or more indexes can be signaled to indicate which DVO within the DVO candidates can be selected as the DVO, where the DVO information can include one or more indexes.

[0151] In one example, the DVO is signaled using the MMVD mode. For example, the DVO is an MVD indicated by a direction index and / or a distance index as described in Tables 2 to 3. For example, two indexes including a distance index indicating the magnitude of the DVO and a direction index indicating the direction of the DVO are signaled to indicate the DVO as described in Tables 2 to 3.

[0152] In one embodiment, the DVO is directly signaled using any signaling method used to signal the MVD, such as, for example, the AMVP mode, the AMVR mode, etc. The DVO can be signaled at different resolutions, such as 1 / 4, 1 / 2, 1, or 4 luma sample resolutions, using the AMVR mode.

[0153] In one embodiment, one or more of the plurality of DVP candidates can be determined based on the spatial candidates (e.g., MVs) of each spatial neighbor of the current block. The spatial neighbors can include the spatial neighbors (or candidate blocks) A0, A1, B0, B1, and B2 in FIG. 9. The spatial neighbors can include the spatial neighbors A0, A1, B0, and B1 in FIG. 20. The spatial neighbors of the current block can be checked in any suitable order as shown in FIGS. 9 and 20. In one example, the reference picture of the spatial neighbor of the current block used to determine the DVP candidate among the plurality of DVP candidates is the same position reference picture.

[0154] In one embodiment, when the number of available spatial candidates of each spatial neighbor is less than a threshold, the plurality of DVP candidates are filled with zero vectors (e.g., (0, 0)). In one example, each of the available spatial candidates is used to determine each DVP candidate among the plurality of DVP candidates. For example, if the number of DVP candidates determined from the available spatial candidates is less than the threshold, zero vectors are added to the plurality of DVP candidates. The spatial neighbors corresponding to the available spatial candidates can be inter-predicted. In one example, the reference picture associated with each of the available spatial candidates is the same position reference picture.

[0155] Referring to FIG. 20, in one example, spatial neighbors A1, B0, and B1 are inter-predicted. A1 is singly predicted using MV1 associated with reference picture 1. B0 is singly predicted using MV2 associated with reference picture 2. B1 is bi-predicted using MV3 to MV4 respectively associated with reference pictures 3 to 4. Reference picture 1 and reference picture 4 are same-position reference pictures, and reference picture 1, reference picture 4, and the same-position reference pictures are the same picture. Reference pictures 2 to 3 are different from the same-position reference pictures. In this example, the number of spatial neighbors (e.g., A0, A1, B0, and B1) is 4. The available spatial candidates are candidates related to A1 and B1. The plurality of DVP candidates includes a first DVP determined based on MV1 and a second DVP determined based on MV4. In one example, the first DVP is MV1 and the second DVP is MV4.

[0156] In one example, when the threshold is 3, the plurality of DVP candidates further includes a zero vector.

[0157] In one embodiment, the plurality of DVP candidates may be derived from MVs or candidates in the merge candidate list. In one example, the merge candidate list is a normal merge candidate list such as a normal merge / skip candidate list. The normal merge candidate list may be different from a sub-block based merge list or a sub-block based merge candidate list. The candidate(s) in the normal merge candidate list can include any suitable candidate(s) used in the normal merge / skip mode. The candidate can include a spatial candidate (e.g., a spatial MVP from a spatially adjacent CU), a temporal candidate (e.g., a TMVP from a same-position CU), an HMVP candidate, a pairwise average candidate (e.g., a pairwise average MVP), and / or a zero MV. The pairwise average MVP can be generated using two existing candidates in the normal merge candidate list. The normal merge / skip mode may be different from additional merge / skip modes such as the MMVD mode, the CIIP mode, and the GPM mode.

[0158] In one example, a plurality of DVP candidates are derived from a subset of candidates in the merge candidate list, where the subset of candidates does not include (one or more) temporal candidates (e.g., (one or more) TMVP).

[0159] In one embodiment, a default MV including but not limited to an MV pointing to one or more positions of the same-position block in the same-position reference picture from the upper left corner of the current block may be included in the plurality of DVP candidates. One or more positions of the same-position block can include the central position, lower right position, right position, left position, lower left position, and / or upper right position of the same-position block in the same-position reference picture.

[0160] The sub-block based merge list (or sub-block based merge candidate list) of the current block can include sub-block based merge candidates or sub-block merge candidates. In one example, the sub-block based merge candidates include a plurality of SbTMVP merge candidates. Each of the plurality of SbTMVP merge candidates can be determined based on one of the plurality of DVP candidates. For example, the updated DV' is determined based on one of the plurality of DVP candidates and DVO. The corresponding SbTMVP merge candidate (e.g., the motion information of each sub-block in the current block corresponding to one of the plurality of DVP candidates) can be determined using the SbTMVP mode described in FIGS. 20-21. Thus, each of the plurality of SbTMVP merge candidates can correspond to one of the plurality of DVP candidates.

[0161] In one example, the sub-block based merge list is a combined sub-block based merge list that further includes, for example, (one or more) affine merge candidates used in the affine merge / skip mode. The affine merge candidates can include affine inheritance candidates and affine composition candidates.

[0162] In one embodiment, the sub-block based merge list includes a plurality of SbTMVP merge candidates and affine merge candidates. To indicate that DVO can be applied to a first subset of the sub-block based merge candidates, a flag (e.g., DVO flag) can be signaled. In one example, the first subset includes a plurality of SbTMVP merge candidates and does not include affine merge candidates. Which SbTMVP merge candidate is selected from the first subset of the sub-block based merge candidates can be indicated by a base index. Since each of the plurality of SbTMVP merge candidates corresponds to one of each of the plurality of DVP candidates, the base index indicates the selected DVP corresponding to the selected SbTMVP merge candidate.

[0163] In one example, the DVO flag is true, indicating that DVO should be used with the selected DVP among the plurality of DVP candidates. In one example, the DVO flag is the MMVD flag in the MMVD mode, thus indicating that the MMVD mode is used to determine DVO from DVO information. The first subset of the sub-block based merge candidates can be specified such that DVO can be applied to the first subset of the sub-block based merge candidates. In one example, DVO is indicated using the MMVD mode as described above. Signaling a base index (e.g., subblock_mmvd_base_index) can indicate which candidate within the first subset of the sub-block based merge candidates is used.

[0164] In one embodiment, the sub-block-based merge list includes a second subset of sub-block-based merge candidates. The second subset of sub-block-based merge candidates includes K0 SbTMVP merge candidates (e.g., a plurality of SbTMVP merge candidates) and K1 affine merge candidates. The base index can indicate which sub-block-based merge candidate within the second subset of sub-block-based merge candidates is selected to code the current block, and the selected sub-block-based merge candidate corresponds to the selected DVP and each SbTMVP merge candidate within the plurality of SbTMVP merge candidates.

[0165] In one example, the first K0 SbTMVP merge candidates from the SbTMVP merge candidates of the current block are included in the second subset. In one example, the first K1 affine merge candidates of the current block are included in the second subset. K0 and K1 are positive integers. K0 can be signaled in a high-level (e.g., higher than the CU level or block level) syntax such as a sequence parameter set (SPS), a picture parameter set (PPS), a tile header, a tile group header, a slice header, a picture header, etc. In one example, K0 is 2. K1 can be signaled in a high-level syntax such as SPS, PPS, tile header, tile group header, slice header, picture header, etc.

[0166] In one example, the selection of the SbTMVP merge candidate is determined by which DVP candidate in the list from which the DV is derived.

[0167] In one embodiment, two sub-block-based merge lists including an affine merge candidate list and an SbTMVP merge candidate list are used for the current block. The affine merge candidate list and the SbTMVP merge candidate list may be independent and may be configured separately. For example, the SbTMVP merge candidate list includes a plurality of SbTMVP merge candidates and does not include affine merge candidates. The base index can indicate which SbTMVP merge candidate in the SbTMVP merge candidate list is selected when the selected SbTMVP merge candidate corresponds to the selected DVP among the plurality of DVP candidates.

[0168] A flag indicating the merge mode or the MMVD mode, indicating that the merge mode or the MMVD mode is applied to the current block, may be signaled. The affine flag indicating the affine merge mode and the SbTMVP flag indicating the SbTMVP mode may be signaled independently. When the SbTMVP flag is true, the SbTMVP merge candidate list may be configured.

[0169] In one example, a first flag and a second flag are signaled to indicate whether the SbTMVP mode is used. The first flag can indicate whether a sub-block-based merge mode is applied to the current block. When the first flag (e.g., true) indicates that the sub-block-based merge mode is applied to the current block, the second flag is signaled to indicate whether the sub-block-based merge mode is the SbTMVP mode. The second flag also indicates whether an affine merge candidate list or an SbTMVP merge candidate list is configured for the current block. For example, when the second flag indicates that the sub-block-based merge mode is the SbTMVP mode, the SbTMVP merge candidate list is configured for the current block.

[0170] In one embodiment, a sub-block-based merge list includes sub-block-based merge candidates. The sub-block-based merge candidates include only a single SbTMVP candidate with or without an affine merge candidate. The number of SbTMVP candidates in the sub-block-based merge list is 1. A first index may be signaled to indicate which sub-block-based merge candidate in the sub-block-based merge list is selected. The single selected SbTMVP candidate can indicate that the SbTMVP mode is applied to the current block. A second index may be signaled to indicate which DVP among a plurality of DVP candidates is selected to form the single SbTMVP candidate. The second index may be a base index (represented, for example, by the syntax sbtmvp_mmvd_base_index). In one example, the updated DV' is obtained by the vector sum of the selected DVP and DVO. The motion information of each sub-block within the current block can be obtained from the motion information of the corresponding sub-block in the same-position block determined based on the updated DV' as described in FIGS. 20 to 21. Thus, a single SbTMVP candidate is determined, and the single SbTMVP candidate can include the motion information of each sub-block in the current block.

[0171] In one example, the value of the base index (for example, the syntax sbtmvp_mmvd_base_index) is determined by the DVP candidate in the list from which the DV is derived. For example, the value of the base index points to the position of the desired DVP candidate within the configured list.

[0172] FIG. 22 shows a flowchart outlining a process (2200) according to an embodiment of the present disclosure. The process (2200) can be used in a video encoder. The process (2200) can be executed by an apparatus for video coding that can include a processing circuit. In various embodiments, the process (2200) is executed by processing circuits such as processing circuits within terminal devices (310), (320), (330), and (340), a processing circuit that executes the functions of video encoder (403), a processing circuit that executes the functions of video encoder (603), and a processing circuit that executes the functions of video encoder (703). In some embodiments, the process (2200) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (2200). The process begins at (S2201) and proceeds to (S2210).

[0173] (S2210), a displacement vector (DV) predictor (DVP) among a plurality of DVP candidates and a DV offset of a current block in a current picture can be determined. The current block includes a plurality of sub-blocks and can be encoded using a sub-block based temporal motion vector prediction (SbTMVP) mode.

[0174] In one example, the plurality of DVP candidates are determined based on the motion vectors (MVs) of the spatial neighbors of the current block, where each reference picture of the spatial neighbors used to determine one of the plurality of DVP candidates is the same location reference picture.

[0175] If the number of MVs of the spatial neighbors is less than a threshold, zero MVs can be inserted into the plurality of DVP candidates.

[0176] In one example, a plurality of DVP candidates are determined based on candidates in a merge candidate list. The candidates can include at least one of (a) a spatial motion vector predictor (MVP) candidate, (b) a history-based MVP (HMVP) candidate, (c) a pairwise average candidate, or (d) a zero motion vector (MV), and the candidates do not include a temporal MVP (TMVP) candidate. Each reference picture of the candidates used to determine one of the plurality of DVP candidates can be the same location reference picture.

[0177] At S2220, the DV of the current block can be determined based on the DVP and DV offset of the current block. The DV indicates a block in a collocated reference picture. The block is collocated with the current block and is called the same location block of the current block. In one example, the DV is determined to be the vector sum of the DVP and the DV offset.

[0178] At S2230, the motion information of a sub-block among a plurality of sub-blocks is determined based on the motion information of the corresponding sub-block in the same location block.

[0179] At S2240, that sub-block among the plurality of sub-blocks is encoded based on the motion information of the sub-block among the plurality of sub-blocks.

[0180] Then, the process (2200) proceeds to (S2299) and ends.

[0181] The process (2200) can be appropriately adapted to various scenarios, and the steps of the process (2200) can be adjusted accordingly. One or more of the steps during the process (2200) can be adapted, omitted, repeated, and / or combined. Any suitable order can be used to implement the process (2200). Additional steps can be added.

[0182] In one example, a sub-block based merge list for the current block is constructed. The sub-block based merge candidates in the sub-block based merge list can include a plurality of SbTMVP merge candidates and at least one affine merge candidate. Each of the plurality of SbTMVP merge candidates can correspond to one of each of the plurality of DVP candidates.

[0183] In one example, a base index indicating the DVP among the plurality of DVP candidates for the current block is encoded and included in the bitstream.

[0184] In one example, the base index indicates which SbTMVP merge candidate within a subset of the sub-block based merge candidates that are the plurality of SbTMVP merge candidates should be selected, and the selected SbTMVP merge candidate corresponds to the DVP.

[0185] In one example, a first number K0 of the plurality of SbTMVP merge candidates and a second number K1 of the at least one affine merge candidate are signaled in the high level syntax. K0 and K1 are positive integers. The base index indicates which sub-block based merge candidate within a subset of the sub-block based merge candidates that includes (i) the plurality of SbTMVP merge candidates and (ii) the at least one affine merge candidate should be selected. The selected sub-block based merge candidate corresponds to the DVP and one of the plurality of SbTMVP merge candidates.

[0186] According to one embodiment, an SbTMVP merge candidate list for the current block including a plurality of SbTMVP merge candidates is constructed. Each of the plurality of SbTMVP merge candidates corresponds to one of each of the plurality of DVP candidates. The SbTMVP merge candidate list does not include affine merge candidates.

[0187] In one embodiment, a merge mode or a merge motion vector difference (MMVD) mode is applied to the current block.

[0188] The flag is encoded and signaled in the bitstream, and the flag indicates that the SbTMVP mode is currently applied to the block.

[0189] In one example, a first flag is encoded and signaled in the bitstream, where the first flag indicates that the sub-block based merge mode is currently applied to the block. A second flag is encoded and signaled in the coded video bitstream, where the second flag indicates that the sub-block based merge mode is the SbTMVP mode.

[0190] In one example, a sub-block based merge list for the current block is constructed. The sub-block based merge list includes one SbTMVP merge candidate and at least one affine merge candidate, where one SbTMVP merge candidate corresponds to DVP. An index pointing to one SbTMVP merge candidate in the sub-block based merge list may be encoded and signaled in the bitstream, where the index indicates that the SbTMVP mode is currently applied to the block.

[0191] In one embodiment, the current block including a plurality of sub-blocks is encoded in the SbTMVP mode. A plurality of DVP candidates for the current block may be determined. Each DVP candidate is derived from one or more MVs. The DV of the current block may be determined based on the DVP among the plurality of DVP candidates of the current block and the DV offset of the current block. The DV indicates a block at the same position as the current block in the same position reference picture. The sub-blocks among the plurality of sub-blocks may be encoded based on the motion information of the corresponding sub-blocks in the same position block.

[0192] FIG. 23A shows a flowchart outlining a process (2300A) according to an embodiment of the present disclosure. The process (2300A) can be used in a video decoder. The process (2300A) can be executed by an apparatus for video encoding that can include a receiving circuit and a processing circuit. In various embodiments, the process (2300A) is executed by a processing circuit such as a processing circuit in terminal devices (310), (320), (330), and (340), a processing circuit that executes the functions of video decoder (410), a processing circuit that executes the functions of video decoder (510), etc. In some embodiments, the process (2300A) is implemented with software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (2300A). The process starts at (S2301) and proceeds to (S2310).

[0193] (S2310), the base index and displacement vector (DV) offset information of the current block in the current picture can be received from the coded video bitstream. The current block can include a plurality of sub-blocks that are reconstructed using sub-block based temporal motion vector prediction (SbTMVP) mode. The base index can indicate a DVP among a plurality of DVP candidates of the current block.

[0194] In one example, the plurality of DVP candidates are determined based on the motion vectors (MVs) of the spatial neighbors of the current block, where each reference picture of the spatial neighbors used to determine one of the plurality of DVP candidates is a collocated reference picture.

[0195] If the number of MVs of the spatial neighbors is less than a threshold, zero MVs can be inserted into the plurality of DVP candidates.

[0196] In one example, multiple DVP candidates are determined based on candidates in a merge candidate list. The candidates can include at least one of (a) a spatial motion vector predictor (MVP) candidate, (b) a history-based MVP (HMVP) candidate, (c) a pairwise average candidate, or (d) a zero motion vector (MV), and the candidates do not include a temporal MVP (TMVP) candidate. Each reference picture of the candidates used to determine one of the multiple DVP candidates can be the same-position reference picture.

[0197] At S2320, the DV of the current block can be determined based on the DVP of the current block and the DV offset of the current block indicated by the DV offset information. The DV can indicate a block within a reference picture at the same position. The block is at the same position as the current block and is called the same-position block. In one example, the DV is determined to be the vector sum of the DVP and the DV offset.

[0198] At S2330, the motion information of a sub-block among multiple sub-blocks is determined based on the motion information of the corresponding sub-block within the same-position block.

[0199] At S2340, based on the motion information of a sub-block among multiple sub-blocks, that sub-block among the multiple sub-blocks can be reconstructed.

[0200] Then, the process (2300A) proceeds to (S2399) and ends.

[0201] The process (2300A) can be appropriately adapted to various scenarios, and the steps of the process (2300A) can be adjusted accordingly. One or more of the steps of the process (2300A) can be adapted, omitted, repeated, and / or combined. Any suitable order can be used to implement the process (2300A). Additional steps can be added.

[0202] In one example, a sub-block based merge list of the current block is constructed. The sub-block based merge candidates in the sub-block based merge list can include a plurality of SbTMVP merge candidates and at least one affine merge candidate. Each of the plurality of SbTMVP merge candidates can correspond to one of each of the plurality of DVP candidates.

[0203] In one example, the base index indicates which SbTMVP merge candidate within a subset of the sub-block based merge candidates that are a plurality of SbTMVP merge candidates should be selected, and the selected SbTMVP merge candidate corresponds to the DVP.

[0204] In one example, a first number K0 of the plurality of SbTMVP merge candidates and a second number K1 of the at least one affine merge candidate are signaled in the high level syntax. K0 and K1 are positive integers. The base index indicates which sub-block based merge candidate within a subset of the sub-block based merge candidates that includes (i) the plurality of SbTMVP merge candidates and (ii) the at least one affine merge candidate should be selected, and the selected sub-block based merge candidate corresponds to the DVP and one of the plurality of SbTMVP merge candidates.

[0205] According to one embodiment, an SbTMVP merge candidate list of the current block including a plurality of SbTMVP merge candidates is constructed. Each of the plurality of SbTMVP merge candidates corresponds to one of each of the plurality of DVP candidates. The SbTMVP merge candidate list does not include affine merge candidates.

[0206] In one embodiment, a merge mode or a merge motion vector difference (MMVD) mode is applied to the current block.

[0207] A flag is received from the coded video bitstream, and the flag indicates that the SbTMVP mode is applied to the current block.

[0208] In one example, a first flag is received from a coded video bitstream, where the first flag indicates that a sub-block based merge mode is applied to a current block. A second flag is received from the coded video bitstream, where the second flag indicates that the sub-block based merge mode is the SbTMVP mode.

[0209] In one example, a sub-block based merge list for a current block is constructed. The sub-block based merge list includes one SbTMVP merge candidate and at least one affine merge candidate, where one SbTMVP merge candidate corresponds to DVP. An index pointing to one SbTMVP merge candidate in the sub-block based merge list may be received, where the index indicates that the SbTMVP mode is applied to the current block.

[0210] FIG. 23B shows a flowchart outlining a process (2300B) according to an embodiment of the present disclosure. The process (2300B) is a variation of the process (2300A). The process (2300B) can be used in a video decoder. The process (2300B) can be executed by an apparatus for video coding that can include a receiving circuit and a processing circuit. In various embodiments, the process (2300B) is executed by a processing circuit such as a processing circuit in terminal devices (310), (320), (330), and (340), a processing circuit that executes the functions of video decoder (410), a processing circuit that executes the functions of video decoder (510), etc. In some embodiments, the process (2300B) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (2300B). The process begins at (S2302) and proceeds to (S2312).

[0211] At S2312, a coded video bitstream including a current picture is received. The current picture includes a current block. The current block includes a plurality of sub-blocks.

[0212] In S2322, it is determined that a current block including a plurality of sub-blocks is coded in a subblock-based temporal motion vector prediction (SbTMVP) mode based on syntax elements in a coded video bitstream.

[0213] In S2332, a plurality of displacement vector (DV) predictor (DVP) candidates for a current block can be determined. Each DVP candidate can be derived from one or more motion vectors (MVs).

[0214] In S2342, a base index and a DV offset of a current block can be received. The base index indicates a DVP among a plurality of DVP candidates of the current block.

[0215] In S2352, the DV of a current block can be described based on the DVP indicated in a plurality of DVP candidates of the current block and the DV offset of the current block. The DV indicates a block at the same position as the current block within the same position reference picture.

[0216] In S2362, the sub-blocks among the plurality of sub-blocks are reconstructed based on motion information of corresponding sub-blocks within the same position block.

[0217] Then, the process (2300B) proceeds to (S2392) and ends.

[0218] The process (2300B) can apply various embodiments used in the process (2300A). The process (2300B) can be appropriately adapted to various scenarios, and the steps of the process (2300B) can be adjusted accordingly. One or more of the steps in the process (2300B) can be adapted, omitted, repeated, and / or combined. Any suitable order can be used to implement the process (2300B). Additional steps can be added.

[0219] Embodiments of the present disclosure may be used separately or combined in any order. Further, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.

[0220] The above-described technology can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, FIG. 24 shows a computer system (2400) suitable for implementing certain embodiments of the disclosed subject matter.

[0221] The computer software can be coded using any suitable machine code or computer language, which can be used to create code containing instructions that can be executed directly, or via interpretation, microcode execution, etc., by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., and can undergo assembly, compilation, linking, or similar mechanisms.

[0222] The instructions can be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, game devices, Internet of Things devices, etc.

[0223] The components shown in FIG. 24 for the computer system (2400) are essentially exemplary and are not intended to suggest any limitation regarding the use or functionality range of the computer software implementing the embodiments of the present disclosure. Also, the configuration of the components should not be construed as having any dependency or requirement regarding any one or combination of the components shown in the exemplary embodiments of the computer system (2400).

[0224] The computer system (2400) may include a human interface input device. Such a human interface input device can respond to input by one or more human users, for example, through tactile input (keystrokes, swipes, movements of a data glove, etc.), voice input (voice, applause, etc.), visual input (gestures, etc.), olfactory input (not shown), etc. The human interface device can also be used to capture certain media such as audio (speech, music, ambient sound, etc.), images (scanned images, photographic images obtained from a still image camera, etc.), video (2D video, 3D video including stereoscopic video, etc.) that are not necessarily directly related to conscious input by humans.

[0225] The input human interface device can include one or more (only one of each is shown) of a keyboard (2401), a mouse (2402), a trackpad (2403), a touch screen (2410), a data glove (not shown), a joystick (2405), a microphone (2406), a scanner (2407), a camera (2408).

[0226] The computer system (2400) may also include certain human interface output devices. Such human interface output devices may stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices include, for example, tactile output devices (e.g., tactile feedback by a touch screen (2410), a data glove (not shown), or a joystick (2405), although there may also be tactile feedback devices that do not function as input devices), audio output devices (such as speakers (2409), headphones (not shown), etc.), visual output devices (screens (2410) including CRT screens, LCD screens, plasma screens, OLED screens, etc., with or without touch screen input capabilities, with or without tactile feedback capabilities, some of which may be capable of outputting 2D visual output or output exceeding 3D, such as through means like stereoscopic output), virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), and may also include printers (not shown).

[0227] The computer system (2400) may include a memory device accessible by humans and its associated media. Associated media can include, for example, optical media such as CD / DVD ROM / RW (2420) including CD / DVD or similar media (2721), thumb drives (2422), removable hard drives or solid state drives (2423), legacy magnetic media such as tapes and floppy (registered trademark) disks (not shown), and associated media such as dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0228] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not include a transmission medium, a carrier wave, or other transient signals.

[0229] The computer system (2400) can also include an interface (2454) to one or more communication networks (2455). The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet (registered trademark), wireless LANs, cellular networks including GSM (registered trademark), 3G, 4G, 5G, LTE, etc., and TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicle and industrial including CANBus, etc. Some networks generally require an external network interface adapter attached to a certain general-purpose data port or peripheral bus (2449) (such as a USB port of the computer system (2400)), and other networks are generally integrated into the core of the computer system (2400) by attachment to the system bus as described below (such as an Ethernet (registered trademark) interface to a personal computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (2400) can communicate with other entities. Such communication can be unidirectional receive-only (such as broadcast TV), unidirectional transmit-only (such as from CANbus to a specific CANbus device), or bidirectional to other computer systems using, for example, local or wide area digital networks. Certain protocols and protocol stacks can be used on each of those networks and network interfaces as described above.

[0230] The aforementioned human interface device, human-accessible memory device, and network interface can be attached to the core (2440) of the computer system (2400).

[0231] The core (2440) can include one or more central processing units (CPUs) (2441), a graphics processing unit (GPU) (2442), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (2443), a hardware accelerator for specific tasks (2444), a graphics adapter (2450), etc. These devices can be connected via a system bus (2448) together with a read-only memory (ROM) (2445), a random access memory (2446), an internal mass storage device such as an internal hard drive not accessible to users, an SSD (2447). In some computer systems, the system bus (2448) can be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be connected directly to the system bus (2448) of the core or via a peripheral bus (2449). In one example, a screen (2410) can be connected to a graphics adapter (2450). Architectures of peripheral buses include PCI, USB, etc.

[0232] The CPU (2441), GPU (2442), FPGA (2443), and accelerator (2444) can execute specific instructions that can, in combination, constitute the aforementioned computer code. The computer code can be stored in the ROM (2445) or RAM (2446). Also, transient data may be stored in the RAM (2446), while persistent data may be stored, for example, in the internal mass storage device (2447). Fast storage and retrieval to any of the memory devices can be enabled through the use of cache memory that can be closely associated with one or more CPUs (2441), GPUs (2442), mass storage devices (2447), ROM (2445), RAM (2446), etc.

[0233] A computer-readable medium can have computer code for performing various computer-implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of this disclosure, or they may be of the kind well known and available to those of ordinary skill in the computer software arts.

[0234] By way of example and not limitation, a computer system having an architecture (2400), specifically a core (2440), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software incorporated in one or more tangible computer-readable media. Such computer-readable media can be a user-accessible mass storage device as introduced above, as well as media associated with a specific storage device of the core (2440) that is of a non-transitory nature, such as a mass storage device (2447) inside the core or a ROM (2445). The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (2440). The computer-readable media can include one or more memory devices or chips, depending on specific requirements. The software can cause the core (2440), specifically the processor (including a CPU, GPU, FPGA, etc.) therein, to define a data structure stored in the RAM (2446) and modify such data structure according to a process defined by the software, thereby executing a specific process or a specific part of a specific process described herein. Additionally, or alternatively, the computer system can provide functionality as a result of circuitry (e.g., an accelerator (2444)) wired or otherwise embodied with logic that can operate instead of or in conjunction with software to execute a specific process or a specific part of a specific process described herein. References to software can, as appropriate, include logic, and vice versa if appropriate. References to computer-readable media can, as appropriate, include circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0235] Appendix A: Acronyms.

[0236] JEM: joint exploration model (Joint Exploration Model) VVC: versatile video coding (Versatile Video Coding) BMS: benchmark set (Benchmark Set) MV: Motion Vector (Motion Vector) HEVC: High Efficiency Video Coding (High Efficiency Video Coding) SEI: Supplementary Enhancement Information (Supplementary Enhancement Information) VUI: Video Usability Information (Video Usability Information) GOPs: Groups of Pictures (Groups of Pictures) TUs: Transform Units (Transform Units) PUs: Prediction Units (Prediction Units) CTUs: Coding Tree Units (Coding Tree Units) CTBs: Coding Tree Blocks (Coding Tree Blocks) PBs: Prediction Blocks (Prediction Blocks) HRD: Hypothetical Reference Decoder (Hypothetical Reference Decoder) SNR: Signal Noise Ratio (Signal Noise Ratio) CPUs: Central Processing Units (Central Processing Units) GPUs: Graphics Processing Units (Graphics Processing Units) CRT: Cathode Ray Tube (Cathode Ray Tube) LCD: Liquid-Crystal Display (Liquid-Crystal Display) OLED: Organic Light-Emitting Diode (Organic Light-Emitting Diode) CD: Compact Disc (Compact Disc) DVD: Digital Video Disc (Digital Video Disc) ROM: Read-Only Memory (Read-Only Memory) RAM: Random Access Memory (Random Access Memory) ASIC: Application-Specific Integrated Circuit (Application-Specific Integrated Circuit) PLD: Programmable Logic Device (Programmable Logic Device) LAN: Local Area Network (Local Area Network) GSM: Global System for Mobile communications (Global System for Mobile Communications) LTE: Long-Term Evolution (Long-Term Evolution) CANBus: Controller Area Network Bus (Controller Area Network Bus) USB: Universal Serial Bus (Universal Serial Bus) PCI: Peripheral Component Interconnect (Peripheral Component Interconnect) FPGA: Field Programmable Gate Areas (Field Programmable Gate Areas) SSD: solid-state drive (Solid-State Drive) IC: Integrated Circuit (Integrated Circuit) CU: Coding Unit (Coding Unit) Although several exemplary embodiments have been described, there are changes, substitutions, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, those skilled in the art will appreciate that many systems and methods can be devised that embody the principles of the present disclosure, yet are not explicitly shown or described herein, and thus are within the spirit and scope of the present disclosure.

Claims

1. A method for video decoding in a decoder, comprising: receiving a coded video bitstream including a current picture, the current picture including a current block, the current block including a plurality of sub-blocks; determining that the current block including the plurality of sub-blocks is coded in a sub-block based temporal motion vector prediction (SbTMVP) mode based on a syntax element in the coded video bitstream; determining a plurality of displacement vector (DV) predictor (DVP) candidates for the current block, each DVP candidate being derived from one or more motion vectors (MVs); receiving a base index and a DV offset of the current block, the base index indicating a DVP among the plurality of DVP candidates of the current block; determining a DV of the current block based on the DVP indicated among the plurality of DVP candidates of the current block and the DV offset of the current block, the DV indicating the same position block as the current block in the same position reference picture; reconstructing sub-blocks among the plurality of sub-blocks based on motion information of corresponding sub-blocks in the same position block and a method.

2. The method according to claim 1, wherein determining the DV includes determining that the DV is a vector sum of the DVP and the DV offset.

3. The method according to claim 1, further comprising determining the plurality of DVP candidates based on MVs of spatial neighbors of the current block, each reference picture of the spatial neighbors used to determine one of the plurality of DVP candidates being the same position reference picture.

4. The method according to claim 3, wherein determining the plurality of DVP candidates includes inserting a zero MV into the plurality of DVP candidates in response to the number of the MVs of the spatial neighbors being less than a threshold.

5. Determining the plurality of DVP candidates based on candidates in a merge candidate list, wherein the candidates include at least one of: (a) a spatial motion vector predictor (MVP) candidate, (b) a history-based MVP (HMVP) candidate, (c) a pairwise average candidate, or (d) a zero motion vector (MV), and the candidates do not include a temporal MVP (TMVP) candidate, and further including that each reference picture of each of the candidates used to determine one of the plurality of DVP candidates is an identical position reference picture, the method according to claim 1.

6. Constructing a sub-block based merge list for the current block, wherein the sub-block based merge candidates in the sub-block based merge list include a plurality of SbTMVP merge candidates and at least one affine merge candidate, and each of the plurality of SbTMVP merge candidates corresponds to one of the plurality of DVP candidates, the method according to claim 1.

7. The base index indicates which SbTMVP merge candidate in a subset of the sub-block based merge candidates that are the plurality of SbTMVP merge candidates should be selected, and the selected SbTMVP merge candidate corresponds to the DVP, the method according to claim 6.

8. A first number K0 of the plurality of SbTMVP merge candidates and a second number K1 of the at least one affine merge candidate are signaled in a high level syntax, and K0 and K1 are positive integers, The base index indicates which sub-block based merge candidate in a subset of the sub-block based merge candidates that includes (i) the plurality of SbTMVP merge candidates and (ii) the at least one affine merge candidate should be selected, and the selected sub-block based merge candidate corresponds to the DVP and one of the plurality of SbTMVP merge candidates, the method according to claim 6.

9. Constructing an SbTMVP merge candidate list for the current block that includes a plurality of SbTMVP merge candidates, wherein each of the plurality of SbTMVP merge candidates corresponds to one of the plurality of DVP candidates, and further including that the SbTMVP merge candidate list does not include an affine merge candidate, the method according to claim 1.

10. A merge mode or a merge motion vector difference (MMVD) mode is applied to the current block, The method includes receiving a flag from the coded video bitstream, the flag indicating that the SbTMVP mode is applied to the current block, The method according to claim 9.

11. A merge mode or a merge motion vector difference (MMVD) mode is applied to the current block, The method Receiving a first flag from the coded video bitstream, the first flag indicating that a sub-block based merge mode is applied to the current block, Receiving a second flag from the coded video bitstream, the second flag indicating that the sub-block based merge mode is the SbTMVP mode, The method according to claim 9.

12. Constructing a sub-block based merge list for the current block, the sub-block based merge list including one SbTMVP merge candidate and at least one affine merge candidate, the one SbTMVP merge candidate corresponding to the DVP, Receiving an index indicating the one SbTMVP merge candidate in the sub-block based merge list, the index indicating that the SbTMVP mode is applied to the current block The method according to claim 1, further comprising.

13. An apparatus for video decoding, Having a processing circuit configured to execute the method steps according to any one of claims 1 to 12, Apparatus.

14. A computer program that, when executed by at least one processor, causes the at least one processor to execute the method steps according to any one of claims 1 to 12.

15. A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform the method steps according to any one of claims 1 to 12 Non-transitory computer-readable storage medium.

Citation Information

Patent Citations

  • Video encoding / decoding method, video encoder, video decoder, video encoding device, video decoding device, and computer-readable storage medium

    JP2022509743A

  • Systems, methods, and non-transitory computer-readable storage media for signaling motion merge modes in video coding

    JP2022515914A

  • Method and apparatus for video coding

    US20200236383A1

  • Device and method for processing video signal by using inter prediction

    US20220078408A1