Methods, apparatus, and computer programs for video coding

Sub-block based temporal motion vector prediction with motion vector offsets addresses inefficiencies in video coding by refining motion information at the sub-block level, enhancing compression efficiency and reducing bandwidth and storage needs.

JP2026062850APending Publication Date: 2026-04-10TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2025-12-26
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing video coding technologies face inefficiencies in reducing redundancy and compression ratio, particularly in intra-prediction and motion vector prediction, due to the increasing number of possible directions and the need for more bits to represent less likely directions, which affects bandwidth and storage requirements.

Method used

Implementing sub-block based temporal motion vector prediction (SbTMVP) with a motion vector offset, using displacement vector (DV) and motion vector (MV) offsets signaled in the encoded video bitstream to refine motion information at the sub-block level, allowing for more efficient compression.

Benefits of technology

Enhances compression efficiency by reducing the data required to represent motion vectors, thereby improving bandwidth utilization and storage efficiency in video encoding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062850000001_ABST
    Figure 2026062850000001_ABST
Patent Text Reader

Abstract

This provides a method for video decoding that a decoder will perform. [Solution] The method determines, based on syntax elements in the encoded video bitstream, that the current block, which includes multiple subblocks, is coded in subblock-based time motion vector prediction (SbTMVP) mode, and receives MVO information indicating the motion vector offset (MVO). The MVO indicates the motion offset of the displacement vector (DV) used to adjust the position of the collated block in the collated reference picture. The method also determines the updated DV of the current block based on the DV and MVO, derives the SbTMVP information for each of the multiple subblocks based on the motion information of the corresponding subblock in the collated block indicated by the updated DV, and reconstructs the multiple subblocks in SbTMVP mode based on the SbTMVP information of the multiple subblocks.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to U.S. Provisional Application No. 63 / 332,131, “SUBBLOCK BASED MOTION VECTOR PREDICTOR WITH MOTION VECTOR OFFSET,” filed on 18 April 2022, and to U.S. Patent Application No. 17 / 984,123, “A SUB-BLOCK BASED TEMPORAL MOTION VECTOR PREDICTOR WITH AN MOTION VECTOR OFFSET,” filed on 9 November 2022. The disclosures of these prior applications are incorporated herein by reference in their entirety.

[0002] This disclosure describes embodiments relating to video coding in general. [Background technology]

[0003] The background information presented herein is intended to provide a general overview of the circumstances relating to the disclosure. Within the scope described in this background section, the work of the inventors listed herein, and any descriptions that might otherwise not qualify as prior art at the time of filing, are not, expressly or implicitly, considered prior art to this disclosure.

[0004] Uncompressed digital images and / or videos consist of a series of pictures, each picture having spatial dimensions of, for example, 1920×1080 luminance samples and associated chrominance samples. A series of pictures can have a fixed or variable picture rate (informally also known as frame rate), for example, 60 pictures per second, i.e., a picture rate of 60Hz. Uncompressed images and / or videos have specific bitrate requirements. For example, 1080p60 4:2:0 video with 8 bits per sample (1920×1080 luminance sample resolution at a 60Hz frame rate) requires a bandwidth of nearly 1.5 Gbit / s. One hour of such video requires more than 600 GByte of storage space.

[0005] One purpose of encoding and decoding images and / or video may be to reduce the redundancy of input image and / or video signals through compression. Compression can help reduce the aforementioned bandwidth and / or storage space requirements by, in some cases, orders of magnitude or more. While this explanation uses video encoding / decoding as an example, the same techniques can be applied to image encoding / decoding in a similar manner without departing from the spirit of this disclosure. Both lossless and lossy compression, as well as combinations thereof, can be used. Lossless compression refers to a technique that allows an exact copy of the original signal to be reconstructed from the compressed original signal. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signals is small enough to make the reconstructed signal useful for its intended purpose. In the case of video, lossy compression is widely used. The amount of distortion that can be tolerated depends on the application; for example, users of a particular consumer streaming application may tolerate higher distortion than users of a television distribution application. The achievable compression ratio reflects this, with higher tolerable distortion resulting in a higher compression ratio.

[0006] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transformation processing, quantization, and entropy coding.

[0007] Video codec technology may include a technique known as intra coding. In intra coding, sample values ​​are represented without referencing samples or other data from a previously reconstructed reference picture. In some video codecs, a picture is spatially subdivided into multiple blocks of samples. If all block samples are coded in intra mode, the picture can be an intra picture. Intra pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and thus can be used as the first picture in the video bitstream and video session being coded, or as a still image. Samples in intra blocks can be subjected to a transformation, and the transformation coefficients can be quantized before entropy coding. Intra prediction can be a technique that minimizes the sample values ​​in the pre-transformation domain. In some cases, a smaller DC value and smaller AC coefficients after transformation result in fewer bits being required to represent the block with a given quantization step size after entropy coding.

[0008] For example, traditional intra-coding used in MPEG-2 generation coding techniques does not use intra-prediction. However, some newer video compression techniques include techniques that attempt to make predictions based on surrounding sample data and / or metadata obtained during the encoding and / or decoding of blocks of data. Such techniques will hereafter be referred to as “intra-prediction” techniques. In at least some cases, intra-prediction uses only reference data from the current picture being reconstructed, and not from the reference picture.

[0009] Numerous different forms of intra-prediction can exist. If two or more such techniques can be used in a given video coding technique, the specific technique used may be coded as a specific intra-prediction mode that employs that particular technique. In a particular case, an intra-prediction mode may have sub-modes and / or parameters, which may be coded individually or included in a mode codeword that defines the prediction mode used. The choice of codeword for a given combination of mode, sub-mode, and / or parameter can affect the coding efficiency gain through intra-prediction, as can the entropy coding technique used to convert the codeword into a bitstream.

[0010] Certain modes of intra-prediction were introduced in H.264, improved in H.265, and further refined with newer coding techniques such as Joint Exploration Models (JEM), Versatile Video Coding (VVC), and Benchmark Sets (BMS). Predictor blocks can be formed using adjacent sample values ​​of already available samples. The sample values ​​of adjacent samples are copied to the predictor block according to direction. A reference to the direction to be used may be coded into the bitstream or predicted itself.

[0011] Referring to Figure 1A, the lower right shows a subset of nine predictor directions known from the 33 possible predictor directions defined in H.265 (corresponding to 33 of the 35 intra-modes, or angular modes). The point where the arrows converge (101) represents the sample being predicted. The arrows indicate the direction from which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples to the upper right at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples to the lower left at an angle of 22.5 degrees from the horizontal.

[0012] Referring again to Figure 1A, a 4x4 sample square block (104) (shown by a thick dashed line) is depicted in the upper left. The square block (104) contains 16 samples, each labeled with "S" and its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample (from the top) in the Y dimension and the first sample (from the left) in the X dimension. Similarly, sample S44 is the fourth sample in both the Y and X dimensions within block (104). Since this block is 4x4 sample size, S44 is in the lower right. Furthermore, a reference sample is shown following a similar numbering scheme. The reference sample is labeled with R and its Y position (e.g., row index) and X position (column index) relative to block (104). In both H.264 and H.265, the predicted samples are adjacent to the block being reconstructed, and therefore, there is no need to use negative values.

[0013] Intra-picture prediction works by copying the reference sample value from an adjacent sample indicated by the signaled prediction direction. For example, suppose the encoded video bitstream contains signaling for this block that the prediction direction coincides with arrow (102), i.e., the sample is predicted from a sample 45 degrees to the upper right of the horizontal. In that case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05, and sample S44 is predicted from reference sample R08.

[0014] In certain cases, particularly when the direction cannot be divided equally into 45-degree intervals, the values ​​of multiple reference samples may be combined, for example, by interpolation, in order to calculate a reference sample.

[0015] As video coding technology advances, the number of possible directions increases. H.264 (2003) could represent nine different directions. This increased to 33 in H.265 (2013). Currently, JEM / VVC / BMS can support up to 65 directions. Experiments are being conducted to identify the most likely directions, and certain techniques in entropy coding are used to represent those less likely directions with fewer bits, accepting a penalty for them. Furthermore, these directions themselves can sometimes be predicted from the adjacent directions used in adjacent, already decoded blocks.

[0016] Figure 1B shows a schematic diagram (110) illustrating 65 intra-prediction directions according to JEM, illustrating the gradually increasing number of prediction directions.

[0017] The mapping of intra-predicted direction bits, which represent direction within an encoded video bitstream, can vary depending on the video coding technique. Such mappings can range from simple direct mapping to complex adaptive schemes involving codewords, most likely modes, and similar techniques. However, in most cases, video content may have certain directions that are statistically less likely to occur than other particular directions. Since the goal of video compression is to reduce redundancy, a well-functioning video coding technique will represent these less likely directions with more bits than the more likely directions.

[0018] Image and / or video encoding and decoding can be performed using interpicture prediction with motion compensation. Motion compensation may be a lossy compression technique and may relate to a technique used to predict a newly reconstructed picture or picture portion after blocks of sample data from a previously reconstructed picture or part thereof (reference picture) are spatially shifted in the direction indicated by a motion vector (MV). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or three dimensions, with a third being an indication pointing to the reference picture to be used (the latter being indirectly a time dimension).

[0019] Some video compression techniques allow predicting a motion vector (MV) applicable to a specific region of sample data from another MV, for example, from an MV that precedes that MV in decoding order and relates to a different region of sample data spatially adjacent to the region being reconstructed. Doing so can significantly reduce the amount of data required to code that MV, thereby eliminating redundancy and improving compression. MV prediction can work effectively because, for example, when coding an input video signal originating from a camera (known as natural video), there is a statistical likelihood that a region larger than the region to which a single MV is applicable will move in a similar direction and, therefore, can be predicted using similar motion vectors derived from the MVs of adjacent regions. This results in the MV found for a given region being similar to or identical to the MV predicted from the surrounding MVs, and thus being representable with fewer bits than would be used if that MV were coded directly after entropy coding. In some cases, MV prediction can be an example of lossless compression of the signal (i.e., MV) derived from the original signal (i.e., the sample stream). In other cases, for example, the MV prediction itself may be irreversible due to rounding errors when calculating the predictor from several surrounding MVs.

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Recommendation H.265, “High Efficiency Video Coding,” December 2016). Of the many MV prediction mechanisms provided by H.265, the technique referred to as “spatial merging” will be explained below with reference to Figure 2.

[0021] Referring to FIG. 2, the current block (201) has samples found by the encoder during the motion search process that it is predictable from a spatially shifted previous block of the same size. Instead of directly coding the MV, the MV can be derived from metadata related to one or more reference pictures, such as from the immediately previous reference picture (in decoding order), using the MV associated with any one of five surrounding samples denoted as A0, A1, and B0, B1, B2 (202 to 206 respectively). In H.265, MV prediction can use predictors from the same reference picture that adjacent blocks are using. Summary of the Invention

[0022] Aspects of the present disclosure provide methods and apparatuses for video encoding / decoding. In some examples, an apparatus for video decoding includes a processing circuit. In one embodiment, the processing circuit receives displacement vector (DV) offset information for a current block within a current picture from an encoded video bitstream. The current block includes a plurality of sub-blocks that are reconstructed using a sub-block based temporal motion vector prediction (SbTMVP) mode. Based on the DV of the current block and the DV offset of the current block, an updated DV of the current block can be determined. The DV offset is indicated by the DV offset information. The updated DV of the current block points to a collocated block within a collocated reference picture. The collocated block is collocated with the current block. The processing circuit determines motion information of a sub-block among the plurality of sub-blocks based on motion information of a corresponding sub-block within the collocated block, and reconstructs the sub-block among the plurality of sub-blocks based on the motion information of the sub-block among the plurality of sub-blocks.

[0023] In one embodiment, the processing circuit determines the updated DV to be a vector sum of the DV and the MV offset.

[0024] In one embodiment, the DV offset information includes a DV offset signaled within the encoded video bitstream.

[0025] In one embodiment, the DV offset information indicates at least one index indicating the magnitude and direction of the DV offset.

[0026] In one example, the at least one index includes a distance index indicating the magnitude of the DV offset, which is one of a set of predetermined distances, and a direction index indicating the direction of the DV offset, which is one of a set of predetermined directions.

[0027] In one example, the set of predetermined distances and the set of predetermined directions are used in a merge motion vector difference (MMVD) mode.

[0028] In one example, the updated DV is constrained such that the collocated block is within a constrained area within the collocated reference picture. The constrained area includes a collocated area corresponding to the current CTU in the current picture, and the current CTU includes the current block.

[0029] In one embodiment, an encoded video bitstream having a current picture is received. The current picture includes a current block. The current block includes a plurality of subblocks. A processing circuit determines, based on syntax elements in the encoded video bitstream, that the current block, including the plurality of subblocks, is encoded in SbTMVP mode. The processing circuit obtains MVO information for the current block, which indicates the motion vector offset (MVO). The MVO indicates the motion offset of the displacement vector (DV) used to adjust the position of the collated block in the collated reference picture. Based on the DV and MVO of the current block, the processing circuit determines the updated DV of the current block. The updated DV indicates the adjusted position of the collated block in the collated reference picture. The processing circuit derives SbTMVP information for each of the plurality of subblocks, based at least on the motion information of the corresponding subblock in the collated block indicated by the updated DV, and reconstructs the plurality of subblocks in SbTMVP mode based on the SbTMVP information of the plurality of subblocks.

[0030] In one embodiment, a processing circuit receives motion vector (MV) offset information for the current block in the current picture from an encoded video bitstream. The current block includes a plurality of subblocks that are reconstructed using subblock-based time motion vector prediction (SbTMVP) mode. The processing circuit determines the displacement vector (DV) of the current block, which represents the collated block in the collated reference picture that is collated with the current block. The processing circuit determines the motion information of one of the plurality of subblocks based on the motion information of the corresponding subblock in the collated block, and determines updated motion information of the subblock based on the motion information of the subblock and the MV offset of the current block indicated by the MV offset information. The processing circuit reconstructs the subblock based on the updated motion information.

[0031] In one example, MV offset information includes the MV offset signaled within the encoded video bitstream.

[0032] In one example, the MV offset information indicates the magnitude of the MV offset and at least one index indicating the direction of the MV offset.

[0033] In one example, at least one index includes a distance index indicating the magnitude of the MV offset, which is one of a set of predetermined distances, and a direction index indicating the direction of the MV offset, which is one of a set of predetermined directions.

[0034] A set of predetermined distances and a set of predetermined directions are used in Merge Motion Vector Difference (MMVD) mode.

[0035] In one example, the motion information of one of the multiple subblocks includes a first motion vector (MV) associated with a first reference picture from the first reference picture list L0. The processing circuit determines an updated first MV, which is the vector sum of the first MV and the MV offset, and the updated motion information includes the updated first MV.

[0036] In one example, the motion information of the subblock among the multiple subblocks mentioned above includes a second MV associated with a second reference picture from a second reference picture list L1. The processing circuit determines an updated second MV, which is either (i) the vector difference between the second MV and the MV offset, or (ii) the vector sum of the second MV and the scaled MV offset. The scaled MV offset is based on the MV offset, the picture order count (POC) of the current picture, the POC of the first reference picture, and the POC of the second reference picture.

[0037] Aspects of this disclosure also provide a non-temporary computer-readable medium that stores instructions causing a computer to perform a method for decoding video when executed by the computer for video decoding. [Brief explanation of the drawing]

[0038] Further characteristics, nature, and various advantages of the matters to be disclosed will become even clearer from the following detailed description and attached drawings. [Figure 1A] A schematic representation of an exemplary subset of intra-predictive modes is shown. [Figure 1B] This shows an example of an intra-prediction direction. [Figure 2] An example of the current block (201) and surrounding samples is shown. [Figure 3] A schematic block diagram of an exemplary communication system (300) is shown. [Figure 4] A schematic block diagram of an exemplary communication system (400) is shown. [Figure 5] A schematic block diagram illustrating a decoder is shown. [Figure 6] A schematic block diagram illustrating an example of an encoder is shown. [Figure 7] A block diagram of an exemplary encoder is shown. [Figure 8] An illustrative block diagram of a decoder is shown. [Figure 9] The position of a spatial merge candidate according to one embodiment of the present invention is shown. [Figure 10] This shows candidate pairs considered for redundancy testing of spatial merge candidates according to one embodiment of the present invention. [Figure 11] This shows an example of motion vector scaling for time merge candidates. [Figure 12] This shows an example of a candidate location for a time merge in the current coding unit. [Figure 13]Figures 13-14 show an example of the search process in merge motion vector difference (MMVD) mode. [Figure 14] Figures 13-14 show an example of the search process in merge motion vector difference (MMVD) mode. [Figure 15] This shows additional refinement positions along multiple oblique angles in MMVD mode. [Figure 16] Figures 16-17 show an exemplary subblock-based time-motion vector prediction (SbTMVP) process used in SbTMVP mode. [Figure 17] Figures 16-17 show an exemplary subblock-based time-motion vector prediction (SbTMVP) process used in SbTMVP mode. [Figure 18] A flowchart outlining an encoding process according to some embodiments of this disclosure is shown. [Figure 19A] A flowchart outlining the decryption process according to some embodiments of this disclosure is shown. [Figure 19B] A flowchart outlining the decryption process according to some embodiments of this disclosure is shown. [Figure 20] A flowchart outlining an encoding process according to some embodiments of this disclosure is shown. [Figure 21] A flowchart outlining the decryption process according to some embodiments of this disclosure is shown. [Figure 22] This is a schematic diagram of a computer system according to one embodiment. [Modes for carrying out the invention]

[0039] Figure 3 shows an exemplary block diagram of a communication system (300). The communication system (300) includes a plurality of terminal devices that can communicate with each other, for example, via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of Figure 3, the first pair of terminal devices (310) and (320) perform one-way transmission of data. For example, terminal device (310) may code video data (e.g., a stream of video pictures captured by terminal device (310)) for transmission to other terminal devices (320) via the network (350). The coded video data may be transmitted in the form of one or more coded video bitstreams. Terminal device (320) may receive coded video data from the network (350), decode the coded video data to restore video pictures, and display video pictures according to the restored video data. One-way data transmission can be common in media service delivery applications and similar systems.

[0040] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) for bidirectional transmission of encoded video data, for example, during a video conference. In the bidirectional transmission of data, in one example, each terminal device of terminal devices (330) and (340) may encode video data (e.g., a stream of video pictures captured by that terminal device) for transmission to the other terminal device of terminal devices (330) and (340) via the network (350). Each terminal device of terminal devices (330) and (340) may also receive encoded video data transmitted by the other terminal device of terminal devices (330) and (340), decode the encoded video data to restore video pictures, and display the video pictures on an accessible display device according to the restored video data.

[0041] In the example in Figure 3, terminal devices (310), (320), (330), and (340) are shown as a server, a personal computer, and a smartphone, respectively, but the principles of this disclosure may not be limited thereto. Embodiments of this disclosure find applications in laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network (350) represents any number of networks that transmit encoded video data between terminal devices (310), (320), (330), and (340), including, for example, wired communication networks and / or wireless communication networks. Communication networks (350) may exchange data over circuit-switched channels and / or packet-switched channels. Typical networks include far-field communication networks, local area networks, wide area networks, and / or the Internet. For the purposes of this description, the architecture and topology of network (350) may not be important to the operation of this disclosure unless described below.

[0042] Figure 4 shows an example of an application relating to the disclosures, illustrating the arrangement of a video encoder and video decoder in a streaming environment. The disclosures can be equally applied to other applications where video is usable, including, for example, video conferencing, digital TV, streaming services, and the storage of compressed video on digital media including CDs, DVDs, memory sticks, and similar devices.

[0043] The streaming system may include a capture subsystem (413) which may include a video source (401), such as a digital camera, that produces, for example, a stream (402) of uncompressed video pictures. In one example, the stream (402) of video pictures includes a sample captured by the digital camera. The stream (402) of video pictures is drawn as a thick line to emphasize that it has a higher data volume compared to encoded video data (404) (or encoded video bitstream) and may be processed by an electronic device (420) which includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof to enable or implement aspects of the matters disclosed in more detail later. The encoded video data (404) (or encoded video bitstream) is drawn as a thin line to emphasize that it has a lower data volume compared to the stream (402) of video pictures and may be stored in a streaming server (405) for later use. For example, one or more streaming client subsystems, such as client subsystems (406) and (408) in Figure 4, can access a streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include a video decoder (410), for example, within an electronic device (430). The video decoder (410) can decode the incoming copy of the encoded video data (407) to produce an outgoing stream of video pictures (411), which may be rendered on a display (412) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., video bitstreams) may be encoded according to a specific video coding / compression standard.Examples of these standards include ITU-T Recommendation H.265. For example, a video coding standard under development is informally known as Versatile Video Coding (VVC). The disclosures may be used in the context of VVC.

[0044] The electronic devices (420) and (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).

[0045] Figure 5 shows an exemplary block diagram of a video decoder (510). The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) can be used in place of the video decoder (410) in the example of Figure 4.

[0046] A receiver (531) can receive one or more encoded video sequences to be decoded by a video decoder (510). In one embodiment, one encoded video sequence is received at a time, and the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. Encoded video sequences can be received from a channel (501) which may be a hardware / software link to a storage device that stores encoded video data. The receiver (531) may receive encoded video data together with other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective user entities (not shown). The receiver (531) may isolate the encoded video sequence from other data. To combat network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser 520 (hereinafter, “Parser (520)”). In certain applications, the buffer memory (515) is part of the video decoder (510). In other cases, it may be located outside the video decoder (510) (not shown). In yet other cases, for example to counter network jitter, a buffer memory (not shown) may exist outside the video decoder (510), and further, for example to handle playback timing, another buffer memory (515) may exist inside the video decoder (510). When the receiver (531) is receiving data from a storage / transfer device with sufficient bandwidth and controllability or from an isosynchronous network, the buffer memory (515) may not be necessary or may be made small. For example, in use on a best-effort packet network such as the Internet, the buffer memory (515) may be necessary and may be made relatively large and, advantageously, of an adaptable size, and may be implemented at least in part in an operating system or similar element (not shown) outside the video decoder (510).

[0047] The video decoder (510) may include a parser (520) for reconstructing symbols (521) from the encoded video sequence. These symbol categories may include information used to manage the operation of the video decoder (510) and, possibly, information for controlling rendering devices, such as a renderer (512) (e.g., a display screen), which are not part of the electronic device (530) but can be coupled to the electronic device (530), as shown in Figure 5. Control information for (one or more) rendering devices may take the form of Supplementary Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not shown). The parser (520) may parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence may be by video coding techniques or standards and may follow a variety of principles, including variable-length coding, Huffman coding, and context-independent or non-context-independent arithmetic coding. The parser(520) can extract from the encoded video sequence a set of subgroup parameters relating to at least one of the subgroups of pixels in the video decoder, based on at least one parameter corresponding to a group. Subgroups may include groups of pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), and prediction units (PU). The parser(520) can also extract information from the encoded video sequence information such as transform coefficients, quantization parameter values, and motion vectors.

[0048] The parser (520) may perform entropy decoding / parsing on the video sequence received from buffer memory (515) to produce symbols (521).

[0049] The reconstruction of symbol (521) may involve multiple different units, depending on the type of encoded video picture or part thereof and other factors (e.g., interpicture and intrapicture, interblock and intrablock, etc.). Which units are involved and how they are involved can be controlled by subgroup control information parsed from the encoded video sequence by the parser (520). The flow of such subgroup control information between the parser (520) and the following multiple units is not illustrated for clarity.

[0050] Beyond the functional blocks described above, the video decoder (510) can be conceptually subdivided into numerous functional units, as described below. In a practical implementation operating under commercial constraints, many of these units may interact closely with each other and be at least partially integrated. However, for the purpose of illustrating the matters to be disclosed, the following conceptual subdivision into functional units is appropriate.

[0051] The first unit is the scaler / inverse unit (551). The scaler / inverse unit (551) receives quantized transformation coefficients as symbols (521) (one or more) from the parser (520), along with control information including block size, quantization coefficients, and quantization scaling matrix, indicating which transformation should be used. The scaler / inverse unit (551) can output a block containing sample values ​​that can be input to the aggregator (555).

[0052] In some cases, the output samples of the scaler / inverse unit (551) may relate to intracoded blocks. Intracoded blocks are blocks that do not use prediction information from previously reconstructed pictures, but can use prediction information from previously reconstructed portions of the current picture. Such prediction information can be provided by the intrapicture prediction unit (552). In some cases, the intrapicture prediction unit (552) generates a block of the same size and shape as the block being reconstructed, using surrounding already reconstructed information fetched from the current picture buffer (558). The current picture buffer (558) buffers, for example, partially reconstructed current pictures and / or fully reconstructed current pictures. In some cases, the aggregator (555) adds the prediction information generated by the intraprediction unit (552) to the output sample information provided by the scaler / inverse unit (551) for each sample.

[0053] In other cases, the output samples of the scaler / inverse unit (551) may relate to an intercoded, potentially motion-compensated block. In such cases, a motion-compensated prediction unit (553) can access a reference picture memory (557) to fetch samples to be used for prediction. After the fetched samples are motion-compensated according to symbols (521) related to the block, these samples can be appended by an aggregator (555) to the output of the scaler / inverse unit (551) (in this case, called residual samples or residual signals) to generate output sample information. From there, the addresses in the reference picture memory (557) from which the motion-compensated prediction unit (553) fetches prediction samples can be controlled by a motion vector, which may be available to the motion-compensated prediction unit (553) in the form of symbols (521) having, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​fetched from reference picture memory (557) when the precise motion vector of a subsample is used, or motion vector prediction mechanisms.

[0054] The output samples from the aggregator (555) can be subjected to various loop filtering techniques in the loop filter unit (556). The video compression technique may include in-loop filtering, which is controlled by parameters included in the encoded video sequence (also called the encoded video bitstream) and made available to the loop filter unit (556) as symbols (521) from the parser (520). The video compression can also respond to metadata obtained during the decoding of preceding portions (in decoding order) of the encoded picture or encoded video sequence, as well as to previously reconstructed and loop-filtered sample values.

[0055] The output of the loop filter unit (556) can be a sample stream that can be output to the renderer (512), which can also be stored in reference picture memory (557) for use in future interpicture prediction.

[0056] A particular encoded picture, once fully reconstructed, can be used as a reference picture for future predictions. For example, when the encoded picture corresponding to the current picture is fully reconstructed and that encoded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) may become part of the reference picture memory (557), and a new current picture buffer may be reallocated before the reconstruction of the next encoded picture begins.

[0057] The video decoder (510) may perform decoding according to a specified video compression technique or standard, such as ITU-T Recommendation H.265. The encoded video sequence may conform to the syntax defined by the video compression technique or standard used, in the sense that it faithfully adheres to both the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. Specifically, a profile may select a particular tool from all the tools available in the video compression technique or standard, such that only that tool is available for use under that profile. Also, for compliance, the complexity of the encoded video sequence must be within the range defined by the level of the video compression technique or standard. Depending on the case, the level may constrain the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limitations set by the level may, depending on the case, be further constrained through the Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the encoded video sequence.

[0058] In one embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of one or more encoded video sequences. The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may take the form of, for example, a temporal, spatial, or signal-to-noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, or forward error correction codes.

[0059] Figure 6 shows an exemplary block diagram of a video encoder (603). The video encoder (603) is included in an electronic device (620). For example, the electronic device (620) includes a transmitter (640) (e.g., a transmitting circuit). The video encoder (603) can be used in place of the video encoder (403) in the example of Figure 4.

[0060] The video encoder (603) may receive video samples from a video source (601) (not part of the electronic device (620) in the example of Figure 6) that can capture (one or more) video images to be encoded by the encoder (603). In another example, the video source (601) is part of the electronic device (620).

[0061] The video source (601) may provide a source video sequence encoded by a video encoder (603) in the form of a digital video sample stream, which can have any preferred bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any preferred sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service delivery system, the video source (601) may be a storage device storing pre-prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a series of individual pictures that convey motion when viewed sequentially. These pictures themselves can be organized as a spatial array of pixels, and each pixel may have one or more samples depending on the sampling structure, color space, etc., used. Those skilled in the art will immediately understand the relationship between pixels and samples. The following description focuses on samples.

[0062] According to one embodiment, the video encoder (603) can code and compress pictures from a source video sequence into an encoded video sequence (643) in real time or under other required time constraints. One function of the controller (650) is to enforce an appropriate coding speed. In some embodiments, the controller (650) controls and is functionally coupled to other functional units, such as those described later. The coupling is not illustrated for clarity. Parameters set by the controller (650) may include rate control-related parameters (picture skip, quantizer, lambda value of rate distortion optimization technique, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) may be configured to have other preferred functions related to the video encoder (603) that are optimized for a particular system design.

[0063] In some embodiments, the video encoder (603) is configured to operate in a coding loop. For an oversimplified explanation, in one example, the coding loop may include a source coder (630) (responsible for creating symbols, such as a symbol stream, based on, for example, the input picture to be encoded and one or more reference pictures) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to generate sample data, similar to how a (remote) decoder would create them. The reconstructed sample stream (sample data) is input to the reference picture memory (634). Since decoding the symbol stream yields bit-accurate results independent of the decoder location (local or remote), the contents of the reference picture memory (634) are also bit-accurate between the local and remote encoders. In other words, the predictive portion of the encoder "sees" the exact same sample values ​​as the reference picture samples that the decoder "sees" when using predictions during decoding. The fundamental principle of this reference picture synchronization (and the resulting drift when synchronization cannot be maintained, for example, due to channel errors) is also used in some related technologies.

[0064] The operation of the “local” decoder (633) can be the same as that of a “remote” decoder, such as a video decoder (510), which has already been described in detail above in relation to Figure 5. However, also briefly referring to Figure 5, since symbols are available and the encoding / decoding of symbols to an encoded video sequence by the entropy coder (645) and parser (520) can be reversible, the entropy decoding portion of the video decoder (510), including the buffer memory (515) and parser (520), does not need to be fully implemented in the local decoder (633).

[0065] In one embodiment, the decoder technology, excluding parsing / entropy decoding present in the decoder, is present in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosure focuses on the decoder operation. A description of the encoder technology can be omitted, as it is the inverse of the thoroughly described decoder technology. More detailed descriptions are provided below for certain specific areas.

[0066] During operation, in some examples, the source coder (630) may perform motion-compensated predictive coding, which predictively codes the input picture against one or more previously coded pictures from a video sequence designated as “reference pictures”. Thus, the coding engine (632) codes the difference between the pixel blocks of the input picture and the pixel blocks of one or more reference pictures that may be selected as prediction criteria for the input picture.

[0067] The local video decoder (633) can decode the encoded video data of a picture that may be designated as a reference picture based on symbols created by the source coder (630). The operation of the coding engine (632) can, advantageously, be a lossy process. When the encoded video data can be decoded by a video decoder (not shown in Figure 6), the reconstructed video sequence may typically be a replica of the source video sequence with some error. The local video decoder (633) can replicate the decoding process that may be performed by the video decoder on the reference picture and cause the reconstructed reference picture to be stored in the reference picture memory (634). Thus, the video encoder (603) can locally store a copy of the reconstructed reference picture that has content common to the reconstructed reference picture that will be obtained by the far-end video decoder.

[0068] The predictor (635) may perform a predictive search for the coding engine (632). That is, with respect to a new picture to be coded, the predictor (636) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors or block shapes that can serve as appropriate predictive criteria for the new picture. The predictor (635) may operate pixel block by pixel to find appropriate predictive criteria. The input picture may have predictive criteria drawn from multiple reference pictures stored in the reference picture memory (634), as determined by the search results obtained by the predictor (635).

[0069] The controller (650) may manage the coding process of the source coder (630), including, for example, setting parameters and subgroup parameters used to encode video data.

[0070] The outputs of all the aforementioned functional units can be subjected to entropy coding in the entropy coder (645). The entropy coder (645) converts the symbols generated by the various functional units into encoded video sequences by applying lossless compression to the symbols according to techniques such as Huffman coding, variable-length coding, and arithmetic coding.

[0071] The transmitter (640) may buffer (one or more) encoded video sequences generated by the entropy coder (645) and prepare them for transmission over the communication channel (660). The communication channel (660) may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (640) may merge the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).

[0072] The controller (650) may manage the operation of the video encoder (603). In coding, the controller (650) may assign each encoded picture a specific encoded picture type that may influence the coding techniques that may be applied to that picture. For example, a picture may often be assigned one of the following picture types:

[0073] An intra-picture (I-picture) may be one that can be encoded and decoded without using any other picture in the sequence as a source for prediction. Some video codecs allow several different types of intra-pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art know these variations of I-pictures, as well as their respective uses and characteristics.

[0074] A prediction picture (P-picture) may be encoded and decoded using intra-prediction or inter-prediction, with at most one motion vector and a reference index to predict the sample values ​​of each block.

[0075] A bidirectional predictive picture (B-picture) may be one that can be encoded and decoded using intra-prediction or inter-prediction, using up to two motion vectors and a reference index to predict the sample values ​​of each block. Similarly, a multiple predictive picture may use three or more reference pictures and associated metadata for the reconstruction of a single block.

[0076] A source picture can generally be subdivided spatially into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each), and each block can be coded. Blocks can be coded predictively by referring to other (already coded) blocks, determined by the coding assignment applied to each picture in those blocks. For example, blocks of picture I can be coded unpredictably, or they can be coded predictively by referring to already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of picture P can be coded unpredictably, or via spatial or temporal prediction by referring to one previously coded reference picture. Blocks of picture B can be coded unpredictably, or via spatial or temporal prediction by referring to one or two previously coded reference pictures.

[0077] The video encoder (603) may perform coding operations in accordance with a predetermined video coding technique or standard, such as ITU-T Recommendation H.265. In its operation, the video encoder (603) may perform various compression operations, including predictive coding operations that take advantage of temporal and spatial redundancy in the input video sequence. The coded video data may therefore conform to the syntax defined by the video coding technique or standard being used.

[0078] In one embodiment, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, and the like.

[0079] Video can be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) utilizes spatial correlations within a given picture, while inter-picture prediction utilizes (temporal or other) correlations between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into multiple blocks. When a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, that block in the current picture can be coded by a vector called a motion vector. The motion vector points to a reference block in the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0080] In some embodiments, a dual prediction technique can be used in interpicture prediction. According to the dual prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture, both of which are earlier in the decoding order than the current picture in the image (but may be past and future in the display order, respectively). A block in the current picture can be coded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. That block can be predicted by a combination of the first and second reference blocks.

[0081] Furthermore, merge mode techniques can be used to improve coding efficiency in interpicture prediction.

[0082] According to some embodiments of this disclosure, predictions such as interpicture prediction and intrapicture prediction are performed in units of blocks. For example, according to the HEVC standard, a picture in a sequence of video pictures is divided into multiple coding tree units (CTUs) for compression, and these CTUs in a picture have the same size, for example, 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU contains three coding tree blocks (CTBs), which are one luma CTB and two chroma CTBs. Each CTU can be recursively quadtree-divided into one or more coding units (CUs). For example, a 64x64 pixel CTU can be divided into one 64x64 pixel CU, or four 32x32 pixel CUs, or sixteen 16x16 pixel CUs. In one example, each CU is analyzed to determine the prediction type of that CU, for example, inter-prediction type or intra-prediction type. A CU is divided into one or more prediction units (PUs) depending on its temporal and / or spatial predictability. Generally, each PU includes a Luma prediction block (PB) and two Chroma PBs. In one embodiment, the prediction operation during coding (encoding / decoding) is performed in units of prediction blocks. Using a Luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values ​​(e.g., Luma values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, and similar.

[0083] Figure 7 shows an exemplary diagram of a video encoder (703). The video encoder (703) is configured to receive sample values ​​of a processing block (e.g., a prediction block) in the current video picture within a sequence of video pictures, and to encode the processing block into an encoded picture which is part of an encoded video sequence. In one example, the video encoder (703) is used instead of the video encoder (403) in the example of Figure 4.

[0084] In the HEVC example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as an 8x8 sample of a prediction block. The video encoder (703) determines, for example using rate-distortion optimization, whether the processing block is best coded using intra-mode, inter-mode, or bi-prediction mode. If the processing block is coded in intra-mode, the video encoder (703) can encode the processing block into an encoded picture using the intra-prediction technique; if the processing block is coded in inter-mode or bi-prediction mode, the video encoder (703) can encode the processing block into an encoded picture using the inter-prediction technique or the bi-prediction technique, respectively. In certain video coding techniques, the merge mode may be an inter-picture prediction submode in which motion vectors are derived from one or more motion vector predictors without benefiting from the encoded motion vector components outside the predictors. In certain other video coding techniques, there may be motion vector components applicable to the target block. In one example, the video encoder (703) includes other components, such as a mode determination module (not shown) for determining the mode of a processing block.

[0085] In the example shown in Figure 7, the video encoder (703) includes an interencoder (730), an intraencoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725), all coupled together as shown in Figure 7.

[0086] The interencoder (730) is configured to receive a sample of the current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a preceding picture and a subsequent picture), generate interprediction information (e.g., a description of redundant information according to the intercoding technique, motion vectors, merge mode information), and then use some preferred technique to compute an interprediction result (e.g., a predicted block) based on the interprediction information. In some examples, the reference picture is a reference picture decoded based on encoded video information.

[0087] The intra encoder (722) receives a sample of the current block (e.g., a processing block), and in some cases compares the block to an already coded block in the same picture to generate transformed quantization coefficients, and in some cases also generates intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques). In one example, the intra encoder (722) also calculates an intra prediction result (e.g., a prediction block) based on the intra prediction information and a reference block in the same picture.

[0088] The general controller (721) is configured to determine general control data and to control other components of the video encoder (703) based on the general control data. In one example, the general controller (721) determines the mode of a block and provides control signals to the switch (726) based on that mode. For example, when the mode is intra-mode, the general controller (721) controls the switch (726) to select intra-mode results for use by the residual calculator (723) and controls the entropy encoder (725) to select intra-prediction information and include it in the bitstream. When the mode is inter-mode, the general controller (721) controls the switch (726) to select inter-prediction results for use by the residual calculator (723) and controls the entropy encoder (725) to select inter-prediction information and include it in the bitstream.

[0089] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra-encoder (722) or inter-encoder (730). The residual encoder (724) operates on the residual data and is configured to encode the residual data to generate conversion coefficients. In one example, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain and generate conversion coefficients. The conversion coefficients are then subjected to a quantization process to obtain quantized conversion coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be suitably used by the intra-encoder (722) and inter-encoder (730). For example, an interencoder (730) can generate a decoded block based on decoded residual data and interprediction information, and an intraencoder (722) can generate a decoded block based on decoded residual data and intraprediction information. The decoded block is suitably processed to generate a decoded picture, which can be buffered in a memory circuit (not shown) and, in some examples, can be used as a reference picture.

[0090] The entropy encoder (725) is configured to format the bitstream to include the encoded blocks. The entropy encoder (725) is configured to include various information in the bitstream according to a preferred standard, such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other preferred information in the bitstream. According to the present disclosure, when coding blocks in either inter-mode or bi-prediction mode merge submode, residual information is not present.

[0091] Figure 8 shows an exemplary diagram of a video decoder (810). The video decoder (810) is configured to receive an encoded picture, which is part of an encoded video sequence, and to decode the encoded picture to produce a reconstructed picture. In one example, the video decoder (810) is used instead of the video decoder (410) in the example of Figure 4.

[0092] In the example shown in Figure 8, the video decoder (810) includes an entropy decoder (871), an interdecoder (880), a residual decoder (873), a reconfiguration module (874), and an intradecoder (872), all coupled together as shown in Figure 8.

[0093] The entropy decoder (871) may be configured to reconstruct specific symbols from the encoded picture that represent the syntax elements constituting the encoded picture. Such symbols may include, for example, the mode in which the block is coded (e.g., intra-mode, inter-mode, bi-prediction mode, merge sub-mode, or two of the latter in other sub-modes), as well as prediction information (e.g., intra-prediction information or inter-prediction information) that can identify specific samples or metadata used for prediction by the intra-decoder (872) or inter-decoder (880), respectively. The symbols may also include residual information in the form of quantized transformation coefficients, and similar information. In one example, when the prediction mode is inter-mode or bi-prediction mode, inter-prediction information is provided to the inter-decoder (880), and when the prediction type is intra-prediction type, intra-prediction information is provided to the intra-decoder (872). The residual information may be subjected to inverse quantization and provided to the residual decoder (873).

[0094] The interdecoder (880) is configured to receive interprediction information and generate interprediction results based on the interprediction information.

[0095] The intra decoder (872) is configured to receive intra prediction information and generate prediction results based on the intra prediction information.

[0096] The residual decoder (873) is configured to perform inverse quantization to extract the dequantized transformation coefficients, and then process the dequantized transformation coefficients to convert the residual information from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (including quantizer parameters (QP)), which may be provided by the entropy decoder (871) (the data path is not shown as this may only be low-volume control information).

[0097] The reconstruction module (874) is configured to combine the residual information output by the residual decoder (873) and the prediction results (output by the inter or intra prediction module, as applicable) in the spatial domain to form a reconstruction block. The reconstruction block can be part of a reconstruction picture, and conversely, the reconstruction picture can be part of a reconstruction video. In addition, other suitable processes, such as deblocking and similar processes, can be performed to improve visual quality.

[0098] The video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) may be implemented using any preferred technology. In one embodiment, the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) may be implemented using one or more processors that execute software instructions.

[0099] VVC allows the use of various interpretation modes. For interpreted CUs, motion parameters can include (one or more) MVs, one or more reference picture indices, a reference picture list usage index, and additional information about specific coding features used for sample generation by interpretation. Motion parameters can be signaled explicitly or implicitly. When a CU is coded in skip mode, the CU can be associated with a PU and does not need to have significant residual coefficients, coded motion vector delta or MV difference (e.g., MVD), or reference picture indices. A merge mode can be specified, in which case the motion parameters for the current CU are derived from (one or more) adjacent CUs, including spatial and / or temporal candidates, and optionally from additional information, such as that introduced in VVC. The merge mode can be applied not only to skip mode but also to interpreted CUs. In one example, an alternative to merge mode is the explicit transmission of motion parameters, with each CU explicitly signaling (one or more) MVs, the corresponding reference picture index for each reference picture list, a reference picture list usage flag, and other information.

[0100] For example, in one embodiment such as in VVC, the VVC test model (VTM) reference software supports extended merge prediction, merge motion vector difference (MMVD) mode, adaptive motion vector prediction (AMVP) mode using symmetric MVD signaling, affine motion compensation prediction, subblock-based temporal motion vector prediction (SbTMVP), adaptive motion vector resolution (AMVR), motion field storage (1 / 16 luma sample MV storage and 8x8 motion field compression), bi-prediction with CU-level weights (BCW), bi-directional optical flow (BDOF), prediction refinement using optical flow (PROF), decoder side motion vector refinement (DMVR), combined inter and intra prediction (CIIP), and geometric partitioning mode. This includes one or more refined interpredictive coding tools, including partitioning mode (GPM) and similar methods. Interpretation and related methods are described in detail below.

[0101] Extended merge prediction may be used in some cases. For example, in VTM4, the merge candidate list is constructed by including, in order, the following five types of candidates: (one or more) spatial motion vector predictors (MVPs) from (one or more) spatially adjacent CUs, (one or more) temporal MVPs from (one or more) collated CUs, (one or more) historical-based MVPs (HMVPs) from first-in, first-out (FIFO) tables, (one or more) pairwise-mean MVPs, and (one or more) zero MVs.

[0102] The size of the merge candidate list can be signaled within the slice header. In one example, the maximum allowable size of the merge candidate list is 6 in VTM4. For each CU coded in merge mode, the index of the best merge candidate (e.g., the merge index) can be encoded using truncated unary binarization (TU). The first bin of the merge index can be coded using context (e.g., context-adaptive binary arithmetic coding (CABAC)), while bypass coding may be used for the other bins.

[0103] Several examples of the generation process for each category of merge candidates are provided below. In one embodiment, (one or more) spatial candidates are derived as follows. The derivation of spatial merge candidates in VVC may be the same as that in HEVC. In one example, up to four merge candidates are selected from the candidates located in the positions shown in Figure 9. Figure 9 shows the positions of spatial merge candidates according to one embodiment of the present disclosure. Referring to Figure 9, the derivation order is B1, A1, B0, A0, and B2. Position B2 is considered only if any of the CUs in positions A0, B0, B1, and A1 are unavailable (for example, because that CU belongs to a different slice or a different tile) or are intracoded. After the candidate for position A1 is added, the addition of the remaining candidates is subjected to redundancy checks to ensure that candidates with the same motion information are excluded from the candidate list in order to improve coding efficiency.

[0104] To reduce computational complexity, the redundancy check described above does not consider all possible candidate pairs. Instead, only pairs linked by the arrows in Figure 10 are considered, and a candidate is added to the candidate list only if the corresponding candidate used for the redundancy check does not have the same motion information. Figure 10 shows candidate pairs considered for redundancy checking of spatial merge candidates according to one embodiment of the present disclosure. Referring to Figure 10, the pairs linked by the respective arrows are A1 and B1, A1 and A0, A1 and B2, B1 and B0, and B1 and B2. Thus, candidates for positions B1, A0, and / or B2 can be compared with the candidate for position A1, and candidates for positions B0 and / or B2 can be compared with the candidate for position B1.

[0105] In one embodiment, (one or more) time candidates are derived as follows. In one example, only one time merge candidate is added to the candidate list. Figure 11 shows exemplary motion vector scaling for time merge candidates. To derive a time merge candidate for the current CU (1111) in the current picture (1101), a scaled MV (1121) (e.g., shown as a dotted line in Figure 11) may be derived based on a collated CU (1112) belonging to a collated reference picture (1104). In one example, a collated reference picture (also referred to as a collated picture) is a specific reference picture used, for example, for time motion vector prediction. A collated reference picture used for time motion vector prediction can be indicated by a reference index in syntax, such as high-level syntax (e.g., picture header, slice header).

[0106] The list of reference pictures used to derive the collated CU(1112) can be explicitly signaled within the slice header. A scaled MV(1121) for the time merge candidate can be obtained, as shown by the dotted line in Figure 11. The scaled MV(1121) can be scaled from the MV of the collated CU(1112) using picture order count (POC) distances tb and td. The POC distance tb can be defined as the POC difference between the current reference picture (1102) and the current picture (1101) of the current picture (1101). The POC distance td can be defined as the POC difference between the collated reference picture (1104) and the collated reference picture (1103) of the collated reference picture (1103). The reference picture index of the time merge candidate can be set to zero.

[0107] Figure 12 shows exemplary candidate positions (e.g., C0 and C1) for the current CU time merge candidate. A time merge candidate position can be selected from candidate positions C0 and C1. Candidate position C0 is located in the lower right corner of the current CU's collated CU(1210). Candidate position C1 is located in the center of the current CU's collated CU(1210). If the CU at candidate position C0 is unavailable, intracoded, or outside the current row of the CTU, candidate position C1 is used to derive the time merge candidate. Otherwise, for example, if the CU at candidate position C0 is available, intracoded, and within the current row of the CTU, candidate position C0 is used to derive the time merge candidate.

[0108] The Merge Motion Vector Difference (MMVD) mode can be used in merge or skip modes with a motion vector representation method. For example, merge candidates (one or more) used in VVC can be reused in MMVD mode. Candidates can be selected from the merge candidates as starting points (e.g., MV predictors (MVPs)) and can be further extended by MMVD mode. MMVD mode can provide a new motion vector representation using simplified signaling. This motion vector representation method includes a starting point and an MV difference (MVD). In one example, the MVD is indicated by the magnitude of the MVD (or magnitude of motion) and the direction of the MVD (e.g., direction of motion).

[0109] The MMVD mode can use a merge candidate list, such as the one used in VVC. In one embodiment, only (one or more) candidates of the default merge type (e.g., MRG_TYPE_DEFAULT_N) are considered in MMVD mode. The starting point can be indicated or defined by a base candidate index (IDX). The base candidate index can indicate a candidate (e.g., the best candidate) among multiple candidates (e.g., multiple base candidates) in the merge candidate list. Table 1 shows an exemplary relationship between the base candidate index and the corresponding starting point. A base candidate index of 0, 1, 2, or 3 indicates that the corresponding starting point is the 1st MVP, 2nd MVP, 3rd MVP, or 4th MVP. In one example, if the number of base candidates is equal to 1, the base candidate IDX is not signaled. [Table 1]

[0110] The distance index can indicate information about the magnitude of motion of an MVD, such as the size of the MVD. For example, the distance index indicates the distance (e.g., a predetermined distance) from a starting point (e.g., an MVP indicated by the base candidate index). In one example, that distance is one of several predetermined distances, such as those shown in Table 2. Table 2 shows an exemplary relationship between the distance index and the corresponding distance (in units of samples or pixels). In Table 2, 1 pel is equal to 1 sample or 1 pixel. For example, a distance index of 1 indicates a distance of 1 / 2 pel, or 1 / 2 sample. [Table 2]

[0111] The direction index can represent the direction of the MVD relative to the starting point. The direction index can represent one of several directions, such as the four directions shown in Table 3. For example, a direction index of 00 indicates that the direction of the MVD is along the positive x-axis. [Table 3]

[0112] The MMVD flag can be signaled after sending the skip and merge flags. If the skip and merge flags are true, the MMVD flag may be parsed. For example, if the MMVD flag is equal to 1, the MMVD syntax (e.g., including distance and / or direction indices) may be parsed. If the MMVD flag is not equal to 1, the AFFINE flag may be parsed. If the AFFINE flag is equal to 1, the AFFINE mode is used to code the current block. If the AFFINE flag is not equal to 1, the skip / merge index may be parsed for the skip / merge mode, for example, as used in VTM.

[0113] Figures 13-14 show an example of the search process in MMVD mode. By performing the search process, it is possible to determine the index, including the base candidate index, direction index, and / or distance index, for the current block (1300) within the current picture (or referred to as the current frame) (1301).

[0114] The first motion vector (MV) (1311) and the second MV (1321) belonging to the first merge candidate are shown. The first merge candidate can be a merge candidate in the merge candidate list currently constructed for block (1300). The first and second MVs (1311) and (1321) can be associated with two reference pictures (1302) and (1303) in reference picture lists L0 and L1, respectively. Thus, the two starting points (1411) and (1421) in Figures 13-14 can be determined within reference pictures (1302) and (1303), respectively.

[0115] In one example, based on the starting points (1411) and (1421), a number of predetermined points extending vertically (represented by +Y or -Y) or horizontally (represented by +X and -X) from the starting points (1411) and (1421) within the reference pictures (1302) and (1303) may be evaluated. In one example, using pairs of points that are mirror images of each other with respect to the respective starting points (1411) or (1421), such as the pair of points (1414) and (1424) or the pair of points (1415) and (1425), a pair of MV(1314) and (1324) or a pair of MV(1315) and (1325) may be determined to form a candidate MV predictor (MVP) for the current block (1300). MVP candidates determined based on predetermined points surrounding the starting points (1411) and / or (1421) may be evaluated. Referring to Figure 13, the MVD(1312) between the first MV(1311) and MV(1314) has a magnitude of 1S. The MVD(1322) between the second MV(1321) and MV(1324) has a magnitude of 1S. Similarly, the MVD between the first MV(1311) and MV(1315) has a magnitude of 2S. The MVD between the second MV(1321) and MV(1325) has a magnitude of 2S.

[0116] In addition to the first merge candidate, other available or valid merge candidates in the current block (1300) merge candidate list may also be evaluated. In one example, for a single-prediction merge candidate, only one prediction direction associated with one of the two reference picture lists is evaluated.

[0117] In one example, the best MVP candidate can be determined based on the evaluation. Therefore, the optimal merge candidate corresponding to the optimal MVP candidate can be selected from the merge list, and the direction and distance of movement can also be determined. For example, the base candidate index can be determined based on the selected merge candidate and Table 1. Based on the selected MVP, the direction and distance (e.g., 2S) of point (1415) relative to the starting point (1411) can be determined, for example, corresponding to a given point (1415) (or (1425)). The direction index and distance index can be determined accordingly according to Tables 2 and 3.

[0118] As mentioned above, the MVD in MMVD mode can be represented using two indexes, such as a distance index and a direction index. Alternatively, the MVD in MMVD mode can be represented using a single index, for example, by using a table that pairs a single index with the MVD.

[0119] For example, in some prediction modes such as MMVD mode and affine MMVD mode, template matching (TM)-based candidate sorting may be used. In one embodiment, the MMVD offset is extended for MMVD mode and affine MMVD mode. Figure 15 shows additional refinement positions along multiple oblique angles, such as an oblique angle of k × π / 8, where k is an integer from 0 to 15. The number of directions can be increased, for example from 4 directions (e.g., +X, -X, +Y, and -Y) to 16 directions (e.g., k=0,1,2,…,15). In one example, each of these 16 directions is represented by the angle between the +X direction and the direction indicated by the center point (1500) and one of points 1 through 16. For example, point 1 indicates the +X direction with an angle of 0 (i.e., k=0), point 2 indicates the direction along an angle of 1 × π / 8 (i.e., k=1), and so on.

[0120] TM can be performed in MMVD mode. For example, for each MMVD refinement position, the TM cost may be determined based on the current template of the current block and one or more reference templates. The TM cost may be determined using any method, such as the sum of absolute differences (SAD) (e.g., SAD cost), the sum of absolute transformation differences (SATD), the sum of squared errors (SSE), mean-removed SAD / SATD / SSE, variance, partial SAD, partial SSE, partial SATD, or similar.

[0121] The current template of the current block may include some preferred samples, such as the top row of samples in the current block and / or the left column of samples in the current block. Based on the TM cost (e.g., SAD cost) between the current template for the refinement position and the corresponding reference template, the MMVD refinement positions for each base candidate (e.g., MVP) can be sorted, such as all possible MMVD refinement positions (e.g., 16x6 representing 16 directions and 6 sizes). In one example, the top MMVD refinement positions with the lowest TM cost (e.g., the lowest SAD cost) are retained as MMVD refinement positions available for MMVD index coding. For example, a subset (e.g., 8) of MMVD refinement positions with the lowest TM cost is used for MMVD index coding. For example, the MMVD index indicates which one of the subset of MMVD refinement positions with the lowest TM cost is selected to code the current block. In one example, an MMVD index of 0 indicates that the MVD (e.g., the MMVD refined position) corresponding to the smallest TM cost is used to code the current block. The MMVD index can be binarized, for example, by a rice code with a parameter equal to 2.

[0122] In one embodiment, in addition to the above-described MMVD offset extension, such as in Figure 15, the affine MMVD sorting is extended to add further refinement positions along a diagonal angle of k × π / 4. After sorting, the top half of the refinement positions having the minimum TM cost (e.g., SAD cost) are held for coding the current block.

[0123] To improve coding efficiency and reduce the transmission overhead of (one or more) MVs, subblock-level MV refinement may be applied to extend CU-level time-motion vector prediction (TMVP). For example, the subblock-based TMVP (SbTMVP) mode allows inheritance of subblock-level motion information from collated reference pictures. As mentioned above, collated reference pictures can be indicated by reference indices in syntax, such as high-level syntax (e.g., picture header, slice header). Each subblock of the current CU in the current picture (e.g., a current CU with a large size) can have its own motion information without explicitly transmitting the block partition structure or its respective motion information. In SbTMVP mode, the motion information for each subblock can be obtained in, for example, three steps as follows: In the first step, the displacement vector (DV) of the current CU can be derived. The DV can indicate a block in the collated reference picture, for example, the DV points from the current block in the current picture to a block in the collated reference picture. Therefore, the block indicated by the DV is considered to be a colcate of the current block and is referred to as the colcate block of the current block. In the second step, the availability of SbTMVP candidates can be checked and the central motion (e.g., the central motion of the current CU) may be derived. In the third step, subblock motion information may be derived from the corresponding subblocks in the colcate block using the DV. These three steps may be combined into one or two steps, and / or the order of these three steps may be adjusted.

[0124] Unlike TMVP candidate derivation, which derives the time MV from a collated block in a reference frame or reference picture, SbTMVP mode allows the application of DV (e.g., DV derived from the MV of the adjacent CU to the left of the current CU) to locate the corresponding subblock in the collated reference picture for each subblock in the current CU within the current picture. If the corresponding subblock is not encoded, the motion information of the current subblock may be set to the central motion of the collated block.

[0125] SbTMVP mode can be supported by various video coding standards, including VVC, for example. Similar to TMVP mode, HEVC, for example, can use motion fields (also referred to as motion information fields or MV fields) in a collated reference picture to improve MV prediction and merging modes for CUs in the current picture. In one example, the same collated reference picture used by TMVP mode is used in SbTMVP mode. In one example, SbTMVP mode differs from TMVP mode in the following respects: (i) TMVP mode predicts motion information at the CU level, whereas SbTMVP mode predicts motion information at the subCU level; and (ii) TMVP mode fetches time MV from collated blocks in the collated reference picture (e.g., the collated block is the bottom right or center block relative to the current CU), whereas SbTMVP mode can apply motion shifts before fetching time motion information from the collated reference picture. In one example, the motion shift used in SbTMVP mode is currently obtained from one of the spatially adjacent blocks of the CU.

[0126] Figures 16-17 show an exemplary SbTMVP process used in SbTMVP mode. This SbTMVP process can predict the motion shift (MV) of a sub-CU (e.g., sub-block) within the current CU (e.g., current block) (1601) of the current picture (1711) in, for example, two steps. In the first step, spatially adjacent elements (spatial neighbors) (e.g., A1) to the current block (1601) in Figures 16-17 are examined. If the spatial neighbor (e.g., A1) has an MV (1721) that uses a collated reference picture (1712) as the reference picture of the spatial neighbor (e.g., A1), then the MV (1721) may be selected to be the motion shift (or DV) that should be applied to the current block (1601). If no such MV (e.g., an MV using collated reference picture (1712) as the reference picture) is identified, the motion shift or DV may be set to zero MV (e.g., (0,0)). In some examples, if no such MV is identified for spatial neighbor A1, the MVs (one or more) of further spatial neighbors such as A0, B0, and B1 are checked.

[0127] In the second step, the motion shift or DV(1721) identified in the first step can be applied to the current block (1601) (for example, by adding DV(1721) to the coordinates of the current block) to obtain sub-CU level motion information (e.g., including MV and reference index) from the collated reference picture (1712). In the example shown in Figure 17, the motion shift or DV(1721) is set to be the MV of the spatial neighbor A1 (e.g., block A1) of the current block (1601). For each sub-CU or sub-block (1731) within the current block (1601), the motion information of the sub-CU or sub-block (1731) can be derived using the motion information of the corresponding collated block (1701) in the collated reference picture (1712) (e.g., motion information of the minimum motion grid covering the central sample of the collated block (1701)). After the motion information of the collated subCU (1732) within the collated block (1701) is identified, the motion information of the collated subCU (1732) can be converted into motion information of the current subCU (1731) (e.g., MV and one or more reference indices) using a scaling method, such as the TMVP process used in HEVC, where time motion scaling is applied to align the time MV reference picture with the current CU reference picture.

[0128] The motion field of the current block (1601), derived based on DV(1721), may include motion information for each subblock (1731) within the current block (1601), such as (one or more) MVs and one or more associated reference indices. The motion field of the current block (1601) may also be referred to as the SbTMVP candidate and corresponds to DV(1721).

[0129] Figure 17 shows an example of a motion field or SbTMVP candidate for block 1601. For example, motion information for a bipredicted subblock (1731(1)) includes a first MV and a first index indicating a first reference picture in reference picture list 0 (L0), and a second MV and a second index indicating a second reference picture in reference picture list 1 (L1). In one example, motion information for a unipredicted subblock (1731(2)) includes an MV and an index indicating a reference picture in L0 or L1.

[0130] In one example, DV(1721) is applied to the center position of the current block (1601) to locate the displaced center position in the collated reference picture (1712). If the block containing the displaced center position is not intercoded, the SbTMVP candidate is considered unavailable. Rather, if the block containing the displaced center position (e.g., collated block (1701)) is intercoded, the motion information of the center position of the current block (1601), referred to as the center motion of the current block (1601), can be derived from the motion information of the block containing the displaced center position in the collated reference picture (1712). In one example, a scaling process may be used to derive the center motion of the current block (1601) from the motion information of the block containing the displaced center position in the collated reference picture (1712). When an SbTMVP candidate is available, DV(1721) can be applied to each subblock(1732) of the current block(1601) to find the corresponding subblock(1731) in the collated reference picture(1712). Using the motion information of the corresponding subblock(1732), the motion information of the subblock(1731) within the current block(1601) can be derived, for example, using the same method used to derive the central motion of the current block(1601). In one example, if the corresponding subblock(1732) is not encoded, the motion information of the current subblock(1731) is set to the central motion of the current block(1601).

[0131] In some cases, such as in VVC, a combined subblock-based merge list containing SbTMVP candidates and (one or more) affine merge candidates is used in signaling for subblock-based merge mode. SbTMVP mode can be enabled or disabled by a Sequence Parameter Set (SPS) flag. When SbTMVP mode is enabled, an SbTMVP candidate (or SbTMVP predictor) is added as the first entry in a subblock-based merge list containing subblock-based merge candidates, which can then be followed by (one or more) affine merge candidates. The size of the subblock-based merge list can be signaled within the SPS. In one example, the maximum allowable size of a subblock-based merge list is 5 in VVC. In one example, multiple SbTMVP candidates may be included in the subblock-based merge list.

[0132] For example, in some cases, such as in VVC, the sub-CU size used in SbTMVP mode is fixed to 8x8, as used in affine merge mode. In one example, SbTMVP mode is only applicable to CUs where both the width and height are 8 or greater. The sub-block size (e.g., 8x8) may be configured to other sizes, such as 4x4 in the use of ECM software models for exploration after VVC. In one example, multiple collated reference pictures, such as two collated frames, are used to provide time-motion information for SbTMVP and / or TMVP in AMVP mode.

[0133] For example, in some examples of SbTMVP modes, such as in VVC and ECM, the current CU's DV (e.g., DV(1721) in Figure 17) is derived solely from the MV of the adjacent CUs of the current CU. However, the SbTMVP candidates derived using the DV may not be an exact match.

[0134] In SbTMVP mode, a DV offset (DVO) can be used. For example, to obtain a more accurate match, the DV (e.g., the original DV) can be modified by the DV offset to determine the updated DV'. For example, the updated DV' is the vector sum of the DV (e.g., the original DV) and the DVO. The original DV can be determined using any method, such as those described in Figures 16-17. For example, the original DV is determined based on the MV of the adjacent blocks of the current block. For the current block, the DVO can be signaled and analyzed to indicate an additional motion offset of the original DV. The DVO can be indicated, for example, by signaling an index that points to the DVO from among several DVO candidates. For example, the DVO is signaled. For example, MMVD mode is used to indicate the DVO, where the DVO is an MVD indicated by a directional index and / or distance index, such as those listed in Tables 2-3. By using DVO, the position of a collated CU (or collated block) within a collated reference picture can be adjusted, and therefore the MV field of the collated CU (or collated block) can be changed based on the DVO. When DVO is non-zero, the updated DV' can be used as a displacement vector indicating the position of the collated CU (or collated block) to perform the SbTMVP process. Referring to Figure 17, instead of using the original DV (e.g., DV(1721)), which is the MV of spatial neighbor A1, the updated DV' can be used to determine the collated block for the current block. The updated DV' (e.g., the vector sum of the original DV and DVO) can be used to derive SbTMVP candidates for the current block.

[0135] In one embodiment, the DVO is directly signaled using any signaling method used to signal the MVD, for example, in AMVP mode, AMVR mode, and / or similar modes. In AMVR mode, the MVD of a block can be signaled at different resolutions, such as 1 / 4, 1 / 2, 1, or 4 luma sample resolutions. The DVO can be signaled at different resolutions using AMVR mode.

[0136] A given DVO list may contain multiple DVO candidates (e.g., DVOs that are currently available for use by the block). One or more indices may be signaled to indicate which of these DVO candidates will be selected as the DVO.

[0137] In one example, DVOs are signaled using MMVD mode. For example, as shown in Tables 2-3, two indices are signaled to indicate a DVO candidate, including a first index indicating the size of the DVO candidate (e.g., a distance index or step index) and a second index indicating the direction of the DVO candidate (e.g., a direction index).

[0138] Referring again to Figure 14, the distance index (or step index) and direction index can be predetermined as described above with reference to the MMVD mode. The distance index indicates motion magnitude information, such as the magnitude of the DVO. For example, the distance index indicates a predetermined distance from the starting point (e.g., initial DV). In one example, the available predetermined distances are shown in Table 2. The direction index represents the direction of the DVO relative to the starting point (e.g., initial DV). The direction index can indicate one of several directions, such as the four directions shown in Table 3.

[0139] In one example, the collated CTU in the collated reference picture is the current CTU and collated, which includes the current block. The current CTU is located in the current picture. In one embodiment, the location of the collated block corresponding to the updated DV' is restricted to being within a first area in the collated reference picture. In one example, the first area in the collated reference picture includes the collated CTU. In one example, the first area in the collated reference picture includes the collated CTU, plus one column of the 4x4 blocks on the right boundary of the collated CTU. The updated DV' may be restricted so that the collated block corresponding to the updated DV' is within the first area in the collated reference picture. In one example, DVO (e.g., the horizontal component DVO of DVO) x and / or the vertical component of DVO y ) is constrained to ensure that the updated DV' satisfies the positional constraints of the collocate block described above.

[0140] In one example, the maximum vertical component of the updated DV' is H. DVO (for example, the vertical component of DVO) y ) is made less than or equal to the amount obtained by subtracting the vertical component DVy of DV from H, for example, DVO y ≤H-DV y Therefore, the maximum horizontal component of the updated DV' is W, and DVO is the horizontal component of DV from W. x It is reduced to less than or equal to the amount obtained by subtracting, for example, DVO x ≤W-DV x That is the case.

[0141] As illustrated in Figures 16-17, a collated block (e.g., (1701)) within a collated reference picture (e.g., (1712)) can be determined based on the DV (e.g., (1721)) of the current block (e.g., (1601)) within the current picture (e.g., (1711)). Thus, the motion information (e.g., TMVP) of each subblock within the current block can be obtained based on the motion information of the corresponding subblock within the collated block. According to one embodiment of the present disclosure, the updated motion information (e.g., updated TMVP) of each subblock within the current block can be determined based on the motion information (e.g., TMVP) of the subblock within the current block and the motion vector offset (MVO) of the current block.

[0142] In one embodiment, the MVO is added to each of the derived subblock level TMVPs of each subblock in the current block to generate an updated subblock level TMVP. The MVO may be signaled and analyzed to show the additional motion offset of each subblock base TMVP of each subblock in the current block determined using SbTMVP mode.

[0143] In one embodiment, the MVO is directly signaled using any signaling method used to signal the MVD, for example, in AMVP mode, AMVR mode, and / or similar modes. In AMVR mode, the MVD of a block can be signaled at different resolutions, such as 1 / 4, 1 / 2, 1, or 4 luma sample resolutions. The MVO can be signaled at different resolutions using AMVR mode.

[0144] A given MVO list may contain multiple MVO candidates (e.g., MVOs that are currently available for use by a block). One or more indices may be signaled to indicate which of the DVO candidates in the given MVO list could be selected as the MVO.

[0145] In one example, MVOs are signaled using MMVD mode. For example, as shown in Tables 2-3, two indices are signaled to indicate MVO candidates, including a first index indicating the size of the MVO candidate (e.g., a distance index or step index) and a second index indicating the direction of the MVO candidate (e.g., a direction index).

[0146] Referring again to Figure 14, the distance index (or step index) and direction index of the MVO can be predetermined as described above with reference to the MMVD mode. The distance index indicates motion magnitude information, such as the magnitude of the MVO, and indicates a predetermined distance from the starting point (e.g., the DV used to determine the position of the collated block in the collated reference picture). In one example, the available predetermined distances are shown in Table 2. The direction index represents the direction of the MVO relative to the starting point (e.g., the DV used to determine the position of the collated block in the collated reference picture). The direction index can indicate one of several directions, such as the four directions shown in Table 3.

[0147] In one example, a subblock within the current block is bipredicted. Referring to Figure 17, the subblock (1731(1)) is bipredicted and has a first MV associated with a first reference picture in reference list L0 and a second MV associated with a second reference picture in reference list L1. An MVO may be applied to the first MV associated with reference list L0. The updated first MV can be the vector sum of the first MV and the MVO. The following embodiment may be applied to the second MV associated with reference list L1.

[0148] In one example, the MVO does not apply to the second MV associated with reference list L1. For example, no MVO applies to the second MV associated with reference list L1. Therefore, the updated motion information of subblock (1731(1)) includes the updated first MV and the second MV.

[0149] In one example, the mirror MVO of the MVO is applied to a second MV associated with the reference list L1. The mirror MVO and the MVO can have the same magnitude and opposite directions. -1 is multiplied by the horizontal and vertical components of the MVO (e.g., those signaled) to obtain the horizontal and vertical components of the mirror MVO, respectively. The updated second MV can be the vector sum of the second MV and the mirror MVO, or the vector difference between the second MV and the MVO. Thus, the updated motion information of the sub-block (1731(1)) includes the updated first MV (e.g., the first MV + MVO) and the updated second MV (e.g., the second MV - MVO).

[0150] In one example, a scaled MVO (MVO’) can be applied to a second MV associated with the reference list L1. The value of each component of the MVO (e.g., the horizontal and vertical components) can be scaled based on a first POC difference and a second POC difference as shown in Equation 1. Equation 1 can be applied to the vectors MVO’ and MVO. Equation 1 can be applied to each component of the vectors MVO’ and MVO. The first POC difference is the difference between the POC of the current picture (POC curr ) and the POC of the first reference picture in the reference picture list L0 (POC L0 ). The second POC difference is the difference between the POC of the current picture (POC curr ) and the POC of the second reference picture in the reference picture list L1 (POC L1 ). MVO’ = MVO × (POC L1 - POC curr ) / (POC L0 - POC curr ) Equation 1

[0151] In one example, the scaled MVO (MVO’) is added to the second MV associated with the reference list L1. In one example, the mirror MVO’ of the scaled MVO (MVO’) is added to the second MV associated with the reference list L1.

[0152] As described above, in one example of SbTMVP mode (for example, a variation of SbTMVP mode described in Figures 16-17), applying DVO to DV allows the position of the collated block in the collated reference picture to be adjusted, thereby influencing the motion information of the subblocks within the current block. In another example of SbTMVP mode (for example, another variation of SbTMVP mode described in Figures 16-17), applying MVO allows the motion information of the subblocks within the current block to be directly adjusted. DVO and MVO can be signaled using the same method. For example, a given MVO list is the same as a given DVO list, and DVO and MVO can use the same given MVO list.

[0153] In some cases, applying DVO to DV can allow the position of a collated block within a collated reference picture to be adjusted. Based on the motion information of the corresponding subblock in the collated reference picture, determined based on the updated DV' (e.g., DV + DVO), MVO may be applied to further adjust the motion information of the subblock within the current block after obtaining the motion information of the subblock within the current block. DVO may be the same as or different from MVO.

[0154] Figure 18 shows a flowchart outlining an encoding process (1800) according to one embodiment of the present disclosure. Process (1800) can be used in a video encoder. Process (1800) can be performed by a device for video coding that includes processing circuits. In various embodiments, process (1800) is performed by processing circuits, such as, for example, processing circuits of terminal devices (310), (320), (330) and (340), processing circuits that perform the functions of a video encoder (e.g., (403), (603), (703)), and similar. In some embodiments, process (1800) is implemented by software instructions, and so the processing circuit performs process (1800) when the processing circuit executes the software instructions. The process starts at (S1801) and proceeds to (S1810).

[0155] In (S1810), the updated DV of the current block in the current picture can be determined based on the DV of the current block and the DV offset of the current block (also called the MV offset (MVO)). The DV can be determined as described above, for example in Figures 16-17. The updated DV of the current block indicates the collated block in the collated picture. The collated block is the current block and collated.

[0156] Currently, a block contains multiple subblocks that are encoded using subblock-based time-motion vector prediction (SbTMVP) mode.

[0157] In one example, the DV offset (or MVO) is determined from the DV offset candidates using any preferred method.

[0158] In one example, the updated DV is determined to be the vector sum of the DV and the DV offset.

[0159] In one example, the updated DV is constrained such that the collated block is within a constrained area in the collated reference picture. The constrained area includes the collated area corresponding to the current CTU in the current picture, and the current CTU includes the current block.

[0160] In (S1820), the movement information of one of the above subblocks may be determined based on the movement information of the corresponding subblock within the collated block.

[0161] In (S1830), DV offset information (also called MVO information) indicating the DV offset can be encoded. The above subblock among the multiple subblocks can be encoded based on the motion information of the above subblock among the multiple subblocks.

[0162] In one embodiment, the DV offset information indicates the magnitude of the DV offset and at least one index indicating the direction of the DV offset. In one example, the at least one index includes a distance index indicating the magnitude of the DV offset, which is one of a set of predetermined distances, and a direction index indicating the direction of the DV offset, which is one of a set of predetermined directions.

[0163] In one example, the set of predetermined distances and the set of predetermined directions are used in merge motion vector difference (MMVD) mode.

[0164] In (S1840), the encoded DV offset information can be included in the bitstream and signaled to the decoder.

[0165] In one example, the DV offset information includes the DV offset, which is encoded into the bitstream and signaled.

[0166] Then, process (1800) proceeds to (S1899) and terminates.

[0167] Process (1800) can be suitably adapted to various scenarios, and the steps of Process (1800) can be adjusted accordingly. One or more steps of Process (1800) can be adapted, omitted, repeated, and / or combined. Process (1800) can be carried out in any preferred order. Additional (one or more) steps can be added.

[0168] In one embodiment, the updated displacement vector (DV) of the current block in the current picture is determined based on the current block's DV and MVO (also referred to as the DV offset). The MVO indicates the motion offset of the DV used to adjust the position of the collated block in the collated reference picture. The updated DV indicates the adjusted position of the collated block in the collated reference picture. The current block is coded in SbTMVP mode.

[0169] The SbTMVP information (e.g., motion information) of each of the above multiple subblocks can be derived based on, at least, the motion information of the corresponding subblock in the collated block indicated by the updated DV. The above multiple subblocks can be encoded in SbTMVP mode based on the SbTMVP information of the above multiple subblocks.

[0170] Figure 19A shows a flowchart outlining a decoding process (1900A) according to one embodiment of the present disclosure. Process (1900A) can be used in a video decoder. Process (1900A) can be performed by a device for video coding which may include a receiving circuit and a processing circuit. In various embodiments, process (1900A) is performed by processing circuits such as, for example, processing circuits for terminal devices (310), (320), (330) and (340), processing circuits that perform the functions of a video encoder (403), processing circuits that perform the functions of a video decoder (410), processing circuits that perform the functions of a video decoder (510), processing circuits that perform the functions of a video encoder (603), and similar. In some embodiments, process (1900A) is implemented by software instructions, and therefore the processing circuit performs process (1900A) when the processing circuit executes software instructions. The process starts at (S1901) and proceeds to (S1910).

[0171] In (S1910), displacement vector (DV) offset (also referred to as MV offset (MVO)) information for the current block in the current picture may be received from the encoded video bitstream. The current block contains multiple subblocks that are reconstructed using subblock-based time motion vector prediction (SbTMVP) mode. The DV offset (or MVO) information may indicate a motion offset relative to the DV used to adjust the position of the collated block in the collated reference picture. In one example, the position of the collated block in the collated reference picture is adjusted by the DV offset.

[0172] In one example, the DV offset information includes the DV offset (or MVO) signaled within the encoded video bitstream.

[0173] In one embodiment, the DV offset information indicates the magnitude of the DV offset and at least one index indicating the direction of the DV offset. In one example, the at least one index includes a distance index indicating the magnitude of the DV offset, which is one of a set of predetermined distances, and a direction index indicating the direction of the DV offset, which is one of a set of predetermined directions.

[0174] In one example, the set of predetermined distances and the set of predetermined directions are used in merge motion vector difference (MMVD) mode.

[0175] In (S1920), the updated DV of the current block can be determined based on the DV of the current block and the DV offset of the current block. The DV offset is indicated by the DV offset information. The updated DV of the current block points to a block in the collated reference picture. This block is considered collated with the current block and is referred to as the collated block of the current block.

[0176] In one example, the updated DV is determined to be the vector sum of the DV and the DV offset.

[0177] In one example, the updated DV is constrained such that the collated block is within a constrained area in the collated reference picture. The constrained area includes the collated area corresponding to the current CTU in the current picture, and the current CTU includes the current block.

[0178] In (S1930), the movement information of one of the above subblocks may be determined based on the movement information of the corresponding subblock within the collated block.

[0179] In (S1940), the above subblock can be reconfigured based on the movement information of the above subblock.

[0180] Process (1900A) proceeds to (S1999) and terminates.

[0181] Process (1900A) can be suitably adapted to various scenarios, and the steps of Process (1900A) can be adjusted accordingly. One or more steps of Process (1900A) can be adapted, omitted, repeated, and / or combined. Process (1900A) can be carried out in any preferred order. Additional (one or more) steps can be added.

[0182] Figure 19B shows a flowchart outlining a decoding process (1900B) according to one embodiment of the present disclosure. Process (1900B) is a variation of decoding process (1900A). Process (1900B) can be used in a video decoder. Process (1900B) can be performed by a device for video coding that includes receiving and processing circuits. In various embodiments, process (1900A) is performed by processing circuits such as, for example, processing circuits for terminal devices (310), (320), (330) and (340), processing circuits that perform the functions of a video encoder (403), processing circuits that perform the functions of a video decoder (410), processing circuits that perform the functions of a video decoder (510), processing circuits that perform the functions of a video encoder (603), and similar. In some embodiments, process (1900B) is implemented by software instructions, and therefore the processing circuit performs process (1900B) when the processing circuit executes software instructions. The process starts at (S1902) and proceeds to (S1912).

[0183] In (S1912), an encoded video bitstream containing the current picture is received. The current picture contains the current block. The current block contains multiple subblocks.

[0184] In (S1922), it is determined that the current block, which includes the above-mentioned subblocks, will be coded in SbTMVP mode based on the syntax elements within the encoded video bitstream.

[0185] In (S1932), the motion vector offset (MVO) information of the current block is obtained. The MVO indicates the motion offset of the displacement vector (DV) used to adjust the position of the collated block within the collated reference picture.

[0186] In (S1942), the updated DV of the current block is determined based on the DV and MVO of the current block. The updated DV indicates the adjusted position of the collated block within the collated reference picture.

[0187] In (S1952), the SbTMVP information (e.g., motion information) of each of the above subblocks is derived based on the motion information of the corresponding subblock in the collated block, indicated by the updated DV.

[0188] In (S1962), the above-mentioned subblocks are reconfigured in SbTMVP mode based on the SbTMVP information of the above-mentioned subblocks.

[0189] Process (1900B) proceeds to (S1992) and terminates.

[0190] Process (1900B) can be suitably adapted to various scenarios, and the steps of Process (1900B) can be adjusted accordingly. One or more steps of Process (1900B) can be adapted, omitted, repeated, and / or combined. Process (1900B) can be carried out in any preferred order. Additional (one or more) steps can be added.

[0191] Embodiments of this disclosure may be used separately or in combination in any order. Furthermore, each of these methods (or embodiments), the encoder, and the decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, these one or more processors execute a program stored on a non-temporary computer-readable medium.

[0192] Figure 20 shows a flowchart outlining an encoding process (2000) according to one embodiment of the present disclosure. Process (2000) can be used in a video encoder. Process (2000) can be performed by a device for video coding that includes processing circuits. In various embodiments, process (2000) is performed by processing circuits, such as, for example, processing circuits of terminal devices (310), (320), (330) and (340), processing circuits that perform the functions of a video encoder (e.g., (403), (603), (703)), and similar. In some embodiments, process (2000) is implemented by software instructions, and therefore the processing circuit executes process (2000) when the processing circuit executes software instructions. The process starts at (S2001) and proceeds to (S2010).

[0193] In (S2010), the displacement vector (DV) of the current block in the current picture can be determined, for example, as explained in Figures 16-17. The current block contains multiple subblocks encoded using the subblock-based time motion vector prediction (SbTMVP) mode. The DV represents the current block and the collated block in the collated reference picture, which is the collated block.

[0194] In (S2020), as explained in Figures 16-17, for example, the movement information of one of the multiple subblocks can be determined based on the movement information of the corresponding subblock within the collated block.

[0195] In (S2030), updated motion information for one of the multiple subblocks can be determined based on the motion information of the subblock and the motion vector (MV) offset of the current block.

[0196] In one example, the motion information of one of the subblocks includes a first motion vector (MV) associated with a first reference picture from a first reference picture list L0. The updated first MV may be determined to be the vector sum of the first MV and the MV offset. The updated motion information includes the updated first MV.

[0197] In one example, the motion information of the subblock among the multiple subblocks described above includes a second MV associated with a second reference picture from a second reference picture list L1. The updated second MV may be determined to be either (i) the vector difference between the second MV and the MV offset, or (ii) the vector sum of the second MV and the scaled MV offset. The scaled MV offset may be based on the MV offset, the picture order count (POC) of the current picture, the POC of the first reference picture, and the POC of the second reference picture. The updated motion information includes the updated second MV.

[0198] In (S2040), MV offset information indicating the MV offset may be encoded. The above subblock among the multiple subblocks may be encoded based on the updated motion information. The MV offset information may be included in the bitstream.

[0199] In one example, the MV offset information includes the MV offset.

[0200] In one example, the MV offset information indicates the magnitude of the MV offset and at least one index indicating the direction of the MV offset. This at least one index includes a distance index indicating the magnitude of the MV offset, which is one of a set of predetermined distances, and a direction index indicating the direction of the MV offset, which is one of a set of predetermined directions. The set of predetermined distances and the set of predetermined directions are used in merge motion vector difference (MMVD) mode.

[0201] Then, process (2000) proceeds to (S2099) and terminates.

[0202] Process (2000) can be suitably adapted to various scenarios, and the steps of Process (2000) can be adjusted accordingly. One or more steps of Process (2000) can be adapted, omitted, repeated, and / or combined. Process (2000) can be carried out in any preferred order. Additional (one or more) steps can be added.

[0203] Figure 21 shows a flowchart outlining a decoding process (2100) according to one embodiment of the present disclosure. Process (2100) can be used in a video decoder. Process (2100) can be performed by a device for video coding which may include a receiving circuit and a processing circuit. In various embodiments, process (2100) is performed by processing circuits such as, for example, processing circuits for terminal devices (310), (320), (330) and (340), processing circuits that perform the functions of a video encoder (403), processing circuits that perform the functions of a video decoder (410), processing circuits that perform the functions of a video decoder (510), processing circuits that perform the functions of a video encoder (603), and similar. In some embodiments, process (2100) is implemented by software instructions, and therefore the processing circuit performs process (2100) when the processing circuit executes the software instructions. The process starts at (S2101) and proceeds to (S2110).

[0204] At (S2110), motion vector (MV) offset information for the current block in the current picture may be received from the encoded video bitstream. The current block contains multiple subblocks which are reconstructed using subblock-based time motion vector prediction (SbTMVP) mode.

[0205] In (S2120), the displacement vector (DV) of the current block can be determined, for example, as explained in Figures 16-17. The DV can represent the current block and a block in the collated reference picture that is collated. This block can be referred to as the collated block of the current block.

[0206] In (S2130), as explained in Figures 16-17, for example, the movement information of one of the multiple subblocks can be determined based on the movement information of the corresponding subblock within the collated block.

[0207] In (S2140), updated motion information for one of the multiple subblocks can be determined based on the motion information of the subblock and the MV offset of the current block indicated by the MV offset information.

[0208] In one example, MV offset information includes the MV offset signaled within the encoded video bitstream.

[0209] In one example, the MV offset information indicates the magnitude of the MV offset and at least one index indicating the direction of the MV offset. This at least one index includes a distance index indicating the magnitude of the MV offset, which is one of a set of predetermined distances, and a direction index indicating the direction of the MV offset, which is one of a set of predetermined directions. The set of predetermined distances and the set of predetermined directions are used in merge motion vector difference (MMVD) mode.

[0210] In one example, the motion information of one of the subblocks includes a first motion vector (MV) associated with a first reference picture from a first reference picture list L0. The updated first MV may be determined to be the vector sum of the first MV and the MV offset. The updated motion information includes the updated first MV.

[0211] In one example, the motion information of the subblock among the multiple subblocks described above includes a second MV associated with a second reference picture from a second reference picture list L1. The updated second MV may be determined to be either (i) the vector difference between the second MV and the MV offset, or (ii) the vector sum of the second MV and the scaled MV offset. The scaled MV offset may be based on the MV offset, the picture order count (POC) of the current picture, the POC of the first reference picture, and the POC of the second reference picture. The updated motion information includes the updated second MV.

[0212] In (S2150), the above subblock among the multiple subblocks can be reconfigured based on the updated motion information.

[0213] Then, process (2100) proceeds to (S2199) and terminates.

[0214] Process (2100) can be suitably adapted to various scenarios, and the steps of Process (2100) can be adjusted accordingly. One or more steps of Process (2100) can be adapted, omitted, repeated, and / or combined. Process (2100) can be carried out in any preferred order. Additional (one or more) steps can be added.

[0215] The technologies described above can be implemented as computer software using computer-readable instructions, physically stored on one or more computer-readable media. For example, Figure 22 shows a computer system (2200) suitable for implementing a particular embodiment of the disclosed subject matter.

[0216] Computer software can be coded using any suitable machine code or computer language that can be assembled, compiled, linked, or subjected to similar mechanisms to produce code having instructions that can be executed directly or via interpretation, microcode execution, and similar means by one or more computer central processing units (CPUs), graphics processing units (GPUs), and similar devices.

[0217] Instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, game consoles, Internet of Things devices, and similar devices.

[0218] The components shown in Figure 22 with respect to the computer system (2200) are essentially illustrative and are not intended to imply any limitation on the scope of use or functionality of computer software implementing embodiments of the present disclosure. Nor should the configuration of the components be construed as having any dependency or requirement on any one or combination of the components shown in this exemplary embodiment of the computer system (2200).

[0219] The computer system (2200) may include certain human interface input devices. Such human interface input devices may respond to input from one or more human users, for example, via tactile input (e.g., keystrokes, swipes, moving a data glove), audio input (e.g., voice, applause), visual input (e.g., gestures), or olfactory input (not shown). The human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., conversations, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still camera), or video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0220] The input human interface device may include one or more of the following: keyboard (2201), mouse (2202), trackpad (2203), touchscreen (2210), data glove (not shown), joystick (2205), microphone (2206), scanner (2207), and camera (2208) (only one of each is shown).

[0221] The computer system (2200) may also include certain human interface output devices. Such human interface output devices may stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (e.g., tactile feedback via a touchscreen (2210), data glove (not shown), or joystick (2205), although there may also be tactile feedback devices that do not function as input devices), audio output devices (e.g., speakers (2209), headphones (not shown), etc.), visual output devices (e.g., screens (2210) including CRT screens, LCD screens, plasma screens, and OLED screens (each having or not having touchscreen input functionality; each having or not having tactile feedback functionality; some of these may be able to output two-dimensional visual output, or output of four or more dimensions through means such as stereoscopic output, etc.), virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), etc.), and printers (not shown).

[0222] The computer system (2200) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (2220) having CD / DVD or similar media (2221), thumb drives (2222), removable hard drives or / or solid-state drives (2223), legacy magnetic media such as tapes and floppy disks (registered trademarks, not shown), specialized ROM / ASIC / PLD-based devices such as security dongles (not shown), and similar devices.

[0223] Those skilled in the art will also understand that the term “computer-readable medium” as used in connection with the matters disclosed herein does not include transmission media, carrier waves, or other transient signals.

[0224] The computer system (2200) may also include an interface (2254) to one or more communication networks (2255). The networks may be, for example, wireless, wired, or optical. The networks may further be local, wide-area, metropolitan, vehicle and industrial, real-time, or latency-tolerant. Examples of networks include local area networks such as Ethernet®, cellular networks including wireless LANs, GSM, 3G, 4G, 5G, LTE and similar technologies, wired or wireless wide-area digital TV networks including cable TV, satellite TV, and terrestrial broadcast TV, and vehicle and industrial networks including CANBus. Certain networks generally require an external network interface adapter attached to a specific general-purpose data port or peripheral bus (2249) (e.g., a USB port on the computer system (2200)), while others are generally integrated into the core of the computer system (2200) by attachment to a system bus as described below (e.g., an Ethernet interface to a PC computer system, or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (2200) can communicate with other entities. Such communication may be unidirectional reception only (e.g., broadcast television), unidirectional transmission only (e.g., CANbus to a specific CANbus device), or bidirectional to other computer systems, for example, using local or wide-area digital networks. Specific protocols and protocol stacks may be used on each of the networks and network interfaces as described above.

[0225] The aforementioned human interface device, human-accessible storage device, and network interface can be mounted on the core (2240) of the computer system (2200).

[0226] The core (2240) may include one or more central processing units (CPUs) (2241), graphics processing units (GPUs) (2242), specialized programmable processing units in the form of field-programmable gate arrays (FPGAs) (2243), hardware accelerators for specific tasks (2244), graphics adapters (2250), and the like. These devices may be connected via a system bus (2248) along with read-only memory (ROM) (2245), random access memory (2246), internal mass storage such as internal non-user-accessible hard drives, SSDs (2247), and similar devices (2247). In some computer systems, the system bus (2248) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, and similar devices. Peripheral devices may be attached either directly to the core's system bus (2248) or via a peripheral bus (2249). In one example, a screen (2210) can be connected to a graphics adapter (2250). Peripheral bus architectures include PCI, USB, and similar technologies.

[0227] The CPU (2241), GPU (2242), FPGA (2243), and accelerator (2244) can execute certain instructions that, in combination, constitute the aforementioned computer code. This computer code may be stored in ROM (2245) or RAM (2246). Transient data may also be stored in RAM (2246), while permanent data may be stored, for example, in internal mass storage (2247). Fast storage and retrieval to any of the memory devices may be made possible by the use of cache memory that may be associated with one or more CPUs (2241), GPUs (2242), mass storage (2247), ROM (2245), RAM (2246), and similar devices.

[0228] A computer-readable medium may have computer code thereon for performing various computer implementation processes. The medium and computer code may be specifically designed and constructed for the purposes of this disclosure, or they may be of a type that is well known and available to those skilled in the computer software technology.

[0229] As an example, but not limited to, a computer system having architecture (2200), in particular core (2240), can provide functionality as a result of (one or more) processors (including CPUs, GPUs, FPGAs, accelerators, and similar) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be specific storage of the core (2240) that is non-transient in nature, such as internal mass storage (2247) or ROM (2245) within the core, and media related to user-accessible mass storage as described above. Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (2240). The computer-readable media may include one or more memory devices or chips, depending on the specific needs. The software may cause the core (2240) and, in particular, the processor within it (including CPUs, GPUs, FPGAs, and similar devices) to execute the specific processes or specific parts of the specific processes described herein, including defining data structures to be stored in RAM (2246) and modifying such data structures according to processes defined by the software. In addition, or alternatively, a computer system may provide functionality as a result of logic wired or otherwise embodied in circuits (e.g., accelerators (2244)) that can operate in place of or with the software to execute the specific processes or specific parts of the specific processes described herein. References to software include logic, and vice versa, where appropriate. References to computer-readable media may include circuits containing software for execution (e.g., integrated circuits (ICs), etc.), circuits embodying logic for execution, or both, where appropriate. This disclosure includes preferred combinations of hardware and software. Appendix A: Acronyms JEM: Joint Exploration Model VVC: Versatile Video Coding BMS: benchmark set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplementary Enhancement Information VUI: Video Usability Information GOPs: Groups of Pictures TUs: Transform Units PU: Prediction Units CTUs: Coding Tree Units CTBs: Coding Tree Blocks PBs: Prediction Blocks HRD: Hypothetical Reference Decoder SNR: Signal Noise Ratio CPUs: Central Processing Units GPUs: Graphics Processing Units CRT:Cathode Ray Tube LCD: Liquid-Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Areas SSD: Solid-state drive IC: Integrated Circuit CU: Coding Unit

[0230] While this disclosure describes several exemplary embodiments, there are many variations, substitutions, and equivalent alternatives that fall within the scope of the disclosure. Therefore, it should be understood that those skilled in the art could devise numerous systems and methods, though not explicitly illustrated or described herein, that embody the principles of the disclosure and thus lie within its spirit and scope.

Claims

[Claim 1] A method of video decoding performed by a decoder, The steps include receiving an encoded video bitstream having a current picture, wherein the current picture includes a current block, and the current block includes a plurality of subblocks, The steps include determining, based on syntax elements in the encoded video bitstream, that the current block, which includes the plurality of subblocks, is coded in subblock-based time-motion vector prediction (SbTMVP) mode, The steps include obtaining MVO information for the current block, which indicates the motion vector offset (MVO), the MVO indicating the motion offset of the displacement vector (DV) used to adjust the position of the collated block in the collated reference picture, The steps include determining the updated DV of the current block based on the DV and MVO of the current block, wherein the updated DV indicates the adjusted position of the collated block in the collated reference picture, The steps include: deriving the SbTMVP information of each of the plurality of subblocks based at least on the motion information of the corresponding subblock within the collated block, as indicated by the updated DV; The steps include: reconfiguring the plurality of subblocks in the SbTMVP mode based on the SbTMVP information of the plurality of subblocks; A method of having.