Sub-block based motion vector predictor with MV offset in AMVP mode

The sub-block based motion vector predictor with MV offset in AMVP mode addresses inefficiencies in video coding by enhancing intra-prediction and motion vector prediction, achieving improved compression ratios and reduced redundancy in video data transmission.

JP2026502408APending Publication Date: 2026-01-23TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025515450
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-06
Filing Date
2023-11-07
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing video coding technologies face inefficiencies in reducing redundancy and improving compression ratios, particularly in intra-prediction and motion vector prediction, due to the use of less likely directions requiring more bits and the need for improved motion vector prediction mechanisms.

Method used

The implementation of a sub-block based motion vector predictor with an MV offset in the AMVP mode, which includes a processing circuit that determines motion information for sub-blocks using a co-located block's motion information and a displacement vector offset, and configures candidate lists for MVP selection with precision-adaptive motion vector resolution.

Benefits of technology

Enhances video coding efficiency by reducing redundancy and improving compression ratios through precise motion vector prediction, especially in intra-prediction and motion compensation, thereby optimizing bit usage and decoding accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026502408000001_ABST
    Figure 2026502408000001_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide a method and apparatus, including a processing circuit, that determines, based on syntax elements of a coded video bitstream, that a current block including multiple sub-blocks is coded in sub-block-based temporal motion vector prediction (SbTMVP) mode. Motion vector offset (MVO) information indicating the MVO is received. The MVO indicates a motion offset of a displacement vector (DV) used to adjust the position of the co-located block in the co-located reference picture. An updated DV for the current block is determined based on the DV and the MVO. SbTMVP information for each sub-block of the multiple sub-blocks is derived based on motion information for a corresponding sub-block of the co-located block indicated by the updated DV. The multiple sub-blocks in the SbTMVP mode are reconstructed based on the SbTMVP information for the sub-blocks of the multiple sub-blocks.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Application No. 18 / 387,378, entitled "SUBBLOCK BASED MOTION VECTOR PREDICTOR WITH MV OFFSET IN AMVP MODE," filed November 6, 2023, which claims the benefit of priority to U.S. Patent Application No. 63 / 437,990, entitled "Subblock Based Motion Vector Predictor With MV Offset In AMVP Mode," filed January 9, 2023. The entire disclosure of the prior application is incorporated herein by reference in its entirety.

[0002] This disclosure generally describes embodiments related to video coding. [Background technology]

[0003] The background discussion provided herein is intended to generally present the context for the present disclosure. To the extent described in this background section, the inventors' work, as well as aspects of the description that may not be admitted as prior art at the time of filing, are not admitted expressly or implicitly as prior art to the present disclosure.

[0004] Uncompressed digital images and / or videos can include a series of pictures, each with spatial dimensions of, for example, 1920 x 1080 luma samples and associated chroma samples. The series of pictures can have a fixed or variable picture rate (informally known as a frame rate), for example, 60 pictures per second or 60 Hz. Uncompressed images and / or videos have specific bitrate requirements. For example, 1080p 60 4:2:0 video (1920 x 1080 luma sample resolution at a 60 Hz frame rate) with 8 bits per sample requires a bandwidth approaching 1.5 Gbit / s. One hour of such video requires more than 600 GBytes of storage space.

[0005] One goal of image and / or video coding and decoding may be reducing redundancy in the input image and / or video signal through compression. Compression can help reduce the aforementioned bandwidth and / or storage space requirements, in some cases by two or more orders of magnitude. While the description herein uses video encoding / decoding as an illustrative example, the same techniques may be applied to image encoding / decoding in a similar manner without departing from the spirit of this disclosure. Both lossless and lossy compression, and combinations thereof, may be employed. Lossless compression refers to techniques by which an exact copy of the original signal can be reconstructed from a compressed version of the original signal. When lossy compression is used, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signal is small enough to make the reconstructed signal useful for its intended use. For video, lossy compression is widely adopted. The amount of acceptable distortion depends on the application; for example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio may indicate that the higher the acceptable / tolerable distortion, the higher the compression ratio that can be achieved.

[0006] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transform processing, quantization, and entropy coding.

[0007] Video codec technology can include a technique known as intra-coding. In intra-coding, sample values ​​are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, pictures are spatially subdivided into blocks of samples. If all blocks of samples are coded in intra mode, the picture may be an intra-picture. Intra-pictures and their derivatives, such as independent decoder refresh pictures, are used to reset the decoder state and may therefore be used as the first picture of a coded video bitstream and video session, or as a still image. Samples of intra-blocks are transformed, and the transform coefficients may be quantized before entropy coding. Intra-prediction can be a technique that minimizes sample values ​​in the pre-transform domain. In some cases, the smaller the DC value and AC coefficients after the transform, the fewer bits required for a given quantization step size to represent the block after entropy coding.

[0008] Traditional intra-coding, for example, as used in MPEG-2 generation coding techniques, does not use intra-prediction. However, some newer video compression techniques include techniques that attempt to make predictions based on surrounding sample data and / or metadata obtained during the encoding / decoding of a block of data. Such techniques are hereinafter referred to as "intra-prediction" techniques. Note that, at least in some cases, intra-prediction uses only reference data from the current picture being reconstructed, and not reference data from a reference picture.

[0009] Intra-prediction may take many different forms. When two or more of such techniques may be used in a given video coding technique, the particular technique in use may be coded as a particular intra-prediction mode that uses the particular technique. In certain cases, an intra-prediction mode may have sub-modes and / or parameters that may be coded separately or included in a mode codeword that defines the prediction mode used. The codeword used for a given mode, sub-mode, and / or parameter combination may affect coding efficiency gains via intra-prediction, and therefore may also affect the entropy coding technique used to convert the codeword into a bitstream.

[0010] Specific modes of intra prediction were introduced in H.264, refined in H.265, and further refined in new coding techniques such as the Joint Search Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). Predictor blocks can be formed using neighboring sample values ​​of already available samples. The sample values ​​of the neighboring samples are copied to the predictor block according to the direction. A reference to the direction in use can be coded in the bitstream or can itself be predicted.

[0011] Referring to FIG. 1A, depicted at the bottom right is a subset of nine known predictor directions from the 33 possible predictor directions defined in H.265 (corresponding to the 33 angular modes out of the 35 intra modes). The point where the arrows converge (101) represents the sample being predicted. The arrows represent the direction from which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from the sample to the upper right, at an angle of 45 degrees from horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from the sample to the lower left of sample (101), at an angle of 22.5 degrees from horizontal.

[0012] 1A, a square block (104) of 4x4 samples (indicated by a thick dashed line) is shown in the upper left. The square block (104) contains 16 samples, each labeled with "S," its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample (from the top) in the Y dimension and the first sample (from the left) in the X dimension. Similarly, sample S44 is the fourth sample in the block (104) in both the Y and X dimensions. Because the block is 4x4 samples in size, S44 is located in the lower right. Reference samples, which follow a similar numbering scheme, are also shown. The reference samples are labeled R, their Y position (e.g., row index), and their X position (column index) relative to the block (104). In both H.264 and H.265, the predicted samples are near the block being reconstructed, so there is no need to use negative values.

[0013] Intra-picture prediction can work by copying reference sample values ​​from neighboring samples indicated by the signaled prediction direction. For example, assume that the coded video bitstream includes signaling for this block indicating a prediction direction consistent with the arrow (102), i.e., the sample is predicted from the upper right sample at an angle of 45 degrees from the horizontal. In this case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.

[0014] In certain cases, to calculate a reference sample, especially when the direction is not evenly divisible by 45 degrees, the values ​​of multiple reference samples may be combined, for example by interpolation.

[0015] The number of possible directions has increased as video coding technology has evolved. H.264 (2003) allowed for nine different directions to be represented. This increased to 33 in H.265 (2013). Currently, JEM / VVC / BMS can support up to 65 directions. Experiments have been conducted to identify the most likely directions, and specific entropy coding techniques are used to represent these likely directions with a small number of bits, accepting a certain penalty for less likely directions. Furthermore, the direction itself may be predicted from nearby directions used in nearby, already decoded blocks.

[0016] FIG. 1B shows a schematic diagram (110) depicting 65 intra-prediction directions with JEM to illustrate the increasing number of prediction directions over time.

[0017] The mapping of intra-prediction direction bits to represent directions within an encoded video bitstream may vary depending on the video coding technique. Such mappings can range, for example, from simple direct mappings to complex adaptive schemes including codeword most probable modes and similar techniques. However, in most cases, there may be certain directions that are statistically less likely to occur within the video content than certain other directions. Because the goal of video compression is to reduce redundancy, in a well-performing video coding technique, these less likely directions are represented by more bits than more likely directions.

[0018] Image and / or video coding and decoding may be performed using inter-picture prediction with motion compensation. Motion compensation may be a lossy compression technique and may refer to a technique in which blocks of sample data from a previously reconstructed picture or portion thereof (reference picture) are spatially shifted in a direction indicated by a motion vector (hereinafter, MV) and then used to predict a newly reconstructed picture or portion thereof. In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or three dimensions, with the third dimension being an indication of the reference picture in use (the third dimension may indirectly be a temporal dimension).

[0019] In some video compression techniques, the MV applicable to a particular region of sample data can be predicted from other MVs, e.g., from an MV associated with another region of sample data that is spatially adjacent to the region being reconstructed and precedes that MV in decoding order. Doing so can significantly reduce the amount of data required to code the MV, thereby eliminating redundancy and improving compression ratios. For example, when coding an input video signal derived from a camera (known as natural video), MV prediction can work effectively because there is a statistical likelihood that regions larger than the region to which a single MV is applicable move in similar directions and therefore, in some cases, can be predicted using similar motion vectors derived from MVs of nearby regions. As a result, the MV found for a given region will be similar or identical to the MV predicted from surrounding MVs, and thus, after entropy coding, can be represented with fewer bits than would be used if the MV were coded directly. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., an MV) derived from the original signal (i.e., a sample stream). In other cases, the MV prediction itself can be lossy, for example, due to rounding errors when calculating a predictor from several surrounding MVs.

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms provided by H.265, the one described with reference to Figure 2 is a technique hereafter referred to as "spatial merging".

[0021] Referring to Figure 2, a current block (201) contains samples that the encoder finds during the motion search process to be predictable from a spatially shifted previous block of the same size. Instead of directly coding its MV, the MV can be derived from metadata associated with one or more reference pictures, e.g., the most recent reference picture (in decoding order), using the MV associated with any one of five surrounding samples, denoted A0, A1, and B0, B1, B2 (202-206, respectively). In H.265, MV prediction can use predictors from the same reference picture used by neighboring blocks. Summary of the Invention [Means for solving the problem]

[0022] Aspects of the present disclosure provide methods and apparatuses for video encoding / decoding. In some examples, the apparatus for video decoding includes a processing circuit. In an embodiment, the processing circuit receives displacement vector (DV) offset information for a current block of a current picture from an encoded video bitstream. The current block includes multiple sub-blocks reconstructed using a sub-block-based temporal motion vector prediction (SbTMVP) mode. An updated DV for the current block can be determined based on the DV of the current block and a DV offset for the current block. The DV offset is indicated by the DV offset information. The updated DV for the current block indicates a co-located block of a co-located reference picture. The co-located block is co-located with the current block. The processing circuit determines motion information for a sub-block of the multiple sub-blocks based on motion information of a corresponding sub-block of the co-located block, and reconstructs the sub-block of the multiple sub-blocks based on the motion information of the sub-block of the multiple sub-blocks. In some examples, the coding information of a current block of the coded video bitstream indicates an advanced motion vector prediction (AMVP) mode with motion vector predictor (MVP) information and motion vector offset (MVO) information in the coding information of the current block. The processing circuit selects an MVP from an MVP candidate list based on the MVP information in the coding information of the current block for the AMVP mode, and derives the DV using the MVP as an SbTMVP candidate.

[0023] In some examples, the processing circuit configures an MVP candidate list that includes a sub-block-based merge candidate list, the sub-block-based merge candidate list including one or more SbTMVP candidates. In examples, the sub-block-based merge candidate list includes multiple spatially neighboring blocks of the current block in a predetermined order. In another example, the sub-block-based merge candidate includes 0DV for use as the SbTMVP candidate.

[0024] In some examples, the processing circuit checks one or more spatially neighboring blocks of the current block in a predetermined order for the availability of a central sub-block motion vector. For a spatially neighboring block of the current block, in response to the central sub-block motion vector of the spatially neighboring block being available, the processing circuit adds the spatially neighboring block as a candidate to the MVP candidate list. In response to none of the one or more spatially neighboring blocks having an available central sub-block motion vector, the processing circuit adds 0DV to the MVP candidate list.

[0025] In some examples, the coding information indicates an affine AMVP mode. The processing circuitry configures an affine AMVP candidate list including one or more SbTMVP candidates.

[0026] In an example, the processing circuitry inserts the SbTMVP candidate into the first position of the affine AMVP candidate list.

[0027] In some examples, the processing circuit checks whether an affine-coded block exists in a spatially neighboring block of the current block. In response to any of the spatially neighboring blocks being affine-coded, the processing circuit inserts an SbTMVP candidate in the first position of an affine AMVP candidate list. In response to any of the spatially neighboring blocks being affine-coded, the processing circuit inserts an SbTMVP candidate in the last position of an affine AMVP candidate list.

[0028] In some examples, the processing circuit determines precision of MVO information in coding information of a current block of the coded video bitstream, and the MVO information is coded into the coded video bitstream with precision according to adaptive motion vector resolution (AMVR). The processing circuit determines a DV motion offset based on the precision of the MVO information.

[0029] In some examples, the processing circuitry decodes, from the coded video bitstream, an index that indicates the precision of the MVO information according to the AMVR.

[0030] In some examples, the accuracy of the MVO information is in M ​​pixel units, where M is a positive integer.

[0031] In some examples, the accuracy of the MVO information is one of 1 pixel, 2 pixels, 4 pixels, and 8 pixels.

[0032] In some examples, the processing circuitry also scales the DV motion offset based on the size of the sub-block.

[0033] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method for video decoding.

[0034] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Brief explanation of the drawings]

[0035] [Figure 1A] FIG. 2 is a schematic diagram of an example subset of intra-prediction modes. [Figure 1B] FIG. 1 is a diagram of an exemplary intra-prediction direction. [Figure 2] FIG. 2 shows an example of a current block (201) and surrounding samples. [Figure 3] FIG. 3 is a schematic diagram of an exemplary block diagram of a communication system (300). [Figure 4] FIG. 4 is a schematic diagram of an exemplary block diagram of a communication system (400). [Figure 5] FIG. 2 is a schematic diagram of an exemplary block diagram of a decoder. [Figure 6] FIG. 2 is a schematic diagram of an exemplary block diagram of an encoder. [Figure 7]1 is a block diagram illustrating an exemplary encoder. [Figure 8] FIG. 2 is a block diagram illustrating an exemplary decoder. [Figure 9] FIG. 10 illustrates the locations of spatial merge candidates according to an embodiment of the present disclosure. [Figure 10] FIG. 10 illustrates candidate pairs considered for redundancy check of spatial merge candidates, according to an embodiment of the present disclosure. [Figure 11] FIG. 10 illustrates exemplary motion vector scaling for temporal merge candidates. [Figure 12] FIG. 10 illustrates exemplary candidate positions for temporal merge candidates for a current coding unit. [Figure 13] FIG. 1 illustrates an example of a search process in merged motion vector differential (MMVD) mode. [Figure 14] FIG. 1 illustrates an example of a search process in merged motion vector differential (MMVD) mode. [Figure 15] 10 shows additional refinement positions along multiple diagonals in MMVD mode. [Figure 16] 1 illustrates an exemplary sub-block-based temporal motion vector prediction (SbTMVP) process used in SbTMVP mode. [Figure 17] 1 illustrates an exemplary sub-block-based temporal motion vector prediction (SbTMVP) process used in SbTMVP mode. [Figure 18] 1 shows a flowchart outlining an encoding process according to some embodiments of the present disclosure. [Figure 19A] 1 shows a flowchart outlining a decoding process according to some embodiments of the present disclosure. [Figure 19B] 1 shows a flowchart outlining a decoding process according to some embodiments of the present disclosure. [Figure 20] 1 shows a flowchart outlining the process of encoding according to some embodiments of the present disclosure. [Figure 21]1 shows a flowchart outlining a decoding process according to some embodiments of the present disclosure. [Figure 22] 1 is a flowchart outlining a process according to some embodiments of the present disclosure. [Figure 23] 1 is a flowchart outlining a process according to some embodiments of the present disclosure. [Figure 24] FIG. 1 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0036] Figure 3 shows an exemplary block diagram of a communication system (300). The communication system (300) includes multiple terminal devices capable of communicating with each other, for example, via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of Figure 3, the first pair of terminal devices (310) and (320) perform unidirectional data transmission. For example, the terminal device (310) can code video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The coded video data can be transmitted in the form of one or more coded video bitstreams. The terminal device (320) can receive the coded video data from the network (350), decode the coded video data to reconstruct the video pictures, and display the video pictures according to the reconstructed video data. One-way data transmission may be common, such as in media serving applications.

[0037] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of coded video data, for example, during a video conference. In the case of bidirectional transmission of data, in the example, each of the terminal devices (330) and (340) can code video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) over the network (350). Additionally, each of the terminal devices (330) and (340) can receive coded video data transmitted by the other of the terminal devices (330) and (340), decode the coded video data to recover the video pictures, and display the video pictures on an accessible display device according to the recovered video data.

[0038] In the example of FIG. 3, terminal devices 310, 320, 330, and 340 are shown as a server, a personal computer, and a smartphone, respectively, although the principles of the present disclosure need not be so limited. Embodiments of the present disclosure apply to laptop computers, tablet computers, media players, and / or dedicated videoconferencing equipment. Network 350 represents any number of networks that convey coded video data between terminal devices 310, 320, 330, and 340, including, for example, wired and / or wireless communication networks. Communication network 350 may exchange data over circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of the present discussion, the architecture and topology of network 350 may not be important to the operation of the present disclosure, unless otherwise described herein below.

[0039] 4 shows a video encoder and a video decoder in a streaming environment as an example of an application for the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0040] The streaming system may include a capture subsystem (413) that may include a video source (401), such as a digital camera, that creates a stream of uncompressed video pictures (402). In the example, the stream of video pictures (402) includes samples captured by the digital camera. The stream of video pictures (402), depicted as a thick line to emphasize its large amount of data compared to the encoded video data (404) (or encoded video bitstream), may be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (404) (or encoded video bitstream), depicted as a thin line to emphasize its small amount of data compared to the stream of video pictures (402), may be stored on a streaming server (405) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to obtain copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include a video decoder (410), for example, within an electronic device (430). The video decoder (410) decodes the incoming copy of the encoded video data (407) and creates an outgoing stream of video pictures (411) that can be rendered on a display (412) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., a video bitstream) may be encoded according to a particular video coding / compression standard. Examples of these standards include ITU-T Recommendation H.265. In an example, a video coding standard under development is informally known as Versatile Video Coding (VVC).The subject matter of this disclosure may be used in the context of VVC.

[0041] It should be noted that the electronic devices (420) and (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).

[0042] 5 shows an exemplary block diagram of a video decoder (510). The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., receiving circuitry). The video decoder (510) can be used in place of the video decoder (410) in the example of FIG.

[0043] The receiver (531) can receive one or more coded video sequences to be decoded by the video decoder (510). In an embodiment, one coded video sequence is received at a time, and the decoding of each coded video sequence is independent of the decoding of other coded video sequences. The coded video sequences may be received from a channel (501), which may be a hardware / software link to a storage device that stores the coded video data. The receiver (531) can receive coded video data with other data, such as coded audio data and / or auxiliary data streams, which may be transferred to each other using entities (not shown). The receiver (531) can separate the coded video sequences from other data. To combat network jitter, a buffer memory (515) can be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter, "parser (520)"). In certain applications, the buffer memory (515) is part of the video decoder (510). In other cases, the buffer memory (515) may be external to the video decoder (510) (not shown). In still other cases, there may be a buffer memory (not shown) external to the video decoder (510), for example, to combat network jitter, and another buffer memory (515) internal to the video decoder (510), for example, to handle playback timing. When the receiver (531) receives data from a storage / forwarding device with sufficient bandwidth and controllability, or from an asynchronous network, the buffer memory (515) may be unnecessary or may be small. For use in best-effort packet networks such as the Internet, a buffer memory (515) may be required, and the buffer memory (515) may be relatively large, advantageously adaptively sized, and at least partially implemented in an operating system or similar element (not shown) external to the video decoder (510).

[0044] The video decoder (510) may include a parser (520) that reconstructs symbols (521) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (510) and potential information for controlling a rendering device, such as a render device (512) (e.g., a display screen) that is not an integral part of the electronic device (530) but may be coupled to the electronic device (530), as shown in FIG. 5. The rendering device control information may be in the form of a supplemental enhancement information (SEI) message or a video usability information (VUI) parameter set fragment (not shown). The parser (520) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) may extract a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. A subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. Additionally, the parser (520) may extract information from the coded video sequence, such as transform coefficients, quantization parameter values, motion vectors, etc.

[0045] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).

[0046] The reconstruction of the symbols (521) can involve several different units, depending on the format of the coded video picture or portion thereof (e.g., inter-picture and intra-picture, inter-block and intra-block, etc.), as well as other factors. Which units are involved and how may be controlled by subgroup control information parsed by the parser (520) from the coded video sequence. The flow of such subgroup control information between the parser (520) and the following units is not depicted for clarity.

[0047] In addition to the functional blocks already mentioned, the video decoder (510) may be conceptually subdivided into several functional units, as described below. In actual implementations operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the subject matter of this disclosure, the following conceptual subdivision into functional units is appropriate:

[0048] The first unit is a scalar / inverse transform unit (551), which receives quantized transform coefficients as well as control information from the parser (520) as symbols (521), including which transform to use, block size, quantization coefficients, quantization scaling matrices, etc. The scalar / inverse transform unit (551) can output blocks containing sample values ​​that can be input to an aggregator (555).

[0049] In some cases, the output samples of the scaler / inverse transform unit (551) may relate to intra-coded blocks. Intra-coded blocks are blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates blocks of the same size and shape as the block being reconstructed using surrounding, already reconstructed information fetched from the current picture buffer (558). The current picture buffer (558), for example, buffers partially reconstructed and / or fully reconstructed current pictures. The aggregator (555) may add, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).

[0050] In other cases, the output samples of the scalar / inverse transform unit (551) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion-compensated prediction unit (553) may access a reference picture memory (557) to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols (521) associated with the block, these samples may be added by an aggregator (555) to the output of the scalar / inverse transform unit (551) (in this case, referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (557) from which the motion-compensated prediction unit (553) fetches prediction samples may be controlled by motion vectors available to the motion-compensated prediction unit (553) in the form of symbols (521), which may have, for example, X, Y, and reference picture components. Furthermore, motion compensation may also include interpolation of sample values ​​fetched from the reference picture memory (557) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.

[0051] The output samples of the aggregator (555) can be subjected to various loop filtering techniques in a loop filter unit (556). Video compression techniques can include in-loop filter techniques controlled by parameters contained in the coded video sequence (also called an encoded video bitstream) and available to the loop filter unit (556) as symbols (521) from the parser (520). Video compression can also be responsive to meta-information obtained during decoding of previous portions (in decoding order) of the coded picture or coded video sequence, and to previously reconstructed loop-filtered sample values.

[0052] The output of the loop filter unit (556) may be a sample stream that can be output to a render device (512) as well as stored in a reference picture memory (557) for use in future inter-picture prediction.

[0053] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a fresh current picture buffer can be reallocated before beginning reconstruction of the next coded picture.

[0054] The video decoder (510) may perform decoding operations according to a given video compression technology or standard, such as ITU-T Recommendation H.265. The coded video sequence may conform to the syntax specified by the video compression technology or standard being used, in the sense that the coded video sequence adheres to both the syntax of the video compression technology or standard and the profile documented in the video compression technology or standard. Specifically, a profile may select certain tools from among all tools available in the video compression technology or standard as the only tools available under that profile. Compliance may also require that the complexity of the coded video sequence be within a range defined by the level of the video compression technology or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further constrained by a hypothetical reference decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.

[0055] In embodiments, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0056] 6 shows an example block diagram of a video encoder (603). The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmitting circuit). The video encoder (603) may be used in place of the video encoder (403) in the example of FIG.

[0057] The video encoder (603) may receive video samples from a video source (601) (not part of the electronic device (620) in the example of FIG. 6) that may capture video images to be coded by the video encoder (603). In another example, the video source (601) is part of the electronic device (620).

[0058] The video source (601) may provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream, which may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (601) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures that, when viewed in sequence, impart motion. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art can readily understand the relationship between pixels and samples. The following description focuses on samples.

[0059] According to an embodiment, the video encoder (603) may code and compress pictures of a source video sequence into a coded video sequence (643) in real time or under any other required time constraints. Enforcing an appropriate coding rate is one function of the controller (650). In some embodiments, the controller (650) controls and is operatively coupled to other functional units described below. This coupling is not depicted for clarity. Parameters set by the controller (650) may include rate control-related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, ...), picture size, Group of Pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured with other appropriate functions related to the video encoder (603) optimized for a particular system design.

[0060] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As an overly simplified explanation, in an example, the coding loop may include a source coder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be coded and a reference picture) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to that used by a (remote) decoder. The reconstructed sample stream (sample data) is input to a reference picture memory (634). Because decoding of the symbol stream produces bit-exact results regardless of the location of the decoder (local or remote), the contents of the reference picture memory (634) are also bit-exact between the local and remote encoders. In other words, the predictive portion of the encoder "sees" the exact same sample values ​​as the decoder would "see" when using prediction during decoding. This basic principle of reference picture synchronism (and the resulting drift if synchronism cannot be maintained, for example due to channel errors) is also used in several related techniques.

[0061] The operation of the "local" decoder (633) may be the same as that of a "remote" decoder, such as the video decoder (510), as already described in detail above in connection with Figure 5. Referring also briefly to Figure 5, however, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder (645) and parser (520) may be lossless, the entropy decoding portion of the video decoder (510), including the buffer memory (515), and the parser (520), may not be fully implemented in the local decoder (633).

[0062] In embodiments, decoder technology, excluding analysis / entropy decoding, present in a decoder is present in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on decoder operation. Descriptions of encoder technology may be omitted, as they are the reverse of the decoder technology described generically. In certain areas, more detailed descriptions are provided below.

[0063] In operation, in some examples, the source coder (630) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as “reference pictures.” In this manner, the coding engine (632) codes differences between pixel blocks of the input picture and pixel blocks of reference pictures that may be selected as predictive references for the input picture.

[0064] The local video decoder (633) may decode coded video data of pictures that may be designated as reference pictures based on symbols created by the source coder (630). The operation of the coding engine (632) may advantageously be a lossy process. When the coded video data is decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (633) may replicate the decoding process that may be performed by the video decoder on the reference pictures and store the reconstructed reference pictures in the reference picture memory (634). In this way, the video encoder (603) may locally store copies of reconstructed reference pictures that have common content as reconstructed reference pictures obtained by the far-end video decoder (without transmission errors).

[0065] The predictor (635) may perform predictive searches for the coding engine (632). That is, for a new picture to be coded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) that may serve as suitable predictive references for the new image, or for specific metadata such as reference picture motion vectors, block shapes, etc. The predictor (635) may operate on sample blocks on a pixel block by pixel block basis to find a suitable predictive reference. In some cases, as determined by search results obtained by the predictor (635), the input picture may have predictive references drawn from multiple reference pictures stored in the reference picture memory (634).

[0066] The controller (650) may manage the coding operations of the source coder (630), including, for example, setting parameters and subgroup parameters used to encode the video data.

[0067] The output of all of the aforementioned functional units may be entropy coded in an entropy coder (645), which converts the symbols produced by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0068] The transmitter (640) may buffer the coded video sequence created by the entropy coder (645) in preparation for transmission over a communication channel (660), which may be a hardware / software link to a storage device that stores the coded video data. The transmitter (640) may merge the coded video data from the video encoder (603) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0069] The controller (650) can manage the operation of the video encoder (603). During coding, the controller (650) can assign a particular coded picture format to each coded picture, which can affect the coding technique that may be applied to the respective picture. For example, pictures can often be assigned as one of the following picture types:

[0070] An intra picture (I-picture) may be one that can be coded and decoded without using other pictures in a sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh ("IDR") pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.

[0071] A predictive picture (P picture) may be one that can be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict the sample values ​​of each block.

[0072] A bi-predictive picture (B-picture) may be one that can be coded and decoded using intra- or inter-prediction, which uses at most two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, a multi-predictive picture may use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0073] A source picture is generally spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.

[0074] The video encoder (603) may perform coding operations in accordance with a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In doing so, the video encoder (603) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.

[0075] In an embodiment, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0076] Video may be captured as a time-sequence of multiple source pictures (video pictures). Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or other) between pictures. In an example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. If a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture may be coded by a vector called a motion vector. The motion vector points to a reference block within the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0077] In some embodiments, bi-prediction techniques can be used in inter-picture prediction. According to bi-prediction techniques, two reference pictures, such as a first reference picture and a second reference picture, are used, both of which are before the decoding order of the current picture in the video (but their display orders may be past and future, respectively). A block in the current picture may be coded by a first motion vector pointing to a first reference block in the first reference picture and by a second motion vector pointing to a second reference block in the second reference picture. A block may be predicted by a combination of the first reference block and the second reference block.

[0078] Furthermore, in inter-picture prediction, merge mode techniques can be used to improve coding efficiency.

[0079] According to some embodiments of the present disclosure, prediction, such as inter-picture prediction and intra-picture prediction, is performed block-by-block. For example, according to the HEVC standard, a picture in a sequence of video pictures is divided into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU may be recursively quadtree-divided into one or more coding units (CUs). For example, a 64x64 pixel CTU may be divided into one CU of 64x64 pixels, four CUs of 32x32 pixels, or 16 CUs of 16x16 pixels. In an example, each CU is analyzed to determine a prediction type for the CU, such as an inter-prediction type or an intra-prediction type. A CU is divided into one or more prediction units (PUs) according to temporal predictability and / or spatial predictability. Generally, each PU includes one luma prediction block (PB) and two chroma PBs. In an embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Taking a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values ​​(e.g., luma values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.

[0080] 7 shows an example diagram of a video encoder (703). The video encoder (703) is configured to receive a processed block of sample values ​​(e.g., a predictive block) in a current video picture in a sequence of video pictures and encode the processed block into a coded picture that is part of a coded video sequence. In an example, the video encoder (703) is used in place of the video encoder (403) of the example of FIG. 4.

[0081] In an HEVC example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as a predictive block of 8x8 samples. The video encoder (703) determines whether the processing block is best coded using intra-mode, inter-mode, or bi-predictive mode, for example, using rate-distortion optimization. When the processing block is to be coded in intra-mode, the video encoder (703) may encode the processing block into a coded picture using intra-prediction techniques, and when the processing block is to be coded in inter-mode or bi-predictive mode, the video encoder (703) may encode the processing block into a coded picture using inter-prediction techniques or bi-prediction techniques, respectively. In certain video coding techniques, the merge mode can be an inter-picture prediction sub-mode in which a motion vector is derived from one or more motion vector predictors without the aid of coded motion vector components outside the predictors. In certain other video coding techniques, there may be motion vector components applicable to the current block. In examples, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the mode of the processing blocks.

[0082] In the example of Figure 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725), coupled to each other as shown in Figure 7.

[0083] The inter-encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture), generate inter-prediction information (e.g., a description of redundancy information, motion vectors, merge mode information according to an inter-coding technique), and calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture that is decoded based on the coded video information.

[0084] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), possibly compare the block with previously coded blocks in the same picture, generate transformed quantized coefficients, and possibly also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques). In an example, the intra encoder (722) also calculates intra prediction results (e.g., predicted blocks) based on the intra prediction information and reference blocks in the same picture.

[0085] The general-purpose controller (721) is configured to determine general-purpose control data and control other components of the video encoder (703) based on the general-purpose control data. In an example, the general-purpose controller (721) determines the mode of the block and provides a control signal to the switch (726) based on the mode. For example, if the mode is intra-mode, the general-purpose controller (721) controls the switch (726) to select intra-mode results for use by the residual calculator (723) and controls the entropy encoder (725) to select intra-prediction information and include the intra-prediction information in the bitstream. If the mode is inter-mode, the general-purpose controller (721) controls the switch (726) to select inter-prediction results for use by the residual calculator (723) and controls the entropy encoder (725) to select inter-prediction information and include the inter-prediction information in the bitstream.

[0086] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and a prediction result selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) is configured to operate based on the residual data to encode the residual data to generate transform coefficients. In an example, the residual encoder (724) is configured to transform the residual data from the spatial domain to the frequency domain to generate transform coefficients. The transform coefficients then undergo a quantization process to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data may be used by the intra-encoder (722) and the inter-encoder (730) as appropriate. For example, the inter-encoder (730) can generate decoded blocks based on the decoded residual data and inter-prediction information, and the intra-encoder (722) can generate decoded blocks based on the decoded residual data and intra-prediction information. The decoded blocks are processed appropriately to generate decoded pictures, which, in some examples, may be buffered in a memory circuit (not shown) and used as reference pictures.

[0087] The entropy encoder (725) is configured to format a bitstream to include the encoded blocks. The entropy encoder (725) is configured to include various information in the bitstream in accordance with an appropriate standard, such as the HEVC standard. In an example, the entropy encoder (725) is configured to include in the bitstream general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information. Note that, in accordance with the subject matter of this disclosure, when coding a block in a merged sub-mode of either an inter mode or a bi-prediction mode, residual information is not present.

[0088] 8 shows an example diagram of a video decoder (810). The video decoder (810) is configured to receive coded pictures that are part of a coded video sequence and decode the coded pictures to generate reconstructed pictures. In an example, the video decoder (810) is used in place of the video decoder (410) of the example of FIG. 4.

[0089] In the example of Figure 8, the video decoder (810) includes an entropy decoder (871), an inter-decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-decoder (872), coupled together as shown in Figure 8.

[0090] The entropy decoder (871) may be configured to reconstruct, from a coded picture, specific symbols representing syntax elements that make up the coded picture. Such symbols may include, for example, prediction information (e.g., intra-mode, inter-mode, bi-predictive mode, etc., where inter-mode and bi-predictive mode are in merged or other submodes) that can identify the mode in which the block is coded, as well as specific samples or metadata used for prediction by the intra-decoder (872) or inter-decoder (880), respectively (e.g., intra-predictive information or inter-predictive information). The symbols may also include, for example, residual information in the form of quantized transform coefficients. In an example, when the prediction mode is an inter-mode or bi-predictive mode, the inter-predictive information is provided to the inter-decoder (880); when the prediction type is an intra-predictive type, the intra-predictive information is provided to the intra-decoder (872). The residual information may undergo inverse quantization and be provided to the residual decoder (873).

[0091] The inter decoder (880) is configured to receive the inter prediction information and generate an inter prediction result based on the inter prediction information.

[0092] The intra decoder (872) is configured to receive the intra prediction information and to generate a prediction result based on the intra prediction information.

[0093] The residual decoder (873) is configured to perform inverse quantization to extract inverse quantized transform coefficients and process the inverse quantized transform coefficients to transform the residual information from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (to include quantizer parameters (QP)), which may be provided by the entropy decoder (871) (data path not shown as this may be only low volume control information).

[0094] The reconstruction module (874) is configured to combine, in the spatial domain, the residual information output by the residual decoder (873) and the prediction results (possibly with the prediction results output by the inter- or intra-prediction module) to form reconstructed blocks that may be part of a reconstructed picture, which in turn may be part of a reconstructed video. It should be noted that other suitable operations, such as deblocking operations, may be performed to improve visual quality.

[0095] It should be noted that the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using any suitable technology. In an embodiment, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (603) and the video decoders (410), (510), and (810) may be implemented using one or more processors executing software instructions.

[0096] VVC can use various inter-prediction modes. For an inter-predicted CU, motion parameters can include MVs, one or more reference picture indexes, a reference picture list usage index, and additional information about specific coding features to be used for inter-predicted sample generation. Motion parameters can be signaled explicitly or implicitly. When a CU is coded in skip mode, the CU can be associated with a PU and may have no significant residual coefficients, coded motion vector deltas, MV differences (e.g., MVDs), or reference picture indexes. A merge mode can be specified when the motion parameters of the current CU are obtained from neighboring CUs, including spatial and / or temporal candidates, and optionally additional information such as those introduced in VVC. The merge mode can be applied not only to skip mode but also to inter-predicted CUs. In an example, an alternative to the merge mode is explicit transmission of motion parameters, in which the MVs, the corresponding reference picture indexes of each reference picture list, and a reference picture list usage flag, and other information are explicitly signaled for each CU.

[0097] In embodiments such as VVC, the VVC Test Model (VTM) reference software includes one or more refined inter-prediction coding tools, including enhanced merge prediction, merge motion vector differential (MMVD) mode, adaptive motion vector prediction with symmetric MVD signaling (AMVP) mode, affine motion compensation prediction, sub-block-based temporal motion vector prediction (SbTMVP), adaptive motion vector resolution (AMVR), motion field storage (1 / 16 luma sample MV storage and 8x8 motion field compression), bi-prediction with CU-level weights (BCW), bidirectional optical flow (BDOF), prediction refinement using optical flow (PROF), decoder-side motion vector refinement (DMVR), combined inter and intra prediction (CIIP), geometric partitioning mode (GPM), etc. Inter prediction and related methods are described in detail below.

[0098] In some examples, enhanced merge prediction may be used. In examples such as VTM4, the merge candidate list is constructed by including five types of candidates in order: spatial motion vector predictors (MVPs) from spatially nearby CUs, temporal MVPs from co-located CUs, history-based MVPs (HMVPs) from a first-in-first-out (FIFO) table, pairwise average MVPs, and zero MVs.

[0099] The size of the merge candidate list may be signaled in the slice header. In an example, the maximum allowed size of the merge candidate list is 6 in VTM4. For each CU coded in merge mode, the index of the best merge candidate (e.g., merge index) may be coded using truncated unary binarization (TU). The first bin of the merge index may be coded with context (e.g., context-adaptive binary arithmetic coding (CABAC)), and bypass coding may be used for the other bins.

[0100] Some examples of the generation process for each category of merge candidates are given below. In an embodiment, spatial candidate(s) are derived as follows: The derivation of spatial merge candidates in VVC may be the same as that in HEVC. In an example, up to four merge candidates are selected from the candidates at the positions shown in FIG. 9. FIG. 9 illustrates the positions of spatial merge candidates according to an embodiment of the present disclosure. Referring to FIG. 9, the derivation order is B1, A1, B0, A0, and B2. Position B2 is only considered when none of the CUs at positions A0, B0, B1, and A1 are available or intra-coded (e.g., because the CU belongs to another slice or another tile). After the candidate at position A1 is added, the addition of the remaining candidates is subject to a redundancy check that ensures that candidates with the same motion information are excluded from the candidate list, so that coding efficiency is improved.

[0101] To reduce computational complexity, not all possible candidate pairs are considered in the above-described redundancy check. Instead, only pairs linked by arrows in FIG. 10 are considered, and a candidate is added to the candidate list only if the corresponding candidate used for the redundancy check does not have the same motion information. FIG. 10 illustrates candidate pairs considered for the spatial merge candidate redundancy check according to an embodiment of the present disclosure. Referring to FIG. 10, each arrow-linked pair includes A1 and B1, A1 and A0, A1 and B2, B1 and B0, and B1 and B2. Therefore, a candidate at position B1, A0, and / or B2 can be compared with a candidate at position A1, and a candidate at position B0 and / or B2 can be compared with a candidate at position B1.

[0102] In an embodiment, temporal candidates are derived as follows: In an example, only one temporal merge candidate is added to the candidate list. Figure 11 shows exemplary motion vector scaling for temporal merge candidates. To derive a temporal merge candidate for a current CU (1111) in a current picture (1101), a scaled MV (1121) (e.g., indicated by a dotted line in Figure 11) may be derived based on a collocated CU (1112) belonging to a collocated reference picture (1104). In an example, a collocated reference picture (also referred to as a collocated picture) is a specific reference picture used, for example, for temporal motion vector prediction. The collocated reference picture used for temporal motion vector prediction may be indicated by a reference index in a syntax statement such as a high-level syntax (e.g., a picture header, a slice header).

[0103] The reference picture list used to derive the co-located CU (1112) may be explicitly signaled in the slice header. The scaled MV (1121) for the temporal merge candidate may be obtained as shown by the dotted line in Figure 11. The scaled MV (1121) may be scaled from the MV of the co-located CU (1112) using picture order count (POC) distances tb and td. The POC distance tb may be defined to be the POC difference between the current reference picture (1102) of the current picture (1101) and the current picture (1101). The POC distance td may be defined to be the POC difference between the co-located reference picture (1104) of the co-located reference picture (1103) and the co-located reference picture (1103). The reference picture index of the temporal merge candidate may be set to zero.

[0104] FIG. 12 shows exemplary candidate positions (e.g., C0 and C1) of temporal merge candidates for the current CU. The position of the temporal merge candidate may be selected from candidate positions C0 and C1. Candidate position C0 is located at the bottom right corner of the current CU's co-located CU (1210). Candidate position C1 is located at the center of the current CU's co-located CU (1210). If the CU at candidate position C0 is unavailable, intra-coded, or outside the current row of the CTU, candidate position C1 is used to derive the temporal merge candidate. Otherwise, for example, if the CU at candidate position C0 is available, intra-coded, and located in the current row of the CTU, candidate position C0 is used to derive the temporal merge candidate.

[0105] Merge of motion vector difference (MMVD) may be used in either skip mode or merge mode using the motion vector representation method. Merge candidates such as those used in VVC can be reused in MMVD mode. A candidate can be selected from the merge candidates as a starting point (e.g., MV predictor (MVP)) and further extended by MMVD mode. In MMVD mode, simplified signaling may be used for the new motion vector representation. In an example where the motion vector representation method includes a starting point and MV difference (MVD), the MVD is indicated by the magnitude of the MVD (or motion magnitude) and the direction of the MVD (e.g., motion direction).

[0106] The MMVD mode can use a merge candidate list such as that used in VVC. In an embodiment, only candidates of the default merge type (e.g., MRG_TYPE_DEFAULT_N) are considered for the MMVD mode. The starting point can be indicated or defined by a base candidate index (IDX). The base candidate index can indicate a candidate (e.g., the best candidate) among candidates (e.g., base candidates) in the merge candidate list. Table 1 shows an example relationship between the base candidate index and the corresponding starting point. A base candidate index of 0, 1, 2, or 3 indicates that the corresponding starting point is the first MVP, second MVP, third MVP, or fourth MVP. In an example, if the number of base candidates is equal to 1, the base candidate index is not signaled.

[0107] [Table 1]

[0108] The distance index may indicate information about the magnitude of the motion of the MVD, such as the magnitude of the MVD. For example, the distance index indicates a distance (e.g., a predetermined distance) from a starting point (e.g., an MVP indicated by a base candidate index). In the example, the distance is one of a plurality of predetermined distances as shown in Table 2. Table 2 shows an exemplary relationship between the distance index and the corresponding distance (in units of samples or pixels). One pixel in Table 2 is one sample or one pixel. For example, a distance index of 1 indicates a distance of 1 / 2 pixel or 1 / 2 sample.

[0109] [Table 2]

[0110] The direction index can represent the direction of the MVD relative to the starting point. The direction index can represent one of multiple directions, such as the four directions shown in Table 3. For example, a direction index of 00 indicates that the direction of the MVD is along the positive x-axis.

[0111] [Table 3]

[0112] The MMVD flag can be signaled after sending the skip-and-merge flag. If the skip-and-merge flag is true, the MMVD flag is parsed. In an example, if the MMVD flag is equal to 1, the MMVD syntax (e.g., distance index and / or direction index) can be parsed. If the MMVD flag is not equal to 1, the AFFINE flag can be parsed. If the AFFINE lag is equal to 1, the AFFINE mode is used to code the current block. If the AFFINE flag is not equal to 1, the skip / merge index can be parsed for the skip / merge mode, such as used in VTM.

[0113] 13-14 show an example of a search process in MMVD mode, which can determine indices, including a base candidate index, a direction index, and / or a distance index, for a current block (1300) in a current image (also called a current frame) (1301).

[0114] A first motion vector (MV) (1311) and a second motion vector (MV) (1321) belonging to a first merging candidate are shown. The first merging candidate may be a merging candidate on the merging candidate list configured for the current block (1300). The first MV (1311) and the second MV (1321) may be associated with two reference pictures (1302) and (1303) in the reference picture lists L0 and L1, respectively. Thus, the two starting points (1411) and (1421) in Figures 13-14 may be determined at the reference pictures (1302) and (1303), respectively.

[0115] In an example, based on the starting points (1411) and (1421), a number of predetermined points (1-12 shown in FIG. 14) extending from the starting points (1411) and (1421) in the vertical direction (represented by +Y or −Y) or the horizontal direction (represented by +X and −X) of the reference pictures (1302) and (1303) may be evaluated. In one example, pairs of points that mirror each other with respect to the respective starting points (1411) or (1421), such as pair of points (1414) and (1424) or pair of points (1415) and (1425), may be used to determine a pair of MVs (1314) and (1324) or a pair of MVs (1315) and (1325) that may form candidate MV predictors (MVPs) for the current block (1300). MVP candidates determined based on predetermined points surrounding the starting point (1411) and / or (1421) can be evaluated. Referring to Figure 13, the MVD (1312) between the first MV (1311) and MV (1314) has a magnitude of 1S. The MVD (1322) between the second MV (1321) and MV (1324) has a magnitude of 1S. Similarly, the MVD between the first MV (1311) and MV (1315) has a magnitude of 2S. The MVD between the second MV (1321) and MV (1325) has a magnitude of 2S.

[0116] In addition to the first merge candidate, other available or valid merge candidates in the merge candidate list of the current block (1300) may also be evaluated. In one example, for a uni-predictive merge candidate, only one prediction direction associated with one of the two reference picture lists is evaluated.

[0117] For example, based on these evaluations, the best MVP candidate can be determined. Therefore, the best merge candidate can be selected from the merge list corresponding to the best MVP candidate, and the motion direction and motion distance can also be determined. For example, based on the selected merge candidate and Table 1, a base candidate index can be determined. Based on the selected MVP, such as the one corresponding to a predetermined point (1415) (or (1425)), the direction and distance (e.g., 2S) of point (1415) relative to the starting point (1411) can be determined. According to Tables 2 and 3, the direction index and distance index can be determined accordingly.

[0118] As described above, two indexes, such as a distance index and a direction index, can be used to indicate the MVD in the MMVD mode, or a single index can be used to indicate the MVD in the MMVD mode, for example, using a table that pairs a single index with the MVD.

[0119] Template matching (TM)-based candidate reordering can be used in some prediction modes, such as MMVD mode and affine MMVD mode. In an embodiment, the MMVD offset is extended for MMVD mode and affine MMVD mode. FIG. 15 shows additional refinement positions along multiple diagonals, such as a k×π / 8 diagonal, where k is an integer from 0 to 15. The additional refinement positions along multiple diagonals can increase the number of directions, for example, from four directions (e.g., +X, −X, +Y, and −Y) to 16 directions (e.g., k=0, 1, 2, ..., 15). In the example, each of the 16 directions is represented by an angle between the +X direction and a direction indicated by a center point (1500) and one of points 1 to 16. For example, point 1 indicates the +X direction at angle 0 (i.e., k=0), point 2 indicates a direction along angle 1×π / 8 (i.e., k=1), and so on.

[0120] TM can be performed in MMVD mode. In an example, for each MMVD refinement position, a TM cost can be determined based on the current template of the current block and one or more reference templates. The TM cost can be determined using any method, such as sum of absolute differences (SAD) (e.g., SAD cost), sum of absolute transformed differences (SATD), sum of squared errors (SSE), mean-removed SAD / SATD / SSE, variance, partial SAD, partial SSE, or partial SATD.

[0121] The current template for the current block may include any suitable samples, such as one row of samples above the current block and / or one column of samples to the left of the current block. Based on the TM cost (e.g., SAD cost) between the current template and the corresponding reference template for the refinement position, the MMVD refinement positions, e.g., all possible MMVD refinement positions (e.g., 16×6 represents 16 directions and 6 magnitudes) for each base candidate (e.g., MVP), may be sorted. In an example, the top MMVD refinement positions with the smallest TM costs (e.g., smallest SAD costs) are retained as available MMVD refinement positions for MMVD index coding. For example, a subset (e.g., 8) of the MMVD refinement positions with the smallest TM costs is used for MMVD index coding. For example, the MMVD index indicates which of the subset of MMVD refinement positions with the smallest TM costs is selected for coding the current block. In an example, an MMVD index of 0 indicates that the MVD (e.g., MMVD refinement position) corresponding to the smallest TM cost is used for coding the current block. The MMVD index can be binarized, for example, by a Rice code with a parameter equal to 2.

[0122] In an embodiment, in addition to the above-described MMVD offset extension as in Figure 15, the affine MMVD reordering is extended in which additional refinement positions along the k × π / 4 diagonal are added. After reordering, the top half of refinement positions with the smallest TM cost (e.g., SAD cost) are retained for coding the current block.

[0123] In some examples (e.g., HEVC), AMVP can be used to predict the motion vector of a current block by exploiting the spatial and temporal correlation of neighboring partitions. On the encoder side, a rate-distortion optimization (RDO) process is used to select the best motion vector predictor from a candidate list of candidates. The index of the selected candidate is then encoded and transmitted to the decoder. On the decoder side, the same candidate list as the encoder side can be constructed, for example, in a defined manner. The candidate list can be constructed in three steps. In the first step, the decoder retrieves spatial and temporal motion vectors from a memory buffer to form the candidate list. In the second step, a redundancy check process is utilized to remove duplicate motion vectors from the candidate list. In the third step, a zero motion check process is optionally employed to check for the presence of zero motion in the candidate list. Note that the construction of the candidate list can shorten the candidate list by removing duplicate motion vectors, and fewer bits can be used to signal the index of the selected candidate in the candidate list.

[0124] Note that in the case of AMVP, the encoder also signals a reference picture index to specify the reference picture to which the motion vector predictor specified by the index of the selected candidate in the candidate list points. Furthermore, in the case of AMVP, the encoder can determine a motion vector differential (MVD) for the current block, where the MVD is the difference between the motion vector predictor and the true or disparity motion vector used for the current block. In the case of AMVP, in addition to the reference picture index and the index of the selected candidate in the candidate list, the encoder also signals the MVD of the current block in the bitstream. Due to the signaling of the reference picture index and the predictor vector difference for a given block, AMVP may not be as efficient as merge mode, but it can improve the fidelity of the coded video data.

[0125] In some examples, the number of spatial neighbor candidates in an AMVP candidate list is limited to no more than a threshold number, and the number of temporal candidates is limited to no more than a threshold number. For example, an AMVP candidate list has a maximum of two spatial neighbor candidates and one co-located temporal candidate. Potential AMVP spatial neighbor candidates are located at the bottom-left, left, top-right, top, and top-left positions of the current block. In the examples, potential AMVP spatial neighbor candidates are classified into two classes. The left and bottom-left candidates are classified into a first class, and the top-right, top, and top-left candidates are classified into a second class. In the examples, a scan order is used to place potential AMVP spatial neighbor candidates in the candidate list. For example, the scan order is bottom-to-top for the first class and right-to-left for the second class, respectively.

[0126] In some examples, affine AMVP mode may apply to CUs with both width and height equal to or greater than 16. A CU-level affine flag may be signaled in the bitstream to indicate whether affine AMVP mode is used, and then another flag may be signaled to indicate whether 4-parameter affine or 6-parameter affine is applied. In affine AMVP mode, the difference between the CPMV of the current CU and the predictor of the CPMVP of the current CU may be signaled in the bitstream. In some examples, an affine AMVP candidate list may have any suitable number of candidates. In examples, the size of the affine AMVP candidate list may be 2. In some examples, an affine AMVP candidate list may be generated by using four types of CPMV candidates in the following order:

[0127] (1) Inherited affine AMVP candidates extrapolated from the CPMVs of nearby CUs, (2) constructed affine AMVP candidates with CPMVPs derived using the translational MVs of nearby CUs; (3) translational MVs from nearby CUs, and (4) Zero MV.

[0128] The check order of inherited affine AMVP candidates may be the same as the check order of inherited affine merge candidates. To determine the AMVP candidate, affine CUs with the same reference picture as the current block may be considered. When an inherited affine motion predictor is inserted into the candidate list, the pruning process may not be applied.

[0129] According to aspects of the present disclosure, a technique referred to as adaptive motion vector resolution (AMVR) may be used in video coding. Note that a motion vector differential (MVD) (between a true motion vector and a motion vector predictor of a CU) may be signaled. In some examples (e.g., HEVC), if a flag (e.g., use_integer_mv_flag) is equal to 0 in a slice header, the MVD is signaled in units of 1 / 4 luma samples. In some examples (e.g., VVC), a CU-level adaptive motion vector resolution (AMVR) scheme is used. AMVR enables the MVD of a CU to be coded with different precisions. In some examples, the MVD of a current CU may be adaptively selected depending on the mode of the current CU (normal AMVP mode or affine AMVP mode). For example, in normal AMVP mode, the resolution may include 1 / 4 luma samples, 1 / 2 luma samples, integer luma samples, or 4 luma samples, and in affine AMVP mode, the resolution may include 1 / 4 luma samples, integer luma samples, or 1 / 16 luma samples.

[0130] In some examples, the CU-level MVD resolution indication is conditionally signaled if the current CU has at least one non-zero MVD component. If all MVD components (e.g., which may include both horizontal and vertical MVDs for reference list L0 and reference list L1) are 0, a 1 / 4 luma sample MVD resolution is inferred.

[0131] In some examples, for a CU with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter luma sample MVD precision is used for the CU. If the first flag is 0, no further signaling is required and quarter luma sample MVD precision is used for the current CU. Otherwise (e.g., if the first flag is 1), a second flag is signaled to indicate whether half luma sample or other MVD precision (integer or 4 luma sample) is used for the normal AMVP CU. For half luma samples (e.g., the second flag is 0), a 6-tap interpolation filter is used for half luma sample positions instead of the default 8-tap interpolation filter. Otherwise (e.g., the second flag is 1), a third flag is signaled to indicate whether integer luma sample or 4 luma sample MVD precision is used for the normal AMVP CU. For an affine AMVP CU, the second flag is used to indicate whether integer luma sample or 1 / 16 luma sample MVD precision is used.

[0132] In some examples, to ensure that the reconstructed MV has the intended precision (¼ luma samples, ½ luma samples, integer luma samples, or 4 luma samples), the motion vector predictor of a CU can be rounded to the same precision as the MVD precision before being added together with the MVD. The motion vector predictor is rounded towards 0 (i.e., negative motion vector predictors are rounded towards positive infinity, and positive motion vector predictors are rounded towards negative infinity).

[0133] In some examples, the encoder uses a rate-distortion (RD) check to determine the motion vector resolution of the current CU. To avoid always performing four CU-level RD checks for each MVD, in some examples (e.g., VTM19), the RD check for MVD precisions other than 1 / 4 luma sample precision is only conditionally invoked. For normal AVMP mode, the RD costs for 1 / 4 luma sample MVD precision and integer luma sample MVD precision can be calculated first. Then, the RD cost for integer luma sample MVD precision is compared with the RD cost for 1 / 4 luma sample MVD precision to determine whether the RD cost for 4 luma sample MVD precision needs to be further checked. If the RD cost for 1 / 4 luma sample MVD precision is much smaller than the RD cost for integer luma sample MVD precision, the RD check for 4 luma sample MVD precision is skipped. Next, if the RD cost for integer luma sample MVD precision is significantly larger than the best RD cost of the previously tested MVD precision, the check for 1 / 2 luma sample MVD precision is skipped. For affine AMVP mode, if affine inter mode is not selected after checking the rate-distortion costs of affine merge / skip mode, merge / skip mode, regular AMVP mode with 1 / 4 luma sample MVD precision, and affine AMVP mode with 1 / 4 luma sample MVD precision, the 1 / 16 luma sample MV precision and 1 pixel MV precision affine inter modes are not checked. Furthermore, the affine parameters obtained in the 1 / 4 luma sample MV precision affine inter mode are used as the starting search points for the affine inter modes with 1 / 16 luma sample and 1 / 4 luma sample MV precision.

[0134] To improve coding efficiency and reduce MV transmission overhead, subblock-level MV refinement can be applied to extend CU-level temporal motion vector prediction (TMVP). For example, subblock-based TMVP (SbTMVP) mode enables subblock-level inheritance of motion information from a co-located reference picture. As described above, the co-located reference picture can be indicated by a reference index in a syntax such as a high-level syntax (e.g., a picture header, a slice header). Each sub-block of a current CU (e.g., a large-sized current CU) of a current picture can have its own motion information without explicitly transmitting a block partition structure or its own motion information. In SbTMVP mode, the motion information for each sub-block can be obtained, for example, in three steps as follows: In the first step, a displacement vector (DV) of the current CU can be derived. The DV can indicate a block of the co-located reference picture, for example, the DV points from the current block of the current picture to the block of the co-located reference picture. Therefore, the block indicated by the DV is considered to be co-located with the current block and is called the co-located block of the current block. In the second step, the availability of SbTMVP candidates can be checked and the center motion (e.g., the center motion of the current CU) can be derived. In the third step, the DV can be used to derive sub-block motion information from the corresponding sub-block of the co-located block. The three steps can be combined into one or two steps, and / or the order of the three steps can be adjusted.

[0135] Unlike TMVP candidate derivation, which derives temporal MVs from co-located blocks in a reference frame or picture, SbTMVP mode can apply DVs (e.g., DVs derived from the MVs of the current CU's left neighboring CUs) to locate the corresponding sub-blocks in the co-located reference picture for each sub-block of the current CU in the current picture. If the corresponding sub-block is not inter-coded, the motion information of the current sub-block can be the motion of the center of the co-located block.

[0136] The SbTMVP mode can be supported by various video coding standards, including, for example, VVC. Similar to the TMVP mode, for example, in HEVC, the SbTMVP mode can use the motion fields (also called motion information fields or MV fields) of the co-located reference picture to improve the MV prediction and merging mode of the CU of the current picture. In an example, the SbTMVP mode uses the same co-located reference picture used by the TMVP mode. In an example, the SbTMVP mode differs from the TMVP mode in the following aspects: (i) the TMVP mode predicts motion information at the CU level, while the SbTMVP mode predicts motion information at the sub-CU level; and (ii) the TMVP mode fetches temporal MV from a co-located block of the co-located reference picture (e.g., the co-located block is the bottom-right or center block relative to the current CU), while the SbTMVP mode can apply a motion shift before fetching temporal motion information from the co-located reference picture. In the example, the motion shift used in the SbTMVP mode is obtained from the MV of one of the spatially neighboring blocks of the current CU.

[0137] 16-17 show an exemplary SbTMVP process used in SbTMVP mode. The SbTMVP process can predict the MV of a sub-CU (e.g., a sub-block) within a current CU (e.g., a current block) (1601) of a current picture (1711), for example, in two steps. The first step examines the spatial neighborhood (e.g., A1) of the current block (1601) in FIGS. 16-17. If the spatial neighborhood (e.g., A1) has a MV (1721) that uses a co-located reference picture (1712) as the reference picture for the spatial neighborhood (e.g., A1), the MV (1721) can be selected to be the motion shift (or DV) to be applied to the current block (1601). If no such MV (e.g., an MV that uses the co-located reference picture (1712) as a reference picture) is identified, the motion shift or DV can be set to 0 MV (e.g., (0,0)). In some examples, additional spatially neighboring MVs, such as A0, B0, B1, etc., are checked if no such MVs are identified for spatial neighbor A1.

[0138] In a second step, the motion shift or DV (1721) identified in the first step may be applied to the current block (1601) (e.g., the DV (1721) may be added to the coordinates of the current block) to obtain sub-CU level motion information (e.g., including MV and reference index) from the co-located reference picture (1712). In the example shown in FIG. 17, the motion shift or DV (1721) is set to be the MV of the spatial neighborhood A1 (e.g., block A1) of the current block (1601). For each sub-CU or sub-block (1731) of the current block (1601), the motion information of the corresponding co-located block (1701) in the co-located reference picture (1712) (e.g., the motion information of the minimum motion grid covering the center sample of the co-located block (1701)) may be used to derive the motion information of the sub-CU or sub-block (1731). After the motion information of the co-located sub-CU (1732) of the co-located block (1701) is identified, the motion information of the co-located sub-CU (1732) can be converted into motion information (e.g., MV and one or more reference indices) of the current sub-CU (1731) using a scaling method, such as a method similar to the TMVP process used in HEVC, and temporal motion scaling is applied to align the reference picture of the temporal MV to the reference picture of the current CU.

[0139] The motion field of the current block (1601) derived based on the DV (1721) may include motion information for each sub-block (1731) of the current block (1601), such as the motion vector (MV) and one or more associated reference indices. The motion field of the current block (1601) may also be referred to as an SbTMVP candidate and corresponds to the DV (1721).

[0140] 17 shows an example of a motion field or SbTMVP candidate for a current block (1601). For example, the motion information of a bi-predicted sub-block (1731(1)) includes a first motion vector (MV) and a first index indicating a first reference picture in reference picture list 0 (L0), a second motion vector (MV) and a second index indicating a second reference picture in reference picture list 1 (L1). In the example, the motion information of a uni-predicted sub-block (1731(2)) includes a motion vector and an index indicating a reference picture in L0 or L1.

[0141] In the example, DV (1721) is applied to the center position of the current block (1601) to locate the displaced center position of the co-located reference picture (1712). If the block containing the displaced center position is not inter-coded, SbTMVP candidates are considered unavailable. Alternatively, if the block containing the displaced center position (e.g., the co-located block (1701)) is inter-coded, motion information for the center position of the current block (1601), referred to as center motion of the current block (1601), may be derived from motion information for the block containing the displaced center position in the co-located reference picture (1712). In the example, a scaling process may be used to derive the center motion of the current block (1601) from motion information for the block containing the displaced center position in the co-located reference picture (1712). If an SbTMVP candidate is available, the DV (1721) can be applied to find a corresponding sub-block (1732) in the co-located reference picture (1712) for each sub-block (1731) of the current block (1601). The motion information of the corresponding sub-block (1732) can be used to derive motion information for the sub-block (1731) of the current block (1601), such as in the same manner as used to derive the central motion of the current block (1601). In an example, if the corresponding sub-block (1732) is not inter-coded, the motion information of the current sub-block (1731) is set to the central motion of the current block (1601).

[0142] In some examples, such as VVC, a combined subblock-based merge list containing SbTMVP candidates and affine merge candidates is used in signaling the subblock-based merge mode. SbTMVP mode can be enabled or disabled by a sequence parameter set (SPS) flag. When SbTMVP mode is enabled, the SbTMVP candidate (or SbTMVP predictor) is added as the first entry of a subblock-based merge list containing subblock-based merge candidates, followed by affine merge candidates. The size of the subblock-based merge list can be signaled in the SPS. In an example, the maximum allowed size of a subblock-based merge list is 5 in VVC. In an example, multiple SbTMVP candidates are included in the subblock-based merge list.

[0143] In some examples, such as VVC, the sub-CU size used in SbTMVP mode is fixed at 8x8, such as used for affine merge mode. In examples, SbTMVP mode is only applicable to CUs whose width and height are both 8 or greater. The size of the sub-blocks (e.g., 8x8) may be configurable to other sizes, such as 4x4 in the ECM software model used for search beyond VVC. In examples, multiple co-located reference pictures, such as two co-located frames, are utilized to provide temporal motion information for SbTMVP and / or TMVP in AMVP mode.

[0144] In some examples of SbTMVP modes, such as VVC and ECM, the DV of the current CU (e.g., DV (1721) in Figure 17) is derived only from the MVs of CUs in the vicinity of the current CU. However, the SbTMVP candidates derived from the DV may not be an exact match.

[0145] In the SbTMVP mode, a DV offset (DVO) can be used. In an example, to obtain a more accurate match, a DV (e.g., an initial DV) can be modified by the DV offset to determine an updated DV'. In an example, the updated DV' is a vector sum of a DV (e.g., the initial DV) and a DVO. The initial DV can be determined using any method such as those described in FIGS. 16 and 17. For example, the initial DV is determined based on the MVs of blocks neighboring the current block. For the current block, the DVO can be signaled and analyzed to indicate an additional motion offset of the initial DV. The DVO can be indicated, for example, by signaling an index indicating the DVO from a DVO candidate. In an example, the DVO is signaled. In an example, the MMVD mode is used to indicate the DVO, for example, the DVO is an MVD indicated by a direction index and / or a distance index as described in Tables 2 and 3. By using DV_O, the position of the co-located CU (or co-located block) in the co-located reference picture can be adjusted, and therefore, the MV field of the co-located CU (or co-located block) can vary based on DV_O. If DV_O is not 0, the updated DV' can be used as a displacement vector to indicate the position of the co-located CU (or co-located block) to perform the SbTMVP process. Referring to FIG. 17, instead of using the initial DV (e.g., DV(1721)), which is the MV of the spatial neighborhood A1, the updated DV' can be used to determine the co-located block of the current block. The SbTMVP candidate for the current block can be derived using the updated DV' (e.g., the vector sum of the initial DV and DV_O).

[0146] In an embodiment, DVO is signaled directly using any signaling method used to signal MVD, e.g., in AMVP mode, AMVR mode, etc. In AMVR mode, the MVD of a block can be signaled at different resolutions, such as a resolution of 1 / 4, 1 / 2, 1, or 4 luma samples. DVO can be signaled at different resolutions using AMVR mode.

[0147] The predefined DVO list may include DVO candidates (e.g., possible DVOs used by the current block). One or more indexes may be signaled to indicate which of the DVO candidates should be selected as the DVO.

[0148] In the example, the DVO is signaled using the MMVD mode. For example, as shown in Tables 2 and 3, two indexes are signaled to indicate the DVO candidate, including a first index (e.g., a distance index or a step index) indicating the size of the DVO candidate and a second index (e.g., a direction index) indicating the direction of the DVO candidate.

[0149] Referring back to FIG. 14 , the distance index (or step index) and direction index may be predefined as described above with reference to the MMVD mode. The distance index indicates movement magnitude information, such as the magnitude of the DVO. For example, the distance index indicates a predetermined distance from a starting point (e.g., the initial DV). In an example, available predetermined distances are shown in Table 2. The direction index represents the direction of the DVO relative to the starting point (e.g., the initial DV). The direction index can indicate one of multiple directions, such as the four directions shown in Table 3.

[0150] In an example, a co-located CTU of a co-located reference picture is co-located with a current CTU that includes the current block. The current CTU is located in the current picture. In an embodiment, the location of the co-located block corresponding to the updated DV' is constrained to be within a first region of the co-located reference picture. In an example, the first region of the co-located reference picture includes the co-located CTU. In an example, the first region of the co-located reference picture includes the co-located CTU and one column of 4x4 blocks at the right boundary of the co-located CTU. The updated DV' can be constrained such that the co-located block corresponding to the updated DV' is within the first region of the co-located reference picture. In an example, the DV (e.g., the horizontal component of the DV) is constrained to be within the first region of the co-located reference picture. x and / or the vertical component of DVO y ) is constrained to ensure that the updated DV′ satisfies the co-located block position constraints mentioned above.

[0151] In the example, the maximum vertical component of the updated DV' is H. y ) is the vertical component of DV from H y minus, e.g., DVO y ≦HDV y The maximum horizontal component of the updated DV' is W, and DVO is the horizontal component of DV from W to DV. x minus, e.g., DVO x ≦W-DV x is.

[0152] As shown in Figures 16-17, a co-located block (e.g., 1701) of a co-located reference picture (e.g., 1712) may be determined based on the DV (e.g., 1721) of a current block (e.g., 1601) of a current picture (e.g., 1711). Thus, the motion information (e.g., TMVP) of each sub-block of the current block may be based on the motion information of a corresponding sub-block of the co-located block. According to an embodiment of the present disclosure, updated motion information (e.g., updated TMVP) of each sub-block of the current block may be determined based on the motion information (e.g., TMVP) of the sub-block of the current block and the motion vector offset (MVO) of the current block.

[0153] In an embodiment, the MVO is added to each derived sub-block-level TMVP of each sub-block of the current block to generate an updated sub-block-level TMVP. The MVO can be signaled and analyzed to indicate the additional motion offset of each sub-block-based TMVP of each sub-block of the current block determined using the SbTMVP mode.

[0154] In an embodiment, the MVO is signaled directly using any signaling method used to signal the MVD, e.g., in AMVP mode, AMVR mode, etc. In AMVR mode, the MVD of a block can be signaled at different resolutions, such as a resolution of 1 / 4, 1 / 2, 1, or 4 luma samples. The MVO can be signaled at different resolutions using AMVR mode.

[0155] The predefined MVO list may include MVO candidates (e.g., possible MVOs used by the current block). One or more indexes may be signaled to indicate which MVO candidates in a given MVO list may be selected as an MVO.

[0156] In the example, the MVO is signaled using the MMVD mode. For example, as shown in Tables 2 and 3, two indexes are signaled to indicate the MVO candidate, including a first index (e.g., a distance index or a step index) indicating the size of the MVO candidate and a second index (e.g., a direction index) indicating the direction of the MVO candidate.

[0157] Referring back to FIG. 14 , the distance index (or step index) and direction index of the MVO may be predefined, as described above with reference to the MMVD mode. The distance index indicates motion magnitude information, such as the magnitude of the MVO, and indicates a predetermined distance from a starting point (e.g., a DV used to determine the position of a co-located block in a co-located reference picture). In an example, available predetermined distances are shown in Table 2. The direction index represents the direction of the MVO relative to the starting point (e.g., a DV used to determine the position of a co-located block in a co-located reference picture). The direction index can indicate one of multiple directions, such as the four directions shown in Table 3.

[0158] In the example, a sub-block of a current block is bi-predicted. Referring to Figure 17, a sub-block (1731(1)) is bi-predicted and has a first MV associated with a first reference picture in reference list L0 and a second MV associated with a second reference picture in reference list L1. An MVO may be applied to the first MV associated with reference list L0. The updated first MV may be a vector sum of the first MV and the MVO. The following embodiments may be applied to the second MV associated with reference list L1.

[0159] In the example, no MVO is applied to the second MV associated with reference list L1. In the example, no MVO is applied to the second MV associated with reference list L1. Therefore, the updated motion information of sub-block (1731(1)) includes the updated first MV and second MV.

[0160] In the example, a mirrored MVO of the MVO is applied to a second MV associated with reference list L1. The mirrored MVO and the MVO may have the same magnitude and opposite direction. The horizontal and vertical components of the MVO (e.g., signaled) are multiplied by -1 to obtain the horizontal and vertical components of the mirrored MVO, respectively. The updated second MV may be a vector sum of the second MV and the mirrored MVO, or a vector difference between the second MV and the MVO. Thus, the updated motion information of sub-block (1731(1)) includes an updated first MV (e.g., first MV + MVO) and an updated second MV (e.g., second MV - MVO).

[0161] In an example, a scaled MVO (MVO') may be applied to a second MV associated with reference list L1. The value of each component (e.g., horizontal and vertical components) of the MVO may be scaled based on the first POC difference and the second POC difference, as shown in Equation 1. Equation 1 may be applied to vectors MVO' and MVO. Equation 1 may be applied to each component of vectors MVO' and MVO. The first POC difference is calculated based on the POC (POC) of the current picture in reference picture list L0. curr ) and the POC of the first reference picture (POC L0 ) of the current picture in the reference picture list L1. curr ) and the POC of the second reference picture (POC L1 ) is the difference. MVO' = MVO × (POC L1 -POC curr ) / (POC L0 -POC curr ) Equation 1

[0162] In the example, the scaled MVO (MVO') is added to a second MV associated with reference list L1. In the example, a mirrored MVO' of the scaled MVO (MVO') is added to a second MV associated with reference list L1.

[0163] As described above, in an example of the SbTMVP mode (e.g., the variant of the SbTMVP mode described in Figures 16-17), DVO may be applied to DV, which may adjust the position of the co-located block within the co-located reference picture and thus affect the motion information of the sub-blocks of the current block. In another example of the SbTMVP mode (e.g., another variant of the SbTMVP mode described in Figures 16-17), MVO may be applied to directly adjust the motion information of the sub-blocks of the current block. DVO and MVO may be signaled using the same method. For example, a predetermined MVO list may be identical to a predetermined DVO list, and DVO and MVO may use the same predetermined MVO list.

[0164] In some examples, DVO can be applied to DV to adjust the position of the co-located block within the co-located reference picture. After obtaining the motion information of the sub-block of the current block based on the motion information of the corresponding sub-block of the co-located reference picture determined based on the updated DV' (e.g., DV+DVO), MVO can be applied to further adjust the motion information of the sub-block of the current block. DVO may be the same as or different from MVO.

[0165] Some aspects of the present disclosure provide techniques for a sub-block-based motion vector predictor with a motion vector offset in AMVP mode. For example, the techniques may be used to enable sub-block-based temporal motion vector prediction (SbTMVP) in AMVP mode. A signal in AMVP mode may include MVP information and MVD information of an encoded video bitstream. For SbTMVP, a DV may be derived based on an SbTMVP candidate and an offset, where the SbTMVP candidate is indicated by the MVP information and the offset is indicated by the MVD information. For example, to code a current block in SbTMVP mode, an encoder may generate coding information in AMVP mode with MVP information and MV offset information of the coding information of the current block, where the MVP information indicates an SbTMVP candidate, the MV offset information indicates an offset, and a combination of the SbTMVP candidate and the offset may indicate a final DV that points to a corresponding sub-block of a co-located picture of a sub-block of the current block. On the decoder side, the decoder can determine the offset from the MVP information and MV offset information of the coding information of the SbTMVP candidate and the current block, determine the final DV pointing to the corresponding sub-block in the co-located picture of the sub-block of the current block, and then reconstruct the current block accordingly. Note that allowing SbTMVP in AMVP mode can improve the coding gain of AMVP. Furthermore, in some examples, AMVR is used to code the motion vector (MVD) information to achieve a favorable trade-off between MV accuracy and MVD bit consumption.

[0166] According to aspects of the present disclosure, SbTMVP candidates are derived as MVP candidates in AMVP mode. In some examples, to construct a candidate list of candidates in AMVP mode, SbTMVP candidates are derived and inserted into the candidate list. SbTMVP candidates can be selected in the same way as other candidates in the candidate list, such as by following an index.

[0167] In some embodiments, SbTMVP derivation of a sub-block-based merge candidate list is used to derive SbTMVP candidates as MVP candidates in AMVP mode. In some examples, SbTMVP candidates are derived from the sub-block-based merge candidate list. The sub-block-based merge candidate list is defined to include one or more SbTMVP candidates and / or suitable sub-block-based merge candidates, such as affine merge candidates. In examples, to construct a candidate list in AMVP mode, a sub-block-based merge candidate list is constructed, and candidates from the sub-block-based merge candidate list are added to the candidate list in AMVP mode.

[0168] In a first embodiment, a predetermined order of spatially neighboring coding blocks is used to derive a displacement vector (DV) for SbTMVP. In some examples, the spatially neighboring blocks of a current block are checked according to a predetermined order to find the first spatially neighboring block for which a co-located central sub-block motion vector is available. For example, the motion vector of the first spatially neighboring block points to a corresponding block in a co-located picture, and the center of the corresponding block has available sub-block motion information (also called a co-located central sub-block motion vector). The first spatially neighboring block for which a co-located central sub-block motion vector is available can then be added to a candidate list and used as an SbTMVP candidate to derive a DV.

[0169] In the second embodiment, a zero displacement vector (DV) (e.g., (0,0)) is used to derive SbTMVP. More specifically, if a co-located central sub-block motion vector is available, the co-located sub-block-based motion field data of the co-located picture is used as the SbTMVP candidate. In an example, if the central sub-block motion vector information (also called the co-located central sub-block motion vector) of the current block's co-located block (0DV) is available, 0DV can be added to the candidate list. This allows 0DV to be selected and used as the SbTMVP candidate.

[0170] In some embodiments, the features of the first and second embodiments can be combined. In some examples, the spatially neighboring blocks of the current block are checked according to a predetermined order to find the first spatially neighboring block for which a co-located central sub-block motion vector is available. If the central sub-block motion vectors of all spatially neighboring blocks in the predetermined order are not available, 0DV is added to the candidate list.

[0171] In some embodiments, SbTMVP MVP candidates are added to a sub-block AMVP candidate list, such as an affine AMVP candidate list.

[0172] In some examples, the MVP index is always the first MVP candidate in the subblock (affine) AMVP list, for example, the SbTMVP candidate is inserted as the first candidate in the affine candidate AMVP list.

[0173] In some examples, the adaptive position is determined by checking whether neighboring coding blocks have affine coding blocks. When none of the neighboring coding blocks are coded in affine mode, the SbTMVP MVP candidate is the first candidate in the sub-block (affine) AMVP list. Otherwise, it is placed at the end of the sub-block (affine) AMVP list. For example, all of the spatially neighboring blocks are checked to determine whether any spatially neighboring blocks are affine coded. If none of the spatially neighboring blocks are affine coded, the SbTMVP candidate is inserted as the first candidate in the affine AMVP candidate list. If at least one of the spatially neighboring blocks is affine coded, the SbTMVP candidate is inserted at the end of the affine AMVP candidate list.

[0174] According to an aspect of the present disclosure, the MVD for the AMVR is signaled as an offset of the displacement vector (DV) to derive the SbTMVP of the AMVP. For example, to derive the SbTMVP in the AMVP mode, the MVD is signaled using the AMVR as an offset of the DV.

[0175] In some embodiments, the signaled index of AMVR precision and the signaled MVD are used to derive the offset of the DV of the SbTMVP MVP candidate. In some examples, an index indicating the AMVR precision is signaled. The offset of the DV is determined based on the index of the coded video bitstream and the signaled MVD.

[0176] In some examples, the unit of AMVR accuracy of SbTMVP in AMVP mode is in pixels. The AMVR accuracy of SbTMVP in AMVP mode can be changed to, but is not limited to, 1 pixel, 2 pixels, 4 pixels, and 8 pixels. In examples, the AMVR accuracy of SbTMVP in AMVP mode is a positive integer number of pixels. The AMVR accuracy of SbTMVP in AMVP mode is modified from (1 / 4 pixel, 1 / 2 pixel, 1 pixel, 4 pixels) to (1 pixel, 2 pixels, 4 pixels, 8 pixels).

[0177] In some examples, the value of the offset is derived from Equation 2 below: DV offset (x,y)=N×MVD(x,y)×AMVR[amvr_precision_idx] Equation (2) where N is the width and / or height of the square sub-block, and amvr_precision_idx can be 0, 1, 2, or 3 when used in VVC and ECM. In examples, the AMVR precision (e.g., 1 pixel, 2 pixels, 4 pixels, or 8 pixels) is selected from a lookup table using amvr_precision_idx.

[0178] In some examples, a final displacement vector can be derived by adding the offset value and the displacement vector (DV) of the SbTMVP candidate that points to the corresponding sub-block motion field data of the co-located picture. In some examples, for each sub-block of the current block, the corresponding sub-block of the co-located picture is determined by a final displacement vector that combines the offset value (e.g., according to Equation 2) and the DV of the SbTMVP candidate. The SbTMVP candidate can be determined based on a signaled index that points to the SbTMVP candidate in the candidate list.

[0179] FIG. 18 shows a flowchart outlining an encoding process (1800) according to an embodiment of the present disclosure. The process (1800) can be used in a video encoder. The process (1800) can be performed by an apparatus for video coding, which can include processing circuitry. In various embodiments, the process (1800) is performed by a processing circuit, such as the processing circuitry in terminal devices (310), (320), (330), and (340), or the processing circuitry that performs the functions of a video encoder (e.g., (403), (603), (703)). In some embodiments, the process (1800) is implemented with software instructions, and thus, the processing circuit performs the process (1800) when the processing circuit executes the software instructions. The process begins at (S1801) and proceeds to (S1810).

[0180] At (S1810), an updated DV of the current block of the current picture may be determined based on the DV of the current block and the DV offset of the current block (also referred to as MV offset (MVO)). The DV may be determined as described above, such as in Figures 16-17. The updated DV of the current block indicates a co-located block of the co-located picture. The co-located block is co-located with the current block.

[0181] The current block includes multiple sub-blocks coded using the sub-block-based temporal motion vector prediction (SbTMVP) mode.

[0182] In examples, the DV offset (or MVO) is determined from the DV offset candidates using any suitable method.

[0183] In the example, the updated DV is determined to be the vector sum of the DV and the DV offset.

[0184] In the example, the updated DV is constrained so that the co-located block is within a restricted region of the co-located reference picture, where the restricted region includes the co-located region corresponding to the current CTU of the current picture, and the current CTU includes the current block.

[0185] In (S1820), the motion information of a sub-block among the plurality of sub-blocks may be determined based on the motion information of the corresponding sub-block in the co-located block.

[0186] In (S1830), DV offset information (also referred to as MVO information) indicating a DV offset can be coded. A sub-block among the plurality of sub-blocks can be coded based on motion information of the sub-block among the plurality of sub-blocks.

[0187] In an embodiment, the DV offset information indicates a magnitude of the DV offset and at least one index indicating a direction of the DV offset. In an example, the at least one index includes a distance index indicating the magnitude of the DV offset, which is one of a set of predetermined distances, and a direction index indicating the direction of the DV offset, which is one of a set of predetermined directions.

[0188] In the example, the set of predetermined distances and the set of predetermined directions are used in merged motion vector differential (MMVD) mode.

[0189] At (S1840), the encoded DV offset information can be included in the bitstream and signaled to the decoder.

[0190] In an example, the DV offset information includes a DV offset that is encoded and signaled in the bitstream.

[0191] The process (1800) then proceeds to (S1899) and ends.

[0192] Process 1800 can be appropriately adapted to various scenarios, and steps within process 1800 can be adjusted accordingly. One or more of the steps within process 1800 can be adapted, omitted, repeated, and / or combined. Any suitable order can be used to perform process 1800. Additional steps can be added.

[0193] In an embodiment, an updated displacement vector (DV) of a current block of a current picture is determined based on the DV and MVO (also referred to as DV offset) of the current block. The MVO indicates a motion offset of the DV used to adjust the position of the co-located block of the co-located reference picture. The updated DV indicates the adjusted position of the co-located block of the co-located reference picture. The current block is coded in SbTMVP mode.

[0194] SbTMVP information (e.g., motion information) of each sub-block of the plurality of sub-blocks can be derived based on at least the motion information of the corresponding sub-block of the co-located block indicated by the updated DV. The plurality of sub-blocks can be coded in SbTMVP mode based on the SbTMVP information of the sub-block of the plurality of sub-blocks.

[0195] FIG. 19A is a flowchart outlining a decoding process (1900A) according to an embodiment of the present disclosure. The process (1900A) can be used in a video decoder. The process (1900A) can be performed by an apparatus for video coding, which can include a receiving circuit and a processing circuit. In various embodiments, the process (1900A) is performed by a processing circuit, such as a processing circuit within the terminal devices (310), (320), (330), and (340), a processing circuit that performs the functions of the video encoder (403), a processing circuit that performs the functions of the video decoder (410), a processing circuit that performs the functions of the video decoder (510), or a processing circuit that performs the functions of the video encoder (603). In some embodiments, the process (1900A) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit performs the process (1900A). The process starts at (S1901) and proceeds to (S1910).

[0196] At (S1910), information of a displacement vector (DV) offset (also referred to as an MV offset (MVO)) of a current block of a current picture may be received from a coded video bitstream. The current block includes multiple sub-blocks reconstructed using a sub-block-based temporal motion vector prediction (SbTMVP) mode. The DV offset (or MVO) information may indicate a motion offset relative to the DV used to adjust the position of the co-located block of the co-located reference picture. In an example, the position of the co-located block of the co-located reference picture is adjusted by the DV offset.

[0197] In an example, the DV offset information includes a DV offset (or MVO) signaled in the coded video bitstream.

[0198] In an embodiment, the DV offset information indicates a magnitude of the DV offset and at least one index indicating a direction of the DV offset. In an example, the at least one index includes a distance index indicating the magnitude of the DV offset, which is one of a set of predetermined distances, and a direction index indicating the direction of the DV offset, which is one of a set of predetermined directions.

[0199] In the example, the set of predetermined distances and the set of predetermined directions are used in merged motion vector differential (MMVD) mode.

[0200] In (S1920), the updated DV of the current block can be determined based on the DV of the current block and the DV offset of the current block. The DV offset is indicated by the DV offset information. The updated DV of the current block indicates a block of the co-located reference picture. A block is considered to be co-located with the current block and is called a co-located block of the current block.

[0201] In the example, the updated DV is determined to be the vector sum of the DV and the DV offset.

[0202] In the example, the updated DV is constrained so that the co-located block is within a restricted region of the co-located reference picture, where the restricted region includes the co-located region corresponding to the current CTU of the current picture, and the current CTU includes the current block.

[0203] In (S1930), the motion information of a sub-block of the plurality of sub-blocks may be determined based on the motion information of the corresponding sub-block in the co-located block.

[0204] In (S1940), a sub-block of the plurality of sub-blocks may be reconstructed based on motion information of the sub-block of the plurality of sub-blocks.

[0205] In some examples, an MVP is selected from the MVP candidate list based on the MVP information in the coding information of the current block for AMVP mode, and the MVP is used as the SbTMVP candidate for deriving the final DV.

[0206] In some examples, an MVP candidate list is constructed that includes a sub-block-based merge candidate list. The sub-block-based merge candidate list includes one or more SbTMVP candidates. In an example, the sub-block-based merge candidate list includes multiple spatially neighboring blocks of the current block in a predetermined order. In another example, the sub-block-based merge candidate list includes 0DV for use as an SbTMVP candidate.

[0207] In some examples, one or more spatially neighboring blocks of the current block are checked in a predetermined order for the availability of a central sub-block motion vector. For a spatially neighboring block of the current block, if the central sub-block motion vector of the spatially neighboring block is available, the spatially neighboring block is added as a candidate to the MVP candidate list. If none of the one or more spatially neighboring blocks has an available central sub-block motion vector, 0DV is added to the MVP candidate list.

[0208] In some examples, the coding information indicates affine AMVP mode, and an affine AMVP candidate list is constructed that includes one or more SbTMVP candidates. In examples, the SbTMVP candidate is inserted into the first position of the affine AMVP candidate list.

[0209] In some examples, it is checked whether an affine-coded block exists in the spatially neighboring blocks of the current block. In response to none of the spatially neighboring blocks being affine-coded, the SbTMVP candidate is inserted in the first position of the affine AMVP candidate list. In response to affine-coded blocks existing in the spatially neighboring blocks, the SbTMVP candidate is inserted in the last position of the affine AMVP candidate list.

[0210] In some examples, a precision of MVO information in coding information of a current block of a coded video bitstream is determined, where the MVO information is coded with a precision according to adaptive motion vector resolution (AMVR) in the coded video bitstream.

[0211] In some examples, the offset for deriving the final DV is determined based on the precision of the MVO information. In examples, an index indicating the precision used in the AMVR is decoded from the coded video bitstream. In examples, the precision of the MVO information is in units of M pixels, where M is a positive integer. In examples, the precision of the MVO information is one of 1 pixel, 2 pixels, 4 pixels, and 8 pixels.

[0212] In some examples, the offset for deriving the final DV is scaled based on the size of the sub-block.

[0213] The process (1900A) proceeds to (S1999) and ends.

[0214] The process (1900A) can be appropriately adapted to various scenarios, and the steps of the process (1900A) can be adjusted accordingly. One or more of the steps of the process (1900A) can be adapted, omitted, repeated, and / or combined. Any suitable order can be used to perform the process (1900A). Additional steps can be added.

[0215] FIG. 19B is a flowchart outlining a decoding process (1900B) according to an embodiment of the present disclosure. Process (1900B) is a variation of decoding process (1900A). Process (1900B) can be used in a video decoder. Process (1900B) can be performed by an apparatus for video coding, which can include receiving circuitry and processing circuitry. In various embodiments, process (1900B) is performed by processing circuitry, such as processing circuitry within terminal devices (310), (320), (330), and (340), processing circuitry performing the functions of the video encoder (403), processing circuitry performing the functions of the video decoder (410), processing circuitry performing the functions of the video decoder (510), processing circuitry performing the functions of the video encoder (603), etc. In some embodiments, process (1900B) is implemented by software instructions, and thus, when the processing circuitry executes the software instructions, the processing circuitry performs process (1900B). The process starts at (S1902) and proceeds to (S1912).

[0216] At (S1912), a coded video bitstream including a current picture is received. The current picture includes a current block. The current block includes a plurality of sub-blocks.

[0217] At (S1922), it is determined that the current block including multiple sub-blocks is coded in a sub-block-based temporal motion vector prediction (SbTMVP) mode based on syntax elements of the coded video bitstream.

[0218] At (S1932), motion vector offset (MVO) information of the current block is obtained. The MVO indicates the motion offset of a displacement vector (DV) used to adjust the position of the co-located block in the co-located reference picture.

[0219] At (S1942), an updated DV of the current block is determined based on the DV and MVO of the current block. The updated DV indicates the adjusted position of the co-located block in the co-located reference picture.

[0220] In (S1952), SbTMVP information (eg, motion information) of each sub-block of the plurality of sub-blocks is derived based at least on the motion information of the corresponding sub-block of the co-located block indicated by the updated DV.

[0221] In (S1962), the plurality of sub-blocks are reconstructed in the SbTMVP mode based on the SbTMVP information of the sub-blocks among the plurality of sub-blocks.

[0222] In some examples, an MVP is selected from the MVP candidate list based on the MVP information in the coding information of the current block for AMVP mode, and the MVP is used as the SbTMVP candidate for deriving the final DV.

[0223] In some examples, an MVP candidate list is constructed that includes a sub-block-based merge candidate list. The sub-block-based merge candidate list includes one or more SbTMVP candidates. In an example, the sub-block-based merge candidate list includes multiple spatially neighboring blocks of the current block in a predetermined order. In another example, the sub-block-based merge candidate list includes 0DV for use as an SbTMVP candidate.

[0224] In some examples, one or more spatially neighboring blocks of the current block are checked in a predetermined order for the availability of a central sub-block motion vector. For a spatially neighboring block of the current block, if the central sub-block motion vector of the spatially neighboring block is available, the spatially neighboring block is added as a candidate to the MVP candidate list. If none of the one or more spatially neighboring blocks has an available central sub-block motion vector, 0DV is added to the MVP candidate list.

[0225] In some examples, the coding information indicates affine AMVP mode, and an affine AMVP candidate list is constructed that includes one or more SbTMVP candidates. In examples, the SbTMVP candidate is inserted into the first position of the affine AMVP candidate list.

[0226] In some examples, it is checked whether an affine-coded block exists in the spatially neighboring blocks of the current block. In response to none of the spatially neighboring blocks being affine-coded, the SbTMVP candidate is inserted in the first position of the affine AMVP candidate list. In response to affine-coded blocks existing in the spatially neighboring blocks, the SbTMVP candidate is inserted in the last position of the affine AMVP candidate list.

[0227] In some examples, a precision of MVO information in coding information of a current block of a coded video bitstream is determined, where the MVO information is coded with a precision according to adaptive motion vector resolution (AMVR) in the coded video bitstream.

[0228] In some examples, the offset for deriving the final DV is determined based on the precision of the MVO information. In examples, an index indicating the precision used in the AMVR is decoded from the coded video bitstream. In examples, the precision of the MVO information is in units of M pixels, where M is a positive integer. In examples, the precision of the MVO information is one of 1 pixel, 2 pixels, 4 pixels, and 8 pixels.

[0229] In some examples, the offset for deriving the final DV is scaled based on the size of the sub-block.

[0230] The process (1900B) proceeds to (S1992) and ends.

[0231] Process (1900B) can be appropriately adapted to various scenarios, and the steps of process (1900B) can be adjusted accordingly. One or more of the steps of process (1900B) can be adapted, omitted, repeated, and / or combined. Any suitable order can be used to perform process (1900B). Additional steps can be added.

[0232] The embodiments of the present disclosure may be used separately or combined in any order. Furthermore, each of the methods (or embodiments), encoders, and decoders may be implemented by processing circuitry (e.g., one or more processors, or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium.

[0233] FIG. 20 shows a flowchart outlining an encoding process (2000) according to an embodiment of the present disclosure. The process (2000) may be used in a video encoder. The process (2000) may be performed by an apparatus for video coding, which may include processing circuitry. In various embodiments, the process (2000) is performed by a processing circuit, such as the processing circuitry in the terminal devices (310), (320), (330), and (340), or the processing circuitry that performs the functions of a video encoder (e.g., (403), (603), (703)). In some embodiments, the process (2000) is implemented in software instructions, and thus, the processing circuit performs the process (2000) when the processing circuit executes the software instructions. The process starts at (S2001) and proceeds to (S2010).

[0234] In (S2010), a displacement vector (DV) of a current block of a current picture can be determined, as shown in Figures 16-17. The current block includes multiple sub-blocks coded using a sub-block-based temporal motion vector prediction (SbTMVP) mode. The DV indicates a co-located block in a co-located reference picture that is co-located with the current block.

[0235] In (S2020), the motion information of a sub-block among the plurality of sub-blocks can be determined based on the motion information of the corresponding sub-block of the co-located block, as illustrated in FIGS.

[0236] In (S2030), updated motion information of the sub-block among the plurality of sub-blocks may be determined based on the motion information of the sub-block among the plurality of sub-blocks and a motion vector (MV) offset of the current block.

[0237] In an example, the motion information of a sub-block among the plurality of sub-blocks includes a first motion vector (MV) associated with a first reference picture obtained from a first reference picture list L0. The updated first MV can be determined to be a vector sum of the first MV and an MV offset. The updated motion information includes the updated first MV.

[0238] In an example, the motion information of a sub-block among the plurality of sub-blocks includes a second MV associated with a second reference picture from a second reference picture list L1. The updated second MV can be determined to be one of (i) a vector difference of the second MV and an MV offset, or (ii) a vector sum of the second MV and a scaled MV offset. The scaled MV offset can be based on the MV offset, a picture order count (POC) of the current picture, a POC of the first reference picture, and a POC of the second reference picture. The updated motion information includes the updated second MV.

[0239] In (S2040), MV offset information indicating a MV offset can be coded. A sub-block among the plurality of sub-blocks can be coded based on the updated motion information. The MV offset information can be included in the bitstream.

[0240] In an example, the MV offset information includes an MV offset.

[0241] In an example, the MV offset information indicates at least one index indicating a magnitude of the MV offset and a direction of the MV offset. The at least one index includes a distance index indicating the magnitude of the MV offset, which is one of a set of predetermined distances, and a direction index indicating the direction of the MV offset, which is one of a set of predetermined directions. The set of predetermined distances and the set of predetermined directions are used in merged motion vector differential (MMVD) mode.

[0242] Next, the process (2000) proceeds to (S2099) and ends.

[0243] Process 2000 can be adapted to various scenarios as appropriate, and steps within process 2000 can be adjusted accordingly. One or more of the steps within process 2000 can be adapted, omitted, repeated, and / or combined. Any suitable order can be used to perform process 2000. Additional steps can be added.

[0244] FIG. 21 shows a flowchart outlining a decoding process (2100) according to an embodiment of the present disclosure. The process (2100) can be used in a video decoder. The process (2100) can be performed by an apparatus for video coding, which can include receiving circuitry and processing circuitry. In various embodiments, the process (2100) is performed by a processing circuit, such as a processing circuit within a terminal device (310), (320), (330), or (340), a processing circuit that performs the functions of a video encoder (403), a processing circuit that performs the functions of a video decoder (410), a processing circuit that performs the functions of a video decoder (510), or a processing circuit that performs the functions of a video encoder (603). In some embodiments, the process (2100) is implemented in software instructions, and thus, the processing circuit performs the process (2100) when the processing circuit executes the software instructions. The process begins at (S2101) and proceeds to (S2110).

[0245] At (S2110), motion vector (MV) offset information for a current block in a current picture may be received from a coded video bitstream. The current block includes multiple sub-blocks reconstructed using a sub-block-based temporal motion vector prediction (SbTMVP) mode.

[0246] In (S2120), a displacement vector (DV) of the current block can be determined, as shown in Figures 16-17. The DV can indicate a block of a co-located reference picture that is co-located with the current block. The block can be referred to as a co-located block of the current block.

[0247] In (S2130), the motion information of a sub-block among the plurality of sub-blocks can be determined based on the motion information of the corresponding sub-block of the co-located block, as illustrated in FIGS.

[0248] In (S2140), updated motion information of the sub-block among the plurality of sub-blocks may be determined based on the motion information of the sub-block among the plurality of sub-blocks and the MV offset of the current block indicated by the MV offset information.

[0249] In an example, the MV offset information includes MV offsets signaled in the coded video bitstream.

[0250] In an example, the MV offset information indicates at least one index indicating a magnitude of the MV offset and a direction of the MV offset. The at least one index includes a distance index indicating the magnitude of the MV offset, which is one of a set of predetermined distances, and a direction index indicating the direction of the MV offset, which is one of a set of predetermined directions. The set of predetermined distances and the set of predetermined directions are used in merged motion vector differential (MMVD) mode.

[0251] In an example, the motion information of a sub-block among the plurality of sub-blocks includes a first motion vector (MV) associated with a first reference picture obtained from a first reference picture list L0. The updated first MV can be determined to be a vector sum of the first MV and an MV offset. The updated motion information includes the updated first MV.

[0252] In an example, the motion information of a sub-block among the plurality of sub-blocks includes a second MV associated with a second reference picture from a second reference picture list L1. The updated second MV can be determined to be one of (i) a vector difference of the second MV and an MV offset, or (ii) a vector sum of the second MV and a scaled MV offset. The scaled MV offset can be based on the MV offset, a picture order count (POC) of the current picture, a POC of the first reference picture, and a POC of the second reference picture. The updated motion information includes the updated second MV.

[0253] At (S2150), sub-blocks of the plurality of sub-blocks may be reconstructed based on the updated motion information.

[0254] The process (2100) proceeds to (S2199) and ends.

[0255] The process (2100) can be adapted appropriately for various scenarios, and the steps within the process (2100) can be adjusted accordingly. One or more of the steps of the process (2100) can be adapted, omitted, repeated, and / or combined. Any suitable order can be used to perform the process (2100). Additional steps can be added.

[0256] FIG. 22 shows a flowchart outlining a decoding process (2200) according to an embodiment of the present disclosure. The process (2200) may be used in a video decoder. The process (2200) may be performed by an apparatus for video coding, which may include receiving and processing circuits. In various embodiments, the process (2200) is performed by a processing circuit, such as a processing circuit within a terminal device (310), (320), (330), or (340), a processing circuit that performs the functions of a video encoder (403), a processing circuit that performs the functions of a video decoder (410), a processing circuit that performs the functions of a video decoder (510), or a processing circuit that performs the functions of a video encoder (603). In some embodiments, the process (2200) is implemented by software instructions, and thus, the processing circuit performs the process (2200) when the processing circuit executes the software instructions. Processing begins at (S2201) and proceeds to (S2210).

[0257] At (S2210), coding information for a current block of a current picture is received from an encoded video bitstream, and the coding information indicates an advanced motion vector prediction (AMVP) mode having motion vector predictor (MVP) information and motion vector (MV) offset (MVO) information of the coding information for the current block.

[0258] At (S2220), it is determined that the current block including multiple sub-blocks is coded in sub-block-based temporal motion vector prediction (SbTMVP) mode.

[0259] In (S2230), a final displacement vector (DV) of the current block is determined based on a combination of MVP information and MV offset information of the coding information of the current block, and the final DV indicates multiple corresponding sub-blocks in the co-located picture, which respectively correspond to multiple sub-blocks of the current block.

[0260] In (S2240), the motion information of each of the plurality of sub-blocks of the current block is determined according to the motion information of each of the plurality of corresponding sub-blocks. The motion information of the sub-blocks of the current block is determined according to the motion information of the corresponding sub-block indicated by the final DV.

[0261] In (S2250), a plurality of sub-blocks of the current block are reconstructed respectively based on the motion information of each of the plurality of sub-blocks.

[0262] In some examples, an MVP is selected from the MVP candidate list based on the MVP information in the coding information of the current block for AMVP mode, and the MVP is used as the SbTMVP candidate for deriving the final DV.

[0263] In some examples, an MVP candidate list is constructed that includes a sub-block-based merge candidate list. The sub-block-based merge candidate list includes one or more SbTMVP candidates. In an example, the sub-block-based merge candidate list includes multiple spatially neighboring blocks of the current block in a predetermined order. In another example, the sub-block-based merge candidate list includes 0DV for use as an SbTMVP candidate.

[0264] In some examples, one or more spatially neighboring blocks of the current block are checked in a predetermined order for the availability of a central sub-block motion vector. For a spatially neighboring block of the current block, if the central sub-block motion vector of the spatially neighboring block is available, the spatially neighboring block is added as a candidate to the MVP candidate list. If none of the one or more spatially neighboring blocks has an available central sub-block motion vector, 0DV is added to the MVP candidate list.

[0265] In some examples, the coding information indicates affine AMVP mode, and an affine AMVP candidate list is constructed that includes one or more SbTMVP candidates. In examples, the SbTMVP candidate is inserted into the first position of the affine AMVP candidate list.

[0266] In some examples, it is checked whether an affine-coded block exists in the spatially neighboring blocks of the current block. In response to none of the spatially neighboring blocks being affine-coded, the SbTMVP candidate is inserted in the first position of the affine AMVP candidate list. In response to affine-coded blocks existing in the spatially neighboring blocks, the SbTMVP candidate is inserted in the last position of the affine AMVP candidate list.

[0267] In some examples, a precision of MVO information in coding information of a current block of a coded video bitstream is determined, where the MVO information is coded with a precision according to adaptive motion vector resolution (AMVR) in the coded video bitstream.

[0268] In some examples, the offset for deriving the final DV is determined based on the precision of the MVO information. In examples, an index indicating the precision used in the AMVR is decoded from the coded video bitstream. In examples, the precision of the MVO information is in units of M pixels, where M is a positive integer. In examples, the precision of the MVO information is one of 1 pixel, 2 pixels, 4 pixels, and 8 pixels.

[0269] In some examples, the offset for deriving the final DV is scaled based on the size of the sub-block.

[0270] The process (2200) proceeds to (S2299) and ends.

[0271] Process (2200) can be adapted appropriately for various scenarios, and steps within process (2120) can be adjusted accordingly. One or more of the steps of process (2200) can be adapted, omitted, repeated, and / or combined. Any suitable order can be used to perform process (2200). Additional steps can be added.

[0272] FIG. 23 shows a flowchart outlining a process (2300) according to an embodiment of the present disclosure. The process (2300) may be used in a video encoder. The process (2300) may be performed by an apparatus for video coding, which may include processing circuitry. In various embodiments, the process (2300) is performed by a processing circuit, such as the processing circuitry in the terminal devices (310), (320), (330), and (340), or the processing circuitry that performs the functions of a video encoder (e.g., (403), (603), (703)). In some embodiments, the process (2300) is implemented with software instructions, and thus, the processing circuit performs the process (2300) when the processing circuit executes the software instructions. The process starts at (S2301) and proceeds to (S2310).

[0273] At (S2310), it is determined that a current block including multiple sub-blocks is coded in sub-block-based temporal motion vector prediction (SbTMVP) mode.

[0274] In (S2020), a final DV is determined by combining the MVP information and the MV offset information. The final DV indicates a plurality of corresponding sub-blocks in the co-located picture, which respectively correspond to the plurality of sub-blocks of the current block.

[0275] In (S2030), the motion information of each of the plurality of sub-blocks of the current block is determined according to the motion information of each of the plurality of corresponding sub-blocks. The motion information of the sub-blocks of the current block is determined according to the motion information of the corresponding sub-block indicated by the final DV.

[0276] In (S2340), the plurality of sub-blocks of the current block are respectively reconstructed based on the motion information of each of the plurality of sub-blocks.

[0277] In (S2350), coding information of the current block is generated using an advanced motion vector prediction (AMVP) mode with MVP information and MV offset information of the coding information.

[0278] In some examples, the MVP information indicates an MVP from an MVP candidate list, which is used as an SbTMVP candidate for deriving the final DV.

[0279] In some examples, an MVP candidate list is constructed that includes a sub-block-based merge candidate list. The sub-block-based merge candidate list includes one or more SbTMVP candidates. In an example, the sub-block-based merge candidate list includes multiple spatially neighboring blocks of the current block in a predetermined order. In another example, the sub-block-based merge candidate list includes 0DV for use as an SbTMVP candidate.

[0280] In some examples, one or more spatially neighboring blocks of the current block are checked in a predetermined order for the availability of a central sub-block motion vector. For a spatially neighboring block of the current block, if the central sub-block motion vector of the spatially neighboring block is available, the spatially neighboring block is added as a candidate to the MVP candidate list. If none of the one or more spatially neighboring blocks has an available central sub-block motion vector, 0DV is added to the MVP candidate list.

[0281] In some examples, an affine AMVP candidate list is constructed that includes one or more SbTMVP candidates, and coding information is generated to indicate affine AMVP mode. In examples, the SbTMVP candidate is inserted into the first position of the affine AMVP candidate list.

[0282] In some examples, it is checked whether an affine-coded block exists in the spatially neighboring blocks of the current block. In response to none of the spatially neighboring blocks being affine-coded, the SbTMVP candidate is inserted in the first position of the affine AMVP candidate list. In response to affine-coded blocks existing in the spatially neighboring blocks, the SbTMVP candidate is inserted in the last position of the affine AMVP candidate list.

[0283] In some examples, the MVO information is coded in the coded video bitstream with the precision used in adaptive motion vector resolution (AMVR).

[0284] In an example, an index indicating the precision used in the AMVR is coded into the coded video bitstream. In an example, the precision of the MVO information is in units of M pixels, where M is a positive integer. In an example, the precision of the MVO information is one of 1 pixel, 2 pixels, 4 pixels, and 8 pixels.

[0285] In some examples, the offset for deriving the final DV is scaled based on the size of the sub-block.

[0286] Next, the process (2300) proceeds to (S2399) and ends.

[0287] The process (2300) can be appropriately adapted to various scenarios, and the steps of the process (2300) can be adjusted accordingly. One or more of the steps in the process (2300) can be adapted, omitted, repeated, and / or combined. Any suitable order can be used to implement the process (2300). Additional steps can be added.

[0288] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 24 illustrates a computer system (2400) suitable for implementing certain embodiments of the disclosed subject matter.

[0289] Computer software may be coded using any suitable machine code or computer language that can undergo mechanisms such as assembly, compilation, linking, etc. to produce code containing instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or that can be executed via interpretation, microcode execution, etc.

[0290] The instructions may be executed on various types of computers or computer components including, for example, personal computers, tablet computers, servers, smartphones, gaming consoles, Internet of Things devices, and the like.

[0291] 24 for computer system 2400 are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of the present disclosure. Neither should the arrangement of components be interpreted as having any dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of computer system 2400.

[0292] The computer system (2400) may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, via tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), or olfactory input (not shown). The human interface devices may also be used to capture certain media not necessarily directly associated with conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images, still images obtained from a camera, etc.), and video (two-dimensional video, three-dimensional video, including stereoscopic video, etc.).

[0293] The input human interface devices may include one or more (only one of each is shown) of a keyboard (2401), a mouse (2402), a trackpad (2403), a touchscreen (2410), a data glove (not shown), a joystick (2405), a microphone (2406), a scanner (2407), and a camera (2408).

[0294] The computer system (2400) may also include certain human interface output devices that may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen (2410), data gloves (not shown), or joystick (2405), although haptic feedback devices that do not function as input devices may also be present), audio output devices (e.g., speakers (2409), headphones (not shown)), visual output devices (e.g., screens (2410) including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touchscreen input capability, each with or without haptic feedback capability, some of which may be capable of outputting two-dimensional visual output or three-dimensional or higher-dimensional output via means such as stereo projection output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0295] The computer system (2400) may also include human-accessible storage devices and their associated media, such as optical media (2421) including CD / DVD ROM / RW (2420) along with CD / DVD or similar media, USB memory (2422), external hard drives or external solid state drives (2423), legacy magnetic media such as tape and floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), etc.

[0296] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.

[0297] The computer system (2400) may also include an interface (2454) to one or more communication networks (2455). The network may be, for example, wireless, wired, or optical. The network may further be local, wide-area, metropolitan, vehicular, industrial, real-time, delay-tolerant, or the like. Examples of networks include local area networks such as Ethernet and wireless LAN; mobile communication networks including GSM, 3G, 4G, 5G, LTE, and the like; television wired or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; and vehicle and industrial networks including CAN bus. Certain networks generally require an external network interface adapter attached to a particular general-purpose data port or peripheral bus (2449) (e.g., a USB port on the computer system (2400)), while other networks are generally integrated into the core of the computer system (2400) by attachment to a system bus (e.g., an Ethernet interface to a PC computer system or a mobile communication network interface to a smartphone computer system) as described below. Using any of these networks, the computer system (2400) may communicate with other elements. Such communications may be unidirectional receive only (e.g., broadcast TV), unidirectional transmit only (e.g., a CAN bus to a particular CAN bus device), or bidirectional, for example, to other computer systems using local or wide-area digital networks. Specific protocols and protocol stacks may be used with each of these networks and network interfaces, as described above.

[0298] The aforementioned human interface devices, human accessible storage devices, and network interfaces may be attached to a central portion (2440) of the computer system (2400).

[0299] The core (2440) may include one or more central processing units (CPUs) (2441), graphics processing units (GPUs) (2442), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (2443), task-specific hardware accelerators (2444), graphics adapters (2450), etc. These devices may be connected via a system bus (2448), along with read-only memory (ROM) (2445), random access memory (2446), and internal mass storage devices (2447) such as internal, non-user-accessible hard drives, SSDs, etc. In some computer systems, the system bus (2448) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core system bus (2448) or via a peripheral bus (2449). In an example, a screen (2410) may be connected to the graphics adapter (2450). Architectures for peripheral buses include PCI, USB, and the like.

[0300] The CPU (2441), GPU (2442), FPGA (2443), and accelerator (2444) may execute specific instructions that may combine to constitute the aforementioned computer code. That computer code may be stored in ROM (2445) or RAM (2446). Transient data may also be stored in RAM (2446), while permanent data may be stored, for example, in internal mass storage (2447). Rapid storage and retrieval from any of the memory devices may be enabled through the use of cache memory, which may be closely associated with one or more CPUs (2441), one or more GPUs (2442), one or more mass storage devices (2447), one or more ROMs (2445), one or more RAMs (2446), etc.

[0301] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.

[0302] By way of example and not limitation, the computer system (2400) having the architecture, and in particular its central unit (2440), may provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be user-accessible mass storage devices as described above, as well as media associated with specific storage devices of the central unit (2440) that are non-transitory in nature, such as the central unit's internal mass storage device (2447) or ROM (2445). Software implementing various embodiments of the present disclosure may be stored in such devices and executed by the central unit (2440). The computer-readable media may include one or more memory devices or chips, depending on specific needs. The software may cause the central unit (2440), and in particular the processor (including a CPU, GPU, FPGA, etc.) therein, to perform specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM (2446) and modifying such data structures in response to software-defined operations. Additionally, or alternatively, a computer system may provide functionality as a result of logic embodied in hardwired or otherwise circuitry (e.g., accelerator (2444)), which may operate in place of or in conjunction with software to perform particular operations or portions of particular operations described herein. References to software may encompass logic, and vice versa, where appropriate. References to computer-readable media may encompass circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both, as appropriate. The present disclosure encompasses any appropriate combination of hardware and software.

[0303] Appendix A: Acronyms JEM: Joint exploration model VVC: versatile video coding BMS: Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplementary Enhancement Information VUI: Video Usability Information GOP: Group of Pictures TU: Transform Unit PU: Prediction Unit CTU: Coding Tree Unit CTB: Coding Tree Block PB: Prediction Block HRD: Hypothetical Reference Decoder SNR: Signal-to-Noise Ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: Cathode Ray Tube LCD: Liquid-Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile communications LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Area SSD: Solid-state drive IC: Integrated Circuit CU: Coding Unit

[0304] While this disclosure describes several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure. [Explanation of symbols]

[0305] 101 point, 101 sample, 102 arrow, 103 arrow, 104 square block, 104 block, 201 current block, 300 communication system, 310 terminal device, 320 terminal device, 330 terminal device, 350 communication network, 350 network, 400 communication system, 401 video source, 402 stream, 403 video encoder, 404 video data, 405 streaming server, 406 client subsystem, 407 copy, 407 incoming copy, 410 video decoder, 411 video picture, 412 display, 413 capture subsystem, 420 electronic device, 430 electronic device, 501 channel, 510 video decoder, 512 render device, 515 buffer memory, 520 parser, 521 symbol, 530 electronic device, 531 receiver, 551 inverse transformation unit, 552 Intra prediction unit, 552, intra picture prediction unit, 553, motion compensated prediction unit, 555, aggregator, 556, loop filter unit, 557, reference picture memory, 558, current picture buffer, 601, video source, 603, video encoder, 620, electronic device, 630, source coder, 632, coding engine, 633, local video decoder, 633, decoder, 633, local decoder, 634, reference picture memory, 635, predictor, 640, transmitter, 643, video sequence, 645, entropy coder, 650, controller, 660, communication channel, 703, video encoder, 721, general controller, 722, intra encoder, 723, residual calculator, 724, residual encoder, 725, entropy encoder, 726, switch, 728, residual decoder, 730, inter encoder, 810, video decoder, 871 Entropy decoder, 872 Intra decoder, 873 Residual decoder, 874 Reconstruction module, 880 Inter decoder, 1101 Current picture, 1102 Current reference picture, 1103 Co-located reference picture, 1104 Co-located reference picture, 1112 Co-located CU, 1210 Co-located CU, 1300 Block, 1302 Reference picture, 1311First motion vector (MV), 1311 First MV, 1314 MV pair, 1315 MV pair, 1321 Second motion vector (MV), 1321 Second MV, 1411 Start point, 1414 Paired point, 1415 Paired point, 1421 Start point, 1500 Center point, 1601 Block, 1701 Block, 1711 Picture, 1712 Reference picture, 1731 Sub-block, 1731(1) Sub-block, 1731(2) Sub-block, 1732 Corresponding sub-block, 1800 Encoding process, 1800 Process, 1900A Decoding process, 1900A Process, 1900B Process, 1900B Decoding process, 2000 Encoding process, 2000 Process, 2100 Decoding process, 2100 Process, 2120 Process, 2200 Process, 2200 Decoding process, 2200 Processing, 2300 Process, 2400 Computer system, 2401 Keyboard, 2402 Mouse, 2403 Trackpad, 2405 Joystick, 2406 Microphone, 2407 Scanner, 2408 Camera, 2409 Speaker, 2410 Touchscreen, 2410 Screen, 2410 Screen, 2421 Optical medium, 2422 USB memory, 2423 External solid state drive, 2440 Central part, 2444 Accelerator, 2444 Hardware accelerator, 2446 Random access memory, 2447 Mass storage device, 2448 System bus, 2449 Peripheral bus, 2450 Graphics adapter, 2454 Interface, 2455 Communication network, 2441 Central Processing Unit (CPU), 2443 Field Programmable Gate Array (FPGA), 2442 Graphics Processing Unit (GPU)

Claims

1. receiving a coded video bitstream including a current picture, the current picture including a current block, the current block including a plurality of sub-blocks; determining, based on syntax elements of the coded video bitstream, that the current block including the plurality of sub-blocks is coded in a sub-block-based temporal motion vector prediction (SbTMVP) mode; obtaining motion vector offset (MVO) information of the current block indicating a motion vector offset (MVO), the MVO indicating a motion offset of a displacement vector (DV) used to adjust a position of a co-located block in a co-located reference picture; determining an updated DV of the current block based on the DV and the MVO of the current block, wherein the updated DV indicates the adjusted position of the co-located block in the co-located reference picture; deriving SbTMVP information for each sub-block among the plurality of sub-blocks based on at least motion information of a corresponding sub-block of the co-located block indicated by the updated DV; and reconfiguring the plurality of sub-blocks in the SbTMVP mode based on the SbTMVP information of the sub-blocks among the plurality of sub-blocks; A method of video decoding, comprising:

2. wherein coding information of the current block of the coded video bitstream indicates an advanced motion vector prediction (AMVP) mode with motion vector predictor (MVP) information and motion vector offset (MVO) information of the coding information of the current block, and the method further comprises: selecting an MVP from an MVP candidate list based on the MVP information of the coding information of the current block for the AMVP mode; and deriving the DV using the MVP as a SbTMVP candidate; The method of claim 1 , comprising:

3. 3. The method of claim 2, further comprising: configuring the MVP candidate list to include a sub-block based merge candidate list, the sub-block based merge candidate list including one or more SbTMVP candidates.

4. The method of claim 3 , wherein the sub-block based merge candidate list includes a plurality of spatially neighboring blocks of the current block in a predetermined order.

5. The method of claim 3 , wherein the sub-block-based merge candidate list includes 0DV for use as an SbTMVP candidate.

6. The step of constructing the MVP candidate list comprises: checking one or more spatially neighboring blocks of said current block in a predetermined order for the availability of a central sub-block motion vector; adding, for a spatially neighboring block of the current block, the spatially neighboring block as a candidate to the MVP candidate list according to the availability of a central sub-block motion vector of the spatially neighboring block; adding 0DV to the MVP candidate list in response to none of the one or more spatially neighboring blocks having an available central sub-block motion vector. The method of claim 3 further comprising:

7. the coding information indicates an affine AMVP mode, and the method comprises: constructing an affine AMVP candidate list containing one or more SbTMVP candidates; The method of claim 2 , comprising:

8. The step of constructing the affine AMVP candidate list comprises: inserting an SbTMVP candidate into the first position of the affine AMVP candidate list; The method of claim 7, comprising:

9. The step of constructing the affine AMVP candidate list comprises: checking whether an affine coding block exists in a spatially neighboring block of the current block; inserting an SbTMVP candidate into a first position of the affine AMVP candidate list in response to none of the spatially neighboring blocks being affine coded; and inserting an SbTMVP candidate at the last position of the affine AMVP candidate list in response to the affine coding block being present in the spatially neighboring blocks; The method of claim 7, comprising:

10. The step of obtaining the MVO information of the current block includes: determining precision of the MVO information of coding information of the current block of the coded video bitstream, the MVO information being coded into the coded video bitstream with the precision according to adaptive motion vector resolution (AMVR); and determining the motion offset of the DV based on the accuracy of the MVO information; The method of claim 1 further comprising:

11. The step of obtaining the MVO information of the current block includes: decoding, from the encoded video bitstream, an index indicating the accuracy according to the AMVR; The method of claim 10 further comprising:

12. The method of claim 10 , wherein the precision of the MVO information is in units of M pixels, where M is a positive integer.

13. The method of claim 12 , wherein the precision of the MVO information is one of 1 pixel, 2 pixels, 4 pixels, and 8 pixels.

14. The step of determining the motion offset of the DV comprises: Scaling the motion offset of the DV based on the size of a sub-block. The method of claim 10, comprising:

15. receiving motion vector (MV) offset information for a current block of a current picture from a coded video bitstream, the current block comprising a plurality of sub-blocks reconstructed using a sub-block-based temporal motion vector prediction (SbTMVP) mode; determining a displacement vector (DV) of the current block indicating a co-located block of a co-located reference picture that is co-located with the current block; determining motion information of a sub-block of the plurality of sub-blocks based on motion information of a corresponding sub-block of the co-located block; determining updated motion information of the sub-block among the plurality of sub-blocks based on the motion information of the sub-block among the plurality of sub-blocks and the motion vector offset of the current block indicated by the motion vector offset information; and reconstructing the sub-blocks of the plurality of sub-blocks based on the updated motion information; A method of video decoding, comprising:

16. coding information of the current block of the coded video bitstream indicates an advanced motion vector prediction (AMVP) mode with motion vector predictor (MVP) information and the MV offset information of the coding information of the current block, and the method further comprises: selecting an MVP from an MVP candidate list based on the MVP information of the coding information of the current block for the AMVP mode; and deriving the DV using the MVP as a SbTMVP candidate; 16. The method of claim 15, comprising:

17. 17. The method of claim 16, further comprising: constructing the MVP candidate list to include a sub-block-based merge candidate list, the sub-block-based merge candidate list including at least one of a plurality of spatially neighboring blocks of the current block in a predetermined order and a 0DV.

18. the coding information indicates an affine AMVP mode, and the method comprises: constructing an affine AMVP candidate list containing one or more SbTMVP candidates; 17. The method of claim 16, comprising:

19. The step of receiving the MV offset information of the current block comprises: determining precision of the MV offset information of coding information of the current block of the coded video bitstream, the MV offset information being coded into the coded video bitstream with the precision according to adaptive motion vector resolution (AMVR); and determining the DV based on the accuracy of the MV offset information; 16. The method of claim 15, further comprising:

20. receiving coding information for a current block of a current picture from a coded video bitstream, the coding information indicating an advanced motion vector prediction (AMVP) mode having motion vector predictor (MVP) information and motion vector (MV) offset information of the coding information for the current block; determining that the current block including a plurality of sub-blocks is coded in a sub-block-based temporal motion vector prediction (SbTMVP) mode; determining a final displacement vector (DV) of the current block based on a combination of the MVP information and the MV offset information of the coding information of the current block, wherein the final DV indicates a plurality of corresponding sub-blocks in a collocated picture, which respectively correspond to the plurality of sub-blocks of the current block; determining motion information of each of the plurality of sub-blocks of the current block according to motion information of each of the plurality of corresponding sub-blocks, wherein the motion information of the sub-blocks of the current block is determined according to the motion information of the corresponding sub-block indicated by the final DV; and reconstructing the plurality of sub-blocks of the current block based on the respective motion information of the plurality of sub-blocks; A method of video decoding, comprising: