Video decoding method, apparatus, and program, and video encoding method

SbTMVP addresses inefficiencies in motion vector prediction by deriving sub-block motion vectors from central and neighboring sub-blocks, enhancing coding efficiency and reducing overhead in video encoding/decoding processes.

JP2025534683AActive Publication Date: 2025-10-17TENCENT AMERICA LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025521008
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-08-31
Filing Date
2023-09-01
Publication Date
2025-10-17
Estimated Expiration
2043-09-01

AI Technical Summary

Technical Problem

Existing video coding technologies face inefficiencies in motion vector prediction, particularly at the sub-block level, leading to increased data transmission overhead and reduced coding efficiency.

Method used

The implementation of sub-block-based template matching in temporal motion vector prediction (SbTMVP) mode, which derives motion vectors for sub-blocks within a current template by utilizing a central motion vector and neighboring sub-blocks, reducing the need for explicit transmission of block partition structures and motion information.

Benefits of technology

Enhances coding efficiency and reduces motion vector transmission overhead by allowing each sub-block to inherit motion information from collocated reference pictures, improving prediction accuracy and overall video encoding/decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025534683000001_ABST
    Figure 2025534683000001_ABST
Patent Text Reader

Abstract

A video bitstream is received. The video bitstream includes a current block having a plurality of sub-blocks and a template region of the current block having a plurality of template sub-blocks adjacent to at least one of the upper and left sides of the current block. A motion vector (MV) located at a center position of the current block is determined. The MV is determined based on the MV of at least one of the plurality of sub-blocks of the current block. The MV of each of the plurality of template sub-blocks is determined based on the MV located at the center position of the current block and the MV of each corresponding sub-block among the plurality of sub-blocks adjacent to each template sub-block. The current block is reconstructed based on the determined MVs of the plurality of template sub-blocks.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure describes embodiments generally related to video coding. [Background technology]

[0002] The background description provided herein is intended to generally present the context for the present disclosure. The work of the currently named inventors is not expressly or implicitly admitted as prior art to the present disclosure to the extent that the work is described in the background section, and aspects of the description that may not otherwise qualify as prior art at the time of filing.

[0003] Image / video compression can help transmit image / video data between different devices, storage, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In examples, video codecs can use a technique called intra-prediction, which can compress images based on spatial redundancy. For example, intra-prediction can use reference data from a current picture being reconstructed for sample prediction. In other examples, video codecs can use a technique called inter-prediction, which can compress images based on temporal redundancy. For example, inter-prediction can predict samples in a current picture from a previously reconstructed picture through motion compensation. Motion compensation can be indicated by motion vectors (MVs). Summary of the Invention

[0004] Aspects of the present disclosure include methods and apparatus for video encoding / decoding. In some examples, an apparatus for video decoding includes a processing circuit.

[0005] According to an aspect of the present disclosure, a video decoding method executed in a video decoder is provided. In the method, a video bitstream is received. The video bitstream includes a current block having a plurality of sub-blocks and a template region of the current block having a plurality of template sub-blocks adjacent to at least one of the upper and left sides of the current block. A motion vector (MV) located at a center position of the current block is determined. The MV is determined based on the MV of at least one of the plurality of sub-blocks of the current block. The MV of each of the plurality of template sub-blocks is determined based on the MV located at the center position of the current block and the MV of each corresponding sub-block among the plurality of sub-blocks adjacent to the respective template sub-block. The current block is reconstructed based on the determined MVs of the plurality of template sub-blocks.

[0006] In the example, the MV located at the center position of the current block is determined as the MV from one of the upper left sub-block, the lower left sub-block, the upper right sub-block, and the lower right sub-block of the current block.

[0007] In the example, the MV located at the center position of the current block is determined as the MV from one of the plurality of sub-blocks, and one of the plurality of sub-blocks is selected based on one of the prediction mode and the median sample value of the one of the plurality of sub-blocks.

[0008] In the example, the MV located at the center position of the current block is determined as the average of a subset of the MVs of multiple sub-blocks.

[0009] In one aspect, the MV of each of the plurality of template sub-blocks is determined as a uni-predictive MV based on the MV located at the center position of the current block being a uni-predictive MV. In another aspect, the MV of each of the plurality of template sub-blocks is determined as a bi-predictive MV based on the MV located at the center position of the current block being a bi-predictive MV.

[0010] In the example, based on the fact that (i) the MV of a first subblock among a plurality of subblocks is a unidirectionally predictive MV in a first reference list, (ii) the MV of a first template subblock among a plurality of template subblocks adjacent to the first subblock is a unidirectionally predictive MV in a second reference list, and (iii) the MV located at the center position of the current block is a unidirectionally predictive MV in the second reference list, the MV of the first template subblock is determined to be the MV located at the center position of the current block.

[0011] In the example, based on (i) the MV of a first subblock among the plurality of subblocks is a uni-predictive MV in the first reference list, (ii) the MV of a first template subblock among the plurality of template subblocks adjacent to the first subblock is a bi-predictive MV, and (iii) the MV located at the center position of the current block is a bi-predictive MV including a first component in the first reference list and a second component in the second reference list, it is determined that the MV of the first template subblock includes the MV of the first subblock in the first reference list and the second component of the MV located at the center position of the current block in the second reference list.

[0012] In the example, based on the fact that the MV located at the center position of the current block is a unidirectional predictive MV, the MV of each of the multiple template sub-blocks is determined as the MV of that sub-block among the multiple sub-blocks adjacent to each template sub-block.

[0013] In the example, based on the fact that the MV located at the center position of the current block is a bi-predictive MV, the MV of each of the multiple template sub-blocks is determined as the MV of that sub-block among the multiple sub-blocks adjacent to each template sub-block.

[0014] In the example, based on (i) the MV located at the center position of the current block is a uni-predictive MV in the first reference list, and (ii) the MV of a first sub-block among a plurality of sub-blocks adjacent to the first template sub-block among a plurality of template sub-blocks is a uni-predictive MV in the second reference list, the MV of the first template sub-block is determined as a bi-predictive MV including the MV located at the center position of the current block in the first reference list and the MV of the first sub-block in the second reference list.

[0015] In the example, to reconstruct the current block, a reference block for the current block is determined based on a difference value between a template region of the reference block and a template region of the current block, and the template region of the reference block is indicated by MVs of a plurality of template sub-blocks, each of which is further reconstructed based on each sub-block of the reference block.

[0016] According to another aspect of the present disclosure, an apparatus is provided, the apparatus including a processing circuit, the processing circuit may be configured to perform any of the described methods of video decoding / encoding.

[0017] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method of video decoding / encoding.

[0018] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Brief explanation of the drawings]

[0019] [Figure 1] FIG. 1 is a schematic diagram of an exemplary block diagram of a video processing system (100). [Figure 2] FIG. 2 is a schematic diagram of an exemplary block diagram of a decoder. [Figure 3] FIG. 2 is a schematic diagram of an exemplary block diagram of an encoder. [Figure 4] 1 illustrates exemplary spatially adjacent blocks used in temporal motion vector prediction (TMVP). [Figure 5] Schematic diagram of the sub-block-based TMVP (SbTMVP) process. [Figure 6] 1 shows an example block coded by SbTMVP. [Figure 7] 1 illustrates a first example of a sub-block-based template matching process for SbTMVP, in accordance with some embodiments of the present disclosure. [Figure 8] 10 illustrates a second example of a sub-block-based template matching process for SbTMVP, in accordance with some embodiments of the present disclosure. [Figure 9] 10 illustrates a third example of a sub-block-based template matching process for SbTMVP, in accordance with some embodiments of the present disclosure. [Figure 10] 10 illustrates a fourth example of a sub-block-based template matching process for SbTMVP, in accordance with some embodiments of the present disclosure. [Figure 11] FIG. 10 illustrates a fifth example of a sub-block-based template matching process for SbTMVP, in accordance with some embodiments of the present disclosure. [Figure 12] 1 shows a flowchart illustrating a decoding process according to some embodiments of the present disclosure. [Figure 13] 1 shows a flowchart illustrating an encoding process according to some embodiments of the present disclosure. [Figure 14] FIG. 1 is a schematic diagram of an exemplary computer system according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0020] 1 shows a block diagram of a video processing system 100 in some examples. The video processing system 100 is an example of an application of the disclosed subject matter, namely, a video encoder and video decoder in a streaming environment. The disclosed subject matter can be similarly applicable to other video-enabled applications including, for example, video conferencing, digital TV, streaming services, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0021] The video processing system (100) includes a capture subsystem (113) that may include, for example, a video source (101), such as a digital camera, that generates an uncompressed stream of video pictures (102). By way of example, the stream of video pictures (102) includes samples captured by the digital camera. The stream of video pictures (102) is represented by a bold line to emphasize its high data volume compared to the encoded video data (104) (or coded video bitstream) and may be processed by an electronic device (120) that includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (104) (or coded video bitstream) is represented by a thin line to emphasize its lower data volume compared to the stream of video pictures (102) and may be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as the client subsystems 106 and 108 of FIG. 1, can access the streaming server 105 to retrieve copies 107 and 109 of the encoded video data 104. The client subsystem 106 may include a video decoder 110, for example, in an electronic device 130. The video decoder 110 decodes the incoming copy 107 of the encoded video data and generates an outgoing stream 111 of video pictures that can be rendered on a display 112 (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data 104, 107, and 109 (e.g., a video bitstream) may be encoded according to a particular video coding / compression standard. An example of such a standard is ITU-T Recommendation H.265.By way of example, a video coding standard under development is commonly known as Versatile Video Coding (VVC), and the disclosed subject matter may be used in connection with VVC.

[0022] It should be noted that electronic devices 120 and 130 may include other components (not shown). For example, electronic device 120 may include a video decoder (not shown), and electronic device 130 may similarly include a video encoder (not shown).

[0023] 2 shows an example block diagram of a video decoder (210). The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used in place of the video decoder (110) in the example of FIG. 1.

[0024] The receiver (231) may receive one or more coded video sequences, e.g., contained in a bitstream, to be decoded by the video decoder (210). In an embodiment, one coded video sequence is received at a time, and the decoding of each coded video sequence is independent of the decoding of other coded video sequences. The coded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device storing the coded video data. The receiver (231) may receive the coded video data along with other data, such as coded audio data and / or auxiliary data streams, which may be forwarded to their respective using entities (not shown). The receiver (231) may separate the coded video sequences from other data. To combat network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter "parser (220)"). In certain applications, the buffer memory 215 is part of the video decoder 210. In others, it can be external to the video decoder 210 (not shown). In still other applications, there can be a buffer memory (not shown) external to the video decoder 210, e.g., to combat network jitter, plus another buffer memory 215 within the video decoder 210, e.g., to manipulate playback timing. When the receiver 231 is receiving data from a storage / forwarding device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory 215 may not be required or may be small. For use with best-effort packet networks such as the Internet, the buffer memory 215 may be required, but it can be relatively large and advantageously adaptively sized, and may be implemented at least in part in an operating system or similar element (not shown) external to the video decoder 210.

[0025] The video decoder (210) may include a parser (220) for reconstructing symbols (221) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (210) and, potentially, information for controlling a rendering device, such as a render device (212) (e.g., a display screen) that is not an essential part of the electronic device (230) but may be coupled to the electronic device (230) as shown in FIG. 2. Control information for the rendering device may take the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (220) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, context-dependent or non-context-dependent arithmetic coding, etc. The parser (220) may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroups may include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (220) may also extract information from the coded video sequence, such as transform coefficients, quantization parameter values, motion vectors, etc.

[0026] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to generate symbols (221).

[0027] The reconstruction of the symbols (221) can have many different units depending on the type of coded video picture or portion thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. Which units are included and how may be controlled by subgroup control information parsed by the parser (220) from the coded video sequence. The flow of such subgroup control information between the parser (220) and the following units is not shown for clarity.

[0028] Beyond the functional blocks already described, the video decoder (210) may be conceptually subdivided into a number of functional units, which are described below. In an actual implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.

[0029] The first unit is a scalar / inverse transform unit (251), which receives quantized transform coefficients as symbols (221) from the parser (220) along with control information including which transform to use, block size, quantization coefficients, quantization scaling matrices, etc. The scalar / inverse transform unit (251) can output blocks containing sample values ​​that can be input to an aggregator (255).

[0030] In some cases, the output samples of the scaler / inverse transformer (251) may relate to intra-coded blocks. Intra-coded blocks are blocks that do not use prediction information from a previously reconstructed picture, but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates a block of the same size and shape as the block being reconstructed using surrounding, already reconstructed information fetched from the current picture buffer (258). The current picture buffer (258), for example, buffers a partially reconstructed and / or fully reconstructed current picture. The aggregator (255), in some cases, adds, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transformer unit (251).

[0031] In other cases, the output samples of the scalar / inverse transform unit (251) may relate to an inter-coded, and potentially motion-compensated, block. In such cases, the motion-compensated prediction unit (253) may access the reference picture memory (257) to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols (221) related to the block, the samples may be added by the aggregator (255) to the output of the scalar / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (257) from which the motion-compensated prediction unit (253) fetches prediction samples may be controlled by motion vectors available to the motion-compensated prediction unit (253), for example, in the form of symbols (221), which may have X, Y, and reference picture components. Motion compensation can also include interpolation of sample values ​​fetched from the reference picture memory (257) when sub-sample accurate motion vectors are used, as well as motion vector prediction mechanisms.

[0032] The output samples of the aggregator (255) may undergo various loop filtering techniques in a loop filter unit (256). Video compression techniques may include in-loop filtering techniques controlled by parameters contained in the coded video sequence (also called the coded video bitstream) and made available to the loop filter unit (256) as symbols (221) from the parser (220). Video compression may also respond to meta-information obtained during decoding of previous portions (in decoding order) of the coded picture or coded video sequence, and may also respond to previously constructed loop-filtered sample values.

[0033] The output of the loop filter unit (256) can be a sample stream that can be output to a render device (212) and further stored in a reference picture memory (257) for use in future inter-picture prediction.

[0034] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and any unused current picture buffer can be reallocated before beginning reconstruction of a subsequent coded picture.

[0035] The video decoder (210) may perform decoding operations in accordance with a given video compression technology or standard, such as ITU-T Recommendation H.265. A coded video sequence may conform to the syntax specified by the video compression technology or standard in use, in the sense that the coded video sequence conforms to both the syntax of the video compression technology or standard and a profile documented in the video compression technology or standard. Specifically, a profile may select specific tools from all tools available in the video compression technology or standard as the only tools available for use under that profile. Compliance also requires that the complexity of the coded video sequence be within the boundaries defined by the level of the video compression technology or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further constrained through a Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.

[0036] In embodiments, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may also be included as part of the coded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0037] 3 shows an example block diagram of a video encoder (303). The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmitting circuit). The video encoder (303) can be used in place of the video encoder (303) in the example of FIG. 1.

[0038] The video encoder (303) may receive video samples from a video source (301) (which is not part of the electronic device (320) in the example of FIG. 3) that may capture video images to be coded by the video encoder (303). In other examples, the video source (301) is part of the electronic device (320).

[0039] The video source (301) may provide a source video sequence to be coded by the video encoder (303) in the form of a digital video sample stream, which can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 YCrCB, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (301) may be a storage device storing prepared video. In a video conferencing system, the video source (301) may be a camera capturing local image information as a video sequence. The video data may be provided as multiple individual pictures that, when viewed in sequence, impart motion. The pictures themselves may be organized as a spatial array of pixels, each of which may have one or more samples depending on the sampling structure, color space, etc., in use. This specification focuses hereinafter on samples.

[0040] According to an embodiment, the video encoder (303) may code and compress pictures of a source video sequence into a coded video sequence (343) in real time or under any other time constraints as needed. Imposing an appropriate coding rate is one function of the controller (350). In some embodiments, the controller (350) controls and is operatively coupled to other functional units, as described below. The coupling is not shown for clarity. Parameters set by the controller (350) may include parameters related to rate control (picture skip, quantizer, lambda value for rate-distortion optimization techniques, etc.), picture size, group-of-picture (GOP) layout, maximum motion vector search range, etc. The controller (350) may be configured with other appropriate functions related to the video encoder (303) optimized for a particular system design.

[0041] In some embodiments, the video encoder (303) is configured to operate in a coding loop. As an overly simplified description, in an example, the coding loop can include a source coder (330) (e.g., responsible for generating symbols, such as a symbol stream, based on an input picture to be coded and a reference picture) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to generate sample data in a manner similar to what a (remote) decoder would also generate. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Because decoding of the symbol stream produces bit-exact results independent of the location (local or remote) of the decoder, the contents of the reference picture memory (334) are also bit-perfect between the local and remote encoders. In other words, the predictive portion of the encoder "sees" exactly the same sample values ​​as the decoder would "see" when using prediction during decoding. This basic principle of reference picture synchronicity (and the resulting drift when synchronicity cannot be maintained, for example due to channel errors) is also used in several related techniques.

[0042] The operation of the "local" decoder (333) can be the same as a "remote" decoder, such as the video decoder (210), already described in detail above in conjunction with Figure 2. Referring also momentarily to Figure 2, however, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and parser (220), may not be fully implemented in the local decoder (233), given the availability of symbols and the fact that the encoding / decoding of symbols into a coded video sequence by the entropy coder (345) and parser (220) can be lossless.

[0043] In embodiments, decoder techniques, with the exception of parsing / entropy decoding, present in a decoder are present in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on the operation of the decoder. Descriptions of encoder techniques may be omitted, as they are the inverse of the decoder techniques described generically. To the extent specified, more detailed descriptions are provided below.

[0044] In operation, in some examples, the source coder (330) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as "reference pictures." In this manner, the coding engine (332) codes differences between pixel blocks of the input picture and pixel blocks of the reference pictures that may be selected as predictive references for the input picture.

[0045] The local video decoder (333) may decode coded video data of pictures that may be designated as reference pictures based on symbols generated by the source coder (330). The operation of the coding engine (332) may advantageously be a lossy process. When the coded video data is decoded by a video decoder (not shown in FIG. 3), the reconstructed video sequence is typically a copy of the source video sequence, with some errors. The local video decoder (333) may replicate the decoding process that may be performed by the video decoder on the reference pictures, causing the reconstructed reference pictures to be stored in a reference picture cache (334). In this way, the video encoder (303) may locally store copies of reconstructed reference pictures that have content in common with reconstructed reference pictures that would be obtained by a far-end video decoder (without transmission errors).

[0046] The predictor (335) may perform a prediction search for the coding engine (332). That is, for a new picture to be coded, the predictor (335) may search the reference picture memory (334) for specific metadata, such as reference picture motion vectors, block shapes, or sample data (as candidate reference pixel blocks) that can serve as suitable prediction references for the new picture. The predictor (335) may operate on a sample block-by-pixel block basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (335), the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (334).

[0047] The controller (350) may manage the coding operations of the source coder (330), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0048] The output of all of the above functional units may undergo entropy coding in an entropy coder (345), which converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, or arithmetic coding.

[0049] The transmitter (340) may buffer the coded video sequence produced by the entropy coder (345) to prepare it for transmission over a communication channel (360), which can be a hardware / software link to a storage device that stores the coded video data. The transmitter (340) may also merge the coded video data from the video coder (303) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0050] The controller (350) may manage the operation of the video encoder (303). During coding, the controller (350) may assign a particular coding picture type to each coded picture, which may affect the coding technique that may be applied to each picture. For example, pictures may often be assigned as one of the following picture types:

[0051] Intra pictures (I-pictures) can be coded and decoded without using any other picture in the sequence as a source of prediction. Some video codecs allow various types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures.

[0052] Predictive Pictures (P-pictures) can be coded and decoded by intra-prediction or inter-prediction, which uses motion vectors and reference indices to predict the sample values ​​of each block.

[0053] Bi-directionally Predictive Pictures (B-pictures) can be coded and decoded using intra- or inter-prediction, which uses two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple-predictive pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0054] A source picture is generally spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples, respectively) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to each picture of the blocks. For example, blocks of an I-picture may be non-predictively coded, or they may be predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P-picture may be predictively coded by spatial prediction or temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be predictively coded by spatial prediction or temporal prediction with reference to one or two previously coded reference pictures.

[0055] The video encoder (303) may perform coding operations according to a predetermined video coding technique or standard, such as ITU-T Recommendation H.265. During its operation, the video encoder (303) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancies in the input video sequence. Thus, the coded video data may conform to a syntax defined by the video coding technique or standard being used.

[0056] In embodiments, the transmitter (340) may transmit additional data along with the encoded video. The source coder (330) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0057] Video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or otherwise) between pictures. As an example, a particular picture being encoded / decoded, called the current picture, is partitioned into blocks. If a block in the current picture is similar to a reference block in a previously coded reference picture in the video that is still buffered, the block in the current picture may be coded by a vector called a motion vector. A motion vector points to a reference block within a reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0058] In some embodiments, bi-prediction techniques may be used in inter-picture prediction. According to bi-prediction techniques, two reference pictures are used, e.g., a first reference picture and a second reference picture, both of which precede the current picture in decoding order (but may be past and future, respectively, in display order) in the video. A block in the current picture may be coded with a first motion vector that points to a first reference block in the first reference picture and a second motion vector that points to a second reference block in the second reference picture. The block is predictable by a combination of the first and second reference blocks.

[0059] Furthermore, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.

[0060] According to some embodiments of the present disclosure, prediction, such as inter-picture prediction and intra-picture prediction, is performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and the CTUs within a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU may be recursively quadtree partitioned into one or more coding units (CUs). For example, a 64x64 pixel CTU can be divided into one CU of 64x64 pixels, four CUs of 32x32 pixels, or 16 CUs of 16x16 pixels. By way of example, each CU is analyzed to determine a prediction type for the CU, such as an inter-prediction type or an intra-prediction type. A CU is divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. Generally, each PU includes one luma prediction block (PB) and two chroma PBs. In an embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values ​​(e.g., luma values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.

[0061] It should be noted that the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using any suitable technology. In some embodiments, the video encoders (103) and (203) and the video decoders (110) and (210) may be implemented using one or more integrated circuits. In other embodiments, the video encoders (103) and (203) and the video decoders (110) and (210) may be implemented using one or more processors executing software instructions.

[0062] The present disclosure includes aspects related to motion vector (MV) derivation, such as MV derivation of sub-block templates within a current template based on sub-block-based template matching in sub-block-based temporal motion vector prediction (SbTMVP) mode.

[0063] To improve coding efficiency and reduce motion vector transmission overhead, subblock-level motion vector refinement can be applied to extend CU-level temporal motion vector prediction (TMVP). Subblock-based TMVP (SbTMVP) can enable subblock-level motion information inheritance from collocated reference pictures. Each subblock of a large-sized CU can have its own motion information without explicitly transmitting block partition structure or motion information. In an example, SbTMVP can obtain motion information for each subblock in three steps. In the first step, a displacement vector (DV) of the current CU can be derived. In the second step, the availability of SbTMVP candidates can be checked and the central motion can be derived. In the third step, subblock motion information can be derived from the corresponding subblock via DV. Unlike TMVP candidate derivation, which can derive temporal motion vectors from co-located blocks in a reference frame, SbTMVP can apply DVs, which can be derived from the MVs of the current CU's left neighboring CU, to find corresponding sub-blocks in the co-located picture for each sub-block of the current CU. If the corresponding sub-block is not inter-coded, the motion information of the current sub-block can be set as the central motion.

[0064] SbTMVP can be supported by codecs such as VVC. Similar to TMVP provided in HEVC, SbTMVP can apply motion fields in collocated pictures to improve merge modes and motion vector prediction for CUs in the current picture. Collocated pictures used by TMVP can also be used by SbTMVP. SbTMVP can differ from TMVP in two main aspects: (1) TMVP predicts movements at the CU level, while SbTMVP predicts movements at the sub-CU level; (2) TMVP can fetch temporal motion vectors from adjacent blocks in adjacent pictures (e.g., the adjacent block can be the bottom-right block or the center block relative to the current CU). SbTMVP can apply a motion shift before the temporal motion information is fetched from the adjacent picture, and the motion shift can be obtained from a motion vector from one of the spatially neighboring blocks of the current CU.

[0065] An exemplary SbTMVP process may be illustrated in FIGS. 4 and 5. FIG. 4 illustrates exemplary spatial neighboring blocks (e.g., A0, A1, B0, and B1) of a current block (402) used in TMVP. FIG. 5 illustrates an exemplary SbTMVP process (500). As shown in FIG. 5, a current CU (502) may be included in a current picture (504). The current CU (502) includes multiple sub-CUs (e.g., 506). The current picture (504) may correspond to a collocated picture (508). In an example, SbTMVP may predict motion vectors of sub-CUs (e.g., 506) within the current CU (502) in two steps. In the first step, spatial neighbors, such as A1, of the current CU (502) may be tested. Exemplary candidate spatial neighbors applied to the SbTMVP process (500) may be illustrated in FIG. 4. If a spatial neighbor, such as A1, has a motion vector (510) that uses the collocated picture (508) as a reference picture, the motion vector (510) may be selected as the motion shift (or displacement vector) for the SbTMVP process (500). If no such motion vector is identified, the motion shift may be set to (0,0).

[0066] In the second step, the motion shift (e.g., 510) identified in the first step is applied, e.g., added to the coordinates of the current CU (502), to obtain sub-CU-level motion information (e.g., motion vectors and reference indices) from the collocated picture (508). As shown in FIG. 5, a reference block A1′ in the collocated picture (508) may be identified according to the motion shift derived based on the motion vector (510) of spatial neighbor A1. The reference block A1′ may correspond to a reference CU (512) in the collocated picture (508). Thus, for each sub-CU (e.g., 506) of the current CU (502), motion information of the corresponding block (or corresponding sub-CU (e.g., 514)) in the reference CU (512) of the collocated picture (508) may be used to derive motion information for the sub-CU (e.g., 506). After the motion information of the co-located sub-CU (e.g., 514) is identified, the motion information can be converted into a motion vector and reference index for the current sub-CU (e.g., 506) in a manner similar to the TMVP process in HEVC, where temporal motion scaling can be applied to align the temporal motion vectors of the reference picture (508) with the temporal motion vectors of the current CU (502).

[0067] A composite sub-block-based merge list, containing both SbTMVP candidates and affine merge candidates, can be used in sub-block-based merge mode. SbTMVP can be enabled / disabled by a sequence parameter set (SPS) flag. When SbTMVP mode is enabled, the SbTMVP predictor is added as the first entry in the list of sub-block-based merge candidates, followed by affine merge candidates. The size of the sub-block-based merge list can be signaled in the SPS; for example, in VVC, the maximum allowed size of the sub-block-based merge list can be 5.

[0068] As in VVC, the sub-CU size used in SbTMVP may be fixed at 8x8. Similar to affine merge mode, SbTMVP mode may be applicable to CUs whose width and height are both 8 or greater. Sub-block (or sub-CU) sizes may be explored beyond VVC. For example, in ECM, the sub-CU size may be configurable to other sizes, such as 4x4. Two side-by-side pictures, or frames, may be utilized to provide temporal motion information for SbTMVP and TMVP in AMVP mode.

[0069] To obtain a better (or improved) match, a signaled extra motion vector offset (MVO) can be added to the displacement motion vector (DV). MVO(x o ,y o ), the position of the MV field of the juxtaposed CU can be adjusted. MVO(x o ,y o ) is not a zero motion offset, DV′, which can be the sum of DV and MVO, can be used as a displacement vector to indicate the position of the collocated CU to derive SbTMVP.

[0070] In a related example, DV can be used as a motion vector for template matching for SbTMVP. However, DV for SbTMVP is configured to indicate the position of the motion field in the collocated reference picture. Therefore, using DV in template matching may not be very reliable because DV may not be used as the motion vector of the current CU in SbTMVP and SbTMVP with MMVD.

[0071] In a related example, multiple collocated pictures may be utilized for SbTMVP, but different derivation methods from multiple collocated reference pictures may have different coding performance.

[0072] In a related example, either the central MV or the neighboring sub-block MV from SbTMVP can be used to point to the reference template or the sub-block reference template, but neither the central MV nor the neighboring sub-block MV may be combined to derive the MV for each sub-block template.

[0073] This disclosure provides MV derivation for a sub-block template within a current template. The MV of the sub-block template can be derived based on sub-block-based template matching in SbTMVP mode. In an example, a central MV (e.g., an MV located at the center position of the current block) is used instead of the MV at (0,0) to prevent random initial values ​​that cause mismatches between the encoder and the decoder. The accuracy and prediction of the coding process can be improved.

[0074] In one aspect, when the current CU is coded in SbTMVP mode, the DV is used to indicate the center position of the MV field of the current CU in the collocated reference picture. In the collocated reference picture, MV data can be obtained from the center position of the corresponding MV field in the collocated reference picture, which can be referred to as the "center MV" in SbTMVP. For example, in FIG. 6, the MV data at position (2,2) in the MV field can be referred to as the "center MV" in SbTMVP. In the example, the template is divided into sub-block templates, and the motion vector of each sub-block template is derived not only by using MVs from neighboring sub-blocks in the current coding block, but also by using the center MV described above.

[0075] In an embodiment, when a current CU is coded in SbTMVP mode, DV may be used to indicate the reference position of the current CU to the reference position of a reference CU in a collocated reference picture. The reference position is a central position in this example. In a collocated reference picture, MV data may be obtained from the central position of the corresponding MV field in the collocated reference picture, and the MV data at the central position may be referred to as the "central MV" in SbTMVP. For example, as shown in FIG. 6, the MV data at position (2,2) in the MV field (600) may be referred to as the "central MV" in SbTMVP.

[0076] In this disclosure, a template (or template region) may be divided into multiple sub-block templates (or template sub-blocks), and the motion vectors of each sub-block template may be derived not only by using MVs from neighboring sub-blocks of the current coding block, but also by using a central MV associated with the current coding block.

[0077] In one aspect, the central MV can be derived from any sub-block MV of the SbTMVP, for example, the central MV is an MV from the top-left, bottom-left, top-right, or bottom-right sub-block of the SbTMVP.

[0078] In an embodiment, the central MV associated with the current block coded in SbTMVP mode can be derived from any appropriate sub-block MV of SbTMVP, for example, the central MV can be determined as an MV from one of the top-left sub-block, bottom-left sub-block, top-right sub-block, and bottom-right sub-block of SbTMVP.

[0079] In one aspect, the central MV may be derived from filtering the sub-block MVs of SbTMVP. The filters utilized may be, but are not limited to, a median filter, a mode filter, a weighted average filter, etc.

[0080] In an embodiment, the central MV can be derived by filtering the sub-block MVs of the SbTMVP. The filters used can be, but are not limited to, a median filter, a mode filter, a weighted average filter, etc. In an example, based on the median filter, the central MV can be determined as the MV from the sub-block of the selected current block based on the median sample value of the selected sub-block. In an example, based on the mode filter, the central MV can be determined as the MV from the sub-block of the selected current block based on the prediction mode of the selected sub-block. In an example, based on the weighted average filter, the central MV can be determined as the average (or weighted combination) of a subset of the MVs of multiple sub-blocks.

[0081] In one aspect, the inter prediction direction from L0, L1, or both of the sub-block templates is determined by the inter prediction direction of the central MV of the SbTMVP.

[0082] In an embodiment, the inter-prediction direction can be a first prediction associated with a first reference list (e.g., L0), a second prediction associated with a second reference list (e.g., L1), or a bi-prediction direction associated with both reference lists (e.g., L0 and L1). The inter-prediction direction of a sub-block template can be determined by the inter-prediction direction of a central MV of SbTMVP. For example, if the inter-prediction direction of the central MV is a bi-prediction direction, the inter-prediction direction of the sub-block template is also a bi-prediction direction.

[0083] In one aspect, if the reference index in reference list x of a neighboring sub-block MV of SbTMVP is not valid and the reference index in reference list x of a central MV of SbTMVP is valid, the reference index on the central MV and reference list X can be used as the MV in reference list x for the sub-block template in reference list x. Examples can be shown in Figures 7 and 8.

[0084] In an embodiment, if the reference index in reference list x of an adjacent sub-block MV of SbTMVP is not valid and the reference index in reference list x of a central MV of SbTMVP is valid, the central MV and the reference index in reference list x of the central MV may be used as the MV in reference list x for the sub-block template of reference list x.

[0085] 7, the current block (700) may include multiple sub-blocks, such as sub-block (702) and sub-block (704). The current block (700) may also include template regions above and to the left of the current block (700). The template region may include multiple template sub-blocks (or sub-block templates), such as template sub-blocks (706) and (708). The MV of the first template sub-block (706) may be determined as the central MV (710) based on the following: (i) the MV of a first sub-block (e.g., (704)) among the multiple sub-blocks of the current block (700) is a uni-predictive MV in a first reference list (e.g., L0); (ii) the MV of a first template sub-block (706) among the multiple template sub-blocks adjacent to the first sub-block (e.g., (704)) is a uni-predictive MV in a second reference list (e.g., L1); and (iii) the central MV (710) is a uni-predictive MV in the second reference list.

[0086] 8, a current block (800) may include multiple sub-blocks, such as sub-block (802) and sub-block (804). The current block (800) may include a template region including multiple template sub-blocks, such as template sub-blocks (806), (808), (810), and (812). (i) The MV of a first sub-block (e.g., (804)) among the multiple sub-blocks is a uni-predictive MV in a first reference list (e.g., L0), (ii) the MV of a first template sub-block (e.g., (810)) among the multiple template sub-blocks adjacent to the first sub-block (e.g., (804)) is a bi-predictive MV, and (iii) the central MV (814) is a first component (e.g., MV P-L0 ) and the second component (e.g., MV) in the second reference list (e.g., L1) P-L1 ), the MV of the first template sub-block (e.g., (810)) is based on the MV of the first sub-block (e.g., (804)) in the first reference list (e.g., L0) (e.g., MV C-L0 ) and the second component of the central MV (814) in the second reference list (e.g., L1) (e.g., MV P-L1 ) may be included.

[0087] In one aspect, when the central MV is bi-predictive, the process of Figure 8 can be applied. Otherwise, when the central MV is uni-predictive, the sub-block template region uses only its neighboring sub-block MVs of the SbTMVP block, as shown in Figure 9.

[0088] In an embodiment, when the central MV is bi-predictive, the derivation of the MV for the template sub-block shown in FIG. 8 may be applied. When the central MV is uni-predictive, the derivation of the MV for the template sub-block may use MVs from neighboring sub-blocks of the SbTMVP block. For example, as shown in FIG. 9, the current block (900) may include multiple sub-blocks, such as sub-block (902) and sub-block (904). The current block (900) may include multiple template sub-blocks, such as template sub-blocks (906) and (908). Based on the central MV (910) being uni-predictive, the MV of each of the multiple template sub-blocks may be determined as the MV of the corresponding sub-block among the multiple sub-blocks neighboring each template sub-block. For example, the MV of the template sub-block (906) may be determined as the MV of the sub-block (902), and the MV of the template sub-block (908) may be determined as the MV of the sub-block (904).

[0089] In one aspect, when the central MV is uni-predictive, the process shown in Figure 7 may be applied. Otherwise, when the central MV is bi-predictive, the sub-block template region uses only its neighboring sub-block MVs of the SbTMVP block, as shown in Figure 10.

[0090] In an embodiment, when the central MV is unipredictive, the derivation of the MV for the template sub-block shown in FIG. 7 may be applied. When the central MV is bipredictive, the derivation of the MV for the template sub-block may use MVs from neighboring sub-blocks of the SbTMVP block. For example, as shown in FIG. 10, the current block (1000) may include multiple sub-blocks, such as sub-block (1002) and sub-block (1004). The current block (1000) may include multiple template sub-blocks, such as template sub-blocks (1006) and (1008). Based on the central MV (1010) being a bipredictive MV, the MV of each of the multiple template sub-blocks may be determined as the MV of the corresponding sub-block among the multiple sub-blocks neighboring each template sub-block. For example, the MV of the template sub-block (1006) may be determined as the MV of the sub-block (1002), and the MV of the template sub-block (1008) may be determined as the MV of the sub-block (1004).

[0091] In one aspect, as shown in FIG. 11, when a central MV is uni-predictive with MVs on one reference list (e.g., L0) and adjacent sub-blocks in an SbTMVP block are uni-predictive with MVs on another reference list (e.g., L1), the template sub-block region uses bi-prediction combined with MVs from the central MV and adjacent sub-block MVs.

[0092] In an embodiment, when a central MV is uni-predictive with MVs on one reference list (e.g., L0) and neighboring sub-blocks in an SbTMVP block are uni-predictive with MVs on another reference list (e.g., L1), the template sub-block region can use bi-prediction combined with MVs from the central MV and neighboring sub-block MVs. For example, as shown in FIG. 11, a current block (or SbTMVP block) (1100) can include multiple sub-blocks, such as sub-block (1102) and sub-block (1104). The current block (1100) can include multiple template sub-blocks, such as template sub-blocks (1106) and (1108). (i) The central MV (1110) is a uni-predictive MV in the first reference list (e.g., L0), and (ii) the MV (e.g., MV 1110) of a first sub-block (e.g., MV 1104) of a plurality of sub-blocks that is neighboring a first template sub-block (e.g., MV 1108) of a plurality of template sub-blocks is bi-predictive. C-L1 ) is a uni-predictive MV in the second reference list (e.g., L1), the MV of the first template sub-block (e.g., (1108)) is calculated by dividing the central MV (1110) in the first reference list (e.g., L0) and the MV of the first sub-block (e.g., (1104)) in the second reference list (e.g., L1) (e.g., MV C-L1 Similarly, the MV of the template sub-block (1106) may be determined as a bi-predictive MV including the MV of the sub-block (1102) (e.g., MV E-L1 ) and the central MV (e.g., MV P-L0 ) may be included.

[0093] 12 shows a flowchart illustrating a process (1200) according to an embodiment of the present disclosure. The process (1200) may be used in a video decoder. In various embodiments, the process (1200) is performed by a processing circuit, such as a processing circuit performing the functions of the video decoder (110), a processing circuit performing the functions of the video decoder (210), etc. In some embodiments, the process (1200) is implemented by software instructions, such that the processing circuit performs the process (1200) when the processing circuit executes the software instructions. The process (1200) begins at (S1201) and proceeds to (S1210).

[0094] At (S1210), a video bitstream is received. The video bitstream includes a current block having a plurality of sub-blocks and a template region of the current block having a plurality of template sub-blocks adjacent to at least one of the upper and left sides of the current block. For example, as shown in any of Figures 7 to 1, the current block includes a plurality of sub-blocks and the template region includes a plurality of template sub-blocks.

[0095] At (S1220), a motion vector (MV) located at the center position of the current block is determined. The MV is determined based on the MV of at least one of the sub-blocks of the current block. For example, in FIG. 6, the MV data at position (2,2) in the MV field may be used as the central MV.

[0096] At (S1230), the MV of each of the plurality of template sub-blocks is determined based on the MV located at the center position of the current block and the MV of each of the corresponding sub-blocks among the plurality of sub-blocks adjacent to each template sub-block. For example, as shown in Figures 7 and 8, if the reference index of the reference list x of the adjacent sub-block MV of SbTMVP is invalid and the reference index of the reference list x of the central MV of SbTMVP is valid, the central MV and the reference index of the reference list x of the central MV can be used as the MV of reference list x for the sub-block template in reference list x.

[0097] At (S1240), the current block is reconstructed based on the determined MVs of the plurality of template sub-blocks.

[0098] In the example, the MV located at the center position of the current block is determined as the MV from one of the upper left sub-block, the lower left sub-block, the upper right sub-block, and the lower right sub-block of the current block.

[0099] In the example, the MV located at the center position of the current block is determined as the MV from one of the plurality of sub-blocks, and the one of the plurality of sub-blocks is selected based on one of the prediction mode and the median sample value of the one of the plurality of sub-blocks.

[0100] In the example, the MV located at the center position of the current block is determined as the average of a subset of the MVs of multiple sub-blocks.

[0101] In one aspect, the MV of each of the plurality of template sub-blocks is determined as a uni-predictive MV based on the MV located at the center position of the current block being a uni-predictive MV. In one aspect, the MV of each of the plurality of template sub-blocks is determined as a bi-predictive MV based on the MV located at the center position of the current block being a bi-predictive MV.

[0102] In the example, the MV of the first template subblock is determined as the MV located at the center position of the current block based on the fact that (i) the MV of a first subblock among the plurality of subblocks is a unidirectionally predictive MV in the first reference list, (ii) the MV of a first template subblock among the plurality of template subblocks adjacent to the first subblock is a unidirectionally predictive MV in the second reference list, and (iii) the MV located at the center position of the current block is a unidirectionally predictive MV in the second reference list.

[0103] In the example, based on (i) the MV of a first subblock among a plurality of subblocks is a uni-predictive MV in a first reference list, (ii) the MV of a first template subblock among a plurality of template subblocks adjacent to the first subblock is a bi-predictive MV, and (iii) the MV located at the center position of the current block is a bi-predictive MV including a first component in the first reference list and a second component in the second reference list, the MV of the first template subblock is determined to include the MV of the first subblock in the first reference list and the second component of the MV located at the center position of the current block in the second reference list.

[0104] In the example, based on the fact that the MV located at the center position of the current block is a unidirectional predictive MV, the MV of each of the multiple template sub-blocks is determined as the MV of that sub-block among the multiple sub-blocks adjacent to each template sub-block.

[0105] In the example, based on the fact that the MV located at the center position of the current block is a bi-predictive MV, the MV of each of the multiple template sub-blocks is determined as the MV of that sub-block among the multiple sub-blocks adjacent to each template sub-block.

[0106] In the example, based on (i) the MV located at the center position of the current block is a uni-predictive MV in the first reference list, and (ii) the MV of a first sub-block among a plurality of sub-blocks adjacent to the first template sub-block among a plurality of template sub-blocks is a uni-predictive MV in the second reference list, the MV of the first template sub-block is determined as a bi-predictive MV including the MV located at the center position of the current block in the first reference list and the MV of the first sub-block in the second reference list.

[0107] In the example, to reconstruct the current block, a reference block of the current block is determined based on a difference value between a template region of the reference block and a template region of the current block, and the template region of the reference block is indicated by MVs of a plurality of template sub-blocks, each of which is further reconstructed based on each sub-block of the reference block.

[0108] The process then proceeds to (S1299) and ends.

[0109] Process 1200 may be adapted as appropriate. Steps of process 1200 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0110] 13 shows a flowchart illustrating a process (1300) according to an embodiment of the present disclosure. The process (1300) may be used in a video encoder. In various embodiments, the process (1300) is performed by a processing circuit, such as a processing circuit performing the functions of the video encoder (103), a processing circuit performing the functions of the video encoder (303), etc. In some embodiments, the process (1300) is implemented by software instructions, such that the processing circuit performs the process (1300) when the processing circuit executes the software instructions. The process (1300) begins at (S1301) and proceeds to (S1310).

[0111] In (S1310), an MV located at the center position of the current block is determined. The current block includes multiple sub-blocks, and the MV located at the center position of the current block is determined based on at least one MV of the multiple sub-blocks of the current block. For example, in FIG. 6, the MV data at position (2,2) in the MV field can be used as the central MV of the current block (or the MV located at the center position of the current block).

[0112] At (S1320), the MV of each of the plurality of template sub-blocks is determined based on the MV located at the center position of the current block and the MV of a sub-block among the plurality of sub-blocks adjacent to each template sub-block. The plurality of template sub-blocks are adjacent to at least one of the upper side or the left side of the current block. For example, as shown in Figures 7 and 8, if the reference index of the reference list x of the adjacent sub-block MV of SbTMVP is invalid and the reference index of the reference list x of the central MV of SbTMVP is valid, the central MV and the reference index of the reference list x of the central MV can be used as the MV of the reference list x for the sub-block template in reference list x.

[0113] At (S1330), the samples of the current block are coded based on the determined MVs of the plurality of template sub-blocks.

[0114] The process then proceeds to (S1399) and ends.

[0115] Process 1300 may be adapted as appropriate. Steps of process 1300 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0116] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 14 illustrates a computer system (1400) suitable for implementing certain embodiments of the disclosed subject matter.

[0117] Computer software can be coded in any suitable machine code or computer language that can be subjected to mechanisms such as assembly, compilation, linking, etc. to generate code containing instructions that can be executed by one or more central processing units (CPUs), graphics processing units (GPUs), etc. directly or through interpretation, microcode execution, etc.

[0118] The instructions may be executable by various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming consoles, Internet of Things devices, and the like.

[0119] 14 for computer system 1400 are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of the present disclosure. The arrangement of components should not be interpreted as having any dependency or requirement regarding any one or combination of components described in the exemplary embodiment of computer system 1400.

[0120] The computer system 1400 may include certain human interface input devices. Such human interface input devices may respond to input by one or more users through, for example, tactile input (e.g., keyboard, swipe, dataglove motion), audio input (e.g., voice, claps), visual input (e.g., gestures), or olfactory input (not shown). The human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still camera), and video (e.g., two-dimensional video, three-dimensional video, including stereoscopic video).

[0121] The input human interface devices may include one or more of a keyboard (1401), a mouse (1402), a trackpad (1403), a touchscreen (1410), a data glove (not shown), a joystick (1405), a microphone (1406), a scanner (1407), and a camera (1408) (only one of each is shown).

[0122] The computer system 1400 may also include certain human interface output devices that may stimulate one or more of the user's senses through, for example, tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen (1410), data gloves (not shown), or joystick (1405), although haptic feedback devices that do not function as input devices may also exist), audio output devices (e.g., speakers (1409), headphones (not shown)), visual output devices (e.g., screens (1410) including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input capability, each with or without haptic feedback capability, some of which are capable of outputting two-dimensional visual output or output in more than three dimensions by means of stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0123] The computer system (1400) may also include human-accessible storage devices and their associated media, such as CD / DVD or similar media (1421), CD / DVD ROM / RW (1420), including thumb drives (1422), removable hard disks or solid state drives (1423), legacy magnetic media, such as tape and floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices, such as security dongles (not shown), and the like.

[0124] Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transitory signals.

[0125] The computer system 1400 may also include interfaces 1454 to one or more communications networks 1455. Networks may be, for example, wireless, wireline, or optical. Networks may also be local, wide-area, metropolitan, vehicular, and industrial, real-time, delay-tolerant, and the like. Examples of networks include local area networks such as Ethernet; wireless LANs; cellular networks including GSM, 3G, 4G, 5G, LTE, and the like; TV wireline or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; and vehicle and factory networks including CAN bus. Certain networks generally require an external network interface adapter attached to a particular general-purpose digital port or peripheral bus 1449 (e.g., a USB port on the computer system 1400). Others are generally integrated into the core of the computer system 1400 by attachment to a system bus as described below (e.g., an Ethernet network interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (1400) can communicate with other entities. Such communication can be one-way receive-only (e.g., broadcast TV) or one-way transmit-only (e.g., a CAN bus to a specific CAN bus device), or it can be two-way to other computer systems, for example, using a local or wide-area digital network. Specific protocols or protocol stacks can be used with each of the networks and network interfaces described above.

[0126] The above-mentioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the core 1440 of the computer system 1400 .

[0127] The core (1440) may include one or more central processing units (CPUs) (1441), graphics processing units (GPUs) (1442), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (1443), task-specific hardware accelerators (1444), graphics adapters (1450), etc. These devices may be connected through a system bus (1448), along with read-only memory (ROM) (1445), random access memory (RAM) (1446), internal mass storage devices such as internal non-user-accessible hard drives, SSDs, etc. (1447). In some computer systems, the system bus (1448) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached to the core's system bus (1448) directly or through a peripheral bus (1449). In an example, a display 1410 may be connected to a graphics adapter 1450. Architectures for peripheral buses include PCI, USB, and the like.

[0128] The CPU (1441), GPU (1442), FPGA (1443), and accelerator (1444) can execute specific instructions that, in combination, can constitute the above-mentioned computer code. The computer code can be stored in ROM (1445) or RAM (1446). Temporary data can also be stored in RAM (1446), while persistent data can be stored, for example, in an internal mass storage device (1447). Rapid storage and retrieval from any of the memory devices can be enabled through the use of cache memory. Cache memory can be closely associated with one or more of the CPU (1441), GPU (1442), mass storage device (1447), ROM (1445), RAM (1446), etc.

[0129] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.

[0130] By way of example, and not limitation, a computer system having the architecture (1400), and in particular the core (1440), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with the user-accessible mass storage devices previously introduced, in addition to specific storage of the core (1440) that is non-transitory in nature, such as the core's internal mass storage device (1447) or ROM (1445). Software implementing various embodiments of the present disclosure can be stored on such devices and executable by the core (1440). The computer-readable media can include one or more memory devices or chips, depending on particular needs. Software can cause the cores (1440), and specifically the processors therein (including CPUs, GPUs, FPGAs, etc.), to perform particular processes or portions of particular processes described herein, including defining data structures stored in RAM (1446) and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system can provide functionality as a result of hardwired or otherwise embodied logic in circuitry (e.g., accelerators (1444)) that can operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. References to software can encompass logic, where appropriate, and vice versa. References to computer-readable media can encompass circuitry (e.g., integrated circuits (ICs)) storing software for execution, circuitry embodying logic for execution, or both, where appropriate. The present disclosure encompasses any appropriate combination of hardware and software.

[0131] The use of "at least one of" or "one of" in this disclosure is intended to include any one or combination of the listed elements. For example, reference to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C is intended to include A only, B only, C only, or any combination thereof. Reference to one of A or B, and one of A and B is intended to include A or B or (A and B). The use of "one of" does not exclude any combination of the listed elements, where applicable, such as when the elements are not mutually exclusive.

[0132] While this disclosure has described several exemplary embodiments, alternatives, permutations, and various substitute equivalents exist and are included within the scope of this disclosure. Thus, it will be apparent to those skilled in the art that numerous systems and methods, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within its spirit and scope.

[0133] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 416,461, entitled "Motion Vector Derivation of Subblock-Based Template-Matching for Subblock Based Motion Vector Predictor," filed on October 14, 2022, which in turn claims the benefit of priority to U.S. Provisional Patent Application No. 18 / 241,084, entitled "MOTION VECTOR DERIVATION OF SUBBLOCK-BASED TEMPLATE-MATCHING FOR SUBBLOCK BASED MOTION VECTOR PREDICTOR," filed on August 31, 2023. The disclosures of the prior applications are incorporated herein by reference in their entirety.

Claims

1. 1. A method of video decoding performed by a video decoder, comprising: receiving a video bitstream including a current block having a plurality of sub-blocks and a template region of the current block having a plurality of template sub-blocks adjacent to at least one of an upper side and a left side of the current block; determining a motion vector (MV) located at a center position of the current block, the MV being determined based on the MV of at least one of the plurality of sub-blocks of the current block; determining a motion vector (MV) of each of the plurality of template sub-blocks based on the motion vector located at the center position of the current block and the motion vectors of corresponding sub-blocks among the plurality of sub-blocks adjacent to each of the template sub-blocks; reconstructing the current block based on the determined MVs of the plurality of template sub-blocks; A method having the following.

2. The step of determining the MV located at the center position of the current block includes: determining the motion vector located at the center position of the current block as a motion vector from one of an upper left sub-block, a lower left sub-block, an upper right sub-block, and a lower right sub-block of the current block; The method of claim 1.

3. The step of determining the MV located at the center position of the current block includes: determining the motion vector located at the center position of the current block as a motion vector from one of the plurality of sub-blocks; the one of the plurality of sub-blocks is selected based on one of a prediction mode and a median sample value of the one of the plurality of sub-blocks. The method of claim 1.

4. The step of determining the MV located at the center position of the current block includes: determining the motion vector located at the center position of the current block as an average of a subset of motion vectors of the plurality of sub-blocks; The method of claim 1.

5. The step of determining MVs for each of the plurality of template sub-blocks comprises: determining that the MV located at the center position of the current block is a uni-predictive MV and that the MVs of each of the plurality of template sub-blocks are uni-predictive MVs; determining that the MV located at the center position of the current block is a bi-predictive MV, and Further comprising: The method of claim 1.

6. The step of determining MVs for each of the plurality of template sub-blocks comprises: (i) a motion vector of a first sub-block among the plurality of sub-blocks is a unidirectionally predicted motion vector in a first reference list, (ii) a motion vector of a first template sub-block among the plurality of template sub-blocks adjacent to the first sub-block is a unidirectionally predicted motion vector in a second reference list, and (iii) the motion vector located at the center position of the current block is a unidirectionally predicted motion vector in the second reference list, determining a motion vector of the first template sub-block as the motion vector located at the center position of the current block; The method of claim 1.

7. The step of determining MVs for each of the plurality of template sub-blocks comprises: (i) a motion vector of a first sub-block among the plurality of sub-blocks is a uni-predictive motion vector in a first reference list; (ii) a motion vector of a first template sub-block among the plurality of template sub-blocks adjacent to the first sub-block is a bi-predictive motion vector; and (iii) the motion vector located at the center position of the current block is a bi-predictive motion vector including a first component in the first reference list and a second component in a second reference list. determining that the motion vector of the first template sub-block includes the motion vector of the first sub-block in the first reference list and the second component of the motion vector located at the center position of the current block in the second reference list; The method of claim 1.

8. The step of determining MVs for each of the plurality of template sub-blocks comprises: determining, based on the fact that the motion vector located at the center position of the current block is a one-way prediction motion vector, a motion vector of each of the plurality of template sub-blocks as a motion vector of a corresponding sub-block among the plurality of sub-blocks adjacent to the respective template sub-block; The method of claim 1.

9. The step of determining MVs for each of the plurality of template sub-blocks comprises: determining, based on the fact that the motion vector located at the center position of the current block is a bi-predictive motion vector, a motion vector of each of the plurality of template sub-blocks as a motion vector of a corresponding sub-block among the plurality of sub-blocks adjacent to the respective template sub-block; The method of claim 1.

10. The step of determining MVs for each of the plurality of template sub-blocks comprises: (i) the motion vector located at the center position of the current block is a unidirectionally predicted motion vector in a first reference list, and (ii) the motion vector of a first sub-block of the plurality of sub-blocks adjacent to a first template sub-block of the plurality of template sub-blocks is a unidirectionally predicted motion vector in a second reference list, determining a motion vector of the first template sub-block, the motion vector being a bi-predictive motion vector including the motion vector located at the center position of the current block in the first reference list and a motion vector of the first sub-block in the second reference list; The method of claim 1.

11. The step of reconstructing the current block includes: determining a reference block of the current block based on a difference value between a template region of the reference block and a template region of the current block, the template region of the reference block being indicated by MVs of the plurality of template sub-blocks; reconstructing each of the plurality of sub-blocks based on each sub-block of the reference block; Further comprising: The method of claim 1.

12. 12. Apparatus comprising processing circuitry configured to perform the method of any one of claims 1 to 11 when executing instructions stored in a memory.

13. A program which, when executed by a processing circuit, causes the processing circuit to carry out a method according to any one of claims 1 to 11.

14. 1. A method of video encoding performed by a video encoder, comprising: determining a motion vector (MV) located at a center position of a current block, the MV being determined based on the MV of at least one of a plurality of sub-blocks included in the current block; determining motion vectors of each of a plurality of template sub-blocks adjacent to at least one of the upper and left sides of the current block based on the motion vector located at the center position of the current block and motion vectors of corresponding sub-blocks among the plurality of sub-blocks adjacent to each of the template sub-blocks; encoding the current block based on the determined MVs of the plurality of template sub-blocks; A method having the following.

Citation Information

Patent Citations

  • Encoder, decoder, encoding method, and decoding method

    US20190335181A1

  • Image decoding method based on inter prediction and image decoding apparatus therefor

    US20200154124A1

  • Video decoding method and apparatus, and video encoding method and apparatus for performing inter prediction according to affine model

    US20220248028A1

  • Template matching based affine prediction for video coding

    US20220329823A1