Dual - Directional Optical Flow and Decoder - Side Motion Vector Refinement in Sub - Block - Based Temporal Motion Vector Refinement (SbTMVP)

SbTMVP with BM-based and BDOF modes addresses inefficiencies in video coding by refining motion vectors for sub-blocks, resulting in improved compression and reduced bandwidth/storage requirements.

JP2025523728APending Publication Date: 2025-07-25TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024517396
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-10
Filing Date
2022-11-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in achieving efficient compression and reconstruction of video data, particularly in intra and inter-picture prediction, due to limitations in motion vector refinement and intra prediction modes, leading to suboptimal bandwidth and storage requirements.

Method used

The implementation of sub-block-based temporal motion vector prediction (SbTMVP) with bilateral matching (BM)-based motion vector refinement and bi-directional optical flow (BDOF) modes to refine motion information for sub-blocks, enhancing the decoding process.

Benefits of technology

Improves video coding efficiency by reducing redundancy and enhancing compression ratios, thereby optimizing bandwidth and storage needs while maintaining video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025523728000001_ABST
    Figure 2025523728000001_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide an apparatus and method including a processing circuit, where a current block obtains prediction information indicating whether the current block is coded in sub-block-based temporal motion vector prediction (SbTMVP) mode. When the current block is coded in SbTMVP mode, it is determined whether sub-blocks within a plurality of sub-blocks of the current block are bi-predicted. In response to a sub-block being bi-predicted, motion information for the sub-block is determined based on the SbTMVP mode. At least one of (i) bilateral matching (BM)-based motion vector (MV) refinement and (ii) bi-directional optical flow (BDOF) mode is applied to the sub-block to refine the motion information for the sub-block. Based on the refined motion information corresponding to one or more sub-blocks within the plurality of sub-blocks, the current block is reconstructed. The refined motion information corresponding to one or more sub-blocks includes the refined motion information for the sub-blocks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001]

[0001] Incorporation by Reference This application claims priority to U.S. Patent Application No. 17 / 984,948, filed November 10, 2022, entitled "Dual-Directional Optical Flow and Decoder-Side Motion Vector Refinement in Sub-Block-Based Temporal Motion Vector Precision (SbTMVP)", which claims priority to U.S. Provisional Application No. 63 / 389,657, filed July 15, 2022, entitled "Dual-Directional Optical Flow and Decoder-Side Motion Vector Refinement in SbTMVP". The disclosure of the prior application is hereby incorporated by reference in its entirety.

[0002]

[0002] Technical Field The present disclosure generally describes embodiments related to video coding.

Background Art

[0003]

[0003] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. The work done under the present inventor's name is not admitted as prior art to the present disclosure, either expressly or implicitly, to the extent that the work is described in a manner that does not otherwise render it eligible as prior art at the time of filing, other than in this background section.

[0004]

[0004] Uncompressed digital images and / or videos can include a series of pictures, each picture having spatial dimensions of, for example, 1920×1080 luminance samples and associated chrominance samples. A series of pictures can have a fixed or variable picture rate (informally known as the frame rate), for example, 60 pictures per second, or 60 Hz. Uncompressed images and / or videos have fairly high bitrate requirements. For example, 1080p60 4:2:0 video with 8 bits per sample (luminance sample resolution of 1920×1080 at a frame rate of 60 Hz) requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires a storage space of more than 600 GB.

[0005]

[0005] One of the purposes of image and / or video encoding and decoding can be said to be the reduction of redundancy in the input image and / or video signal by compression. Compression can, in some cases, help reduce the aforementioned bandwidth and / or storage space requirements by a factor of two or more. The description in this case uses video encoding / decoding as an exemplary example, but the same techniques can be applied to image encoding / decoding in a similar manner without departing from the spirit of the present disclosure. Both lossless compression and lossy compression, as well as combinations thereof, can be used. Lossless compression (reversible compression) refers to a technique in which an exact copy of the original signal can be reconstructed from the compressed original signal. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough that the reconstructed signal is useful for the intended application. In the case of video, lossy compression is widely used. The amount of acceptable distortion depends on the application; for example, users of a particular consumer streaming application may be able to tolerate higher distortion than users of a television distribution application. The achievable compression ratio can reflect the fact that higher acceptable / tolerable distortion can result in a higher compression ratio.

[0006]

[0006] Video encoders and decoders can utilize techniques in several broad categories, including, for example, motion compensation, transform processing, quantization, and entropy coding.

[0007]

[0007] Video codec technology can include techniques known as intra coding. In intra coding, sample values are represented without reference to samples from previously reconstructed reference pictures or other data. In some video codecs, a picture is spatially divided into blocks of samples. If all blocks of samples are coded in the intra mode, that picture can be an intra picture. Intra pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and thus can be used as the first picture in a coded video bitstream and video session or as a still image. Samples of an intra block can be entrusted to a transform, and the transform coefficients can be quantized prior to entropy coding. Intra prediction can be a technique for minimizing sample values in the pre-transform domain. In some cases, the smaller the post-transform DC value and the smaller the AC coefficients, the fewer bits are required to represent the block after entropy coding for a given quantization step size.

[0008]

[0008] For example, conventional intra-coding used in MPEG-2 generation coding technology does not use intra prediction. However, some new video compression technologies include techniques that attempt to perform prediction based on surrounding sample data and / or metadata obtained during encoding and / or decoding of data blocks. Such techniques are hereinafter referred to as "intra prediction" techniques. It should be noted that in at least some cases, intra prediction uses only reference data from the current picture being reconstructed and does not use reference data from reference pictures.

[0009]

[0009] There can be many different forms of intra prediction. If more than one such technique may be used in a given video coding technology, the particular technique used may be coded as a particular intra prediction mode that uses that particular technique. In certain cases, the intra prediction mode may have submodes and / or parameters, in which case the submodes and / or parameters may be individually coded or included in a mode codeword, which defines the prediction mode being used. Which codeword to use for a given combination of mode, submode, and / or parameter can affect the coding efficiency gain due to intra prediction and thus can also affect the entropy coding technology used to convert the codeword into the bitstream.

[0010]

[0010] The specific mode of intra prediction was introduced in H.264, improved in H.265, and further refined in new coding technologies such as JEM (joint exploration model), VVC (versatile video coding), and MBS (benchmark set). The prediction block can be formed using the sample values of samples in the vicinity of already available samples. The sample values of the neighboring samples are copied to the predictor block according to a certain direction. The reference for the direction in use can be coded in the bitstream or can be predicted itself.

[0011]

[0011] Referring to FIG. 1A, what is depicted at the lower right is a subset of 9 predictor directions out of the 33 possible predictor directions defined in H.265 (corresponding to 33 of the 35 intra - modes, the angular modes). The point (101) where the arrows converge represents the sample to be predicted. The arrows represent the direction from which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from the sample that goes diagonally up to the right at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from the sample that goes diagonally down to the left of sample (101) at an angle of 22.5 degrees from the horizontal.

[0012]

[0012] Referring further to FIG. 1A, in the top left, a square block (104) of 4×4 samples is depicted (shown by the thick dashed line). The square block (104) contains 16 samples, each labeled with an “S”, its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample from the top in the Y dimension and the first sample from the left in the X dimension. Similarly, sample S44 is the fourth sample within block (104) in both the Y and X dimensions. Since the block size is 4×4 samples, S44 is in the lower right. Further, reference samples following a similar numbering scheme are shown. The reference samples are labeled with an X position (column index) relative to block (104), a Y position (e.g., row index), and an R. In both H.264 and H.265, since the predicted samples are in the neighborhood of the block being reconstructed, there is no need to use negative values.

[0013]

[0013] Intra-picture prediction can be performed by copying the reference sample value from neighboring samples indicated by the signaled prediction direction. For example, assume that the coded video bitstream includes signaling indicating a prediction direction that coincides with arrow (102) for this block, i.e., assume that the samples are predicted from samples that go diagonally up and to the right at a 45-degree angle from horizontal. In this case, samples S41, S32, S23, S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.

[0014]

[0014] In certain cases, the values of multiple reference samples can be combined, particularly when the direction is not evenly divisible by 45 degrees, for example, by interpolation, to calculate the reference sample.

[0015]

[0015] As video coding technology develops, the number of possible directions is increasing. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013). Currently, JEM / VVC / BMS can support up to 65 directions. Experiments have been conducted to identify the most likely directions, using specific techniques in entropy coding to represent these likely directions with fewer bits while accepting a specific penalty for the less likely directions. Furthermore, the directions themselves can sometimes be predicted from neighboring directions that are already decoded in neighboring blocks.

[0016]

[0016] FIG. 1B shows a schematic diagram (110) depicting 65 intra prediction directions according to JEM, showing the increasing number of prediction directions over time.

[0017]

[0017] The mapping of intra prediction direction bits representing directions within a coded video bitstream can vary for each video coding technology. Such mappings can range from a simple direct mapping to complex adaptive schemes including codewords, the most likely modes, and similar techniques. However, in most cases, there may be specific directions (specific directions that occur in the video content with a statistically lower probability than other specific directions). Since the goal of video compression is redundancy reduction, these less likely directions are represented with more bits than the more likely directions in well - operating video coding technologies.

[0018]

[0018] The encoding and decoding of images and / or videos may be performed using inter-picture prediction with motion compensation. Motion compensation may be a lossless compression technique, and a block of sample data from a previously reconstructed picture or a part thereof (reference picture) may be spatially shifted in the direction indicated by a motion vector (hereinafter referred to as MV) and then used for the prediction of a newly reconstructed picture or picture part. In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV can have two dimensions X and Y, or three dimensions, the third being an index of the reference picture used (the latter may indirectly be the time dimension).

[0019]

[0019] In some video compression techniques, the MVs applicable to a given area of sample data can be predicted from other MVs, for example, from MVs related to another area of sample data that is spatially adjacent to the area being reconstructed and that precedes that MV in decoding order. By doing so, the amount of data required to code the MVs can be significantly reduced, thereby eliminating redundancy and increasing the compression ratio. For example, when coding an input video signal derived from a camera (known as natural video), there is a statistical likelihood that areas larger than the area to which a single MV is applicable move in a similar direction, and thus in some cases, it is possible to predict using similar motion vectors derived from the MVs of adjacent areas, so MV prediction can function effectively. This results in MVs that are found to be similar or identical to the MVs predicted from surrounding MVs for a given area, which can be represented with fewer bits than those used when directly coding the MVs after entropy coding. In some cases, MV prediction may be an example of lossless compression of the signal (i.e., the MV) derived from the original signal (i.e., the sample stream). In other cases, MV prediction itself may be lossy due to rounding errors, for example, when calculating predictors from several surrounding MVs.

[0020]

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec.H.265, “High Efficiency Video Coding”, December 2016). Among the many MV prediction mechanisms provided by H.265, the one described with reference to FIG. 2 is a technique hereafter called “spatial merge”.

[0021] Referring to FIG. 2, the current block (201) includes samples discovered by the encoder during the motion search process such that it is predictable from previous blocks of the same size that are spatially shifted. Instead of directly coding the MV, the MV can be derived from the metadata associated with one or more reference pictures, for example, using the MV associated with any of five surrounding samples (202 to 206 respectively) shown as A0, A1, and B0, B1, B2, from the latest reference picture (in the order of decoding). In H.265, MV prediction can use predictors from the same reference picture that neighboring blocks are using.

SUMMARY OF THE INVENTION

[0022]

[0022] Aspects of the present disclosure provide methods and apparatuses for video encoding / decoding. In some examples, an apparatus for video / image decoding includes a processing circuit that decodes prediction information of a current block in a current picture (currently being processed) from a coded bitstream. The current block includes a plurality of sub-blocks that are reconstructed based on a subblock-based temporal motion vector prediction (SbTMVP) mode. The processing circuit determines motion information of sub-blocks within the plurality of sub-blocks based on the SbTMVP mode. The sub-blocks are bi-predicted. The processing circuit can apply at least one of (i) bilateral matching (BM)-based motion vector (MV) refinement and (ii) bi-directional optical flow (BDOF) mode to the sub-blocks to update the motion information of the sub-blocks and reconstruct the sub-blocks based on the updated motion information.

[0023]

[0023] The processing circuit can determine refined motion information of the sub-block by applying BM-based MV refinement. BM-based MV refinement includes decoder-side motion vector refinement (DMVR) or multi-pass decoder-side motion vector refinement (MP-DMVR). The processing circuit can reconstruct the sub-block based on the refined motion information.

[0024]

[0024] In an embodiment, the motion information includes an initial MV pair of the sub-block, and BM-based MV refinement includes DMVR. The processing circuit can apply DMVR to an area within the sub-block to determine a refined MV pair of the area based on the initial MV pair. The area may be smaller than or equal to the area of the sub-block. The processing circuit can reconstruct the area within the sub-block based on the refined MV pair.

[0025]

[0025] In an embodiment, the motion information includes an initial MV pair of the sub-block, and BM-based MV refinement includes MP-DMVR. When the sub-block size of the sub-block is larger than a first threshold, the processing circuit can apply at least one DMVR to the sub-block to determine a first refined MV pair of the sub-block; and apply the BDOF mode to an area within the sub-block to determine a second refined MV pair of the area based on the first refined MV pair. The area may be smaller than or equal to the area of the first sub-block.

[0026]

[0026] In one example, the motion information includes the initial MV pair of the sub-block. The processing circuit can apply the BDOF mode to each sample of the sub-block to determine the refined MV pair for each sample. The BDOF mode may include a sample-based BDOF mode. The processing circuit reconstructs each sample of the sub-block based on the refined MV pair for each sample.

[0027]

[0027] In one example, the refined motion information of the sub-block includes one or more first refined MV pairs for each of one or more areas within the sub-block. After applying BM-based MV refinement, the processing circuit applies a BDOF mode including a sample-based BDOF mode. The refined MV pair for each sample in the area within the one or more areas can be determined based on the first refined MV pair corresponding to the area and the BDOF mode. The processing circuit reconstructs each sample in the area based on the refined MV pair for each sample.

[0028]

[0028] In one example, the prediction information indicates that at least one of (i) BM-based MV refinement or (ii) the BDOF mode is applied to the sub-block.

[0029]

[0029] In one example, the prediction information includes a flag indicating that at least one of (i) BM-based MV refinement or (ii) the BDOF mode is applied to the sub-block.

[0030]

[0030] In one example, BM-based MV refinement or the BDOF mode is applied based on the first reference picture or the second reference picture of the current picture. The first reference picture is before the current picture in the display order, and the second reference picture is after the current picture in the display order. The distances from the first reference picture and the second reference picture to the current picture are the same.

[0031]

[0031] In an embodiment, a current block includes a plurality of sub-blocks that are reconfigured based on the SbTMVP mode. The processing circuit determines motion information for each of one or more sub-blocks within the plurality of sub-blocks based on the SbTMVP mode. One or more sub-blocks can be bi-predicted. The processing circuit can apply BM-based MV refinement to the current block based on an initial MV pair of the current block to determine updated motion information for the current block. The initial MV pair can be the motion information of one of the one or more sub-blocks. The processing circuit can reconfigure the current block based on the updated motion information of the current block.

[0032]

[0032] In one example, the one or more sub-blocks include a plurality of bi-predicted sub-blocks. The processing circuit determines an initial MV pair of the current block by: (i) applying BM-based MV refinement to the current block based on the motion information for each of the plurality of bi-predicted sub-blocks to determine a bilateral matching cost associated with each of the bi-predicted sub-blocks; and (ii) determining the initial MV pair of the current block as the motion information of a sub-block using the minimum bilateral matching cost among the bilateral matching costs of the plurality of bi-predicted sub-blocks. This can be done by.

[0033]

[0033] In one example, the processing circuit determines the initial MV pair of the current block as the motion information of one of the plurality of bi-predicted sub-blocks based on syntax information in the coded bitstream.

[0034]

[0034] In one example, the prediction information indicates that BM-based MV refinement is applied to the current block based on an initial MV pair that is the motion information of one of the one or more sub-blocks.

[0035]

[0035] In an embodiment, a processing circuit receives a coded bitstream including a current block in a current picture, the current block including a plurality of sub-blocks. The processing circuit obtains prediction information indicating whether the current block is coded in a sub-block-based temporal motion vector prediction (SbTMVP) mode. When the current block is coded in the SbTMVP mode, the processing circuit determines whether sub-blocks within the plurality of sub-blocks of the current block are bi-predicted. When a sub-block is bi-predicted, the processing circuit determines motion information of the sub-block based on the SbTMVP mode. The processing circuit applies at least one of (i) bilateral matching (BM)-based motion vector (MV) refinement and (ii) bi-directional optical flow (BDOF) mode to the sub-block to refine the motion information of the sub-block. The processing circuit reconstructs the current block based on the refined motion information corresponding to one or more sub-blocks within the plurality of sub-blocks of the current block. The refined motion information corresponding to one or more sub-blocks includes the refined motion information of the sub-blocks.

[0036]

[0036] Aspects of the present disclosure also provide a non-transitory computer-readable storage medium storing instructions that, when executed by a computer for video decoding, cause the computer to execute a method for video decoding.

Brief Description of the Drawings

[0037]

[0037] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.

Figure 1A

[0038] FIG. 1A is a schematic diagram of an exemplary subset of intra prediction modes.

Figure 1B

[0039] FIG. 1B is a diagram of exemplary intra prediction directions.

Figure 2

[0040] Figure 2 shows an example of a current block and its surrounding samples.

Figure 3

[0041] Figure 3 is a schematic diagram of an exemplary block diagram of a communication system (300).

Figure 4

[0042] Figure 4 is a schematic diagram of an exemplary block diagram of a communication system (400).

Figure 5

[0043] Figure 5 is a schematic diagram of an exemplary block diagram of a decoder.

Figure 6

[0044] Figure 6 is a schematic diagram of an exemplary block diagram of an encoder.

Figure 7

[0045] Figure 7 shows a block diagram of an exemplary encoder.

Figure 8

[0046] Figure 8 shows a block diagram of an exemplary decoder.

Figure 9

[0047] Figure 9 shows the positions of spatial merge candidates according to an embodiment of the present disclosure.

Figure 10

[0048] Figure 10 shows candidate pairs considered for redundant checking of spatial merge candidates according to an embodiment of the present disclosure.

Figure 11

[0049] Figure 11 shows exemplary motion vector scaling for temporal merge candidates.

Figure 12

[0050] Figure 12 shows exemplary candidate positions for temporal merge candidates of a current coding unit.

Figure 13

[0051] Figure 13 shows an example of a search process in the merge motion vector difference (MMVD) mode.

Figure 14

[0051] Figure 14 shows an example of a search process in the merge motion vector difference (MMVD) mode.

Figure 15

[0052] FIG. 15 shows exemplary additional refinement positions along multiple oblique angles in the MMVD mode.

Figure 16

[0053] FIG. 16 shows an exemplary subblock-based temporal motion vector prediction (SbTMVP) process used in the SbTMVP mode.

Figure 17

[0053] FIG. 17 shows an exemplary subblock-based temporal motion vector prediction (SbTMVP) process used in the SbTMVP mode.

Figure 18

[0054] FIG. 18 shows an example of bilateral matching based decoder side motion vector refinement (DMVR).

Figure 19

[0055] FIG. 19 shows an example of an extended coding unit (CU) region for bi-directional optical flow (BDOF).

Figure 20

[0056] FIG. 20 shows an example of a search area including multiple search regions.

Figure 21

[0057] FIG. 21 shows a flowchart for outlining an encoding process according to some embodiments of the present disclosure.

Figure 22A

[0058] FIG. 22A shows a flowchart for outlining a decoding process according to some embodiments of the present disclosure.

Figure 22B

[0059] FIG. 22B shows a flowchart for outlining a decoding process according to some embodiments of the present disclosure.

Figure 23

[0060] FIG. 23 shows a flowchart for outlining an encoding process according to some embodiments of the present disclosure.

Figure 24

[0061] FIG. 24 shows a flowchart for outlining a decoding process according to some embodiments of the present disclosure.

Figure 25

[0062] FIG. 25 is a schematic diagram of a computer system according to an embodiment.

Best Mode for Carrying Out the Invention

[0038]

[0063] FIG. 3 shows an exemplary block diagram of a communication system (300). The communication system (300) includes a plurality of terminal devices capable of communicating with each other via, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via a network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) perform unidirectional data transmission. For example, the terminal device (310) can code video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The coded video data can be transmitted in the form of one or more coded video bitstreams. The terminal device (320) can receive the coded video data from the network (350), decode the coded video data to restore a video picture, and display the video picture according to the restored video data. Unidirectional data transmission may be common in media serving applications and the like.

[0039]

[0064] In another example, a communication system (300) includes, for example, a second pair of terminal devices (330) and (340) that perform bidirectional transmission of coded video data during a video conference. With respect to the bidirectional transmission of data, for example, each of the terminal devices (330) and (340) can code video data (e.g., a stream of video pictures captured by a terminal device) to transmit to the other of the terminal devices (330) and (340) via a network (350). Each of the terminal devices (330) and (340) can also receive the coded video data transmitted by the other of the terminal devices (330) and (340), can decode the coded video data to restore the video picture, and can display the video picture on an accessible display device according to the restored video data.

[0040]

[0065] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) are shown as a server, a personal computer, and smartphones, but the principles of the present disclosure need not be so limited. Embodiments of the present disclosure have applications using laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network (350) represents any number of networks that carry coded video data between the terminal devices (310), (320), (330), and (340), including, for example, wired (wired) and / or wireless communication networks. The communication network (350) can exchange data over circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of the present disclosure, the architecture and topology of the network (350) may not be important for the operation of the present disclosure, unless otherwise described below.

[0041]

[0066] FIG. 4 shows a video encoder and a video decoder in a streaming environment as an application example of the disclosed subject matter. The disclosed subject matter is equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, streaming services, storage of compressed video on digital media (including CDs, DVDs, memory sticks, etc.).

[0042]

[0067] The streaming system can include a video source (401), such as a digital camera, and may include a capture subsystem (413) capable of generating a stream of, for example, uncompressed video pictures (402). In one example, the stream of video pictures (402) includes samples taken by a digital camera. The stream of video pictures (402), depicted as a thick line to emphasize the large amount of data when compared to the encoded video data (404) (or coded video bitstream), can be processed by an electronic device (420) including a video encoder (403) coupled to the video source (401). The video encoder (403) includes hardware, software, or a combination thereof and can be operative to implement or realize aspects of the disclosed subject matter as detailed below. The encoded video data (404) (or encoded video bitstream), depicted as a thin line to emphasize the smaller amount of data when compared to the stream of video pictures (402), can be stored in a streaming server (405) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include, for example, a video decoder (410) within an electronic device (430). The video decoder (410) decodes an incoming copy (407) of the encoded video data and generates an output stream of video pictures (411) that can be rendered on a display (412), such as a display screen, or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., video bitstreams) can be encoded according to a particular video coding / compression standard.Examples of these standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.

[0043]

[0068] Note that the electronic devices (420) and (430) can include other components (not shown). For example, the electronic device (420) can include a video decoder (not shown), and the electronic device (430) can also include a video encoder (not shown).

[0044]

[0069] FIG. 5 shows an exemplary block diagram of a video decoder (510). The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). The video decoder (510) can be used in place of the video decoder (410) in the example of FIG. 4.

[0045]

[0070] Receiver (531) is capable of receiving one or more coded video sequences to be decoded by video decoder (510). In an embodiment, when the decoding of each coded video sequence is independent of other coded video sequences, it is possible to receive one coded video sequence at a time. The coded video sequence can be received from channel (501), which may be a hardware / software link to a storage device storing the encoded video data. Receiver (531) is capable of receiving the encoded video data together with other data, such as coded audio data and / or auxiliary data streams, and these data can be transferred using respective entities (not shown). Receiver (531) can separate the coded video sequence from other data. To handle network jitter, buffer memory (515) may be coupled between receiver (531) and entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In certain applications, buffer memory (515) is part of video decoder (510). In other cases, it may be outside video decoder (510) (not shown). In yet another example, there may be buffer memory (not shown) outside video decoder (510), for example, to handle network jitter, and furthermore, there may be another buffer memory (515) inside video decoder (510), for example, to handle playback timing. If receiver (531) is receiving data from a store-and-forward device with sufficient bandwidth and controllability, or from a synchronous network, buffer memory (515) may not be required or can be made smaller.For use in a best-effort packet network such as the Internet, buffer memory (515) may be required, which may be relatively large and advantageously may be of an adaptable size and may be implemented at least in part in an operating system or similar element (not shown) outside the video decoder (510).

[0046]

[0071] The video decoder (510) can include a parser (520) to reconstruct symbols (521) from the coded video sequence. The categories of these symbols include information used to manage the operation of the video decoder (510) and potentially information for controlling a rendering device (512) (e.g., a display screen) that is not an essential part of the electronic device (530) but can be coupled to the electronic device (530), as shown in FIG. 5. The control information for the rendering device may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (520) can parse / entropy decode the received coded video sequence. The coding of the video sequence to be coded can conform to a video coding technology or standard and can follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context influence, etc. The parser (520) can extract a set of subgroup parameters for at least one subgroup of pixels within the video decoder based on at least one parameter corresponding to a group. The subgroups can include a Group of Pictures (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), etc. The parser (520) can also extract from the coded video sequence information such as transform coefficients, quantization parameter values, motion vectors, etc.

[0047]

[0072] The parser (520) can perform entropy decoding / analysis processing on the video sequence received from the buffer memory (515) to generate symbols (521).

[0048]

[0073] The reconstruction of the symbols (521) can include a plurality of different units depending on the type of the coded video picture or a part thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. How each unit is included can be controlled by subgroup control information analyzed by the parser (520) from the coded video sequence. Such a flow of subgroup control information between the parser (520) and a plurality of subsequent units is not depicted for clarity.

[0049]

[0074] The video decoder (510) can be conceptually subdivided into a plurality of functional units as described below, in addition to the functional blocks already described. In a practical implementation operating under commercial constraints, many of these units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, the conceptual subdivision into the following functional units is appropriate.

[0050]

[0075] The first unit is the scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives, as symbols (521), not only the quantized transform coefficients but also control information (including the transform to be used, block size, quantization factor, quantization scaling matrix, etc.) from the parser (520). The scaler / inverse transform unit (551) can output a block including sample values that can be input to the aggregator (555).

[0051]

[0076] In some cases, the output samples of the scaler / inverse transform unit (551) may be related to intra-coded blocks. An intra-coded block is a block that does not use prediction information from a previously reconstructed picture but may use prediction information from a previously reconstructed portion of the current picture. Such prediction information can be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) uses the already reconstructed surrounding information fetched from the buffer (558) of the current picture to generate blocks of the same size and shape as the block being reconstructed. The current picture buffer (558) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. The aggregator (555) may, in some cases, add, sample by sample, the prediction information generated by the intra prediction unit (552) to the output sample information as provided by the scaler / inverse transform unit (551).

[0052]

[0077] Otherwise, the output samples of the scaler / inverse transform unit (551) can be related to the inter-coded motion-compensable blocks. In such a case, the motion compensation prediction unit (553) can access the reference picture memory (557) to extract the samples used for prediction. According to the symbol (521) related to the block, after motion-compensating the extracted samples, these samples are added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (in this case, called the residual samples or residual signal), generating output sample information. The address in the reference picture memory (557) from which the motion compensation prediction unit (553) extracts the prediction samples can be controlled by the motion vectors available to the motion compensation prediction unit (553) in the form of, for example, X, Y, and the symbol (521) which can have reference picture components. Also, motion compensation can include interpolation of sample values taken from the reference picture memory (557), a motion vector prediction mechanism, etc., when an accurate sub-sample motion vector is used.

[0053]

[0078] The output samples of the aggregator (555) can be affected by various loop filtering techniques within the loop filter unit (556). The video compression technology can include in-loop filter techniques that are included in the coded video sequence (also called the coded video bitstream) and are controlled by parameters made available to the loop filter unit (556) as symbols (521) from the parser (520). Also, video compression can respond to meta information obtained during the decoding of the coded picture or a previous part of the coded video sequence (in the decoding order), and can also respond to previously reconstructed loop-filtered sample values.

[0054]

[0079] The output of the loop filter unit (556) can be not only output to the rendering device (512), but also a sample stream that can be stored in the reference picture memory (557) for use in future inter-picture prediction.

[0055]

[0080] Once a given coded picture is completely reconstructed, it can be used as a reference picture for future prediction. For example, when the coded picture corresponding to the current picture is completely reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a new current picture buffer can be reallocated before starting the reconstruction of subsequent coded pictures.

[0056]

[0081] The video decoder (510) is capable of performing a decoding operation in accordance with a standard such as ITU-T Rec.H.265 or a predetermined video compression technique. The coded video sequence can comply with the syntax specified by the video compression technique or standard being used, in the sense that the coded video sequence conforms to both the syntax of the video compression technique or standard and the profile as documented in the video compression technique or standard. Specifically, the profile can select specific tools as the only tools available under that profile from all the tools available in the video compression technique or standard. Also, for compliance, it is necessary that the complexity of the coded video sequence falls within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further restricted by the Hypothetical Reference Decoder (HRD) specifications and metadata for HRD buffer management signaled in the coded video sequence.

[0057]

[0082] In an embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0058]

[0083] FIG. 6 shows an exemplary block diagram of a video encoder (603). The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used in place of the video encoder (403) in the example of FIG. 4.

[0059]

[0084] The video encoder (603) can receive video samples from a video source (601) (not part of the electronic device (620) in the example of FIG. 6) that can capture video images to be coded by the video encoder (603). In another example, the video source (601) is part of the electronic device (620).

[0060]

[0085] The video source (601) can provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream that can be any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit,...), any color space (e.g., BT.601 YCrCB, RGB,...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (601) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that convey motion when viewed in sequence. The picture itself can be organized as a spatial array of pixels, and each pixel can include one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0061]

[0086] According to an embodiment, the video encoder (603) can encode and compress pictures of a source video sequence into a coded video sequence (643) in real time or under some other required time constraints. Enforcing an appropriate coding speed is one function of the controller (650). In some embodiments, the controller (650) controls other functional units and is functionally coupled to other functional units as described below. The coupling is not drawn for clarity. Parameters set by the controller (650) can include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization techniques, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other appropriate functions related to the video encoder (603) optimized for a particular system design.

[0062]

[0087] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As a very simplified explanation, in one example, the coding loop may include a source coder (630) (which is responsible for generating symbols such as a symbol stream based on an input picture and reference pictures to be coded), and a (local) decoder (633) incorporated in the video encoder (603). The decoder (633) reconstructs the symbols to generate sample data in the same way as a (remote) decoder does. The reconstructed sample stream (sample data) is input to the reference picture memory (634). Since the decoding of the symbol stream results in a bit-exact result regardless of the decoder's location (local or remote), the content in the reference picture memory (634) is also bit-exact between the local encoder and the remote encoder. In other words, the prediction unit of the encoder "sees" exactly the same sample values as the samples that the decoder would "see" when using prediction during decoding, as reference picture samples. This basic principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, e.g., due to channel errors) is also used similarly in some related arts.

[0063]

[0088] It is possible to assume that the operation of the "local" decoder (633) is the same as that of a "remote" decoder such as the video decoder (510) already described in detail above in relation to FIG. 5. However, referring briefly to FIG. 5, since it is possible to assume that the symbols are available and the encoding / decoding of the symbols into the coded video sequence by the entropy coder (645) and the parser (520) is lossless, the entropy decoding part of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully realized in the local decoder (633).

[0064]

[0089] In one embodiment, decoder techniques other than parsing / entropy decoding that exist in the decoder are also present in the corresponding encoder in an ideal or substantially identical functional form. Therefore, the disclosed subject matter focuses on the operation of the decoder. Since the description of encoder techniques is the reverse of the decoder techniques described comprehensively, it can be omitted. In certain areas, more detailed descriptions are given below.

[0065]

[0090] During operation, in some examples, the source coder (630) can perform motion-compensated predictive coding that predicts and encodes an input picture by referring to one or more previously coded pictures from a video sequence designated as a "reference picture". In this way, the coding engine (632) codes the difference between a pixel block of the input picture and a pixel block of a reference picture that can be selected as a prediction reference for the input picture.

[0066]

[0091] The local video decoder (633) can decode the coded video data of a picture that can be specified as a reference picture based on the symbols generated by the source coder (630). The operation of the coding engine (632) may advantageously be a lossless process. If the coded video data may be decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (633) can repeat the decoding process that can be performed by the video decoder in the reference picture, causing the reconstructed reference picture to be stored in the reference picture cache (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference picture having common content as the reconstructed reference picture obtained by the video decoder at the remote end (assuming no transmission errors).

[0067]

[0092] The predictor (635) can perform a prediction search for the coding engine (632). That is, for a new picture to be coded, the predictor (635) can search the reference picture memory (634) for sample data (as a candidate reference pixel block) or predetermined metadata (reference picture motion vector, block shape, etc.), which may serve as an appropriate prediction reference for the new picture. The predictor (635) can operate on a sample block - pixel block basis to find an appropriate prediction reference. In some cases, the input picture may have a prediction reference drawn from a plurality of reference pictures stored in the reference picture memory (634) as determined by the search result obtained by the predictor (635).

[0068]

[0093] The controller (650) can manage the coding operation of the source coder (630), including, for example, the setting of parameters and subgroup parameters used for encoding video data.

[0069]

[0094] All outputs of the aforementioned functional units can undergo entropy coding in the entropy coder (645). The entropy coder (645) converts the symbols generated by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable-length coding, arithmetic coding, etc.

[0070]

[0095] The transmitter (640) can buffer the coded video sequence as created by the entropy coder (645) and prepare it for transmission via the communication channel (660), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (640) can merge the coded video data from the video coder (603) with other data to be transmitted, such as, for example, coded audio data and / or auxiliary data streams (source not shown).

[0071]

[0096] The controller (650) can manage the operation of the video encoder (603). During coding, the controller (650) can assign a specific coded picture type to each of the coded pictures, which may affect the coding technique applicable to each picture. For example, a picture may often be designated as one of the following picture types:

[0097] An Intra Picture (I Picture) can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs, for example, allow different types of Intra Pictures, including Independent Decoder Refresh (IDR) Pictures. Those skilled in the art are aware of these variations of I Pictures, as well as their respective uses and characteristics.

[0072]

[0098] A Predictive Picture (P Picture) can be encoded and decoded using Intra prediction or Inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.

[0073]

[0099] A Bi-directional Predictive Picture (B Picture) can be encoded and decoded using Intra prediction or Inter prediction that uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for the reconstruction of one block.

[0074]

[0100] The source picture is usually spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and can be coded block by block. The blocks can be predicted and coded by referring to other (already coded) blocks as determined by the coding assignment applied to each block of the picture. For example, blocks of an I picture may be non-predicted coded, or they may be predicted and coded by referring to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be predicted and coded by spatial or temporal prediction by referring to one previously coded reference picture. Blocks of a B picture may be predicted and coded by spatial or temporal prediction by referring to one or two previously coded reference pictures.

[0075]

[0101] The video encoder (603) can perform coding operations according to a predetermined video coding technology or standard such as ITU-T Rec.H.266. In this operation, the video encoder (603) can execute various compression operations including predictive coding operations that utilize temporal and spatial redundancies in the input video sequence. The coded video data can thus conform to the syntax specified by the video coding technology or standard being used.

[0076]

[0102] In an embodiment, the transmitter (640) can transmit additional data together with the coded video. The source coder (630) can include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data (redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.).

[0077]

[0103] Video can be captured as a plurality of source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra prediction) utilizes the spatial correlation in a given picture, and inter-picture prediction utilizes the (temporal or other) correlation between pictures. In one example, a particular picture under encoding / decoding, called the current picture, is partitioned into blocks. If a block within the current picture is similar to a reference block within a reference picture that has been previously coded and is still buffered in the video, the block within the current picture can be coded by a vector called a motion vector. The motion vector points to the reference block within the reference picture and can have a third dimension to identify the reference picture when multiple reference pictures are used.

[0078]

[0104] In some embodiments, it is possible to use dual-prediction techniques for inter-picture prediction. According to the dual-prediction technique, two reference pictures such as a first reference picture and a second reference picture that both precede the current picture in decoding order within the video (however, they may be in the past and future respectively in display order) are used. A block within the current picture can be coded by a first motion vector pointing to a first reference block within the first reference picture and a second motion vector pointing to a second reference block within the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.

[0079]

[0105] Further, in order to improve coding efficiency, it is possible to use merge mode techniques for inter-picture prediction.

[0080]

[0106] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and the CTUs within a picture have the same size such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs) which are one luma CTB and two chroma CTBs. Each CTU can be recursively quad-tree partitioned into one or more coding units (CUs). For example, a 64×64 pixel CTU can be partitioned into one 64×64 pixel CU, four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine the prediction type of the CU, such as an inter-prediction type or an intra-prediction type. The CU is partitioned into one or more prediction units (PUs) depending on the temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In an embodiment, the prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of values (e.g., luma values) for pixels such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0081]

[0107] FIG. 7 shows an exemplary diagram of a video encoder (703). The video encoder (703) receives a processing block (e.g., a prediction block) of sample values within a current video picture in a sequence of video pictures and is configured to encode the processing block into a coded picture that is part of a coded video sequence. In one example, the video encoder (703) is used in place of the video encoder (403) of the example of FIG. 4.

[0082]

[0108] In an example of HEVC, a video encoder (703) receives a matrix of sample values of a processing block, such as a prediction block of 8×8 samples. The video encoder (703) uses an intra-mode, an inter-mode, or a bi-prediction mode to determine, for example using rate distortion optimization, whether the processing block is best coded. If the processing block is to be coded in the intra-mode, the video encoder (703) can use an intra prediction technique to code the processing block into the coded picture; if the processing block is to be coded in the inter-mode or the bi-prediction mode, the video encoder (703) can use an inter prediction technique or a bi-prediction technique respectively to code the processing block into the coded picture. In a specific video coding technique, the merge mode may be an inter-picture prediction sub-mode, in which case the motion vector is derived from one or more motion vector predictors without the benefit of coded motion vector components outside the predictor. In some other specific video coding techniques, there may be motion vector components applicable to the target block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown), to determine the mode of the processing block.

[0083]

[0109] In the example of FIG. 7, the video encoder (703) includes an inter encoder (730), an intra encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general-purpose controller (721), and an entropy encoder (725) coupled together as shown in FIG. 7.

[0084]

[0110] The inter-encoder (730) receives samples of a current block (e.g., a processing block), compares the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and blocks in a subsequent picture), generates inter-prediction information (e.g., a description of redundant information by an inter-coding technique, a motion vector, merge mode information), and is configured to calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using some suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on the encoded video information.

[0085]

[0111] The intra-encoder (722) receives samples of a current block (e.g., a processing block), optionally compares the block with blocks already coded in the same picture, generates quantized coefficients after transformation, and is optionally also configured to generate intra-prediction information (e.g., intra-prediction direction information according to one or more intra-coding techniques). In one example, the intra-encoder (722) also calculates an intra-prediction result (e.g., a predicted block) based on the intra-prediction information and reference blocks in the same picture.

[0086]

[0112] The general-purpose controller (721) is configured to determine general control data and control other components of the video encoder (703) based on the general control data. In one example, the general-purpose controller (721) determines the mode of a block and provides a control signal to the switch (726) based on that mode. For example, when the mode is the intra mode, the general-purpose controller (721) controls the switch (726) to select the intra mode result for use by the residual calculator (723) and controls the entropy encoder (725) to select the intra prediction information and include the intra prediction information in the bitstream; also, when the mode is the inter mode, the general-purpose controller (721) controls the switch (726) to select the inter prediction result for use by the residual calculator (723) and controls the entropy encoder (725) to select the inter prediction information and include the inter prediction information in the bitstream.

[0087]

[0113] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) is configured to operate based on the residual data to encode the residual data and generate transform coefficients. In one example, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain and generate transform coefficients. The transform coefficients are then subjected to quantization processing to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be appropriately used by the intra-encoder (722) and the inter-encoder (730). For example, the inter-encoder (730) can generate a decoded block based on the decoded residual data and the inter-prediction information, and the intra-encoder (722) can generate a decoded block based on the decoded residual data and the intra-prediction information. The decoded block is appropriately processed to generate a decoded picture, and the decoded picture is buffered in a memory circuit (not shown) and can be used as a reference picture in some examples.

[0088]

[0114] The entropy encoder (725) is configured to format the bitstream to include the encoded blocks. The entropy encoder (725) is configured to include various information in the bitstream according to a suitable standard such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other suitable information in the bitstream. Note that there is no residual information when coding a block in either the inter-mode or the merge sub-mode of the bi-prediction mode according to the disclosed subject matter.

[0089]

[0115] FIG. 8 shows an exemplary diagram of a video decoder (810). The video decoder (810) is configured to receive a coded picture that is part of a coded video sequence and decode the coded picture to generate a reconstructed picture. In one example, the video decoder (810) is used in place of the video decoder (410) in the example of FIG. 4.

[0090]

[0116] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) coupled together as shown in FIG. 8.

[0091]

[0117] The entropy decoder (871) can be configured to reconstruct from the coded picture certain symbols that represent the syntax elements that make up the coded picture. Such symbols can include, for example, the mode in which a block is coded (e.g., intra mode, inter mode, bi-prediction mode, the merge submode in the latter two, or another submode), prediction information (e.g., intra prediction information or inter prediction information) that can identify specific samples or metadata used for prediction by the intra decoder (872) or the inter decoder (880) respectively. Also, the symbols can include, for example, residual information in the form of quantized transform coefficients. In one example, when the prediction mode is inter or bi-prediction mode, the inter prediction information is provided to the inter decoder (880); when the prediction type is intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information can be inverse quantized and provided to the residual decoder (873).

[0092]

[0118] The inter decoder (880) is configured to receive inter prediction information and generate an inter prediction result based on the inter prediction information.

[0093]

[0119] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0094]

[0120] The residual decoder (873) is configured to perform inverse quantization to extract non-quantized transform coefficients, process the non-quantized transform coefficients, and convert the residual information from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (including quantization parameter (QP)), and that information may be provided by the entropy decoder (871) (since this may be only a small amount of control information, the data path is not depicted).

[0095]

[0121] The reconfiguration module (874) is configured to form a reconfigured block by combining, in the spatial domain, residual information as an output by the residual decoder (873) and a prediction result (which may be output by an inter or intra prediction module in some cases). The block is part of a reconfigured picture, and the picture may be part of a reconfigured video. Note that other appropriate processes such as deblocking processing may be performed to improve visual quality.

[0096]

[0122] Note that the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) can be implemented using any appropriate technology. In an embodiment, the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) can be implemented using one or more processors that execute software instructions.

[0097]

[0123] In VVC, various inter-prediction modes can be used. Regarding the (CU) to be inter-predicted, the motion parameters can include the MV for a predetermined coding feature to be used for generating the inter-predicted samples, one or more reference picture indices, a reference picture list usage index, and additional information. The motion parameters can be signaled in an explicit or implicit manner. When a CU is coded in skip mode, the CU can be associated with a PU and may not have significant residual coefficients, a coded motion vector delta, or an MV residual (e.g., MVD) or a reference picture index. It is possible to specify the merge mode, where the motion parameters for the current CU are obtained from neighboring CU(s) (spatial candidates and / or temporal candidates, optionally including additional information as introduced in VVC). The merge mode can be applied not only for skip mode but also for inter-predicted CUs. In one example, an alternative to the merge mode is the explicit transmission of motion parameters, in which case the MV(s), the corresponding reference picture indices for each reference picture list, the reference picture list usage flag, and other information are explicitly signaled for each CU.

[0098]

[0124] In embodiments such as VVC, the VVC Test model (VTM) reference software includes one or more refined inter-prediction coding tools, which are: extended merge prediction, merge motion vector difference (MMVD) mode, adaptive motion vector prediction (AMVP) mode using symmetric MVD signaling, affine motion compensation prediction, Subblock-based temporal motion vector prediction (SbTMVP), Adaptive motion vector resolution (AMVR), Motion field memory (1 / 16 luma sample MV storage and 8×8 motion field compression), Bi-prediction with CU-level weights (BCW), Bi-directional optical flow (BDOF), Prediction refinement using optical flow (PROF), Decoder side motion vector refinement (DMVR), Combined inter and intra prediction (CIIP), Geometric partitioning mode (GPM) and the like are included. Inter prediction and related methods are described in detail below.

[0099]

[0125] In some examples, it is possible to use extended merge prediction. In one example such as in VTM4, the merge candidate list consists of the following five types of candidates: Motion vector predictor(s) (MVP) from spatial neighboring CUs, Temporal MVP from collocated CU(s), History-based MVP(s) (HMVP) from a first-in-first-out (FIFO) table, Pairwise average MVP, and zero MV(s) are configured by sequentially including them.

[0100]

[0126] The size of the merge candidate list can be signaled in the slice header. In one example, the maximum allowable size of the merge candidate list is 6 in VTM4. For each CU coded in merge mode, the index of the best merge candidate (e.g., the merge index) can be coded using truncated unary binarization (TU). The first bin of the merge index can be coded using context (e.g., context-adaptive binary arithmetic coding (CABAC)), and bypass coding can be used for the other bins.

[0101]

[0127] Some examples of the generation process for each category of merge candidates are given below. In an embodiment, the spatial candidates are derived as follows. It is possible to assume that the derivation of spatial merge candidates in VVC is the same as that in HEVC. In one example, among the candidates at the positions shown in FIG. 9, up to four merge candidates are selected. FIG. 9 shows the positions of spatial merge candidates according to an embodiment of the present disclosure. Referring to FIG. 9, the order of derivation is B1, A1, B0, A0, and B2. The position B2 is considered only if none of the CUs at the positions A0, B0, B1, and A1 are available (e.g., due to the CU belonging to another slice or another tile), or if it is intra-coded. After the candidate at position A1 is added, the addition of the remaining candidates is subject to a redundancy check, which ensures that candidates with the same motion information are excluded from the candidate list, thereby improving the coding effect.

[0102]

[0128] To reduce the computational complexity, not all possible candidate pairs are considered in the above redundancy check. Instead, only the pairs linked by arrows in FIG. 10 are considered, and a candidate is only added to the candidate list when the corresponding candidates used for the redundancy check do not have the same motion information. FIG. 10 shows candidate pairs considered for the redundancy check of spatial merge candidates according to an embodiment of the present disclosure. Referring to FIG. 10, the pairs connected by each arrow are A1 and B1, A1 and A0, A1 and B2, B1 and B0, B1 and B2. Thus, candidates at positions B1, A0, and / or B2 can be compared with candidates at position A1, and candidates at positions B0 and / or B2 can be compared with candidates at position B1.

[0103]

[0129] In one embodiment, the temporal candidates are derived as follows. In one example, only one temporal merge candidate is added to the candidate list. FIG. 11 shows an exemplary motion vector scaling for temporal merge candidates. To derive the temporal merge candidate of the current CU (1111) within the current picture (1101), a scaled MV (1121) (e.g., as indicated by the dotted line in FIG. 11) can be derived based on the co-located CU (1112) belonging to the co-located reference picture (1104). In one example, the co-located reference picture (also referred to as the co-located picture) is, for example, a specific reference picture used for temporal motion vector prediction. The co-located reference picture used for temporal motion vector prediction can be specified by a reference index in the syntax such as high-level syntax (e.g., picture header, slice header).

[0104]

[0130] The reference picture list used to derive the co-located CU (1112) can be explicitly signaled in the slice header. The scaled MV (1121) for the temporal merge candidate can be obtained as shown by the dotted line in FIG. 11. The scaled MV (1121) can be scaled from the MV of the co-located CU (1112) using the picture order count (POC) distances tb and td. The POC distance tb can be defined as the POC difference between the current reference picture (1102) of the current picture (1101) and the current picture (1101). The POC distance td can be defined as the POC difference between the co-located reference picture (1104) of the co-located reference picture (1103) and the co-located reference picture (1103). The reference picture index of the temporal merge candidate can be set to zero.

[0105]

[0131] FIG. 12 shows exemplary candidate positions (e.g., C0 and C1) for the temporal merge candidate of the current CU. The position of the temporal merge candidate can be selected from the candidate positions C0 and C1. The candidate position C0 is located at the lower right corner of the co-located CU (1210) of the current CU. The candidate position C1 is located at the center of the co-located CU (1210) of the current CU. If the CU at the candidate position C0 is not available, is intra-coded, or is outside the current row of the CTU, the candidate position C1 is used to derive the temporal merge candidate. Otherwise, for example, if the CU at the candidate position C0 is available, is intra-coded, and is in the current row of the CTU, the candidate position C0 is used to derive the temporal merge candidate.

[0106]

[0132] The Merge with Motion Vector Difference (MMVD) mode can be used for the skip or merge mode together with the motion vector representation method. Merge candidates as used in VVC may be reused in the MMVD mode. The candidates can be selected from among the merge candidates as the starting point (e.g., MV predictor (MVP)), and can be further extended by the MMVD mode. The MMVD mode can provide a new motion vector representation using simplified signaling. The motion vector representation method includes a starting point and an MV difference (MVD). In one example, the MVD is indicated by the magnitude of the MVD (or the magnitude of the motion) and the direction of the MVD (e.g., the direction of the motion).

[0133] The MMVD mode can use a merge candidate list as used in VVC. In one embodiment, only candidates of the default merge type (e.g., MRG_TYPE_DEFAULT_N) are considered for the MMVD mode. The starting point can be indicated or defined by a base candidate index (IDX). The base candidate index can indicate a candidate (e.g., the best candidate) among the candidates (e.g., base candidates) in the merge candidate list. Table 1 shows an exemplary relationship between the base candidate index and the corresponding starting point. That the base candidate index is 0, 1, 2, or 3 indicates that the corresponding starting point is the first MVP, the second MVP, the third MVP, or the fourth MVP. In one example, when the number of base candidates is equal to 1, the base candidate IDX is not signaled.

[0107] Table 1. Base candidate IDX

[0108]

Table 1

[0134] The distance index can indicate the magnitude of movement of the MVD, such as the magnitude of the MVD. For example, the distance index indicates the distance (e.g., a predefined distance) from the starting point (e.g., the MVP indicated by the base candidate index). In one example, the distance is one of a plurality of predetermined distances as shown in Table 2. Table 2 shows an exemplary relationship between the distance index and the corresponding distance (in samples or pixels). "1-pel" in Table 2 is 1 sample or 1 pixel. For example, a distance index of 1 indicates that the distance is 1 / 2-pel or 1 / 2 sample.

[0109] Table 2. Distance IDX

[0110]

Table 2

[0135] The direction index can represent the direction of the MVD with respect to the starting point. The direction index can represent one of a plurality of directions, such as the four directions shown in Table 3. For example, a direction index of 00 indicates that the direction of the MVD is along the positive direction of the x-axis.

[0111] Table 3. Index IDX

[0112]

Table 3

[0136] After sending the skip and merge flags, the MMVD flag may be signaled. If the skip and merge flags are true, it is possible to analyze the MMVD flag. In one example, if the MMVD flag is equal to 1, it is possible to analyze the MMVD syntax (including, for example, a distance index and / or a direction index). If the MMVD flag is not equal to 1, it is possible to analyze the AFFINE flag. If the AFFINE flag is equal to 1, code the current block using the affine mode. If the AFFINE flag is not equal to 1, it is possible to analyze the skip / merge index for the skip / merge mode as used in VTM.

[0113]

[0137] FIGS. 13-14 show an example of a search process in the MMVD mode. By performing the search process, it is possible to determine an index including a base candidate index, a direction index, and / or a distance index for a current block (1300) within a current picture (or called a current frame) (1301).

[0114]

[0138] A first motion vector (MV) (1311) belonging to a first merge candidate and a second MV (1321) are shown. It is possible to assume that the first merge candidate is a merge candidate within a merge candidate list constructed for the current block (1300). The first MV (1311) and the second MV (1321) can be associated with two reference pictures (1302) and (1303) within the reference picture lists L0 and L1, respectively. Therefore, two starting points (1411) and (1421) in FIGS. 13-14 can be determined in the reference pictures (1302) and (1303), respectively.

[0115]

[0139] In one example, based on start points (1411) and (1421), in the vertical direction (represented by +Y or -Y) or the horizontal direction (represented by +X and -X) in reference pictures (1302) and (1303), a plurality of predefined points (e.g., 1-12 shown in FIG. 14) extending from the start points (1411) and (1421) can be evaluated. In one example, pairs of points that are in a mirroring relationship with respect to each other for each start point (1411) or (1421), such as the pair of points (1414) and (1424), or the pair of points (1415) and (1425), can be used to determine pairs of MVs (1314) and (1324) or MVs (1315) and (1325) that can form an MV predictor (MVP) candidate for the current block (1300). It is possible to evaluate the MVP candidates determined based on predetermined points surrounding the start point (1411) and / or (1421). Referring to FIG. 13, The MVD (1312) between the first MV (1311) and MV (1314) has a magnitude of 1S.

[0116] The MVD (1322) between the second MV (1321) and MV (1324) has a magnitude of 1S. Similarly, The MVD between the first MV (1311) and MV (1315) has a magnitude of 2S.

[0117] The MVD between the second MV (1321) and MV (1325) has a magnitude of 2S.

[0118]

[0140] In addition to the first merge candidate, other available or valid merge candidates in the merge candidate list of the current block (1300) can also be similarly evaluated. In one example, in the case of a uni-predicted merge candidate, only one prediction direction associated with one of the two reference picture lists is evaluated.

[0119]

[0141] In one example, based on the evaluation, it is possible to determine the best MVP candidate. Therefore, the optimal merge candidate corresponding to the optimal MVP candidate can be selected from the merge list, and the movement direction and movement distance can also be determined. For example, based on the selected merge candidate and Table 1, it is possible to determine the base candidate index. Based on the selected MVP corresponding to a predetermined point (1415) (or (1425)), the direction and distance (e.g., 2S) of point (1415) relative to the starting point (1411) can be determined. According to Table 2 and Table 3, the direction index and distance index can be determined accordingly.

[0120]

[0142] As described above, two indexes, such as the distance index and the direction index, can be used to indicate the MVD in the MMVD mode. Alternatively, a single index can be used to indicate the MVD in the MMVD mode, for example, by using a table that pairs the single index with the MVD.

[0121]

[0143] Template matching (TM)-based candidate rearrangement can be used in some prediction modes, such as the MMVD mode and the affine MMVD mode. In an embodiment, the MMVD offset is extended for the MMVD mode and the affine MMVD mode. FIG. 15 shows additional refinement positions along a plurality of diagonal angles, such as k×π / 8 diagonal angles (k is an integer from 0 to 15). The additional refinement positions along the plurality of diagonal angles can increase the number of directions from, for example, 4 directions (e.g., +X, -X, +Y, -Y) to 16 directions (e.g., k = 0, 1, 2,..., 15). In one example, each of the 16 directions is represented by an angle between the +X direction and the direction indicated by the center point (1500) and one of points 1-16. For example, point 1 indicates the +X direction at an angle of 0 (i.e., k = 0), point 2 indicates the direction along an angle of 1×π / 8 (i.e., k = 1), and so on.

[0122]

[0144] TM can be executed in the MMVD mode. In one example, for each MMVD refinement position, it is possible to determine the TM cost based on the current template of the current block and one or more reference templates. The TM cost can be determined using any method, such as the sum of absolute differences (SAD) (e.g., SAD cost), the sum of absolute transformed differences (SATD), the sum of squared errors (SSE), the mean removed SAD / SATD / SSE, variance, partial SAD, partial SSE, partial SATD, or something similar to these.

[0123]

[0145] The current template of the current block can include any suitable samples, such as one row of samples above the current block and / or one column of samples to the left of the current block. Based on the TM cost (e.g., SAD cost) between the current template and the corresponding reference template for the refinement position, it is possible to reorder all possible MMVD refinement positions (e.g., 16×6 representing 16 directions and 6 sizes) for each base candidate (e.g., MVP), such as the MMVD refinement positions. In one example, the top MMVD refinement position that results in the minimum TM cost (e.g., minimum SAD cost) is retained as the MMVD refinement position available for MMVD index coding. For example, a subset (e.g., 8) of the MMVD refinement positions with the minimum TM cost is used for MMVD index coding. For example, the MMVD index indicates which of the subsets of MMVD refinement positions is selected to code the current block along with the minimum TM cost. In one example, an MMVD index of 0 indicates that the MVD (e.g., MMVD refinement position) corresponding to the minimum TM cost is used to code the current block. The MMVD index can be binarized, for example, by a Rice code with a parameter equal to 2.

[0124]

[0146] In one embodiment, in addition to the above-described MMVD offset extension as shown in FIG. 15, the affine MMVD rearrangement is extended and additional refinement positions along the k×π / 4 oblique angle are added. After the rearrangement, the top half of the refinement positions with the minimum TM cost (e.g., SAD cost) are retained to code the current block.

[0125]

[0147] To improve coding efficiency and reduce the transmission overhead of MVs, sub-block level MV refinement can be applied to extend the temporal motion vector prediction (TMVP) at the CU level. In one example, the sub-block-based TMVP (SbTMVP) mode enables inheriting motion information at the sub-block level from the reference picture at the same position. The reference picture at the same position can be indicated by a reference index in the syntax such as high-level syntax (e.g., picture header, slice header). Each sub-block of the current CU (e.g., the current CU having a large size) within the current picture can have its respective motion information without explicitly transmitting the block partition structure or its respective motion information. In the SbTMVP mode, the motion information of each sub-block can be obtained in three steps as follows, for example.

[0126] In the first step, it is possible to derive the displacement vector (DV) of the current CU. The DV can indicate a block within the reference picture at the same position. For example, the DV points to a block within the reference picture at the same position from the current block within the current picture. Therefore, the block indicated by the DV is considered to be at the same position as the current block and is called the equivalent position block of the current block.

[0127] In the second step, it is possible to check the availability of SbTMVP candidates and derive the central motion (e.g., the central motion of the current CU).

[0128] In the third step, the sub-block motion information can be derived from the corresponding sub-blocks at the same position using the DV. The three steps can be combined into one or two steps, and / or the order of the three steps may be changed.

[0129]

[0148] Different from deriving a TMVP candidate by deriving a temporal MV from blocks at equivalent positions in a reference frame or reference picture, in the SbTMVP mode, for each sub-block of a current CU within a current picture, a DV (e.g., a DV derived from the MV of a left neighboring CU of the current CU) may be applied to identify a corresponding sub-block in a reference picture at an equivalent position. If the corresponding sub-block is not inter-coded, the motion information of the current sub-block can be set to the central motion of the equivalent-position block.

[0130]

[0149] The SbTMVP mode can be supported by various video coding standards including, for example, VVC. Similar to the TMVP mode, for example, in HEVC, in the SbTMVP mode, a motion field (also referred to as a motion information field or an MV field) in a reference picture at an equivalent position can be used to improve MV prediction and the merge mode for CUs within a current picture. In one example, the same reference picture at an equivalent position used by the TMVP mode is used in the SbTMVP mode. In one example, the SbTMVP mode is in the following aspects: (i) An aspect where the TMVP mode predicts motion information at the CU level while the SbTMVP mode predicts motion information at the sub-CU level; (ii) An aspect where the TMVP mode fetches a temporal MV from an equivalent-position block in a reference picture at an equivalent position (e.g., the equivalent-position lock is the bottom-right or central block with respect to the current CU), and the SbTMVP mode can apply a motion shift before fetching temporal motion information from the reference picture at an equivalent position. In which it is different from the TMVP mode. In one example, the motion shift used in the SbTMVP mode is obtained from the MV of one of the spatial neighboring blocks of the current CU.

[0131]

[0150] Figs. 16-17 show an exemplary SbTMVP process used in SbTMVP mode. The SbTMVP process can predict the motion vector (MV) of a sub-CU (e.g., sub-block) within a current CU (e.g., current block) (1601) in the current picture (1711) in, for example, two steps. In the first step, the spatial neighborhood (e.g., A1) of the current block (1601) in Figs. 16-17 is examined. If the spatial neighborhood (e.g., A1) has an MV (1721) that uses a reference picture (1712) at the same position as the reference picture of the spatial neighborhood (e.g., A1), the MV (1721) can be selected to be the motion shift (or DV) to be applied to the current block (1601). If no such MV (e.g., an MV that uses a reference picture at the same position as the reference picture) is identified, the motion shift or DV can be set to zero MV (e.g., (0,0)). In some examples, the MV(s) of additional spatial neighborhoods such as A0, B0, B1 and the like are checked if such an MV is not identified for the spatial neighborhood A1.

[0151] In a second step, a motion shift or DV (1721) identified in the first step is applied to a current block (1601) (e.g., the DV (1721) is added to the coordinates of the current block), and sub-CU level motion information (e.g., including an MV and a reference index) is obtained from a reference picture (1712) at an equivalent position. In the example shown in FIG. 17, the motion shift or DV (1721) is set to be the MV of a spatial neighborhood A1 (e.g., block A1) of the current block (1601). For each sub-CU or sub-block (1731) within the current block (1601), it is possible to derive the motion information of the sub-CU or sub-block (1731) using the motion information of a corresponding equivalent position block (1701) within the reference picture (1712) at the equivalent position (e.g., the motion information of the smallest motion grid covering the central sample of the equivalent position block (1701)). After the motion information of an equivalent position sub-CU (1732) within the equivalent position block (1701) is identified, the motion information of the equivalent position sub-CU (1732) can be converted into the motion information (e.g., an MV and one or more reference indexes) of the current sub-CU (1731) using a scaling method such as a method similar to the TMVP process used in HEVC, where temporal motion scaling is applied to align the reference picture of the temporal MV with the reference picture of the current CU.

[0132]

[0152] The motion field of the current block (1601) derived based on the DV (1721) can include the motion information of each sub-block (1731) within the current block (1601), such as an MV and one or more related reference indexes. The motion field of the current block (1601) is also called an SbTMVP candidate and corresponds to the DV (1721).

[0133]

[0153] FIG. 17 shows an example of a motion field of a current block (1601) or an SbTMVP candidate. For example, the motion information of a bi-predicted sub-block (1731(1)) includes a first MV, a first index indicating a first reference picture in reference picture list 0 (L0), a second MV, and a second index indicating a second reference picture in reference picture list 1 (L1). In one example, the motion information of a uni-predicted sub-block (1731(2)) includes an MV and an index indicating a reference picture in L0 or L1.

[0134]

[0154] In one example, DV (1721) is applied to the center position of the current block (1601) to find the displaced center position in the reference picture (1712) at the same position. If the block including the displaced center position is not inter-coded, the SbTMVP candidate is considered unavailable. Otherwise, when the block including the displaced center position (e.g., the same-position block (1701)) is inter-coded, the motion information of the center position of the current block (1601), called the center motion of the current block (1601), can be derived from the motion information of the block including the displaced center position in the reference picture (1712) at the same position. In one example, the center motion of the current block (1601) can be derived from the motion information of the block including the displaced center position in the reference picture (1712) at the same position using a scaling process. When the SbTMVP candidate is available, DV (1721) can be applied to find the corresponding sub-block (1731) in the reference picture (1712) at the same position for each sub-block (1732) of the current block (1601). The motion information of the corresponding sub-block (1732) can be used to derive the motion information of the sub-block (1731) in the current block (1601) in the same way as used to derive the center motion of the current block (1601). In one example, when the corresponding sub-block (1732) is not inter-coded, the motion information of the current sub-block (1731) is set to be the center motion of the current block (1601).

[0135]

[0155] In some examples, such as in VVC, a combined sub-block-based merge list that includes SbTMVP candidates and affine merge candidates is used in the signaling of the sub-block-based merge mode. The SbTMVP mode can be enabled or disabled by a sequence parameter set (SPS) flag. When the SbTMVP mode is enabled, the SbTMVP candidate (or SbTMVP predictor) is added as the first entry in the sub-block-based merge list that includes sub-block-based merge candidates, and then affine merge candidates may follow. The size of the sub-block-based merge list can be signaled in the SPS. In one example, the maximum allowable size of the sub-block-based merge list is 5 in VVC. In one example, multiple SbTMVP candidates are included in the sub-block-based merge list.

[0136]

[0156] In some examples, such as in VVC, the sub-CU size used in the SbTMVP mode is fixed at 8×8, as used for the affine merge mode. In one example, the SbTMVP mode is applicable only to CUs where both the width and height are 8 or more. The sub-block size (e.g., 8×8) can potentially be set to other sizes, such as 4×4, in the ECM software model used for exploration beyond VVC. In one example, multiple reference pictures at the same position, such as two equally-positioned frames, are utilized to provide the temporal motion information for SbTMVP and / or TMVP in the AMVP mode.

[0137]

[0157] In some examples of the SbTMVP mode, such as in VVC and ECM, the DV of the current CU (e.g., DV(1721) in FIG. 17) is derived from the MV of the neighboring CU of the current CU. In some examples, a DV offset (DVO) is used in the SbTMVP mode. In one example, in order to obtain a more accurate matching, the initial DV (e.g., derived from the MV of the neighboring CU of the current CU) can be modified by the DV offset to determine the updated DV'. In one example, the updated DV' is the vector sum of the initial DV and the DVO. As described in FIGS. 16 - 17, it is possible to determine the initial DV using some method. By utilizing the DVO, it is possible to adjust the position of the co - located CU (or co - located block) in the reference picture at the same position, and then the MV field of the co - located CU (or co - located block) may change based on the DVO. Referring to FIG. 17, instead of using the initial DV (e.g., DV(1721)) which is the MV of the spatial neighbor A1, it is possible to use the updated DV' to determine the co - located block of the current block.

[0138]

[0158] The DVO can be indicated, for example, by signaling an index indicating the DVO from among DVO candidates (e.g., the possible DVO to be used by the current block). It is possible to signal one or more indices to indicate which of the DVO candidates should be selected as the DVO.

[0139]

[0159] In one example, DVO is signaled using the MMVD mode. The SbTMVP mode using DVO signaled using the MMVD mode may be referred to as the SbTMVP-MMVD mode. The embodiments described in the present disclosure can be applied with the SbTMVP mode or a variation of the SbTMVP mode (e.g., the SbTMVP-MMVD mode). For example, two indexes including a first index (e.g., a distance index) indicating the size of the DVO candidate and a second index (e.g., a direction index) indicating the direction of the DVO candidate are signaled to indicate the DVO candidate as described in Table 2-3.

[0140]

[0160] Referring back to FIG. 14, the distance index and the direction index can be predefined as described above with respect to the MMVD mode. The distance index indicates motion magnitude information such as the size of the DVO. For example, the distance index indicates a predetermined distance from a starting point (such as an initial DV). In one example, the available predetermined distances are shown in Table 2. The direction index represents the direction of the DVO with respect to the starting point (such as an initial DV). The direction index can indicate one of a plurality of directions such as the four directions shown in Table 3.

[0141]

[0161] In one embodiment, the DVO is directly signaled using some signaling method used to signal MVD, for example, as in the AMVP mode.

[0142]

[0162] For example, by minimizing the distortion between two reference blocks (or two reference sub-blocks) in reference picture list L0 and reference picture list L1, based on matching the reference blocks (or reference sub-blocks) in two respective reference pictures, it is possible to apply bilateral matching (BM)-based motion vector (MV) refinement to refine the motion information of the block (or sub-block).

[0143]

[0163] Examples of BM-based MV refinement include DMVR or its variants, multi-pass (MP) decoder-side motion vector refinement (MP-DMVR) or its variants, etc. In order to improve the accuracy of the MV in the merge mode, it is possible to apply BM-based DMVR as in VVC. In the bi-prediction operation, it is possible to search for refined MVs around the initial MVs in the reference pictures in L0 and L1. The BM method can calculate the distortion between two reference blocks in the reference pictures of L0 and L1.

[0144]

[0164] In one embodiment, it is possible to apply DMVR to a CU (1801) in a current picture (1811) coded in a regular merge mode as in VVC. The initial MV pair (e.g., MV0 - MV1) is an input to the DMVR process and may also be referred to as the input or input MV pair. The initial MV pair can be obtained from regular merge candidates as in the regular merge mode. MV0 indicates the block (1831) of the first reference picture (1812) in L0, and MV1 indicates the block (1832) of the second reference picture (1813) in L1.

[0145]

[0165] BM can be applied in DMVR to refine the input MV pair. It is possible to determine the output of DMVR (or the refined MV pair MV0’-MV1’). The refined MV pair can be associated with the input MV pair as follows.

[0146]

Number

[0166] In Eq.1, the parameters mv L0 and mv L1 represent MV0 and MV1 respectively. The parameters mv refinedL0 and mv refinedL1 represent MV0’ and MV1’ respectively. By using the MVD mirroring characteristic, the motion vector difference (MVD) Δmv can be applied to the input MV pair to obtain the refined MV pair, where Δmv is applied to MV0 and (-Δmv) is applied to MV1. In one example, MV0 and MV1 refer to two different reference pictures (1812)-(1813) having equal differences in the picture order count (POC) with respect to the current picture (1811), and the two reference pictures (1812)-(1813) are in different temporal directions. For example, the first POC difference (ΔPOC1) is “POC of reference picture (1812)” minus “POC of current picture (1811)”, and the second POC difference (ΔPOC2) is “POC of reference picture (1813)” minus “POC of current picture (1811)”. The first POC difference and the second POC difference have the same magnitude but different signs, for example, ΔPOC1=-ΔPOC2 is true.

[0147]

[0167] The refined MV can be used in motion compensation prediction for both the luma component and the chroma component.

[0148]

[0168] In one example, the DMVR shown in FIG. 18 is used to refine the MV pair of a block when the block size meets the condition, such as when the block size is less than or equal to a threshold M D ×N D (e.g., 16×16). M D and N D are positive integers. M D and N D may be the same or different.

[0149]

[0169] In another embodiment, a luma-coded block larger than the threshold M D ×N D (e.g., 16×16) may be divided into sub-blocks of size, for example, M D ×N D (e.g., 16×16) for the MV refinement process. The DMVR as shown in FIG. 18 may be applied to each of the sub-blocks to refine the MV pair of each sub-block. In this case, the blocks shown in FIG. 18 (e.g., (1801), (1831)-(1832), and (1835)-(1836)) represent sub-blocks.

[0150]

[0170] The DMVR can be applied using any suitable bilateral matching method. In one example, the DMVR is applied in multiple steps as described below. Δmv in Eq. 1 can be independently derived for each block or sub-block in the following two steps, including integer-precision motion search (or integer sample offset search) and fractional motion search. The following description uses a sub-block as an example, but this description is also applicable to blocks as appropriate.

[0151]

[0171] Sub-block motion compensation (MC) is based on the refined MV pair {mv refinedL0 , mv refinedL1} can be applied using

[0152]

[0172] In integer sample - offset search in DMVR, the search space may include a plurality (e.g., 25) of MV - pair candidates as described in Eq. 2:

[0153]

Number

[0154]

Number

[0173] In Eq. 3, W and H are the weight and height of the sub-block. If the SAD of the initial MV pair is smaller than the threshold, it is possible to end the integer sample offset search stage of DMVR. Otherwise, the SADs of the remaining 24 points may be calculated and checked, for example, in raster scan order. The point (e.g., (i,j) pair) or MV pair having the minimum SAD may be selected as the output of the integer sample offset search stage. To reduce the penalty of the uncertainty of DMVR refinement, the initial MV pair may work advantageously during the DMVR process. For example, based on Eqs. 3 and 5, the factor K related to the difference (e.g., SAD) between the reference blocks indicated by the initial MV pair (e.g., i = j = 0) is reduced by 1 / 4 from the factor K associated with the SAD between the reference blocks indicated by another motion vector candidate pair (e.g., Δmv = (i,j)).

[0155]

[0174] In addition to SAD, other functions that determine the difference between two reference blocks indicated by a pair of MVs {mv L0(i,j) ,mv L1(i,j)}, such as SATD, SSE, variance, partial SAD, partial SSE, partial SATD, etc., can be used.

[0156]

[0175] In fractional sample offset search in DMVR, the candidate MV pairs selected in the integer sample offset search step may be further refined. To reduce the computational complexity, instead of additional search (e.g., SAD comparison) that compares the differences between two reference blocks, fractional sample refinement can be derived using a parametric error surface equation. Fractional sample refinement may be conditionally activated based on the output of the integer sample search stage. If the integer sample search stage ends with the center having the minimum SAD, fractional sample refinement may be further applied. In parametric error surface-based sub-pixel offset estimation, the center position cost (e.g., E(0,0)) and the costs at four neighboring positions from the center (e.g., E(-1,0), E(1,0), E(0,-1), E(0,1)) can be used to fit a 2-D parabolic error surface equation as follows.

[0157]

Number

[0158]

Number

[0176] The fractional positions x calculated based on Eqs. 7-8 min and y min are from -1 / 2-pel to 1 / 2-pel and may be in. In one example, 1 / 2-pel is a half pixel. The calculated fractional positions (x min , y min ) can be added to the integer-distance refined MVs to obtain the sub-pixel accuracy refinement Δmv.

[0159]

[0177] In some examples, such as in VVC, x min and y min values are automatically constrained to be between -8 and 8 because all cost values are positive and the minimum value is E(0,0), which corresponds to a 1 / 2-pel offset (e.g., from -1 / 2-pel to 1 / 2-pel) with 1 / 16-pel (1 / 16 pixel) MV accuracy.

[0160]

[0178] When DMVR is applied to a block or sub-block, two steps including integer-precision motion search and subsequent fractional motion search may be executed in DMVR.

[0161]

[0179] The bi-directional optical flow (BDOF) in VVC was previously called BIO in JEM. In one example, compared to the JEM version, the BDOF in VVC is a simple version that requires fewer operations, especially with respect to the number of multiplications and the size of the multiplier.

[0162]

[0180] BDOF can be used to refine the bi-prediction signal of a CU at the sub-block (e.g., 4×4 sub-block) level. BDOF can be applied to a CU when the CU meets the following conditions: (1) When the CU is coded using the "true" bi-prediction mode, i.e., when one of the two reference pictures is before the current picture in the display order and the other is after the current picture in the display order, (2) When the distances (e.g., POC differences) from the two reference pictures to the current picture are the same, (3) When both reference pictures are short-term reference pictures, (4) When the CU is not coded using the affine mode or the SbTMVP merge mode, (5) When the CU has more than 64 luma samples, (6) When both the CU height and the CU width are 8 luma samples or more, (7) When the BCW weight index indicates equal weights, (8) When the weighted position (WP) is not enabled for the current CU, and (9) When the CIIP mode is not used for the current CU.

[0163]

[0181] In one example, BDOF may be applied only to the luma component. The BDOF mode can be based on the concept of optical flow that assumes smooth motion of objects. For each sub-block (e.g., 4×4 sub-block), the motion refinement (v x , v y ) can be calculated by minimizing the difference between the L0 and L1 prediction samples. Then, the motion refinement can be used to adjust the bi-prediction sample values within the sub-block. BDOF can include the following steps:

[0182] First, the horizontal and vertical gradients of the two prediction signals from reference list L0 and reference list L1

[0164]

Number

[0165]

Number

[0166]

[0183] Then, the autocorrelation and cross-correlation S1, S2, S3, S5, S6 of the gradient can be calculated according to the following Eqs. (11)-(15):

[0167]

Number

[0168]

Number

[0169]

[0184] Motion refinement (v x , v y ) can then be derived using cross - correlation and auto - correlation as follows using Eqs. (19) and (20):

[0170]

Equation

[0171]

Equation

[0172]

Equation

[0173]

Equation

[0185] Finally, the CU's BDOF sample can be calculated as follows by adjusting the dual - prediction sample of Eq. (22):

[0174]

Equation

[0175]

[0186] To derive the gradient value, some prediction samples I in list k (k = 0, 1) outside the current CU boundary (k)(i,j) needs to be generated. As shown in FIG. 19, for BDOF in VVC, one extended row / column (1902) can be used around the boundary (1906) of CU (1904). To suppress the computational complexity of generating prediction samples outside the boundary, prediction samples in the extended area (e.g., the non-shaded area in FIG. 19) can be directly generated by obtaining reference samples at neighboring integer positions (e.g., using the floor() operation for coordinates) without interpolation, and prediction samples can be generated within the CU (e.g., the shaded area in FIG. 19) using a normal 8-tap motion compensation interpolation filter. The extended sample values can be used only for gradient calculation. For the remaining steps of the BDOF process, when samples and gradient values outside the CU boundary are used, those samples and gradient values can be padded (e.g., repeated) from the nearest neighbors of the samples and gradient values.

[0176]

[0187] In one embodiment, it is possible to apply BDOF such as sample-based BDOF. An example of sample-based BDOF is applied in ECM. In sample-based BDOF, as described in Eqs. 19-20, instead of deriving motion refinement (v x ,v y ) on a block or sub-block basis, it is possible to determine motion refinement (v x ,v y ) for each sample. In one example, a coding block is divided into sub-blocks (e.g., 8×8 sub-blocks). For each sub-block, whether to apply BDOF can be determined by comparing the difference (e.g., SAD) between two reference sub-blocks with a threshold. If it is determined that BDOF is to be applied to a sub-block, it is possible to use a sliding window (e.g., a sliding 5×5 window) for each sample within the sub-block, and the existing BDOF process as described above is applied for each sliding window to obtain motion refinement v xand v y can be derived. The derived motion refinement (v x , v y ) can be applied to adjust the bi-predicted sample values for a sample (e.g., the central sample of a sliding window).

[0177]

[0188] In one embodiment, such as in ECM, it is possible to apply MP-DMVR (multi-pass decoder-side motion vector refinement, MP-DMVR). In the first pass, bilateral matching is applied to the coding block. In the second pass, BM is applied to each M1×N1 sub-block (e.g., 16×16 sub-block) within the coding block. In the third pass, the MVs within each M2×N2 sub-block (e.g., 8×8 sub-block) can be refined by applying BDOF. In one example, M1 is greater than or equal to M2, and N1 is greater than or equal to N2. The refined MVs can be stored for both spatial motion vector prediction and temporal motion vector prediction.

[0178]

[0189] In the first pass, block-based BM MV refinement is applied to the coding block. The refined MV (or refined MV pair) of the coding block can be derived by applying BM to the coding block. Similar to DMVR, in the bi-prediction operation, the refined MVs may be searched around two initial MVs (MV0 and MV1) in reference picture lists L0 and L1 respectively. The refined MVs (MV0_pass1 and MV1_pass1) of the first pass can be derived around the initial MVs based on the minimum BM cost between two reference blocks in L0 and L1.

[0179]

[0190] BM can perform a local search to derive integer sample precision MVs (intDeltaMV). The local search can loop through a search range [-sHor, sHor] in the horizontal direction and a search range [-sVer, sVer] in the vertical direction by applying a search pattern (e.g., a 3×3 square search pattern). The values of sHor and sVer can be determined by the block size. In one example, the maximum values of sHor and sVer are 8.

[0180]

[0191] The BM cost (bilCost) can be calculated based on the difference (e.g., SAD) between two reference blocks in L0 and L1. In one example, bilCost = mvDistanceCost + sadCost where the parameter sadCost represents the difference (e.g., SAD) between two reference blocks in L0 and L1, and the parameter mvDistanceCost represents the cost related to signaling overhead (e.g., rate cost). When the block size cbW×cbH is larger than 64, the mean removed SAD (MRSAD) cost function can be applied to remove the DC effect of the distortion between two reference blocks. cbW and cbH are the block width and block height. When the bilCost at the center point of the search pattern (e.g., a 3×3 square search pattern) has the minimum cost, the local search (e.g., intDeltaMV local search) can end. Otherwise, the current minimum cost search point becomes the new center point of the search pattern (e.g., a 3×3 square search pattern), and the search for the minimum cost can continue until the search reaches the end of the search range.

[0181]

[0192] Existing fractional sample refinement may be further applied to derive the final motion refinement (e.g., the final delta MV). The refined MV after the first pass (e.g., the output of the first pass) can be derived as follows:

[0182]

Number

[0193] In the second pass, sub-block-based BM MV refinement can be performed, and by applying BM to sub-blocks (e.g., 16×16 grid sub-blocks), refined MV pairs can be derived. For each sub-block, the refined MV pairs can be searched around two MVs (MV0_pass1 and MV1_pass1) obtained from the first pass in reference picture lists L0 and L1 respectively. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) can be derived based on the minimum BM cost between two reference sub-blocks in L0 and L1. The parameter sbIdx2 indicates the index of the sub-block

[0194] For each sub-block, BM can perform a full search to derive the integer sample refined MV (intDeltaMV) of the second pass. The full search can have a horizontal search range [-sHor, sHor] and a vertical search range [-sVer, sVer]. The values of sHor and sVer are determined by the block size. In one example, the maximum values of sHor and sVer are 8

[0183]

[0195] The BM cost (bilCost) is bilCost = satdCost × costFactor It can be calculated by applying a cost coefficient (costFactor) to a cost function (e.g., the SATD cost shown as satdCost) between two reference sub-blocks. The search area (e.g., (2×sHor+1)×(2×sVer+1)) (2000) can be divided into a plurality (e.g., 5) of search regions (e.g., 5 diamond-shaped search regions) shown in FIG. 20. In the example of FIG. 20, the search area (2000) has a square or diamond shape. The five search regions (2001)-(2005) include a central region (2001), three diamond-shaped search regions (2002)-(2004), and an outer region (2005).

[0184]

[0196] Each search region can be assigned a cost coefficient (costFactor) that can be determined by the distance (e.g., intDeltaMV) between each search point and the starting MV. The search regions (2001)-(2005) can be processed in order starting from the center of the search area (e.g., (2001)) towards the outer boundary of the search area (e.g., (2005)). For example, the search regions are processed in the order of (2001) (e.g., processed first), (2002), (2003), (2004), and (2005) (e.g., processed last). In each search region, the search points can be processed in a suitable order such as a raster scan order (e.g., from the upper left to the lower right corner of the search region). When the minimum BM cost (bilCost) in the current search region is smaller than a threshold value (e.g., a value equal to the area sbW×sbH of the sub-block), the integer sample refinement (int-pel) full search may end. Otherwise, the int-pel full search continues to the next search region until all search points within the search area (2000) are inspected. In one example, further, when the difference between the previous minimum cost and the current minimum cost in an iteration is smaller than a threshold value (e.g., a value equal to the area of the block), the search process ends.

[0185]

[0197] DMVR fractional sample refinement as used in VVC can be further applied to derive the final deltaMV(sbIdx2). The refined MV in the second pass can be derived as follows:

[0186]

Number

[0198] In the third pass, sub-block-based BDOF MV refinement can be applied, and the refined MV pair can be derived by applying BDOF to M2×N2 sub-blocks (e.g., 8×8 grid sub-blocks). For each M2×N2 sub-block (e.g., 8×8 sub-block), starting from the refined MV pair of the parent sub-block in the second pass, without clipping, scaled v x and v y BDOF refinement can be applied to derive. The derived bioMv(v x , v y ) can be rounded to 1 / 16 pel accuracy and clipped between -32 and 32.

[0187]

[0199] The refined MVs in the third pass (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) are derived as follows:

[0188]

Number

[0200] The parameter sbIdx3 indicates the index of the sub-block in the third pass.

[0189]

[0201] Paths in MP-DMVR can be omitted, modified, or replaced with another BM MV refinement method. For example, when the block size is M1×N1 or less, the first path or the second path is omitted. Additional path(s) may be included in MP-DMVR. For example, sample-based BDOF may be executed after the third path. In one embodiment, MP-DMVR includes a BM MV refinement method and BDOF. In one embodiment, MP-DMVR includes DMVR and BDOF.

[0190]

[0202] In related art, certain coding tools such as BM-based MV refinement and BDOF (e.g., sub-block-based BDOF, sample-based BDOF, combinations of sub-block-based BDOF and sample-based BDOF, or derivatives thereof) are not supported in the SbTMVP mode. The lack of sub-block MV refinement using bi-directional prediction in SbTMVP may have lower coding efficiency compared to other coding tools that use BM-based MV refinement (e.g., including MP-DMVR) and / or BDOF (e.g., including sample-based BDOF).

[0191]

[0203] This disclosure describes embodiments related to applying DMVR (or a variant of DMVR) and pixel-domain BDOF to sub-block-level temporal motion vectors determined using the SbTMVP mode.

[0192]

[0204] In one embodiment, a current block (or coding block) within a current picture is coded using the SbTMVP mode. The current block may include a plurality of sub-blocks. The motion information of each sub-block within the current block can be determined based on the SbTMVP mode as described in FIGS. 16-17. In one example, the SbTMVP mode is executed without DVO. For example, DV (1721) is determined from the MVs of neighboring blocks of the current block. In one example, the SbTMVP mode is executed using DVO, and the MVs of neighboring blocks of the current block are changed by DVO to generate an updated DV as DV (1721) in FIG. 17. The SbTMVP mode may be referred to as the SbTMVP-MMVD mode when DVO is signaled based on the MMVD mode.

[0193]

[0205] The first sub-block within the current block is predicted in bi-prediction, and the motion information of the sub-block can be determined based on the SbTMVP mode. In one example, the motion information includes an initial MV pair such as a first MV and a second MV of the sub-block. According to an embodiment of the present disclosure, BM-based MV refinement and / or BDOF can be applied to the first sub-block to refine the motion information of the first sub-block determined based on the SbTMVP mode. In one example, when a refinement condition (or inspection condition) is satisfied, BM-based MV refinement and / or BDOF are applied to the first sub-block to refine the motion information of the first sub-block.

[0194]

[0206] In one example, the size of the first sub-block is M sub ×N sub where M sub and N sub can be positive integers. In one example, when one of M sub and N sub is 1, the other of M sub and N sub is greater than 1.

[0195]

[0207] BM-based MV refinement can include any MV refinement method based on bilateral matching, or any suitable combination of any BM-based MV refinement methods, for example, DMVR applied to a coding block, DMVR applied to sub-blocks within a coding block, MP-DMVR described above, derivatives of MP-DMVR, where BDOF and at least DMVR are included. BDOF can refer to sub-block-based BDOF applied to each sub-block within a coding block, sample-based BDOF applied to each sample within a sub-block, a combination of sub-block-based BDOF and sample-based BDOF, or a suitable variant form.

[0196]

[0208] In one example, when the refinement condition is satisfied (e.g., true), BM-based MV refinement and / or BDOF is applied to the first sub-block. Otherwise, when the refinement condition is not satisfied (e.g., false), BM-based MV refinement and BDOF are not implicitly applied. For example, when the refinement condition is not satisfied, signaling is not required to indicate that BM-based MV refinement and BDOF should not be applied to the sub-block.

[0197]

[0209] The sub-block can be selected from a plurality of sub-blocks based on the refinement condition.

[0198]

[0210] In one embodiment, when the refinement condition is true for the selected sub-block, BM-based MV refinement can be applied to each of the selected sub-blocks. In one example, the selected sub-block includes a first sub-block. The refined MV pair (e.g., BM-refined MV pair) associated with the first sub-block can be determined based on BM-based MV refinement that uses the motion information (e.g., initial MV pair) of the first sub-block as an input to the BM-based MV refinement. The initial MV pair of the first sub-block is derived based on the SbTMVP mode. The BM-refined MV pair (e.g., the output from the BM-based MV refinement) associated with the first sub-block can be applied as the final MV pair to predict the sub-block using inter prediction.

[0199]

[0211] In one embodiment, the BM-based MV refinement is DMVR. DMVR can be applied to the area of the first sub-block to determine the refined MV pair of the area. This area may be below the first sub-block. If the area is smaller than the first sub-block, DMVR can be applied to each area of the first sub-block to determine the refined MV pair of each area.

[0200]

[0212] In one example, when the sub-block size of the first sub-block is below a threshold (e.g., M D ×N D ), DMVR is applied to the first sub-block (e.g., the entire first sub-block) to determine the refined MV pair of the first sub-block based on the initial MV pair. When the sub-block size is larger than the threshold, the first sub-block includes a plurality of areas. DMVR is applied to each of the areas to determine the refined MV pair of each area.

[0201]

[0213] In one embodiment, BM-based MV refinement is MP-DMVR. In one example, the sub-block size of the first sub-block is larger than a threshold (e.g., M1×N1). The first pass (e.g., DMVR) is applied to the first sub-block to determine the first refined MV pair of the first sub-block. The first sub-block may include a plurality of first areas. The second pass (e.g., DMVR) can be applied to each of the first areas to determine the second refined MV pair of each of the first areas based on the first refined MV pair. The third pass (e.g., BDOF) can be applied to a second area in one of the first areas to determine the third refined MV pair of the second area based on the second refined MV pair. The second area (or its area) may be less than or equal to the first area (or its area).

[0202]

[0214] In one embodiment, BM-based MV refinement is MP-DMVR. The first pass (e.g., DMVR) is applied to the first sub-block to determine the first refined MV pair of the first sub-block. Based on the first refined MV pair, the second pass (e.g., BDOF) can be applied to a first area within the first sub-block to determine the second refined MV pair of the first area. The first area may be less than or equal to the first sub-block.

[0203]

[0215] In one embodiment, if the refinement condition is true for the selected sub-block(s), BDOF may be applied to each of the selected sub-blocks. In one example, the selected sub-block includes a first sub-block. The refined MV pairs (e.g., BDOF-refined MV pairs) associated with the first sub-block can be determined based on the BDOF and the initial MV pairs of the first sub-block. The BDOF-refined MV pairs associated with the first sub-block can be applied as the final MV pairs to predict the first sub-block using inter prediction.

[0204]

[0216] In one example, the BDOF includes sample-based BDOF. In one example, the input to the sample-based BDOF is the initial MV pair of the first sub-block. The sample-based BDOF can be applied to each sample within the first sub-block to refine the initial MV pair of the first sub-block. For example, the motion refinement (v x , v y ) of each sample is determined, and the refined MV pairs of each sample can be determined based on the motion refinement (v x , v y ) and the initial MV pair. The BDOF-refined MV pair(s) associated with the first sub-block may include the refined MV pairs of each sample within the first sub-block. For example, if the first sub-block includes 16 samples, the BDOF-refined MV pairs associated with the first sub-block include 16 refined MV pairs of the respective 16 samples.

[0205]

[0217] Alternatively, the sample-based BDOF is applied after the sub-block-based BDOF, and the input to the sample-based BDOF includes the refined MV pairs associated with the first sub-block.

[0206]

[0218] In one example, the refined MV pairs associated with the first sub-block (e.g., BM-refined MV pairs and / or BDOF-refined MV pairs) can be used, for example, for motion memory for spatial and / or temporal motion vector prediction. In one example, the refined MV pairs for each sample within the first sub-block predicted in sample-based BDOF may be stored. For example, the refined MV pairs associated with the first sub-block can be stored in a motion buffer (e.g., a motion information buffer that stores motion information) as predictors (e.g., spatial candidates, temporal candidates, history-based candidates, and / or the like) for another block or another sub-block within the current picture including the current block or in a different picture.

[0207]

[0219] In one embodiment, when the refinement condition is true for the selected sub-block, the following processes can be sequentially executed. BM-based MV refinement can be applied to the first sub-block to obtain BM-refined MV pairs. Then, BDOF (e.g., sample-based BDOF) can be applied to each sample based on the BM-refined MV pairs to obtain the final MV pairs of the first sub-block. In one example, when both BDOF and BM-based MV refinement are performed, the following order is required: BDOF follows BM-based MV refinement.

[0208]

[0220] In one embodiment, syntax information (e.g., a first flag) is signaled to indicate whether BM-based MV refinement is applied using the SbTMVP mode (e.g., SbTMVP with DVO or SbTMVP without DVO). In one example, the syntax information (e.g., the first flag) is signaled at a high level, e.g., at a level higher than the CU level. The syntax information (e.g., the first flag) may be signaled at the CTU level, slice header, picture header, sequence parameter set (SPS), picture parameter set (PPS), or the like.

[0209]

[0221] Syntax information (e.g., a second flag) may be signaled to indicate whether BDOF (e.g., sample-based BDOF) is applied using the SbTMVP mode (e.g., SbTMVP with DVO or without DVO). In one example, the syntax information (e.g., the second flag) is signaled at a high level, e.g., at a level higher than the CU level. The syntax information (e.g., the second flag) may be signaled at the CTU level, slice header, picture header, sequence parameter set (SPS), picture parameter set (PPS), or the like.

[0210]

[0222] In one example, the first flag and the second flag are the same flag. In one example, the first flag and the second flag are different flags.

[0211]

[0223] In one embodiment, the refinement condition may depend on the POC of the reference picture and the current picture. In one example, the refinement condition includes (i) a first reference picture that is before the current picture in the display order (e.g., related to the first MV of the first sub-block), and a second reference picture that is after the current picture in the display order (e.g., related to the second MV of the first sub-block); and (ii) the distances from the first reference picture and the second reference picture to the current picture are the same. In one example, when the above refinement condition is satisfied, the first sub-block is coded in the true bi-prediction mode. For example, the first sub-block is predicted by bi-directional prediction. The MV pair (e.g., the initial MV pair) indicates two different reference pictures having equal differences in POC with respect to the current picture, and the two reference pictures are in different temporal directions.

[0212]

[0224] In one embodiment, the refinement condition may depend on the POC of the reference picture and the co-located reference picture in the SBTMVP mode. Referring to FIG. 17, in order to apply BM-based MV refinement and / or BDOF using the SbTMVP mode to the first sub-block, the MV fetched from the related sub-block (e.g., (1732)) within the co-located reference picture (e.g., (1712)) in the SbTMVP mode can be examined. Depending on the MV from the related sub-block within the co-located reference picture, it is determined whether to apply BM-based MV refinement and / or BDOF using the SbTMVP mode for the first sub-block.

[0213]

[0225] In one example, the refinement conditions are: (i) the MVs fetched from the relevant sub-blocks in the reference picture at the same position are the MV pairs used for bi-prediction; and (ii) the MV pairs point to two different reference pictures having equal differences in POC with respect to the reference picture at the same position, and the two reference pictures are in different temporal directions of the picture at the same position. For example, the first reference picture is before the reference picture at the same position in the display order, the second reference picture is after the reference picture at the same position in the display order, and the distances from the first reference picture and the second reference picture to the reference picture at the same position are the same.

[0214]

[0226] In one embodiment, the current block (e.g., a complete or whole current block) is refined based on the motion information (e.g., MVs) of one or more available bi-predicted sub-blocks that satisfy the refinement conditions.

[0215]

[0227] As described above, the motion information of each sub-block within the current block can be determined based on the SbTMVP mode as described in FIGS. 16-17. The motion information of the bi-predicted sub-blocks within the current block can be inspected to determine whether the refinement conditions are satisfied for each respective bi-predicted sub-block. One or more bi-predicted sub-blocks that satisfy the refinement conditions are referred to as one or more available bi-predicted sub-blocks. The motion information of one of the one or more available bi-predicted sub-blocks can be used as an input to BM-based MV refinement (e.g., a DMVR process such as in VVC or a variant, MP-DMVR or a variant, etc.) to refine the motion information of the current block (e.g., the entire current block), and thus, to refine the inter-predictor (e.g., the predicted sample value) of the current block.

[0216]

[0228] In one embodiment, one or more available dual-prediction sub-blocks include a plurality of dual-prediction sub-blocks. The motion information of each of the plurality of dual-prediction sub-blocks is applied as the initial motion information of the current block (e.g., the entire current block), and can then be used as an input for BM-based MV refinement. The BM cost associated with the motion information of each dual-prediction sub-block can be determined by applying BM-based MV refinement to the entire current block. For example, each of the available dual-prediction sub-block MVs can be individually applied to the entire current block as the initial MV (e.g., initial MV pair) in the same manner as or in the same way as DMVR or MP-DMVR. The motion information (e.g., MV pair) of the dual-prediction sub-block with the minimum bilateral matching cost may be selected based on the BM cost. The motion information (e.g., MV pair) of the selected dual-prediction sub-block may be applied to the entire current block to determine the inter-predictor.

[0217]

[0229] In one embodiment, syntax information (e.g., block-level syntax) can be signaled to indicate, for example, among the motion information of one or more available dual-prediction sub-blocks, which sub-block motion information (e.g., MV(s)) can be used as the initial motion information (e.g., initial MV(s)) for refining the entire current block. The refinement process may be the same as or identical to a DMVR process such as in VVC or a variant, MP-DMVR or a variant, or the like. The motion information (e.g., MV(s)) of the sub-block with the minimum bilateral matching cost may be applied to the entire current block.

[0218]

[0230] In one embodiment, for all blocks BM-based MV refinement (e.g., DMVR, MP-DMVR) that use the motion information of sub-blocks as the initial MV pair, syntax information (e.g., block-level syntax such as a flag) may be used to indicate whether it is possible to be enabled, for example, for the current block.

[0219]

[0231] FIG. 21 shows a flowchart for outlining an encoding process (2100) according to an embodiment of the present disclosure. The process (2100) can be used in a video / image encoder. The process (2100) can be executed by an apparatus for video / image coding that can include a processing circuit. In various embodiments, the process (2100) is executed by a processing circuit such as a processing circuit in terminal devices (310), (320), (330), and (340), a processing circuit that executes the functions of video encoders (e.g., (403), (603), (703)). In some embodiments, the process (2100) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (2100). The process starts at (S2101) and proceeds to (S2110).

[0220]

[0232] In (S2110), it is possible to determine the motion information of sub-blocks within a plurality of sub-blocks of the current block based on the sub-block-based temporal motion vector prediction (SbTMVP) mode. The sub-blocks are bi-predicted using reference sub-blocks in the first reference picture at L0 and the second reference picture at L1, respectively. The motion information can include an initial pair of MVs indicating the reference sub-blocks.

[0221]

[0233] In (S2120), at least one of (i) bilateral matching (BM)-based motion vector (MV) refinement or (ii) bi-directional optical flow (BDOF) mode can be applied to a sub-block to update the motion information of the sub-block.

[0222]

[0234] In one example, BM-based MV refinement is applied to determine the refined motion information of the sub-block. BM-based MV refinement can include (i) decoder-side motion vector refinement (DMVR) or a variant thereof, or (ii) multi-pass decoder-side motion vector refinement (MP-DMVR) or a variant thereof.

[0223]

[0235] In (S2130), the sub-block can be encoded based on the updated motion information (or refined motion information).

[0224]

[0236] In one example, prediction information indicating that at least one of (i) BM-based MV refinement or (ii) BDOF mode is applied to the sub-block is encoded and included in the bitstream.

[0225]

[0237] In one example, the prediction information includes a flag indicating that at least one of (i) BM-based MV refinement or (ii) BDOF mode is applied to the sub-block. The flag can be signaled in high-level syntax.

[0226]

[0238] The process (2100) then proceeds to (S2199) and ends.

[0227]

[0239] The process (2100) can be appropriately applied to various scenarios, and the steps in the process (2100) can be correspondingly adjusted. One or more of the steps in the process (2100) can be adapted, omitted, repeated, and / or combined. Any appropriate order can be used to implement the process (2100). It is possible to add additional steps.

[0228]

[0240] In one example, the motion information includes the initial MV pair of the sub-block. BM-based MV refinement includes DMVR. DMVR can be applied to an area within the sub-block to determine a refined MV pair for the area based on the initial MV pair. This area can be less than or equal to the area of the sub-block. The area within the sub-block can be reconstructed based on the refined MV pair.

[0229]

[0241] In one example, BM-based MV refinement includes MP-DMVR. When the sub-block size of the sub-block is larger than a first threshold (e.g., M1×N1), at least one DMVR may be applied to the sub-block to determine a first refined MV pair for the sub-block. Subsequently, BDOF can be applied to an area within the sub-block to determine a second refined MV pair for the area based on the first refined MV pair. This area can be less than or equal to the area of the first sub-block.

[0230]

[0242] In one example, the BDOF mode is applied to each sample within the sub-block to determine the refined MV pair for each sample. The BDOF mode includes a sample-based BDOF mode. Each sample within the sub-block can be reconstructed based on the refined MV pair for each sample.

[0231]

[0243] In one example, the refined motion information of the sub-blocks includes one or more first refined MV pairs for each of one or more areas within the sub-blocks. After applying BM-based MV refinement, it is possible to apply a BDOF mode including a sample-based BDOF mode. The refined MV pairs of each sample in the area within one or more areas can be determined based on the BDOF mode and the first refined MV pairs corresponding to the area. Each sample within the area can be reconstructed based on the refined MV pairs of the respective samples.

[0232]

[0244] In one example, BM-based MV refinement and BDOF are applied based on a first reference picture and a second reference picture of the current picture. The first reference picture is before the current picture in the display order, and the second reference picture is after the current picture in the display order. The distances from the first reference picture and the second reference picture to the current picture are the same.

[0233]

[0245] FIG. 22A shows a flowchart for outlining a decoding process (2200A) according to an embodiment of the present disclosure. The process (2200A) can be used in a video / image decoder. The process (2200A) can be executed by an apparatus for video / image coding that can include a receiving circuit and a processing circuit. In various embodiments, the process (2200A) is executed by a processing circuit such as a processing circuit in terminal devices (310), (320), (330), and (340), a processing circuit that executes the function of video encoder (403), a processing circuit that executes the function of video decoder (410), a processing circuit that executes the function of video decoder (510), a processing circuit that executes the function of video encoder (603), and the like. In some embodiments, the process (2200A) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (2200A). The process starts at (S2201) and proceeds to (S2210).

[0234]

[0246] In (S2210), the prediction information of the current block in the current picture can be decoded from the coded bitstream (e.g., the coded video bitstream). The current block includes a plurality of sub-blocks that are reconstructed based on the sub-block-based temporal motion vector prediction (SbTMVP) mode.

[0235]

[0247] In (S2220), based on the SbTMVP mode, it is possible to determine the motion information of the sub-blocks within the plurality of sub-blocks. The sub-blocks are bi-predicted as described above.

[0236]

[0248] In (S2230), at least one of (i) bilateral matching (BM)-based motion vector (MV) refinement or (ii) the bi-directional optical flow (BDOF) mode can be applied to the sub-blocks to update the motion information of the sub-blocks.

[0237]

[0249] BM-based MV refinement can be applied to determine the refined motion information of the sub-blocks. BM-based MV refinement can include decoder-side motion vector refinement (DMVR) or multi-pass decoder-side motion vector refinement (MP-DMVR).

[0238]

[0250] In (S2240), the sub-blocks can be reconstructed based on the updated motion information (or refined motion information).

[0239]

[0251] The process (2200A) proceeds to (S2299) and ends.

[0240]

[0252] The process (2200A) can be appropriately applied to various scenarios, and the steps in the process (2200A) can be correspondingly adjusted. One or more of the steps in the process (2200A) can be adapted, omitted, repeated, and / or combined. Any appropriate order can be used to implement the process (2200A). It is possible to add additional steps.

[0241]

[0253] The motion information can include the initial MV pair of the sub-block.

[0242]

[0254] In one example, BM-based MV refinement includes DMVR. DMVR can be applied to an area within a sub-block to determine a refined MV pair for the area based on the initial MV pair. This area may be less than or equal to the area of the sub-block. The area within the sub-block can be reconstructed based on the refined MV pair.

[0243]

[0255] In one example, BM-based MV refinement includes MP-DMVR. When the sub-block size of the sub-block is larger than the first threshold M1×N1, at least one DMVR is applied to the sub-block to determine the first refined MV pair of the sub-block. BDOF can be applied to an area within the sub-block to determine a second refined MV pair for the area based on the first refined MV pair. This area is less than or equal to the area of the first sub-block.

[0244]

[0256] In one example, in the BDOF mode, a BDOF mode including a sample-based BDOF mode is applied to each sample within the sub-block to determine a refined MV pair for each sample. Each sample within the sub-block is reconstructed based on the refined MV pair for each sample.

[0245]

[0257] In one example, the refined motion information of the sub-blocks includes one or more first refined MV pairs for each of one or more areas within the sub-blocks. After applying BM-based MV refinement, a BDOF mode including a sample-based BDOF mode is applied to determine the refined MV pairs of each sample in the areas within the one or more areas based on the first refined MV pairs corresponding to the areas. Each sample within the area can be reconstructed based on the refined MV pairs of the respective samples.

[0246]

[0258] The prediction information can indicate that at least one of (i) BM-based MV refinement or (ii) the BDOF mode is applied to the sub-blocks.

[0247]

[0259] The prediction information can include a flag indicating that at least one of (i) BM-based MV refinement or (ii) the BDOF mode is applied to the sub-blocks. The flag can be signaled in a high-level syntax (e.g., SPS, PPS).

[0248]

[0260] BM-based MV refinement and BDOF are applied based on a first reference picture or a second reference picture of the current picture. The first reference picture is before the current picture in the display order, and the second reference picture is after the current picture in the display order. The distances from the first reference picture and the second reference picture to the current picture are the same.

[0249]

[0261] Figure 22B shows a flowchart for outlining a process (e.g., a decoding process) (2200B) according to an embodiment of the present disclosure. The process (2200B) is a variation of the process (2200A). The process (2200B) can be used in a video / image decoder. The process (2200B) can be executed by a device for video / image coding that can include a receiving circuit and a processing circuit. In various embodiments, the process (2200B) is executed by a processing circuit such as a processing circuit in terminal devices (310), (320), (330), and (340), a processing circuit that executes the function of video encoder (403), a processing circuit that executes the function of video decoder (410), a processing circuit that executes the function of video decoder (510), a processing circuit that executes the function of video encoder (603), and the like. In some embodiments, the process (2200B) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (2200B). The process starts at (S2202) and proceeds to (S2212).

[0250]

[0262] At (S2212), a coded bitstream including a current block in a current picture is received, and the current block includes a plurality of sub-blocks.

[0251]

[0263] At (S2222), prediction information indicating whether the current block is coded in a sub-block-based temporal motion vector prediction (SbTMVP) mode is obtained.

[0252]

[0264] At (S2232), when the current block is coded in the SbTMVP mode, it is determined whether sub-blocks within the plurality of sub-blocks of the current block are bi-predicted. When a sub-block is bi-predicted, motion information of the sub-block is determined based on the SbTMVP mode.

[0253]

[0265] In (S2242), apply at least one of (i) bilateral matching (BM)-based motion vector (MV) refinement and (ii) bi-directional optical flow (BDOF) mode to the sub-block to refine the motion information of the sub-block.

[0254]

[0266] In (S2252), it is possible to reconstruct the current block based on the refined motion information corresponding to one or more sub-blocks within the plurality of sub-blocks of the current block. The refined motion information corresponding to one or more sub-blocks includes the refined motion information of the sub-blocks.

[0255]

[2667] The process (2200B) proceeds to (S2292) and ends.

[0256]

[2668] The process (2200B) can be appropriately applied to various scenarios, and the steps in the process (2200B) can be correspondingly adjusted. One or more of the steps in the process (2200B) can be adapted, omitted, repeated, and / or combined. It is possible to use any appropriate order to implement the process (2200B). It is possible to add additional steps.

[0257]

[0269] FIG. 23 shows a flowchart for outlining an encoding process (2300) according to an embodiment of the present disclosure. The process (2300) can be used in a video / image encoder. The process (2300) can be executed by an apparatus for video / image coding that can include a processing circuit. In various embodiments, the process (2300) is executed by a processing circuit such as a processing circuit within terminal devices (310), (320), (330), and (340), a processing circuit that executes the functions of video encoders (e.g., (403), (603), (703)). In some embodiments, the process (2300) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (2300). The process starts at (S2301) and proceeds to (S2310).

[0258]

[0270] At (S2310), the motion information of each of one or more sub-blocks can be determined based on a sub-block-based temporal motion vector prediction (SbTMVP) mode. The one or more sub-blocks can be bi-predicted sub-blocks among a plurality of sub-blocks within a current block.

[0259]

[0271] At (S2320), based on an initial motion vector (MV) pair of a current block, bilateral matching (BM)-based motion vector (MV) refinement can be applied to the current block to determine updated (or refined) motion information of the current block. The initial MV pair can be the motion information of one of the one or more sub-blocks.

[0260]

[0272] At (S2330), the current block can be encoded based on the updated motion information of the current block.

[0261]

[0273] Then, the process (2300) proceeds to (S2399) and ends.

[0262]

[0274] Process (2300) can be appropriately applied to various scenarios, and the steps in process (2300) can be adjusted accordingly. One or more of the steps in process (2300) can be adapted, omitted, repeated, and / or combined. Any appropriate order can be used to implement process (2300). It is possible to add additional steps.

[0263]

[0275] In one embodiment, one or more sub-blocks include a plurality of bi-prediction sub-blocks. In one example, (i) based on the motion information of each of the plurality of bi-prediction sub-blocks, apply BM-based MV refinement to the current block to determine the bilateral matching cost associated with each bi-prediction sub-block; and (ii) determine the initial MV pair of the current block by determining the motion information of the sub-block having the minimum bilateral matching cost among the bilateral matching costs of each of the plurality of bi-prediction sub-blocks.

[0264]

[0276] In one example, syntax information such as an index is encoded and included in the bitstream to specify the sub-block having the minimum bilateral matching cost.

[0265]

[0277] In one example, prediction information indicating that BM-based MV refinement is applied to the current block based on the initial MV pair, which is the motion information of one of the one or more sub-blocks, is encoded and included in the bitstream.

[0266]

[0278] Figure 24 shows a flowchart for outlining a decoding process (2400) according to an embodiment of the present disclosure. The process (2400) can be used in a video / image decoder. The process (2400) can be executed by an apparatus for video / image coding that can include a receiving circuit and a processing circuit. In various embodiments, the process (2400) is executed by a processing circuit such as a processing circuit in terminal devices (310), (320), (330), and (340), a processing circuit that executes the function of video encoder (403), a processing circuit that executes the function of video decoder (410), a processing circuit that executes the function of video decoder (510), a processing circuit that executes the function of video encoder (603), and the like. In some embodiments, the process (2400) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (2400). The process starts at (S2401) and proceeds to (S2410).

[0267]

[0279] In (S2410), the prediction information of the current block in the current picture can be decoded from the coded bitstream (e.g., the coded video bitstream). The current block includes a plurality of sub-blocks that are reconstructed based on the sub-block-based temporal motion vector prediction (SbTMVP) mode.

[0268]

[0280] In (S2420), based on the SbTMVP mode, it is possible to determine the motion information of each of one or more sub-blocks within the plurality of sub-blocks. One or more sub-blocks can be bi-predicted using, for example, an MV pair.

[0269]

[0281] In (S2430), bilateral matching (BM)-based motion vector (MV) refinement can be applied to a current block based on the initial MV pair of the current block in order to determine updated motion information (or refined motion information) of the current block. The initial MV pair may be motion information of one sub-block among one or more sub-blocks.

[0270]

[0282] In (S2440), the current block can be reconstructed based on the updated motion information of the current block.

[0271]

[0283] Process (2400) proceeds to (S2499) and ends. Process (2400A) can be appropriately applied to various scenarios, and the steps in process (2400) can be adjusted accordingly. One or more of the steps in process (2400) can be adapted, omitted, repeated, and / or combined. Any suitable order can be used to implement process (2400). Additional steps can be added.

[0272]

[0285] In one embodiment, one or more sub-blocks include a plurality of bi-prediction sub-blocks. In one example, the initial MV pair of the current block can be determined by: (i) applying BM-based MV refinement to the current block based on the motion information of each of the plurality of bi-prediction sub-blocks in order to determine the bilateral matching cost associated with each bi-prediction sub-block; and (ii) determining the motion information of the sub-block having the minimum bilateral matching cost among the bilateral matching costs of each of the plurality of bi-prediction sub-blocks as the initial MV pair of the current block.

[0273]

[0286] In one example, the initial MV pair of the current block is determined as the motion information of one sub-block among a plurality of dual-prediction sub-blocks based on the syntax information in the coded bitstream.

[0274]

[0287] In one example, the prediction information indicates that BM-based MV refinement is applied to the current block based on the initial MV pair that is the motion information of one sub-block among one or more sub-blocks.

[0275]

[0288] The embodiments in the present disclosure may be used separately or combined in any order. Further, each of the method (or embodiment), encoder, and decoder can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.

[0276]

[0289] The above-described technology can be implemented as computer software using computer-readable instructions and can be physically stored in one or more computer-readable media. For example, FIG. 25 shows a computer system (2500) suitable for implementing a particular embodiment of the disclosed subject matter.

[0277]

[0290] The computer software can be coded using any suitable machine code or computer language that can be the subject of assembly, compilation, linking, or similar mechanisms to create code including instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or instructions that go through interpretation or microcode execution.

[0278]

[0291] The command can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0279]

[0292] The components shown in FIG. 25 for the computer system (2500) are essentially exemplary and are not intended to suggest any limitation as to the scope or functionality of the computer software for implementing the embodiments of the present disclosure. Also, the component configuration should not be construed as having any dependency or requirement regarding any one or combination of the components shown in the exemplary embodiment of the computer system (2500).

[0280]

[0293] The computer system (2500) can include a specific human interface input device. Such a human interface input device can respond to input by one or more human users via, for example, tactile input (e.g., keystrokes, swipes, movement of a data glove), auditory input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not shown). Also, the human interface device can be used to capture certain media that is not necessarily directly related to conscious human input, such as audio (e.g., conversation, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., 2D video, 3D video including stereoscopic pictures).

[0281]

[0294] The input human interface device may include one or more of a keyboard (2501), a mouse (2502), a trackpad (2503), a touch screen (2510), a data glove (not shown), a joystick (2505), a microphone (2506), a scanner (2507), and a camera (2508) (although only one of each is depicted).

[0282]

[0295] The computer system (2500) may also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by a touch screen (2510), a data glove (not shown), a joystick (2505), although there may also be tactile feedback devices that do not serve as input devices), auditory output devices (e.g., speakers (2509), headphones (not shown)), visual output devices (e.g., a screen (2510) including a CRT screen, an LCD screen, a plasma screen, an OLED screen, each of which may or may not have a touch screen input function, each of which may or may not have a tactile feedback function, and some of which may be capable of outputting three-dimensional or higher-dimensional output by means such as two-dimensional visual output, stereoscopic output; virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and a printer (not shown).

[0283]

[0296] The computer system (2500) can also include human-accessible storage devices and associated media such as an optical medium including a CD / DVD ROM / RW (2520) using a medium (2521) such as a CD / DVD, a thumb drive (2522), a removable hard drive or solid state drive (2523), legacy magnetic media such as tapes and floppy disks (not shown), and specialized ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0284]

[0297] One of ordinary skill in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not include a transmission medium, a carrier wave, or other transient signals.

[0285]

[0298] The computer system (2500) can also include an interface to one or more communication networks (2555). The network can be, for example, wireless, wired, or optical. The network can further be related to local, wide area, metropolitan, vehicle industry, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), wired or wireless wide area digital networks for TV (including cable TV, satellite TV, and terrestrial broadcast TV), vehicle industries including CANBus, etc. Certain networks generally require an external network interface adapter attached to a specific general-purpose data port or peripheral bus (2549) (e.g., the USB port of the computer system (2500)); others are generally integrated commonly into the core of the computer system (2500) by attaching to a system bus as described below (e.g., an Ethernet interface is integrated within a PC computer system, and a cellular network interface is integrated within a smartphone computer system). Using any of these networks, the computer system (2500) can communicate with other entities. Such communication can be one-way receive-only (e.g., broadcast TV), one-way transmit-only (e.g., CANbus for certain CANbus devices), or two-way, e.g., for other computer systems using local or wide area digital networks. Specific protocols and protocol stacks can be used for each of those networks and network interfaces as described above.

[0286]

[0299] The foregoing human interface device, human accessible storage device, and network interface can be attached to the core (2540) of a computer system (2500).

[0287]

[0300] The core (2540) can include one or more central processing units (CPUs) (2541), a graphics processing unit (GPU) (2542), a special programmable processing device in the form of a field programmable gate array (FPGA) (2543), a hardware accelerator for specific tasks (2544), a graphics adapter (2550), etc. These devices can be connected via a system bus (2548) together with a read only memory (ROM) (2545), a random access memory (2546), and an internal mass storage device (e.g., an internal non-user accessible hard drive, SSD, etc.) (2547). In some computer systems, the system bus (2548) may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the system bus (2548) of the core or via a peripheral bus (2549). In one example, a screen (2510) can be connected to a graphics adapter (2550). The architecture of the peripheral bus includes PCI, USB, etc.

[0288]

[0301] The CPU (2541), GPU (2542), FPGA (2543), and accelerator (2544) can be combined to execute specific instructions capable of constituting the aforementioned computer code. The computer code can be stored in the ROM (2545) or RAM (2546). Temporary data can be stored in the RAM (2546), while persistent data can be stored, for example, in the internal mass storage (2547). Fast storage and retrieval for any memory device may be made possible by using cache memory, which can be closely associated with one or more CPUs (2541), GPUs (2542), mass storage (2547), ROM (2545), RAM (2546), etc.

[0289]

[0302] A computer-readable medium can have therein computer code for performing various computer-implemented operations. The medium and the computer code can be those specially designed and constructed for the purposes of this disclosure, or they can be of the kind well-known and available to those of ordinary skill in the computer software art.

[0290]

[0303] By way of example, and not limitation, a computer system having an architecture (2500), specifically a core (2540), can provide functionality by a processor (including a CPU, GPU, FPGA, accelerator, etc.) that executes software embodied on one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage as described above, as well as specific storage of the core (2540) of a non-transitory nature such as mass storage (2547) inside the core or ROM (2545). The software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core (2540). The computer-readable media can include one or more memory devices or chips depending on specific needs. The software includes defining data structures stored in RAM (2546) and modifying such data structures according to processes defined by the software, and causing the core (2540) and in particular the processors (including a CPU, GPU, FPGA, etc.) therein to execute specific processes or specific parts of specific processes described herein. Further or alternatively, the computer system can provide functionality as a result of logic wired or otherwise incorporated within a circuit (e.g., an accelerator (2544)), which circuit can execute a specific process or specific part of a specific process described herein instead of or in addition to software. References to software include logic and, if necessary, vice versa. References to computer-readable media can include circuits (such as integrated circuits (ICs)) that store software for execution, circuits that embody logic for execution, or both, as appropriate. The present disclosure encompasses any suitable combination of hardware and software.

[0291] Appendix A: Acronyms JEM: joint exploration model VVC: Versatile Video Coding BMS: Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplementary Enhancement Information VUI: Video Usability Information GOPs: Groups of Pictures TUs: Transform Units PUs: Prediction Units CTUs: Coding Tree Units CTBs: Coding Tree Blocks PBs: Prediction Blocks HRD: Hypothetical Reference Decoder SNR: Signal Noise Ratio CPUs: Central Processing Units GPUs: Graphics Processing Units CRT: Cathode Ray Tube LCD: Liquid-Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile communications LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Areas SSD: solid-state drive IC: Integrated Circuit CU: Coding Unit

[0304] Although the present disclosure has described several exemplary embodiments, there are modifications, substitutions, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it will be understood that although not explicitly illustrated or described herein, those skilled in the art will be able to devise many systems and methods that embody the principles of the present disclosure and thus fall within its spirit and scope.

[0292]

[0305] Supplementary Note (Supplementary Note 1) A video decoding method in a decoder, comprising: receiving a coded bitstream including a current block in a current picture, the current block including a plurality of sub-blocks; obtaining prediction information indicating whether the current block is coded in a sub-block-based temporal motion vector prediction (SbTMVP) mode; determining whether sub-blocks within the plurality of sub-blocks of the current block are bi-predicted in response to the current block being coded in the SbTMVP mode; Determining motion information of the sub-block based on the SbTMVP mode according to the sub-block being bi-predicted; Applying at least one of (i) bilateral matching (BM)-based motion vector (MV) refinement and (ii) bi-directional optical flow (BDOF) mode to the sub-block to refine the motion information of the sub-block; and Reconstructing the current block based on the refined motion information corresponding to one or more sub-blocks within the plurality of sub-blocks of the current block, wherein the refined motion information corresponding to the one or more sub-blocks includes the refined motion information of the sub-block; A method comprising.

[0293] (Appendix 2) In the method according to Appendix 1, The applying step includes determining refined motion information of the sub-block by applying the BM-based MV refinement, and the BM-based MV refinement includes decoder-side motion vector refinement (DMVR) or multi-pass decoder-side motion vector refinement (MP-DMVR); and The reconstructing step includes reconstructing the sub-block based on the refined motion information, a method.

[0294] (Appendix 3) In the method according to Appendix 2, The motion information includes an initial MV pair of the sub-block; The BM-based MV refinement includes the DMVR; The applying step includes applying the DMVR to an area within the sub-block to determine a refined MV pair of the area based on the initial MV pair, and the area is smaller than or equal to the area of the sub-block; and The reconstructing step includes reconstructing an area within the sub-block based on the refined MV pair, a method.

[0295] (Appendix 4) In the method described in Appendix 2, the motion information includes the initial MV pair of the sub-block; the BM-based MV refinement includes the MP-DMVR; when the sub-block size of the sub-block is larger than a first threshold M1×N1, apply at least one DMVR to the sub-block to determine a first refined MV pair of the sub-block; and apply the BDOF mode to an area within the sub-block to determine a second refined MV pair of the area based on the first refined MV pair, where the area is smaller than or equal to the area of the sub-block, method.

[0296] (Appendix 5) In the method described in Appendix 1, the motion information includes the initial MV pair of the sub-block; the applying step includes applying the BDOF mode to each sample of the sub-block to determine a refined MV pair for each sample, where the BDOF mode includes a sample-based BDOF mode; and the reconstructing step includes reconstructing each sample of the sub-block based on the refined MV pair for each sample, method.

[0297] (Appendix 6) In the method described in Appendix 2, the refined motion information of the sub-block includes one or more first refined MV pairs for each of one or more areas within the sub-block; the applying step further includes applying a BDOF mode including a sample-based BDOF mode after applying the BM-based MV refinement, and the refined MV pair for each sample in an area within the one or more areas is determined based on the first refined MV pair corresponding to the area and the BDOF mode; and The step of reconstructing includes the step of reconstructing each sample in the area based on the refined MV pairs of the respective samples, method.

[0298] (Appendix 7) In the method according to Appendix 1, the prediction information indicates that at least one of (i) the BM-based MV refinement or (ii) the BDOF mode is applied to the sub-block, method.

[0299] (Appendix 8) In the method according to Appendix 1, the prediction information includes a flag indicating that at least one of (i) the BM-based MV refinement or (ii) the BDOF mode is applied to the sub-block, method.

[0300] (Appendix 9) In the method according to Appendix 1, the BM-based MV refinement or the BDOF mode is applied based on a first reference picture or a second reference picture of the current picture; the first reference picture is before the current picture in the display order, and the second reference picture is after the current picture in the display order; and the distances from the first reference picture and the second reference picture to the current picture are the same, method.

[0301] (Appendix 10) A video decoding method in a decoder, comprising: decoding prediction information of a current block in a current picture from a coded bitstream, wherein the current block includes a plurality of sub-blocks reconstructed based on a sub-block-based temporal motion vector prediction (SbTMVP) mode, step; determining motion information of each of one or more sub-blocks within the plurality of sub-blocks based on the SbTMVP mode, wherein the one or more sub-blocks are bi-predicted, step; Based on the initial motion vector (MV) pair of the current block, applying bilateral matching (BM)-based MV refinement to the current block to determine updated motion information of the current block, wherein the initial MV pair of the current block is motion information of one of the one or more sub-blocks; and Reconstructing the current block based on the updated motion information of the current block; A method comprising the steps of.

[0302] (Appendix 11) In the method according to Appendix 10, The one or more sub-blocks include a plurality of bi-predicted sub-blocks; and The method includes determining the initial MV pair of the current block Based on the motion information of each of the plurality of bi-predicted sub-blocks, applying the BM-based MV refinement to the current block to determine the bilateral matching cost associated with each of the bi-predicted sub-blocks; and Using the minimum bilateral matching cost among the bilateral matching costs of each of the plurality of bi-predicted sub-blocks to determine the initial MV pair of the current block as the motion information of the sub-block; A method comprising the steps performed by.

[0303] (Appendix 12) In the method according to Appendix 10, The one or more sub-blocks include a plurality of bi-predicted sub-blocks; and The method includes determining the initial MV pair of the current block as the motion information of one of the plurality of bi-predicted sub-blocks based on the syntax information in the coded bitstream.

[0304] (Appendix 13) In the method according to Appendix 10, The method, wherein the prediction information indicates that the BM-based MV refinement is applied to the current block based on an initial MV pair which is motion information of one of the one or more sub-blocks.

[0305] (Appendix 14) A video decoding apparatus including a processing circuit, comprising: The processing circuit is configured to receive a coded bitstream including a current block in a current picture, the current block including a plurality of sub-blocks; The processing circuit is configured to obtain prediction information indicating whether the current block is coded in a sub-block-based temporal motion vector prediction (SbTMVP) mode; The processing circuit is configured to determine whether sub-blocks within the plurality of sub-blocks of the current block are bi-predicted in response to the current block being coded in the SbTMVP mode; The processing circuit is configured to determine motion information of the sub-blocks based on the SbTMVP mode in response to the sub-blocks being bi-predicted; The processing circuit is configured to apply at least one of (i) bilateral matching (BM)-based motion vector (MV) refinement and (ii) bi-directional optical flow (BDOF) mode to the sub-blocks to refine motion information of the sub-blocks; and The processing circuit is configured to reconstruct the current block based on refined motion information corresponding to one or more sub-blocks within the plurality of sub-blocks of the current block, the refined motion information corresponding to the one or more sub-blocks including the refined motion information of the sub-blocks.

[0306] (Appendix 15) In the apparatus according to Appendix 14: The processing circuit is configured to determine refined motion information of the sub-block by applying the BM-based MV refinement, and the BM-based MV refinement includes decoder-side motion vector refinement (DMVR) or multi-pass decoder-side motion vector refinement (MP-DMVR); and The apparatus, wherein the processing circuit is configured to reconstruct the sub-block based on the refined motion information.

[0307] (Appendix 16) In the apparatus according to Appendix 15, the motion information includes an initial MV pair of the sub-block; the BM-based MV refinement includes the DMVR; the processing circuit is configured to apply the DMVR to an area within the sub-block to determine a refined MV pair of the area based on the initial MV pair, the area being smaller than or equal to the area of the sub-block; and The apparatus, wherein the processing circuit is configured to reconstruct an area within the sub-block based on the refined MV pair.

[0308] (Appendix 17) In the apparatus according to Appendix 15, the motion information includes an initial MV pair of the sub-block; the BM-based MV refinement includes the MP-DMVR; in response to the sub-block size of the sub-block being larger than a first threshold M1×N1, applying at least one DMVR to the sub-block to determine a first refined MV pair of the sub-block; and The apparatus, wherein the BDOF mode is applied to an area within the sub-block to determine a second refined MV pair of the area based on the first refined MV pair, the area being smaller than or equal to the area of the sub-block.

[0309] (Supplementary Note 18) In the apparatus described in Supplementary Note 14, the motion information includes the initial MV pair of the sub-block; the processing circuit is configured to apply the BDOF mode to each sample of the sub-block to determine a refined MV pair for each sample, the BDOF mode including a sample-based BDOF mode; and the processing circuit is configured to reconstruct each sample of the sub-block based on the refined MV pair for each sample. Apparatus

[0310] (Supplementary Note 19) In the apparatus described in Supplementary Note 15, the refined motion information of the sub-block includes one or more first refined MV pairs for each of one or more areas within the sub-block; the processing circuit is configured to apply a BDOF mode including a sample-based BDOF mode after applying the BM-based MV refinement, and the refined MV pair for each sample in the area within the one or more areas is determined based on the corresponding first refined MV pair for the area and the BDOF mode; and the processing circuit is configured to reconstruct each sample in the area based on the refined MV pair for each sample. Apparatus

[0311] (Supplementary Note 20) In the apparatus described in Supplementary Note 14, the prediction information indicates that at least one of (i) the BM-based MV refinement or (ii) the BDOF mode is applied to the sub-block. Apparatus

Claims

1. A video decoding method in a decoder, comprising: receiving a coded bitstream including a current block in a current picture, wherein the current block includes a plurality of sub-blocks; obtaining prediction information indicating whether the current block is coded in a sub-block-based temporal motion vector prediction (SbTMVP) mode; determining whether sub-blocks within the plurality of sub-blocks of the current block are bi-predicted in response to the current block being coded in the SbTMVP mode; determining motion information of the sub-blocks based on the SbTMVP mode in response to the sub-blocks being bi-predicted; applying at least one of (i) bilateral matching (BM)-based motion vector (MV) refinement and (ii) bi-directional optical flow (BDOF) mode to the sub-blocks to refine the motion information of the sub-blocks; and reconstructing the current block based on the refined motion information corresponding to one or more sub-blocks within the plurality of sub-blocks of the current block, wherein the refined motion information corresponding to the one or more sub-blocks includes the refined motion information of the sub-blocks; A method comprising the above steps.

2. The method according to claim 1, wherein the applying step includes applying the BM-based MV refinement to determine the refined motion information of the sub-blocks, and the BM-based MV refinement includes decoder-side motion vector refinement (DMVR) or multi-pass decoder-side motion vector refinement (MP-DMVR); and the reconstructing step includes reconstructing the sub-blocks based on the refined motion information.

3. The method according to claim 2, wherein the motion information includes an initial MV pair of the sub-blocks; the BM-based MV refinement includes the DMVR; the applying step includes applying the DMVR to an area within the sub-blocks to determine a refined MV pair of the area based on the initial MV pair, and the area is smaller than or equal to an area of the sub-blocks; and The method of the reconstructing step includes the step of reconstructing an area within the sub-block based on the refined MV pair. **Claim 4** In the method according to claim 2, the motion information includes an initial MV pair of the sub-block; the BM-based MV refinement includes the MP-DMVR; The sub-block size of the sub-block is greater than the first threshold M 1 ×N 1 depending on whether it is greater applying at least one DMVR to the sub-block to determine a first refined MV pair of the sub-block; and applying the BDOF mode to an area within the sub-block to determine a second refined MV pair of the area based on the first refined MV pair, the area being smaller than or equal to the area of the sub-block. **Claim 5** In the method according to claim 1, the motion information includes an initial MV pair of the sub-block; the applying step includes applying the BDOF mode to each sample of the sub-block to determine a refined MV pair for each sample, the BDOF mode including a sample-based BDOF mode; and the reconstructing step includes reconstructing each sample of the sub-block based on the refined MV pair for each sample. **Claim 6** In the method according to claim 2, the refined motion information of the sub-block includes one or more first refined MV pairs for each of one or more areas within the sub-block; the applying step further includes applying a BDOF mode including a sample-based BDOF mode after applying the BM-based MV refinement, and the refined MV pair for each sample in an area within the one or more areas is determined based on the first refined MV pair corresponding to the area and the BDOF mode; and the reconstructing step includes reconstructing each sample in the area based on the refined MV pair for each sample. **Claim 7** In the method according to claim 1, the prediction information indicates that at least one of (i) the BM-based MV refinement or (ii) the BDOF mode is applied to the sub-block. **Claim 8** The method according to claim 1, wherein the prediction information includes a flag indicating that at least one of (i) the BM-based MV refinement or (ii) the BDOF mode is applied to the sub-block.

9. In the method according to claim 1, the BM-based MV refinement or the BDOF mode is applied based on a first reference picture or a second reference picture of the current picture; the first reference picture is before the current picture in the display order, and the second reference picture is after the current picture in the display order; and the distances from the first reference picture and the second reference picture to the current picture are the same.

10. A video decoding method in a decoder, comprising: decoding prediction information of a current block in a current picture from a coded bitstream, wherein the current block includes a plurality of sub-blocks reconfigured based on a sub-block-based temporal motion vector prediction (SbTMVP) mode; determining motion information of each of one or more sub-blocks within the plurality of sub-blocks based on the SbTMVP mode, wherein the one or more sub-blocks are bi-predicted; applying bilateral matching (BM)-based motion vector (MV) refinement to the current block based on an initial motion vector (MV) pair of the current block to determine updated motion information of the current block, wherein the initial MV pair of the current block is motion information of one of the one or more sub-blocks; and reconstructing the current block based on the updated motion information of the current block; A method comprising the above steps.

11. In the method according to claim 10, the one or more sub-blocks include a plurality of bi-predicted sub-blocks; and the method includes determining the initial MV pair of the current block, applying the BM-based MV refinement to the current block based on the motion information of each of the plurality of bi-predicted sub-blocks to determine a bilateral matching cost associated with each of the bi-predicted sub-blocks; and Determining an initial MV pair of the current block as motion information of the sub-blocks by using the minimum bilateral matching cost among the bilateral matching costs of the plurality of bilaterally predicted sub-blocks; A method including the steps performed thereby.

12. In the method according to claim 10, the one or more sub-blocks include a plurality of bilaterally predicted sub-blocks; and the method includes determining an initial MV pair of the current block as motion information of one of the plurality of bilaterally predicted sub-blocks based on syntax information in the coded bitstream.

13. In the method according to claim 10, the prediction information indicates that based on an initial MV pair which is motion information of one of the one or more sub-blocks, the BM-based MV refinement is applied to the current block.

14. A computer program for causing a computer to execute the method according to any one of claims 1-13.

15. A video decoding apparatus including a processing circuit configured to execute the method according to any one of claims 1-13.

16. A video processing method in an encoder, comprising: transmitting a coded bitstream including a current block in a current picture to a decoder, wherein the coded bitstream includes prediction information of the current block in the current picture, the current block includes a plurality of sub-blocks reconstructed based on a sub-block based temporal motion vector prediction (SbTMVP) mode, the prediction information indicates whether the current block is coded in a sub-block based temporal motion vector prediction (SbTMVP) mode, when the current block is coded in the SbTMVP mode, it is determined whether sub-blocks in the plurality of sub-blocks of the current block are bilaterally predicted, when the sub-blocks are bilaterally predicted, motion information of the sub-blocks is determined based on the SbTMVP mode. Apply at least one of (i) bilateral matching (BM)-based motion vector (MV) refinement and (ii) bi-directional optical flow (BDOF) mode to the sub-block, so that the motion information of the sub-block is refined, Based on the refined motion information corresponding to one or more sub-blocks within a plurality of sub-blocks of the current block, the current block is reconstructed, and the refined motion information corresponding to the one or more sub-blocks includes the refined motion information of the sub-blocks. Method.