Method, apparatus, and computer program for signaling for composite interpredictive modes

JP7917258B2Active Publication Date: 2026-09-08TENCENT AMERICA LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025523077
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-09-08
Filing Date
2023-09-21
Publication Date
2026-09-08
Estimated Expiration
2043-09-21

AI Technical Summary

Benefits of technology

【0026】 開示の主題のさらなる特徴、性質、および様々な利点は、以下の詳細な説明および添付の図面からより明らかになるであろう。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007917258000007
    Figure 0007917258000007
  • Figure 0007917258000008
    Figure 0007917258000008
  • Figure 0007917258000009
    Figure 0007917258000009
Patent Text Reader

Abstract

This disclosure relates generally to video coding, and more particularly to methods and systems for signaling composite inter-prediction modes. For example, to design an exemplary scheme for signaling an optical flow refinement flag and composite inter-prediction mode of a video block, the correlation between the use of optical flow refinement and various composite inter-prediction modes is explored. Specifically, the signaling of the inter-prediction mode of a block may depend on whether optical flow refinement is applied to the block.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-Reference to Related Applications This application claims the benefit of and priority to U.S. Non-Provisional Patent Application No. 18 / 463,508, filed on September 8, 2023, entitled "Method and Apparatus for Signaling for Compound Inter Prediction Mode", which claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 442,721, filed on February 1, 2023, entitled "Method and Apparatus for Improved Signaling for Compound Inter Prediction Mode", which is incorporated herein by reference in its entirety.

[0002] The present disclosure relates generally to video coding, and in particular, to methods and systems for signaling of compound inter prediction modes. [Background Art]

[0003] Uncompressed digital video may include a sequence of pictures and may include specific bitrate requirements for storage, data processing, and transmission bandwidth in streaming applications. One object of video coding and decoding may be to reduce the redundancy of an uncompressed input video signal through various compression techniques. [Summary of the Invention] [Means for Solving the Problems]

[0004] The present disclosure relates generally to video coding, and in particular, to methods and systems for signaling of compound inter prediction modes. To design example schemes for signaling an optical flow refinement flag and a compound inter prediction mode for a video block, for example, the correlation between the use of optical flow refinement and different compound inter prediction modes is explored. Specifically, signaling of an inter prediction mode for a block may depend on whether optical flow refinement is applied to the block.

[0005] An exemplary implementation discloses a method for decoding video blocks in a video bitstream. The method includes the steps of: receiving a first syntactic element signaled in the video bitstream indicating whether optical flow refinement is applied to the video block; determining whether optical flow refinement is applied to the video block based on the value of the first syntactic element; receiving a second syntactic element from the video bitstream indicating the composite interpretation mode of the video block, depending on whether optical flow refinement is applied after receiving the first syntactic element; determining the composite interpretation mode of the video block based on the value of the second syntactic element; and predicting the video block based on the determined composite interpretation mode.

[0006] In the exemplary implementation described above, the value of the second syntactic element is determined by decoding the second syntactic element using a coding context that depends on whether optical flow refinement is applied.

[0007] In any of the above implementations, if optical flow refinement is applied, the first context is used to decode the second syntactic element that indicates the composite interpretation mode of the video block; if optical flow refinement is not applied, a second context different from the first context is used to decode the second syntactic element that indicates the composite interpretation mode of the video block.

[0008] In any of the above implementations, when optical flow refinement is applied to a video block, only composite interprediction modes are permitted that require signaling of a single motion vector difference, or that do not require signaling of a motion vector difference.

[0009] In any of the above implementation forms, when optical flow refinement is applied to a video block, the syntactic value space of the video block's composite interprediction mode includes a subset of the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, JOINT_NEWMV, and JOINT_AMVDNEWMV modes.

[0010] In any of the above implementations, when optical flow refinement is applied to a video block, only composite interpretation modes that include or do not include a single motion vector difference are permitted.

[0011] In any of the above implementation forms, if optical flow refinement is applied to a video block, the syntactic value space of the video block's composite interpretation mode includes a subset of the NEAR_NEARMV, NEAR_NEWMV, and NEW_NEARMV modes.

[0012] In any one of the above implementation configurations, with respect to joint motion vector difference composite interpretation modes, if optical flow refinement is not applied to the video block, two joint motion vector difference composite interpretation modes are allowed; if optical flow refinement is applied to the video block, only one of the two joint motion vector difference composite interpretation modes is allowed.

[0013] In any one of the above implementation forms, the two joint motion vector difference composite interpretation modes include JOINT_NEWMV mode and JOINT_AMVDNEWMV mode, and when optical flow refinement is applied to the video block, only one of the two allowed joint motion vector difference composite interpretation modes is JOINT_AMVDNEWMV mode.

[0014] In any one of the above implementations, the step of determining the composite interpretation mode of a video block may include the steps of determining the mapping between possible values ​​of a second syntactic element and a plurality of composite interpretation modes, based on whether optical flow refinement of the video block is applied, and determining the composite interpretation mode of the video block based on the values ​​of the second syntactic element and the mapping.

[0015] In any one of the above implementations, the mapping between the possible values ​​of the second syntactic element and the multiple composite interpretation modes differs depending on whether optical flow refinement is applied or disabled.

[0016] Another exemplary implementation discloses a method for decoding a video block in a video bitstream. This method may include: receiving a first syntactic element signaled in the video bitstream that indicates a composite interprediction mode of a video block among a plurality of composite interprediction modes; determining the composite interprediction mode of the video block based on the value of the first syntactic element; determining whether a second syntactic element of the video block is included in the video bitstream, or receiving the second syntactic element from the video bitstream in a manner dependent on the composite interprediction mode indicated by the first syntactic element, wherein the second syntactic element indicates whether optical flow refinement is applied to the video block.

[0017] In the exemplary implementation described above, the step of extracting a second syntactic element includes the step of decoding the second syntactic element using a coding context that depends on the compound interpretation mode indicated by the first syntactic element.

[0018] In any one of the above implementations, the second syntactic element of the video block is included in the video bitstream only for a subset of multiple composite interprediction modes.

[0019] In any one of the above implementations, the second syntax element of a video block is included in the video bitstream only when the compound inter prediction mode of the video block requires at most one signaled motion vector difference.

[0020] In any one of the above implementations, the subset of the plurality of compound inter prediction modes includes NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, JOINT_NEWMV, and JOINT_AMVDNEWMV modes.

[0021] In any one of the above implementations, the second syntax element of a video block is included in the video bitstream only when the compound inter prediction mode of the video block does not involve different motion vectors, or uses adaptive motion vector resolution, or has a motion vector accuracy coarser than a predetermined accuracy threshold.

[0022] In any one of the above implementations, the second syntax element of a video block is included in the video bitstream only when the compound inter prediction mode of the video block involves at most one motion vector difference.

[0023] In any one of the above implementations, the subset of the plurality of compound inter prediction modes includes NEAR_NEARMV, NEAR_NEWMV, and NEW_NEARMV modes.

[0024] Aspects of the present disclosure also provide an electronic device or electronic apparatus comprising a circuit or a processor configured to perform any of the above method implementations.

[0025] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by an electronic device, cause the electronic device to perform any one of the above method implementations.

[0026] Further features, properties and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] [Figure 1] FIG. 1 is a schematic diagram showing a simplified block diagram of a communication system (100) according to an exemplary embodiment. [Figure 2] FIG. 2 is a schematic diagram showing a simplified block diagram of a communication system (200) according to an exemplary embodiment. [Figure 3] FIG. 3 is a schematic diagram showing a simplified block diagram of a video decoder according to an exemplary embodiment. [Figure 4] FIG. 4 is a schematic diagram showing a simplified block diagram of a video encoder according to an exemplary embodiment. [Figure 5] FIG. 5 is a block diagram of a video encoder according to another exemplary embodiment. [Figure 6] FIG. 6 is a block diagram of a video decoder according to another exemplary embodiment. [Figure 7] FIG. 7 is a diagram illustrating a coding block splitting scheme according to an exemplary embodiment of the present disclosure. [Figure 8] FIG. 8 is a diagram illustrating another coding block splitting scheme according to an exemplary embodiment of the present disclosure. [Figure 9] FIG. 9 is a diagram illustrating another coding block splitting scheme according to an exemplary embodiment of the present disclosure. [Figure 10] FIG. 10 is a diagram illustrating an exemplary logic flow for a method of signaling an optical flow refinement flag and a combined inter prediction mode. [Figure 11] FIG. 11 is a schematic diagram of a computer system according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF EMBODIMENTS

[0028] Throughout this specification and the claims, terms may have nuances implied or suggested in context beyond their expressly stated meanings. The phrases "in one embodiment / implementation" or "in several embodiments / implementation" used in this disclosure do not necessarily refer to the same embodiment / implementation, and the phrases "in another embodiment / implementation" or "in other embodiments" used in this disclosure do not necessarily refer to different embodiments. For example, the claimed subject matter is intended to include all or some combinations of exemplary embodiments / implementation.

[0029] In general, technical terms can be understood, at least partially, from their usage in context. For example, terms such as “and,” “or,” or “and / or” as used herein can have various context-dependent meanings. Typically, when “or” is used to relate a list such as A, B, or C, it is intended to mean A, B, and C in an inclusive sense, as well as A, B, or C in an exclusive sense. Furthermore, the terms “one or more,” “at least one,” “a,” “an,” or “the” as used herein can be used in a singular or plural sense, at least partially depending on the context. In addition, the terms “based on” or “determined by” may be understood not necessarily to convey an exclusive set of factors, but instead, at least partially depending on the context, may allow for the existence of further factors that are not necessarily explicitly described.

[0030] Figure 1 shows a simplified block diagram of a communication system (100) according to one embodiment of the present disclosure. The communication system (100) includes a plurality of terminal devices, e.g., 110, 120, 130, and 140, which can communicate with each other, for example, over a network (150). In the example of Figure 1, a first pair of terminal devices (110) and (120) may perform unidirectional data transmission. For example, terminal device (110) may encode video data in the form of one or more encoded bitstreams (e.g., a stream of video pictures captured by terminal device (110)) for transmission over the network (150). Terminal device (120) may receive the encoded video data from the network (150), decode the encoded video data to restore the video pictures, and display the video pictures according to the restored video data. Unidirectional data transmission may be performed for media serving applications, etc.

[0031] In another example, a second pair of terminal devices (130) and (140) may perform bidirectional transmission of encoded video data, for example, during video conferencing. For bidirectional transmission of data, in one example, each of terminal devices (130) and (140) may encode video data (e.g., a stream of video pictures captured by a terminal device) for transmission to another device of terminal devices (130) and (140) in order to restore and display a video picture, and may receive the encoded video data from another device of terminal devices (130) and (140).

[0032] In the example in Figure 1, the terminal devices may be implemented as servers, personal computers, and smartphones, but the applicability of the fundamental principles of this disclosure is not limited thereto. Embodiments of this disclosure may be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing equipment, and the like. Network (150) represents any number or any type of network that transmits coded video data between terminal devices, including, for example, wired and / or wireless communication networks. Communication network (150) may exchange data via circuit switching, packet switching, and / or other types of channels. Typical networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0033] Figure 2 shows an example of an application of the disclosed subject matter, illustrating the arrangement of a video encoder and video decoder in a video streaming environment. The disclosed subject matter may be equally applicable to other video applications, including, for example, video conferencing, digital television broadcasting, games, virtual reality, and the storage of compressed video on digital media, including CDs, DVDs, and memory sticks.

[0034] As shown in Figure 2, the video streaming system may include a video capture subsystem (213) which may include a video source (201), for example, a digital camera, for creating an uncompressed video picture or stream of pictures (202). In one example, the stream of video pictures (202) includes samples recorded by the digital camera of the video source 201. The stream of video pictures (202), drawn as a thick line to emphasize the high data volume compared to encoded video data (204) (or encoded video bitstream), can be processed by an electronic device (220) which includes a video encoder (203) coupled to the video source (201). The video encoder (203) may include hardware, software, or a combination thereof, enabling or implementing embodiments of the disclosed subject matter as will be described in more detail below. The encoded video data (204) (or encoded video bitstream (204)), which is drawn as a thin line to emphasize its smaller data size compared to the uncompressed video picture stream (202), can be stored on a streaming server (205) for future use or directly on a downstream video device (not shown). One or more streaming client subsystems, such as client subsystems (206) and (208) in Figure 2, can access the streaming server (205) to retrieve copies (207) and (209) of the encoded video data (204). The client subsystem (206) may include, for example, a video decoder (210) in an electronic device (230). The video decoder (210) decodes the incoming copy (207) of the encoded video data and creates an outgoing stream of an uncompressed video picture (211) that can be rendered on a display (212) (e.g., a display screen) or other rendering device (not shown).

[0035] Figure 3 shows a block diagram of a video decoder (310) of an electronic device (330) according to any embodiment of the present disclosure below. The electronic device (330) may include a receiver (331) (e.g., a receiving circuit). The video decoder 310 may be used in place of the video decoder (210) in the example of Figure 2.

[0036] As shown in Figure 3, the receiver (331) may receive one or more coded video sequences from the channel (301). To address network jitter, a buffer memory (315) may be placed between the receiver (331) and the entropy decoder / parser (320) (hereinafter, "Parser (320)"). The Parser (320) may be configured to reconstruct symbols (321) from the coded video sequences. The categories of those symbols include information used to manage the operation of the video decoder (310), and potentially information for controlling rendering devices such as a display (312) (e.g., a display screen). The Parser (320) may parse / entropy decode the coded video sequences. The Parser (320) may extract from the coded video sequences a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder. Subgroups may include picture groups (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), and predictive units (PUs). The parser (320) may also extract information from the coded video sequence, such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, and motion vectors. The reconstruction of symbols (321) may involve multiple different processing or function units. The units involved and how they are involved may be controlled by the parser (320) by subgroup control information analyzed from the coded video sequence.

[0037] The first unit may include a scaler / inverse unit (351). The scaler / inverse unit (351) may receive control information from the parser (320) including the quantized transformation coefficients, as well as information indicating which type of inverse transformation should be used, block size, quantization coefficients / parameters, quantization scaling matrix, and state as symbol (321). The scaler / inverse unit (351) may output a block containing sample values ​​that can be input to the aggregator (355).

[0038] In some cases, the output samples of the scaler / inverse transform (351) may relate to intracoded blocks, i.e., blocks that do not use prediction information from previously reconstructed pictures but can use prediction information from previously reconstructed portions of the current picture. Such prediction information can be provided by the intrapicture prediction unit (352). In some cases, the intrapicture prediction unit (352) can generate a block of the same size and shape as the block being reconstructed using surrounding block information that has already been reconstructed and is stored in the current picture buffer (358). The current picture buffer (358) buffers, for example, partially reconstructed current pictures and / or fully reconstructed current pictures. In some implementations, the aggregator (355) can add the prediction information generated by the intraprediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351) on a sample-by-sample basis.

[0039] In other cases, the output samples of the scaler / inverse unit (351) may relate to an intercoded, potentially motion-compensated block. In such cases, the motion-compensated prediction unit (353) can access the reference picture memory (357) to fetch samples to be used for interpicture prediction. After motion-compensating the reference samples fetched according to the symbols (321) associated with the block, these samples may be added by the aggregator (355) to the output of the scaler / inverse unit (351) (the output of unit 351 is sometimes called residual samples or residual signal) to generate output sample information.

[0040] The output samples of the aggregator (355) can undergo various loop filtering techniques in the loop filter unit (356), including several types of loop filters. The output of the loop filter unit (356) can be a sample stream that can be output to the rendering device (312) as well as stored in a reference picture memory (357) for use in future interpicture prediction.

[0041] Figure 4 shows a block diagram of a video encoder (403) according to an exemplary embodiment of the present disclosure. The video encoder (403) may be included in an electronic device (420). The electronic device (420) may further include a transmitter (440) (e.g., a transmitting circuit). The video encoder (403) may be used instead of the video encoder (403) in the example of Figure 4.

[0042] The video encoder (403) may receive video samples from the video source (401). According to some exemplary embodiments, the video encoder (403) may encode and compress the pictures of the source video sequence into an encoded video sequence (443) in real time or under other temporal constraints required by the application. Implementing an appropriate coding speed constitutes one function of the controller (450). In some embodiments, the controller (450) can be functionally coupled with and controlled by other functional units, as described below. Parameters set by the controller (450) may include rate control-related parameters (picture skip, quantizer, lambda value of rate distortion optimization technique, ...), picture size, picture group (GOP) layout, maximum motion vector search range, etc.

[0043] In some exemplary embodiments, the video encoder (403) may be configured to operate in a coding loop. The coding loop may include a source coder (430) and a (local) decoder (433) built into the video encoder (403). (Since any compression between symbols and the coded video bitstream in entropy coding may be reversible in the video compression techniques considered in the disclosed subject) the decoder (433) reconstructs the symbols to create sample data in a manner similar to that created by a (remote) decoder, even if the built-in decoder 433 processes the coded video bitstream by the source coder 430 without entropy coding. At this point, it can be said that any decoder technique other than parsing / entropy decoding, which may only exist within the decoder, may also necessarily exist in substantially the same functional form in the corresponding encoder. For this reason, the disclosed subject may focus on decoder operation, which is similar to the decoding portion of the encoder. Thus, a description of encoder technique can be omitted, as it is the inverse of a comprehensive description of decoder technique. A more detailed description of the encoder is provided below, but only in specific areas or embodiments.

[0044] In operation in some exemplary implementations, the source coder (430) may perform motion-compensated predictive coding, predictively coding the input picture by referencing one or more previously coded pictures from a video sequence designated as a “reference picture”.

[0045] The local video decoder (433) may decode the encoded video data of a picture that may be designated as a reference picture. The local video decoder (433) may replicate the decoding process that the video decoder may perform on the reference picture and store the reconstructed reference picture in the reference picture cache (434). In this way, the video encoder (403) can locally store a copy of the reconstructed reference picture that has content in common with the reconstructed reference picture obtained by the far-end (remote) video decoder (without transmission error).

[0046] The predictor (435) can perform predictive searches on the coding engine (432). That is, for a new picture to be coded, the predictor (435) can search the reference picture memory (434) for specific metadata such as sample data (as candidate reference pixel blocks) or reference picture motion vectors, block shapes, etc., which can serve as appropriate predictive references for the new picture.

[0047] The controller (450) can manage the coding operations of the source coder (430), including, for example, setting parameters and subgroup parameters used to encode video data.

[0048] The outputs of all the aforementioned functional units can be entropy-coded by an entropy coder (445). A transmitter (440) may buffer the coded video sequence created by the entropy coder (445) and prepare it for transmission over a communication channel (460), which may be a hardware / software link to a storage device that stores the coded video data. The transmitter (440) can merge the coded video data from the video coder (403) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0049] The controller (450) can manage the operation of the video encoder (403). During coding, the controller (450) may assign a specific coded picture type to each coded picture, which may affect the coding technique that can be applied to each picture. For example, a picture may often be assigned as one of the following picture types: intra-picture (I-picture), predictive picture (P-picture), or bidirectional predictive picture (B-picture), or multi-predictive picture. A source picture can generally be spatially subdivided into multiple sample coding blocks, as will be described in more detail below.

[0050] Figure 5 shows a diagram of a video encoder (503) according to another exemplary embodiment of the present disclosure. The video encoder (503) is configured to receive a processing block (e.g., a prediction block) of sample values ​​in the current video picture within a sequence of video pictures, and to encode the processing block into a coded picture which is part of a coded video sequence. The exemplary video encoder (503) may be used instead of the video encoder (403) in the example of Figure 4.

[0051] For example, the video encoder (503) receives a matrix of sample values ​​for a processing block. The video encoder (503) then determines whether the processing block is best coded using, for example, rate distortion optimization (RDO) in intra-mode, inter-mode, or bi-predictive mode.

[0052] In the example shown in Figure 5, the video encoder (503) includes an interencoder (530), an intraencoder (522), a residual calculator (523), a switch (526), ​​a residual encoder (524), a master controller (521), and an entropy encoder (525), all coupled together as shown in the exemplary configuration of Figure 5.

[0053] The interencoder (530) is configured to receive a sample of the current block (e.g., a processing block), compare the block with one or more reference blocks in the reference picture (e.g., blocks in the previous and subsequent pictures in display order), generate interprediction information (e.g., a description of redundant information by the interencoding technique, motion vectors, merge mode information), and compute an interprediction result (e.g., a predicted block) based on the interprediction information using any appropriate technique.

[0054] The intra encoder (522) is configured to receive a sample of the current block (e.g., a processing block), compare the block with a block already coded within the same picture, generate quantization coefficients after the transformation, and optionally generate intra prediction information (e.g., intra prediction direction information by one or more intra encoding techniques).

[0055] The master controller (521) may be configured, for example, to determine the prediction mode of a block, to determine general-purpose control data to provide control signals to the switch (526) based on the prediction mode, and to control other components of the video encoder (503) based on the general-purpose control data.

[0056] A residual calculator (523) may be configured to calculate the difference (residual data) between the received block and the predicted result of a block selected from an intra encoder (522) or interencoder (530). A residual encoder (524) may be configured to encode the residual data to generate transformation coefficients. The transformation coefficients are then quantized to obtain quantized transformation coefficients. In various exemplary embodiments, the video encoder (503) also includes a residual decoder (528). The residual decoder (528) is configured to perform an inverse transform to produce decoded residual data. An entropy encoder (525) may be configured to format the bitstream to include the encoded blocks and perform entropy coding.

[0057] Figure 6 shows a diagram of an exemplary video decoder (610) according to another embodiment of the present disclosure. The video decoder (610) is configured to receive an encoded picture, which is part of an encoded video sequence, and to decode the encoded picture to produce a reconstructed picture. In one example, the video decoder (610) may be used instead of the video decoder (410) in the example of Figure 4.

[0058] In the example shown in Figure 6, the video decoder (610) includes an entropy decoder (671), an interdecoder (680), a residual decoder (673), a reconstruction module (674), and an intradecoder (672), all coupled together as shown in the exemplary configuration of Figure 6.

[0059] An entropy decoder (671) can be configured to reconstruct specific symbols representing the syntactic elements that make up a coded picture from the coded picture. An interdecoder (680) can be configured to receive interprediction information and generate interprediction results based on the interprediction information. An intradecoder (672) can be configured to receive intraprediction information and generate prediction results based on the intraprediction information. A residual decoder (673) can be configured to perform inverse quantization to extract inversely quantized transformation coefficients and process the inversely quantized transformation coefficients to convert the residual from the frequency domain to the spatial domain. A reconstruction module (674) can be configured to combine the residual output by the residual decoder (673) and the prediction results (optionally output by the interprediction module or intraprediction module) in the spatial domain to form a reconstructed block that forms part of the reconstructed picture as part of the reconstructed video.

[0060] It should be noted that the video encoders (203), (403), and (503), and the video decoders (210), (310), and (610) can be implemented using any suitable technique. In some embodiments, the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610) can be implemented using one or more integrated circuits. In other embodiments, the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610) can be implemented using one or more processors that execute software instructions.

[0061] Focusing on the block partitioning used for coding and decoding, a typical partition may begin with a base block and follow a predefined set of rules, a specific pattern, a partition tree, or some partition structure or scheme. The partitioning may be hierarchical and recursive. After separating or partitioning the base block according to one of the exemplary partitioning procedures or other procedures, or a combination thereof, described below, a final set of partitions or coding blocks may be obtained. Each of these partitions may be at one of the various partitioning levels in the partitioning hierarchy and may be of various shapes. Each of the partitions may be called a coding block (CB). In the various exemplary partitioning implementations described further below, each resulting CB may be any CB of an acceptable size and partitioning level. Such partitions are called coding blocks because some basic coding / decoding decisions may be made for them, and they can form units for which coding / decoding parameters can be optimized, determined, and signaled in the encoded video bitstream. The highest or deepest level in the final partition represents the depth of the coding block partitioning structure of the tree. Coding blocks may be luma coding blocks or chroma coding blocks. The CB tree structure for each color is sometimes called a coding block tree (CBT). The coding blocks for all color channels are sometimes collectively called a coding unit (CU). The hierarchical structure for all color channels is sometimes collectively called a coding tree unit (CTU). The division patterns or structures of the various color channels within a CTU may or may not be the same.

[0062] In some implementations, the partition tree schemes or structures used for lumar channels and chroma channels do not necessarily have to be the same. In other words, lumar channels and chroma channels may have separate coding tree structures or patterns. Furthermore, whether lumar channels and chroma channels use the same or different coding partition tree structures, and the actual coding partition tree structures used, may depend on whether the coded slice is a P, B, or I slice. For example, in an I slice, chroma channels and lumar channels may have separate coding partition tree structures or coding partition tree structure modes, while in a P or B slice, lumar channels and chroma channels may share the same coding partition tree scheme. When separate coding partition tree structures or modes are applied, a lumar channel may be partitioned into CBs by one coding partition tree structure, and a chroma channel may be partitioned into chroma CBs by another coding partition tree structure.

[0063] Figure 7 shows 10 exemplary predefined partition structures / patterns that allow recursive partitioning to form a partition tree. The root block can start from a predefined level (e.g., a base block at a 128x128 level or 64x64 level). The exemplary partition structures in Figure 7 include various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. In some exemplary implementations, none of the rectangular partitions in Figure 7 can be further subdivided. A coding tree depth may be further defined to indicate the partition depth from the root node or root block. For example, the coding tree depth relative to the root node or root block may be set to 0, and after the root block is partitioned one more time according to Figure 7, the coding tree depth increases by 1. In some implementations, only 710 square partitions can allow recursive partitioning to the next level of the partition tree following the pattern in Figure 7.

[0064] In some other exemplary implementations for coding block partitioning, a quadtree structure may be used. Such quadtree partitioning may be applied hierarchically and recursively to any square partition. Whether the base block or intermediate block or partition is further quadtree-partitioned can be adapted to various local characteristics of the base block or intermediate block / partition.

[0065] In some other examples, a ternary pattern may be used to partition a base block or any intermediate block, as shown in Figure 8. The ternary pattern may be implemented vertically, as shown in 802, or horizontally, as shown in 804. The exemplary partition ratio in Figure 8 is shown as 1:2:1, but other ratios may be predefined. In some implementations, two or more different ratios may be predefined. In some implementations, the width and height of the partitions in the exemplary ternary tree are always powers of 2 to avoid further transformations.

[0066] The above partitioning schemes can be combined in any way at different partitioning levels. For example, the quadtree and binary partitioning schemes described above may be combined to partition a base block into a quadtree-binary (QTBT) structure. In such a scheme, the base block or intermediate block / partition may be either quadtree partitioned or binary partitioned, if specified, according to a set of predefined conditions. A particular example is shown in Figure 9, where the base block is initially quadtree partitioned into four partitions, as indicated by 902, 904, 906, and 908. Each of the resulting partitions is then quadtree partitioned into four further partitions (such as 908), or binary partitioned into two further partitions at the next level (e.g., both symmetric, either horizontal or vertical, such as 902 or 906), or not partitioned at all (such as 904). Binary or quadtree partitioning may be recursively possible for square partitions, as shown by the overall exemplary partition pattern in 910 and the corresponding tree structure / representation in 920, where solid lines represent quadtree partitioning and dashed lines represent binary partitioning. A flag may be used for each binary node (non-leaf binary partition) to indicate whether the binary is horizontal or vertical. For example, as shown in 920, which matches the partition structure in 910, a flag "0" may represent horizontal binary and a flag "1" may represent vertical binary. In the case of quadtree partitioning, there is no need to specify the partition type, as quadtree partitioning always divides a block or partition both horizontally and vertically to produce four subblocks / partitions of the same size. In some implementations, a flag "1" may represent horizontal binary and a flag "0" may represent vertical binary.

[0067] In some exemplary implementations of QTBT, the quadtree and binary rule set may be represented by the following predefined parameters and their associated corresponding functions. -CTU size: The size of the root node of the quadtree (the size of the base block). -MinQTSize: Minimum allowable quadtree leaf node size -MaxBTSize: Maximum allowable binary tree root node size -MaxBTDepth: Maximum allowable binary tree depth -MinBTSize: Minimum allowable binary tree leaf node size

[0068] In some exemplary implementations of the QTBT partitioning structure, the CTU size may be set to a 128x128 chroma sample with corresponding 64x64 blocks of chroma samples (when used assuming typical chroma subsampling), MinQTSize may be set to 16x16, MaxBTSize may be set to 64x64, MinBTSize (both width and height) may be set to 4x4, and MaxBTDepth may be set to 4. The quadtree partitioning may be applied to the CTU first to generate quadtree leaf nodes. The quadtree leaf nodes may have sizes ranging from their minimum allowable size of 16x16 (i.e., MinQTSize) to 128x128 (i.e., CTU size). If a node is 128x128, it will not be partitioned by the binary tree first because its size exceeds MaxBTSize (i.e., 64x64). Otherwise, nodes that do not exceed MaxBTSize may be partitioned by the binary tree. In the example in Figure 9, the base block is 128x128. The base block can only be quadtree-partitioned according to a predefined set of rules. The base block has a partitioning depth of 0. Each of the four resulting partitions is 64x64, not exceeding MaxBTSize, and may be further quadtree-partitioned or binary-partitioned at level 1. The process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), further partitioning is not considered. When the width of a binary tree node is equal to MinBTSize (i.e., 4), further horizontal partitioning is not considered. Similarly, when the height of a binary tree node is equal to MinBTSize, further vertical partitioning is not considered.

[0069] In some exemplary implementations, the above QTBT scheme may be configured to support the flexibility for lumens and chromians to have the same QTBT structure or separate QTBT structures. For example, in the case of P-slice and B-slice, the lumens CTB and chromens CTB within one CTU may share the same QTBT structure. However, in the case of I-slice, the lumens CTB may be divided into CBs by a QTBT structure, and the chromens CTB may be divided into chromens CBs by another QTBT structure. This means that CUs may be used to refer to different color channels within an I-slice, for example, an I-slice may consist of a coding block for the lumens component or a coding block for two chromens components, and a CU in a P-slice or B-slice may consist of a coding block for all three color components.

[0070] The various CB partitioning schemes described above, as well as the further partitioning of the CB into PBs, may be combined in any way. The following specific implementations are provided as non-limiting examples.

[0071] Interpretation can be performed, for example, in single-reference mode or composite-reference mode. In some implementations, a skip flag may be initially included in the bitstream of the current block (or at a higher level) to indicate whether the current block is intercoded and will not be skipped. If the current block is intercoded, another flag may be further included in the bitstream as a signal indicating whether single-reference mode or composite-reference mode is being used to predict the current block. In single-reference mode, one reference block may be used to generate the predicted block for the current block. In composite-reference mode, two or more reference blocks may be used, for example, by a weighted average to generate the predicted block. One or more reference blocks may be identified using one or more reference frame indices, and further using one or more corresponding motion vectors indicating the shift between the one or more reference blocks and the current block in terms of its position relative to the frame, for example, in horizontal and vertical pixels. For example, the interpretation block of the current block may be generated from a single reference block identified by a single motion vector in the reference frame, as a prediction block in single reference mode. However, in composite reference mode, the prediction block may be generated by a weighted average of two reference blocks in two reference frames, indicated by two reference frame indices and two corresponding motion vectors. Motion vectors can be coded in various ways and included in the bitstream.

[0072] In some exemplary implementations, one or more reference picture lists, including the identification of short-term and long-term reference frames for inter-prediction, may be formed based on information in a reference picture set (RPS). For example, a single picture reference list may be formed for unidirectional inter-prediction, denoted as L0 reference (or reference list 0), and two picture reference lists may be formed for bidirectional inter-prediction, denoted as L0 (or reference list 0) and L1 (or reference list 1) for each of the two prediction directions. The reference frames included in the L0 and L1 lists may be ordered in various predetermined ways. The lengths of the L0 and L1 lists may be signaled in the video bitstream. Unidirectional inter-prediction can be either single-reference mode or composite-reference mode, provided that the multiple references for generating prediction blocks by weighted averaging in composite-prediction mode are on the same side of the frame in which the block to be predicted is located. Bidirectional inter-prediction can only be composite mode, in that bidirectional inter-prediction includes at least two reference blocks.

[0073] In some implementations, a merge mode (MM) for interpretation may be implemented. Generally, in merge mode, one or more motion vectors in a single reference prediction or a composite reference prediction of the current PB may be derived from other motion vectors rather than being computed and signaled independently. For example, in an encoding system, the current motion vector of the current PB can be represented by the difference between the current motion vector and one or more other already encoded motion vectors (called reference motion vectors). Such a difference of motion vectors, rather than the entire current motion vector, may be encoded and included in the bitstream and linked to the reference motion vectors. Correspondingly, in a decoding system, the motion vector corresponding to the current PB may be derived based on the decoded motion vector difference and the decoded reference motion vector linked to it. As a specific form of general merge mode (MM) interpretation, such interpretation based on motion vector differences is sometimes called merge mode with motion vector differences (MMVD). Thus, general MM, or MMVD in particular, may be implemented to improve coding efficiency by leveraging correlations between motion vectors associated with different PBs. For example, neighboring PBs may have similar motion vectors, and therefore their MVDs may be small and can be coded efficiently. In another example, motion vectors can be correlated temporally (between frames) for blocks that are similarly positioned / placed in space.

[0074] In some exemplary implementations of MMVD, a list of reference motion vectors (RMVs) or MV predictor candidates can be formed for a predicted block. The list of RMV candidates can contain a predetermined number (e.g., two) of MV predictor candidate blocks whose motion vectors could be used to predict the current motion vector. RMV candidate blocks can include blocks selected from neighboring blocks and / or time blocks within the same frame (e.g., blocks that are identically located in the current or subsequent frames). These options represent blocks that are spatially or temporally located relative to the current block and are likely to have a motion vector similar to or identical to the current block. The size of the list of MV predictor candidates may be predetermined. For example, the list may contain two or more candidates. In order to be on the list of RMV candidates, a candidate block may need to have, and must exist, the same reference frame (or more frames) as the current block (e.g., boundary checking must be performed if the current block is near the edge of a frame), and must have already been encoded during the encoding process and / or decoded during the decoding process. In some implementations, the list of merge candidates may be filled first with spatially neighboring blocks (traversed in a specific predefined order) if available and satisfying the above conditions, and then with time blocks if space is still available in the list. Neighboring RMV candidate blocks can be selected, for example, from the blocks to the left and above the current block. The list of RMV predictor candidates can be dynamically formed as a dynamic reference list (DRL) at various levels (sequence, picture, frame, slice, superblock, etc.). The DRL can be signaled with a bitstream.

[0075] In some implementations, the actual MV predictor candidate being used as the reference motion vector for predicting the motion vector of the currently coded block may be signaled. If the RMV candidate list contains two candidates, a one-bit flag called the merge candidate flag may be used to indicate the selection of the reference merge candidate. For the currently coded block being predicted in composite mode, each of the multiple motion vectors predicted using the MV predictor may be associated with a reference motion vector from the merge candidate list. The encoder can determine which RMV candidate more closely predicts the MV of the currently coded block and signal the selection as an index to the DRL.

[0076] In some exemplary implementations of MMVD, an RMV candidate is selected and used as the base motion vector predictor for the motion vector to be predicted, after which the motion vector difference (MVD or deltaMV representing the difference between the motion vector to be predicted and the reference candidate motion vector) can be calculated in the encoding system. Such an MVD may contain information representing the magnitude and direction of the MV difference, both of which can be signaled in the bitstream in various ways.

[0077] In some exemplary implementations of MMVD, a distance index can be used to specify the magnitude of the motion vector difference, indicating one of a set of predefined offsets that represent a predefined motion vector difference from a starting point (reference motion vector). The MV offset corresponding to the signaled index can then be added to either the horizontal or vertical component of the starting (reference) motion vector. Exemplary predefined relationships between distance indices and predefined offsets are specified in Table 1.

[0078] [Table 1]

[0079] In some exemplary implementations of MMVD, a direction index may be further signaled and used to represent the direction of the MVD relative to the reference motion vector. In some implementations, the direction may be restricted to either the horizontal or vertical direction. Exemplary 2-bit direction indices are shown in Table 2. In the examples in Table 2, the interpretation of the MVD may vary depending on the information of the start / reference MV. For example, if the start / reference MV corresponds to a single predictive block, or if both reference frame lists correspond to two predictive blocks pointing to the same side of the current picture (i.e., the POCs of both reference pictures are either greater than the POC of the current picture, or both are less than the POC of the current picture), the sign in Table 2 may specify the sign (direction) of the MV offset applied to the start / reference MV. If the start / reference MV corresponds to a biprediction block having two reference pictures on different sides of the current picture (i.e., the POC of one reference picture is greater than the POC of the current picture, and the POC of the other reference picture is smaller than the POC of the current picture), and the difference between the reference POC in picture reference list 0 and the current frame is greater than the difference between the reference POC in picture reference list 1 and the current frame, then the sign in Table 2 can specify the sign of the MV offset applied to the reference MV corresponding to the reference picture in picture reference list 0, and the sign of the offset of the MV corresponding to the reference picture in picture reference list 1 can have the opposite value (opposite sign of the offset). Otherwise, if the difference between the reference POC in picture reference list 1 and the current frame is greater than the difference between the reference POC in picture reference list 0 and the current frame, then the sign in Table 2 can specify that the sign of the MV offset applied to the reference MV associated with picture reference list 1 and the sign of the offset to the reference MV associated with picture reference list 0 have opposite values.

[0080] [Table 2]

[0081] In some exemplary implementations, the MVD may be scaled according to the difference in POCs in each direction. If the difference in POCs in both lists is the same, scaling is not necessary. If, instead, the difference in POCs in reference list 0 is greater than the difference in reference list 1, the MVD of reference list 1 is scaled. If the difference in POCs in reference list 1 is greater than that of list 0, the MVD of list 0 may be scaled similarly. If the initial MV is single predicted, the MVD is added to the available or reference MVs.

[0082] In some exemplary implementations of MVD coding and signaling for bidirectional composite prediction, in addition to coding and signaling two MVDs separately, or instead, symmetric MVD coding may be implemented such that only one MVD requires signaling and the other MVD can be derived from the signaled MVD. In such implementations, motion information, including the reference picture indices of List 0 and List 1, is not signaled. Specifically, at the slice level, a flag called "mvd_l1_0_flag" may be included in the bitstream to indicate whether reference list 1 is not signaled in the bitstream. If this flag is 1, indicating that reference list 1 is equal to 0 (and therefore not signaled), then a bidirectional prediction flag called "BiDirPredFlag" may be set to 0, which means there is no bidirectional prediction. If mvd_l1_0_flag is 0, BiDirPredFlag may be set to 1 if the nearest reference picture in List 0 and the nearest reference picture in List 1 form a forward-reverse or reverse-reverse-reference picture pair, and both reference pictures in List 0 and List 1 are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. A BiDirPredFlag of 1 may indicate that a symmetric mode flag is additionally signaled in the bitstream. The decoder may extract the symmetric mode flag from the bitstream if BiDirPredFlag is 1. The symmetric mode flag may be signaled, for example, at the CU level (if necessary) to indicate whether a symmetric MVD coding mode is being used for the corresponding CU.If the symmetric mode flag is 1, it indicates the use of the symmetric MVD coding mode, where only the reference picture indices in both List 0 and List 1 (called "mvp_l0_flag" and "mvp_l1_flag") are signaled by the MVD associated with List 0 (called "MVD0"), and the other motion vector difference "MVD1" should be derived rather than signaled. For example, MVD1 might be derived as -MVD0. Thus, in the exemplary symmetric MVD mode, only one MVD is signaled.

[0083] In several other exemplary implementations for MV prediction, harmonic schemes may be used to implement the general merge-mode MMVD for both single-reference-mode and compound-reference-mode MV prediction, and for several other types of MV prediction. Various syntactic elements may be used to signal how the MV of the current block is predicted. For example, in single-reference mode, the following MV prediction modes may be signaled:

[0084] NEARMV uses one of the motion vector predictors (MVPs) in a list directly indicated by a DRL (Dynamic Reference List) index without MVD.

[0085] NEWMV uses one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference, and applies the delta to the MVP (e.g., use MVD).

[0086] GLOBALMV - Uses motion vectors based on global motion parameters at the frame level.

[0087] Similarly, in the case of a composite reference interpretation mode that uses two reference frames corresponding to the two MVs to be predicted, the following MV prediction modes may be signaled:

[0088] NEAR_NEARMV - For each of the two motion vectors to be predicted, use one of the motion vector predictors (MVPs) in a list signaled by a DRL index without MVD.

[0089] NEAR_NEWMV - To predict the first of two motion vectors, use one of the motion vector predictors (MVPs) in the list signaled by a DRL index without MVD as the reference MV, and to predict the second of the two motion vectors, use one of the motion vector predictors (MVPs) in the list signaled by a DRL index as the reference MV, along with an additionally signaled delta MV (MVD).

[0090] NEW_NEARMV - To predict the second of two motion vectors, one of the motion vector predictors (MVPs) in the list signaled by a DRL index without MVD is used as the reference MV, and to predict the first of two motion vectors, one of the motion vector predictors (MVPs) in the list signaled by a DRL index is used as the reference MV in conjunction with an additionally signaled delta MV (MVD).

[0091] NEW_NEWMV uses one of the motion vector predictors (MVPs) in a list signaled by the DRL index as the reference MV, and uses it in conjunction with an additionally signaled delta MV to make predictions for each of the two MVs.

[0092] GLOBAL_GLOBALMV - Uses MV from each reference based on frame-level global motion parameters.

[0093] Therefore, the term "NEAR" above refers to MV prediction using a reference MV without using any MVD, as in the general merge mode, while the term "NEW" refers to MV prediction using a referenced MV and involving offsetting it with a signaled or derived MVD, as in the MMVD mode. In the case of composite interpretation, both the reference-based motion vector and motion vector delta above may generally differ or be independent between the two references or between the two MVDs, even if the two MVDs are correlated, for example, and such correlation can be used to reduce the amount of information required to signal the two motion vector deltas. To take advantage of such correlations, joint signaling of the two MVDs may be performed and shown in a bitstream, as will be described in more detail below.

[0094] In some implementations of MVD, a predefined pixel resolution for the MVD may be acceptable. For example, a motion vector precision (or precision) of 1 / 8 of a pixel may be acceptable. The MVD described above can be constructed and signaled in various ways with various MV prediction modes. In some implementations, various syntactic elements can be used to signal the above motion vector difference in reference frame list 0 or list 1.

[0095] For example, the syntax element called "mv_joint" can specify which components of the associated motion vector difference are non-zero. A mv_joint with a value of 0 can indicate that there are no non-zero MVDs along either the horizontal or vertical direction. 1 can be shown to indicate that there is a non-zero MVD only along the horizontal direction. 2 can be shown that there is a non-zero MVD only along the vertical direction. 3 can be shown to have a non-zero MVD along both the horizontal and vertical directions.

[0096] If the "mv_joint" syntax element for MVD signals that there are no non-zero MVD components, no further MVD information can be signaled. However, if the "mv_joint" syntax signals that there are one or two non-zero components, additional syntax elements can further signal each of the non-zero MVD components, as described below.

[0097] For example, a syntax element called "mv_sign" may be used to further specify whether the corresponding motion vector difference component is positive or negative.

[0098] In another example, a syntactic element called "mv_class" can be used to specify the class of motion vector differences between a predefined set of classes for corresponding non-zero MVD components. These predefined classes for motion vector differences can be used, for example, to divide a continuous size space of motion vector differences into non-overlapping ranges of classes. Thus, the signaled MVD classes indicate the size ranges of the corresponding MVD components. In the exemplary implementation shown in Table 3 below, higher classes correspond to motion vector differences with larger size ranges. The symbol (n,m) is used to represent the range of motion vector differences greater than n pixels and less than or equal to m pixels.

[0099] [Table 3]

[0100] In some other implementations, a syntactic element called “mv_bit” may be used to specify the integer part of the offset between the non-zero motion vector difference component and the magnitude of the start of the MV class size range that is signaled in correspondence. In some other implementations, a syntactic element called “mv_fr” may be used to specify the first two fractional bits of the motion vector difference of the corresponding non-zero MVD component, and a syntactic element called “mv_hp” may be used to specify the third fractional bit (high resolution bit) of the motion vector difference of the corresponding non-zero MVD component. Two “mv_fr” bits essentially provide a quarter-pixel MVD resolution, but the “mv_hp” bits can further provide a resolution of 1 / 8 pixels. In some other implementations, two or more “mv_hp” bits may be used to provide MVD pixel resolutions finer than 1 / 8 pixels. In some exemplary implementations, additional flags may be signaled at one or more of various levels to indicate whether MVD resolutions of 1 / 8 pixels or higher are supported. If an MVD resolution does not apply to a particular coding unit, the above syntactic elements for the corresponding unsupported MVD resolution may not be signaled.

[0101] However, in some other exemplary implementations, the resolution of motion vector differences across different MVD size classes may be differentiated or adaptive. Specifically, a high-resolution MVD for larger MVD sizes in higher MVD classes may not result in a statistically significant improvement in compression efficiency or coding gain. Therefore, an MVD may be coded with a reduced resolution (integer pixel resolution or fractional pixel resolution) or without increasing resolution for larger MVD size ranges corresponding to higher MVD size classes. The term “resolution” is sometimes further referred to as “pixel resolution.”

[0102] In some exemplary implementations, each MVD class may be associated with a single allowed resolution. In some other implementations, one or more MVD classes may each be associated with two or more optional MVD pixel resolutions. For example, adaptively allowed MVD pixel resolutions may include, but are not limited to, 1 / 64pel (pixel), 1 / 32pel, 1 / 16pel, 1 / 8pel, 1-4pel, 1 / 2pel, 1pel, 2pel, 4pel… (in descending order of resolution).

[0103] In some other exemplary implementations, for MV classes above a threshold MV class, only a single MVD value may be permitted. For example, such a threshold MV class may be MV_CLASS 2. Therefore, for MV_CLASS_2 and above, only having a single MVD value and not having fractional pixel resolution may be permitted.

[0104] Looking at the various composite interprediction modes in which each MV is predicted by a reference motion vector and can be coded by an MVD, the two MVDs can be signaled separately in the bitstream or together, as described above. Thus, in some exemplary implementations, in addition to the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes described above, another interprediction mode called JOINT_NEWMV can be introduced for a mode in which the MVDs are for joint signaling of reference lists 0 and 1. Specifically, when the interprediction mode is indicated as NEW_NEWMV, the MVDs of reference lists 0 and 1 are signaled separately, but when the interprediction mode is indicated as JOINT_NEWMV mode, the MVDs of reference lists 0 and 1 are signaled together. In particular, in the case of joint MVD, there may be cases where only one MVD called joint_delta_mv needs to be signaled and transmitted in the bitstream, and the MVDs in reference lists 0 and 1 can be derived from joint_delta_mv. The derived MVD can then be combined with a reference motion vector in reference list 0 or 1 to generate two motion vectors for locating the reference block for composite interpretation.

[0105] In some implementations of composite interpretation, the JOINT_NEWMV mode may be signaled together with the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes. In such implementations, the syntax may be included in the bitstream to indicate any one of these alternative composite interpretation modes at any of the various signaling levels (e.g., sequence level, picture level, frame level, slice level, tile level, superblock level, etc.). Alternatively, the JOINT_NEWMV mode may be implemented as a submode of the NEW_NEWMV mode. In other words, under the NEW_NEWMV mode, two MVDs of two reference blocks are either signaled together (hence the JOINT_NEWMV submode) or not signaled together (another submode of the NEW_NEWMV mode). In such an implementation, a first syntactic element may be included in the bitstream to indicate one of the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes, and if the first syntactic element indicates that the NEW_NEWMV mode has been selected for a coding block, then a second syntactic element may be further included in the bitstream, which may be extractable by the decoder to indicate whether the MVD of the coding block is signaled separately or together.

[0106] In some exemplary implementations, if the JOINT_NEWMV mode is signaled and the POC distances between two reference frames and the current frame are different, the MVD may be scaled for reference list 0 or reference list 1 based on the POC distance. Specifically, the distance between reference frame list 0 and the current frame may be denoted as td0, and the distance between reference frame list 1 and the current frame may be denoted as td1. If td0 is greater than or equal to td1, joint_mvd may be used directly for reference list 0, and the MVD for reference list 1 may be derived from joint_mvd based on equation (1).

number

[0107] Instead, if td1 is greater than or equal to td0, joint_mvd is used directly in reference list 1, and the MVD of reference list 0 is derived from joint_mvd based on equation (2).

number

[0108] In some exemplary implementations, another intercoding mode called AMVDMV may be added to a single reference case. When the AMVDMV mode is selected, it indicates that AMVD (Adaptive Motion Vector Difference) is applied to the signal MVD. To indicate whether AMVD is applied to the Joint MVD coding mode, a flag may be added under the JOINT_NEWMV mode, for example, named amvd_flag. When Adaptive MVD Resolution is applied to the Joint MVD coding mode, the MVDs of the two reference frames are signaled together, and the precision of the MVD may be implicitly determined by the magnitude of the MVD. Otherwise, the MVDs of two (or more) reference frames are signaled together, and conventional MVD coding without Adaptive MVD Resolution may be applied.

[0109] Turning to the composite intermodes, as shown in Figure 10, these modes are two different reference frames F i-1 and F i+1 By combining two hypotheses about the motion vectors MV0 and MV1 from the current frame F i This generates predictions for the blocks within. Therefore, two motion information components (e.g., motion vectors) can be signaled in the bitstream for each block.

[0110] Alternatively, as shown in Figure 11, interpolation can be used to create two reference frames F i-1 and F i+1 The information is combined, and currently frame F i Interpolated frames may be generated by projecting to the same time. Multiple TIP modes may be supported. In one TIP mode, the interpolated frame may be used as an additional reference frame. Current frame F i The coding block directly references the interpolated frame and can utilize information from two different references with only the overhead cost of a single interprediction mode. In another TIP mode, the interpolated frame is currently frame F, skipping any other conventional coding steps. i It can be directly assigned as the output of the decoding process for that purpose. This mode can offer considerable coding and complexity advantages, especially for low-bitrate applications.

[0111] In some implementations, Adaptive Motion Vector Resolution (AMVR) may be supported. In certain exemplary implementations, a total of seven MV accuracies (8, 4, 2, 1, 1 / 2, 1 / 4, 1 / 8) may be supported. For each prediction block, the AMVR encoder may search all supported accuracy values ​​and signal the decoder to the best accuracy.

[0112] To reduce encoder execution time, two MV precision sets are supported. Each precision set can contain four predefined precisions. The precision set can be adaptively selected at the frame level based on the maximum precision value of the frame. The maximum precision can be signaled in the frame header. Table 4 summarizes exemplary supported precision values ​​based on exemplary frame-level maximum precision.

[0113] [Table 4]

[0114] In some exemplary implementations of AMVR, there may be a frame-level flag indicating whether the frame's MV includes sub-perfect precision. AMVR can only be enabled if the value of the cur_frame_force_integer_mv flag is 0. In AMVR, if the block precision is lower than the maximum precision, the motion model and interpolation filter may not be signaled. If the block precision is lower than the maximum precision, the motion mode is inferred to translational motion, and the interpolation filter is inferred to the REGULAR interpolation filter. Similarly, if the block precision is either 4 pixels or 8 pixels, the inter-intra mode may remain unsignalized and may be inferred to be 0.

[0115] In some examples, optical flow-based methods may be employed to refine motion vectors (MVs) at the subblock level for composite prediction. In particular, optical flow equations can be applied to formulate a least-squares problem from which fine-grained motions can be derived from the gradients of composite interprediction samples. These fine-grained motions can be used to refine the MVs for each subblock within the prediction block, thus improving the interprediction quality. This coding feature may be an extension of the well-known bidirectional optical flow (BDOF) concept, as it supports MV refinement when the two reference blocks have an arbitrary temporal distance to the current block. In some implementations, additional intercomposite modes, e.g., as listed below, • NEAR_NEARMV_OPTFLOW, • NEAR_NEWMV_OPTFLOW, ·NEW_NEARMV_OPTFLOW, For example, four additional intercomplex modes, such as NEW_NEWMV_OPTFLOW, may be added.

[0116] In some implementations, signaling for composite interpretation modes now depends on the optical flow refinement flag for the block. For example, one flag called use_optflow may be signaled as a syntactic element in the bitstream before the syntactic element indicating the composite interpretation mode. The flag, e.g., use_optflow, can indicate whether an optical flow-based composite mode is used. If use_optlow is set to 1 (or true), these composite interpretation modes are called optical flow modes, and the reference MV type is defined similarly to conventional composite modes (e.g., NEAR_NEWMV_OPTFLOW has the same reference MV type as NEAR_NEWMV), but the composite prediction is based on a subblock-wise refined MV rather than the original MV. If, instead, use_optflow is set to 0 (false), optical flow refinement is not applied to the block now. Specifically, according to the disclosed method, the decoder may receive a first syntactic element signaled in the video bitstream indicating whether or not optical flow refinement is applied to a video block, and determine whether or not optical flow refinement is applied to the video block based on the value of the first syntactic element. After receiving the first syntactic element, the decoder may receive a second syntactic element from the video bitstream indicating the composite inter prediction mode of the video block depending on whether or not optical flow refinement is applied, and determine the composite inter prediction mode of the video block based on the value of the second syntactic element. The decoder may then predict the video block based on the determined composite inter prediction mode.

[0117] Therefore, in the above implementation, one flag is signaled as a syntactic element in the bitstream before signaling the composite interprediction mode among multiple composite interprediction modes to indicate whether the composite interprediction mode is without optical flow refinement or with optical flow refinement. The usage and distribution / statistics of the composite interprediction mode may differ when the optical flow refinement flag is on compared to when it is off. Such correlations between the use of optical flow refinement and composite interprediction modes can be utilized in signaling to improve coding efficiency.

[0118] The following various exemplary implementations for signaling schemes for optical flow refinement and composite interpretation may be used separately or combined in any order. Furthermore, the methods, encoders, and decoders in these implementations may each be implemented by processing circuits (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-temporary computer-readable medium. Hereinafter, the term "block" may be interpreted as a prediction block, coding block, or coding unit, i.e., CU.

[0119] In this document, the direction of a reference frame may be determined by whether the reference frame is before or after the current frame in the display order.

[0120] Regarding these implementations, in composite reference mode, if the POCs of both reference frames in a motion vector pair are greater than or less than the POC of the current frame, then the directions of the two reference frames may be the same. However, if the POC of one reference frame is greater than the POC of the current frame, and the POC of the other reference frame is smaller than the POC of the current frame, then the directions of the two reference frames will be different.

[0121] In some exemplary implementations, signaling of composite interprediction modes within a bitstream may depend on the optical flow refinement flag of the current block and / or the reference frame of the current block and / or the interprediction mode of the neighboring block and / or the reference frame of the neighboring block.

[0122] In some exemplary implementations, the context for encoding / decoding the signaling of composite interprediction modes may depend on the optical flow refinement flag of the current block and / or the reference frame of the current block and / or the interprediction mode of the neighboring block and / or the reference frame of the neighboring block.

[0123] In some exemplary implementations, if the optical flow refinement flag for the current block is false, one set of contexts may be used to code / decode the composite interprediction mode signaling. Otherwise, a different set of contexts is used instead.

[0124] In some exemplary implementations, the context for signaling a composite interprediction mode may depend on whether the optical flow refinement mode is permitted for the current block, and / or the optical flow refinement flag of the current block, and / or the reference frame of the current block, and / or the interprediction mode of the neighboring block, and / or the reference frame of the neighboring block. Optical flow refinement is not permitted if the corresponding flag should not be signaled in the bitstream. Whether optical flow refinement is permitted may be predefined or signaled at a higher level. If optical flow refinement is permitted, whether it is used for a particular block or at another level is signaled by the optical flow refinement flag described above.

[0125] In some exemplary implementations, when the optical flow refinement flag is on for a single block, only a subset of composite interprediction modes requiring signaling for at most one motion vector difference (MVD) are permitted. In other words, composite interprediction modes requiring signaling for two or more MVDs are not permitted. This constraint improves signaling efficiency by reducing the signaling space when the optical flow refinement flag is on.

[0126] In some exemplary implementations, when the optical flow refinement flag is on for a single block, only a subset of the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, JOINT_NEWMV, and JOINT_AMVDNEWMV composite interprediction modes are allowed and signaled. These allowed composite interprediction modes require, for example, that at most one MVD is signaled. Therefore, when optical flow refinement is applied to a block, the syntactic value space of composite interprediction modes for the video block may not span all possible composite interprediction modes, but instead may include a subset of the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, JOINT_NEWMV, and JOINT_AMVDNEWMV modes. Such constraints improve signaling efficiency by reducing the signaling space when the optical flow refinement flag is on.

[0127] In some exemplary implementations, when the optical flow refinement flag is on for a single block, only a subset of composite interprediction modes with a maximum of one MVD for all reference frames are permitted. In other words, composite interprediction modes containing two or more MVDs are not permitted and therefore not signaled. This constraint improves signaling efficiency by reducing the signaling space when the optical flow refinement flag is on.

[0128] In some exemplary implementations, when the optical flow refinement flag is on for a single block, a subset of the NEAR_NEARMV, NEAR_NEWMV, and NEW_NEARMV composite interprediction modes are allowed and signaled. Other composite interprediction modes are not allowed and therefore not signaled. These allowed composite interprediction modes contain at most one MVD. Therefore, when optical flow refinement is applied to a block, the syntactic value space of composite interprediction modes for the video block may not span all possible composite interprediction modes, but may instead contain a subset of the NEAR_NEARMV, NEAR_NEWMV, and NEW_NEARMV modes. Such constraints improve signaling efficiency by reducing the signaling space when the optical flow refinement flag is on.

[0129] In some exemplary implementations, there may be two permitted joint MVD coding modes, JOINT_NEWMV and JOINT_AMVDNEWMV, when the optical flow refinement flag is off. However, when the optical flow refinement flag is on for a block, at most one of these two joint MVD coding modes may be permitted. Such constraints improve signaling efficiency by reducing the signaling space when the optical flow refinement flag is on.

[0130] In some exemplary implementations, when the optical flow refinement flag is on for a single block, the only permitted joint MVD coding mode may be JOINT_AMVDNEWMV. This constraint improves signaling efficiency by reducing the signaling space when the optical flow refinement flag is on.

[0131] In some exemplary implementations, when the optical flow refinement flag for a block is currently on, the mapping between parsed syntactic values ​​and composite interprediction modes (i.e., the semantics of syntactic values) differs compared to when the optical flow refinement flag for a block is currently off. For example, the same syntactic value extracted for a composite interprediction mode syntactic element maps to different composite interprediction modes when the optical flow refinement flag is on and when it is off. In other words, the same syntactic value can dynamically point to different composite interprediction modes depending on the value of the optical flow refinement flag received / extracted from the bitstream. This implementation can be used to leverage the correlation between the optical flow flag and various composite interprediction modes. For example, different interprediction modes are likely to be invoked when the optical flow refinement flag is on and when it is off. Therefore, more efficient syntactic values ​​(e.g., fewer signaling bits) can be used for coding to map different optical flow refinement flag values ​​to more likely composite interprediction modes.

[0132] In some exemplary implementations, the optical flow refinement flag may be signaled after the composite interpretation mode, and the signaling of the optical flow refinement flag may depend on the composite interpretation mode. Such constraints improve signaling efficiency.

[0133] In some exemplary implementations, the context for signaling the optical flow refinement flag may depend on the composite interpretation mode. Such constraints improve signaling efficiency.

[0134] In some exemplary implementations, the optical flow refinement flag may be signaled only for a subset of composite interpretation modes. In this way, the overall amount of signaling for the optical flow refinement flag can be reduced. The subset of composite interpretation modes can be determined in such a way that overall coding efficiency is not significantly impaired.

[0135] In some exemplary implementations, the optical flow refinement flag may be signaled only for composite interprediction modes that have at most one signaled MVD for multiple reference frames, such as NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, JOINT_NEWMV, and a subset of the JOINT_AMVDNEWMV composite interprediction modes. Such constraints improve signaling efficiency.

[0136] In some exemplary implementations, the optical flow refinement flag may be signaled only for composite interprediction modes involving at most one MVD, such as a subset of the NEAR_NEARMV, NEAR_NEWMV, and NEW_NEARMV composite interprediction modes. Such constraints improve signaling efficiency.

[0137] In some exemplary implementations, the optical flow refinement flag may only be signaled for composite interprediction modes that either lack MVD, use adaptive MVD resolution, or have MV precision coarser than a threshold (e.g., MV precision coarser than 1 / 8 or 1 / 4 MV precision).

[0138] Figure 10 shows an exemplary logic flow 1000 based on the above implementation. The logic flow 1000 starts at S1001. At S1010, a first syntactic element is received, signaled in the video bitstream, indicating whether optical flow refinement is applied to the video block. At S1020, it is determined whether optical flow refinement is applied to the video block based on the value of the first syntactic element. At S1030, after receiving the first syntactic element, a second syntactic element is received from the video bitstream, indicating the composite inter prediction mode of the video block, depending on whether optical flow refinement is applied. At S1040, the composite inter prediction mode of the video block is determined based on the value of the second syntactic element. At S1050, the video block is predicted based on the determined composite inter prediction mode. The logic flow 1000 stops at S1099.

[0139] The following is example pseudocode for reading the optical flow refinement and compound interpretation syntax described above. #if CONFIG_OPTFLOW_REFINEMENT int use_optical_flow = 0; if (cm->features.opfl_refine_type == REFINE_SWITCHABLE && is_opfl_refine_allowed(cm, mbmi)) { use_optical_flow = aom_read_symbol(r, xd->tile_ctx->use_optflow_cdf[ctx], 2, ACCT_INFO(“use_optical_flow”)); } #endif / / CONFIG_OPTFLOW_REFINEMENT const int mode = #if CONFIG_OPTFLOW_REFINEMENT aom_read_symbol(r, xd->tile_ctx->inter_compound_mode_cdf[ctx], INTER_COMPOUND_REF_TYPES, ACCT_INFO(“inter_compound_mode_cdf”)); #else aom_read_symbol(r, xd->tile_ctx->inter_compound_mode_cdf[ctx], INTER_COMPOUND_MODES, ACCT_INFO(“inter_compound_mode_cdf”)); #endif / / CONFIG_OPTFLOW_REFINEMENT #if CONFIG_OPTFLOW_REFINEMENT if (use_optical_flow) { assert(is_inter_compound_mode(comp_idx_to_opfl_mode[mode])); return comp_idx_to_opfl_mode[mode]; } #endif / / CONFIG_OPTFLOW_REFINEMENT assert(is_inter_compound_mode(NEAR_NEARMV + mode)); return NEAR_NEARMV + mode; }

[0140] The operations described above can be combined or arranged in any quantity or order as needed. Two or more steps and / or operations may be performed in parallel. Embodiments and implementations of the disclosure may be used individually or in any order. Furthermore, each of the methods (or embodiments), encoders and decoders may be implemented by processing circuits (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-temporary computer-readable medium. Embodiments of the disclosure may be applied to luma blocks or chroma blocks. The term block may be interpreted as a prediction block, coding block, or coding unit, i.e., CU. The term block may also be used to refer to a transformation block. In the following sections, when we say block size, it may mean the width or height of the block, or the maximum width and height, or the minimum width and height, or the size of the area (width * height), or the aspect ratio of the block (width:height, or height:width).

[0141] The techniques described above can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, Figure 11 shows a computer system (1100) suitable for carrying out a particular embodiment of the disclosed subject matter.

[0142] Computer software can be coded using any suitable machine code or computer language that can undergo mechanisms such as assembly, compilation, and linking to create code that contains instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or that can be executed via interpretation, microcode execution, etc.

[0143] Instructions can be executed on various types of computers or computer components, including, for example, personal computers, tablet computers, servers, smartphones, game consoles, and Internet of Things devices.

[0144] The components shown in Figure 11 with respect to the computer system (1100) are essentially illustrative and are not intended to imply any limitation on the scope of use or functionality of the computer software implementing embodiments of this disclosure. Furthermore, the configuration of the components should not be construed as having any dependencies or requirements relating to any one or combination of components shown in the exemplary embodiments of the computer system (1100).

[0145] The computer system (1100) may include certain human interface input devices. The input human interface devices may include one or more of the following (only one of each is depicted): a keyboard (1101), a mouse (1102), a trackpad (1103), a touchscreen (1110), a data glove (not shown), a joystick (1105), a microphone (1106), a scanner (1107), and a camera (1108).

[0146] The computer system (1100) may also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (e.g., tactile feedback via a touchscreen (1110), data glove (not shown), or joystick (1105), although there may also be tactile feedback devices that do not function as input devices), audio output devices (e.g., speakers (1109), headphones (not shown)), visual output devices (e.g., screens (1110) including CRT screens, LCD screens, plasma screens, OLED screens, etc., each with or without touchscreen input functionality and with or without tactile feedback functionality, some of which can output three or more dimensions by means such as two-dimensional visual output or stereoscopic output, virtual reality glasses (not shown), holographic displays, smoke tanks (not shown)), and printers (not shown).

[0147] The computer system (1100) may also include human-accessible storage devices and associated media, such as optical media including CD / DVD ROM / RW (1120) with media such as CD / DVD (1121), thumb drives (1122), removable hard drives or solid-state drives (1123), legacy magnetic media such as tapes and floppy disks (not shown), and special ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0148] Those skilled in the art should also understand that the term “computer-readable medium” as used in connection with the subject matter of this disclosure does not include transmission media, carriers, or other transient signals.

[0149] The computer system (1100) may also include an interface (1154) to one or more communication networks (1155). The networks may be, for example, wireless, wired, or optical. The networks may further be local, wide-area, metropolitan, automotive, and industrial, real-time, latency-tolerant, etc. Examples of networks include local area networks such as Ethernet, cellular networks including wireless LAN, GSM, 3G, 4G, 5G, LTE, etc., wired television or wireless wide-area digital networks including cable television, satellite television, and terrestrial broadcast television, and automotive and industrial networks including CAN bus.

[0150] The aforementioned human interface device, human-accessible storage device, and network interface may be mounted on the core (1140) of the computer system (1100).

[0151] The core (1140) may include one or more central processing units (CPUs) (1141), graphics processing units (GPUs) (1142), dedicated programmable processing units in the form of field-programmable gate areas (FPGAs) (1143), hardware accelerators for specific tasks (1144), and graphics adapters (1150). These devices may be connected via a system bus (1148) along with read-only memory (ROM) (1145), random-access memory (1146), internal hard drives inaccessible to the user, and internal mass storage devices such as SSDs (1147). In some computer systems, the system bus (1148) can be accessed in the form of one or more physical plugs to enable expansion with additional CPUs, GPUs, etc. Peripheral devices may be connected directly to the core's system bus (1148) or via a peripheral bus (1149). For example, a screen (1110) may be connected to a graphics adapter (1150). The architecture for peripheral buses includes PCI, USB, and others.

[0152] Computer-readable media may contain computer code for performing various computer operations. The media and computer code may be specifically designed and constructed for the purposes of this disclosure, or they may be of a type that is well known and available to persons skilled in computer software technology.

[0153] While this disclosure has described several typical embodiments, variations, substitutions, and various alternative equivalents exist and are included within the scope of this disclosure. Therefore, those skilled in the art will understand that numerous systems and methods not expressly shown or described herein can be devised to embody the principles of this disclosure and thus fall within the spirit and scope of this disclosure. [Explanation of Symbols]

[0154] 0 Picture reference list, 1 Picture reference list, reference frame list, 100 Communication system, 110, 120, 130, 140 Terminal device, 150 Network, 200 Communication system, 201 Video source, 202 Stream, 203 Video encoder, 204 Encoded video data, 205 Streaming server, 206, 208 Client subsystem, 207, 209 Copy, 210 Video decoder, 211 Renderable video picture, 212 Display, 213 Video capture system, 220, 230 Electronic device, 301 Channel, 310 Video decoder, 312 Display, 315 Buffer memory, 319 Symbol, 320 Parser, 330 Electronic device, 331 Receiver, 351 Scaler / Inverse unit, 352 Intra-picture prediction unit, 353 Motion compensation prediction unit, 355 Aggregator, 356 Loop filter, 357 Reference picture memory, 358 Current picture buffer, 401 Video source, 403 Video encoder, 420 Electronic device, 430 Source coder, 432 Coding engine, 433 Decoder, 434 Reference picture memory, 435 Predictor, 440 Transmitter, 443 Video sequence, 445 Entropy coder, 450 Controller, 460 Channel, 503 Video encoder, 521 General controller, 522 Intra encoder, 523 Residual calculator, 524 Residual encoder, 525 Entropy encoder, 526 Switch, 528 Residual decoder, 530 Intercoder, 610 Video decoder, 671 Entropy decoder, 672 Intra decoder, 673 Residual decoder, 674 Reconstruction module, 680 Intercoder, 710 Square partition, 802 Vertical tripartition scheme, 804 Horizontal three-part partitioning, 910 Overall exemplary partition patterns, 902, 904, 906, 908 Base block, 920 Tree structure, 1000 Logical flow, 1100 Computer system, 1101 Keyboard, 1102 Mouse, 1103 Trackpad, 1105 Joystick, 1106 Microphone, 1107 Scanner, 1108 Camera, 1109 Speaker, 1110Touchscreen, 1120 CD / DVD ROM / RW, 1121 CD / DVD and other media, 1122 Thumb drive, 1123 Removable hard drive or solid-state drive, 1140 Core, 1141 Central Processing Unit, 1142 Graphics Processing Unit, 1143 Field-programmable gate area, 1144 Hardware accelerator, 1145 Graphics adapter, 1146 Random access memory, 1147 Internal mass storage device, 1148 System bus, 1149 Peripheral bus, 1150 Graphics adapter, 1154 Network interface, 1155 Communication network

Claims

1. A method for decoding a video block in a video bitstream, The steps include receiving a first syntactic element signaled in the video bitstream that indicates whether optical flow refinement is applied to the video block, A step of determining whether the optical flow refinement is applied to the video block based on the value of the first syntactic element, After receiving the first syntactic element, the step of receiving a second syntactic element from the video bitstream indicating the composite interpretation mode of the video block, depending on whether the optical flow refinement is applied, The steps include determining the composite interpretation mode of the video block based on the value of the second syntactic element, A step of predicting the video block based on the determined composite interpretation mode, Methods that include...

2. The method according to claim 1, wherein the value of the second syntactic element is determined by decoding the second syntactic element using a coding context that depends on whether the optical flow refinement is applied.

3. The method according to claim 2, wherein, when optical flow refinement is applied, a first context is used to decode the second syntactic element indicating the composite interpretation mode of the video block, and, when optical flow refinement is not applied, a second context different from the first context is used to decode the second syntactic element indicating the composite interpretation mode of the video block.

4. The method according to claim 1, wherein when the optical flow refinement is applied to the video block, only composite interprediction modes that require signaling of one motion vector difference or do not require signaling of a motion vector difference are permitted.

5. The method according to claim 4, wherein when the optical flow refinement is enabled for the video block, the syntactic value space of the composite interprediction mode of the video block includes a subset of the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, JOINT_NEWMV, and JOINT_AMVDNEWMV modes.

6. The method according to claim 1, wherein, when the optical flow refinement is applied to the video block, only composite interprediction modes including or not including a motion vector difference are permitted.

7. The method according to claim 6, wherein when the optical flow refinement is applied to the video block, the syntactic value space of the composite interprediction mode of the video block includes a subset of the NEAR_NEARMV, NEAR_NEWMV, and NEW_NEARMV modes.

8. Regarding the joint motion vector difference composite interpretation mode, If the optical flow refinement is not applied to the video block, a two-joint motion vector difference composite interpretation mode is permitted. The method according to claim 1, wherein when the optical flow refinement is applied to the video block, only one of the two joint motion vector difference composite interpretation modes is permitted.

9. The two joint motion vector difference composite interpretation modes include the JOINT_NEWMV mode and the JOINT_AMVDNEWMV mode, When the optical flow refinement is applied to the video block, only one of the two permitted joint motion vector difference composite interpretation modes is the JOINT_AMVDNEWMV mode. The method according to claim 8.

10. The step of determining the composite interpretation mode of the video block is: The steps include determining a mapping between possible values ​​of the second syntactic element and a plurality of composite interprediction modes, based on whether the optical flow refinement of the video block is applied, The method according to claim 1, further comprising the step of determining the composite interpretation mode of the video block based on the value of the second syntactic element and the mapping.

11. The method according to claim 10, wherein the mapping between the possible values ​​of the second syntactic element and the plurality of composite interprediction modes differs when the optical flow refinement is applied and when the optical flow refinement is not applied.

12. A method for decoding a video block in a video bitstream, The steps include receiving a first syntactic element signaled in the video bitstream that indicates the composite interpretation mode of the video block among a plurality of composite interpretation modes, The steps include determining the composite interpretation mode of the video block based on the value of the first syntactic element, A step of determining whether a second syntactic element of the video block is included in the video bitstream, wherein the second syntactic element indicates whether optical flow refinement is applied to the video block, and the second syntactic element is included in the video bitstream only if the composite interpretation mode of the video block requires at most one signaled motion vector difference, Methods that include...

13. The method according to claim 12, wherein the subset of the plurality of composite interprediction modes includes the modes NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, JOINT_NEWMV, and JOINT_AMVDNEWMV.

14. The method according to claim 12, wherein the second syntactic element of the video block is included in the video bitstream only if the composite interpretation mode of the video block has at most one motion vector difference.

15. The method according to claim 12, wherein the second syntactic element of the video block is included in the video bitstream only if the composite interpretation mode of the video block does not involve different motion vectors, or uses adaptive motion vector resolution, or has a motion vector precision coarser than a predetermined precision threshold.

16. An electronic device comprising a memory for storing instructions and a processor for executing the stored instructions and carrying out the method according to any one of claims 1 to 11.

17. An electronic device comprising a memory for storing instructions and a processor for executing the stored instructions and carrying out the method according to any one of claims 12 to 15.

18. A computer program that, when executed by a processor, includes a computer instruction causing the processor to perform the method according to any one of claims 1 to 11.

19. A computer program that, when executed by a processor, includes a computer instruction causing the processor to perform the method according to any one of claims 12 to 15.

Citation Information

Patent Citations

  • Optical flow estimation for motion-compensated prediction in video coding

    JP2020522200A

  • Symmetric motion vector differential coding

    JP2022516433A

  • BDOF-based inter prediction method and device

    US20220078439A1

  • Signaling for motion vector refinement

    WO2021061023A1