Joint signaling method for motion vector differences

JP2026143700APending Publication Date: 2026-09-08TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026097427
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-03-22
Filing Date
2026-06-11
Publication Date
2026-09-08

Smart Images

  • Figure 2026143700000001_ABST
    Figure 2026143700000001_ABST
Patent Text Reader

Abstract

This provides an improved signaling method for joint motion vector difference for interpreting video blocks. [Solution] The method extracts a flag from the video bitstream received by the device indicating whether the interprediction mode is JOINT_NEWMV mode with respect to the current block of the current frame. JOINT_NEWMV mode indicates that a first delta motion vector (MV) for a first reference frame from reference list 0 and a second delta MV for a second reference frame from reference list 1 are signaled together. If it is JOINT_NEWMV mode, the first delta MV and the second delta MV are derived based on the extracted joint delta motion vector (MV), and the current block is decoded based on the first delta MV and the second delta MV.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001]

[0001] Reference This application claims priority to U.S. Provisional Application No. 63 / 280,506 filed on November 17, 2021 and U.S. Provisional Application No. 63 / 289,140 filed on December 14, 2021, the entire contents of both of which are incorporated herein by reference. This application also claims priority to U.S. Application No. 17 / 700,729, which is not a provisional application, filed on March 22, 2022, the entire content of which is incorporated herein by reference.

[0002]

[0002] Technical field The present disclosure relates to video coding and / or decoding techniques, and in particular, to signaling and improved design of joint motion vector difference for coding and / or decoding.

Background Art

[0003]

[0003] This background description provided herein is for the purpose of generally presenting the context of the present disclosure. Work carried out under the name of the current inventor, to the extent that such work is described not only in this background section but also in any other aspects of the description that may not otherwise qualify as prior art as of the filing date of the present application, is not expressly or implicitly admitted as prior art against the present disclosure.

[0004]

[0004] Video coding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can contain a series of pictures, each picture having spatial dimensions of, for example, 1920x1080 luminance samples and associated full or subsampled chrominance samples. The series of pictures can have a fixed or variable picture rate (alternatively called a frame rate), for example, 60 pictures / second or 60 frames per second. Uncompressed video has specific bitrate requirements for streaming or data processing. For example, video with 1920x1080 pixel resolution, a frame rate of 60 frames / second, and 4:2:0 chroma subsampling in 8 bits / pixel / color channel requires a bandwidth of nearly 1.5 Gbit / s. One hour of such video requires more than 600 GB of storage space.

[0005]

[0005] One of the purposes of video coding and decoding is to reduce the redundancy of uncompressed input video signals through compression. Compression can, in some cases, help reduce the aforementioned bandwidth and / or storage space requirements by more than two orders of magnitude. Both lossless and non-lossless compression, as well as combinations thereof, can be used. Lossless compression is a technique in which an exact copy of the original signal can be reconstructed from the compressed original signal by the decoding process. Non-lossless compression is a coding / decoding process in which the original video information is not fully preserved during coding and is not fully recoverable during decoding. When using non-lossless compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signals is small enough that the reconstructed signal is useful for its intended use despite some information loss. In the case of video, non-lossless compression is widely used in many applications. The amount of distortion that is acceptable depends on the application. For example, users of certain consumer video streaming applications may be able to tolerate higher distortion than users of film or television broadcast applications. The compression ratio achievable by a particular coding algorithm can be selected or adjusted to reflect various distortion tolerances: higher tolerance distortion generally allows for coding algorithms that result in higher losses and higher compression ratios.

[0006]

[0006] Video encoders and decoders can utilize techniques from a wide range of categories and steps, including, for example, motion compensation, Fourier transform, quantization, and entropy coding.

[0007]

[0007] Video codec technology may include a technique known as intra coding. In intra coding, sample values ​​are represented without referencing samples or other data from a previously reconstructed reference picture. In some video codecs, the picture is spatially subdivided into blocks of samples. If all blocks of samples are coded in intra mode, the picture may be referred to as an intra picture. Intra pictures and their derivatives (e.g., independent decoder refresh pictures) can be used to reset the decoder state and thus can be used as the first picture in a coded video bitstream and video session, or as a still image. Samples of a block after intra prediction may be subject to conversion to the frequency domain, and the conversion coefficients thus produced can be quantized before entropy coding. Intra prediction represents a technique to minimize the sample values ​​in the domain before conversion. In some cases, the smaller the converted DC value and the smaller the AC coefficient, the fewer bits are required for a given quantization step size to represent a block after entropy coding.

[0008]

[0008] Traditional intra-coding, such as that known from MPEG-2 generation coding techniques, does not use intra-prediction. However, some newer video compression techniques include techniques that attempt to code / decode blocks based on, for example, surrounding sample data and / or metadata obtained during encoding / decoding that are spatially adjacent and precede the block of data being coded or decoded in the decoding order. Such techniques are hereafter referred to as “intra-prediction” techniques. Note that in at least some cases, intra-prediction uses only reference data from the current picture being reconstructed, and not from other reference pictures.

[0009]

[0009] There may be many different forms of intra-prediction. If more than one such technique can be used in a given video coding technique, the technique used may be referred to as an intra-prediction mode. One or more intra-prediction modes may be provided in a particular codec. In some cases, a mode may have submodes and / or be associated with various parameters, and the mode / submode information and intra-coding parameters for a block of video may be coded individually or collectively included in a mode codeword. The codeword used for a given combination of mode, submode, and / or parameters may affect the coding efficiency gain by intra-prediction and therefore may similarly affect the entropy coding technique used to convert the codeword into a bitstream.

[0010]

[0010] A certain mode of intra prediction was introduced in H.264, improved in H.265, and further improved with new coding techniques such as Joint Search Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). Generally, in intra prediction, a prediction block can be formed using available neighboring sample values. For example, available values ​​from a particular set of neighboring samples along a particular direction and / or line can be copied into a predictor block. A reference to the direction to be used can be coded in the bitstream or may be predicted itself.

[0011]

[0011] Referring to Figure 1A, the lower right shows a subset of nine predictor directions specified in H.265 (corresponding to 33 of the 35 intra-modes specified in H.265, corresponding to 33 angular modes). The point where the arrows converge (101) represents the predicted sample. The arrows indicate a direction, and nearby samples from that direction are used to predict the sample at 101. For example, arrow (102) indicates that sample (101) is predicted from one or more nearby samples pointing upwards and to the right at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more nearby samples pointing downwards and to the left of sample (101) at an angle of 22.5 degrees from the horizontal.

[0012]

[0012] Referring further to Figure 1A, a 4x4 sample square block (104) is shown in the upper left (indicated by a thick dashed line). The square block (104) contains 16 samples, each labeled with "S" and its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample (from the top) in the Y dimension and the first sample (from the left) in the X dimension. Similarly, sample S44 is the fourth sample in block (104) in both the Y and X dimensions. Since the block is the size of 4x4 samples, S44 is in the lower right. Furthermore, an exemplary reference sample following a similar numbering scheme is shown. The reference sample is labeled with R and its Y position (e.g., row index) and X position (column index) relative to block (104). In both H.264 and H.265, predicted samples adjacent to the block being reconstructed are used.

[0013]

[0013] Intra-picture prediction in block 104 may begin by copying a reference sample value from a nearby sample according to the signaled prediction direction. For example, suppose the coded video bitstream includes signaling with respect to this block 104 that indicates the prediction direction of arrow (102) - i.e., the sample is predicted from one or more prediction samples that are angled 45 degrees to the upper right from the horizontal. In that case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05, and sample S44 is predicted from reference sample R08.

[0014]

[0014] In some cases, in order to calculate a reference sample, the values ​​of multiple reference samples may be combined, for example, by interpolation; in particular when the direction is not evenly divisible by 45 degrees.

[0015]

[0015] The number of possible directions is increasing as video coding technology continues to develop. For example, in H.264 (2003), nine different directions were available for intra-prediction. This increased to 33 in H.265 (2013), and as of the present disclosure, JEM / VVC / BMS can support up to 65 directions. Experimental studies have been conducted to help identify the most appropriate intra-prediction direction, and using certain techniques in entropy coding, the most appropriate direction can be encoded with fewer bits while accepting a certain bit penalty for the direction. Furthermore, the direction itself can often be predicted from the nearby directions used in the intra-prediction of the nearby block being decoded.

[0016]

[0016] Figure 1B shows a schematic diagram (180) of 65 intra prediction directions according to JEM, illustrating the increasing number of prediction directions in various coding techniques that have evolved over time.

[0017]

[0017] The way bits representing intra-prediction directions are mapped to prediction directions in a coded video bitstream may differ from video coding technique to video coding technique; for example, it may range from a simple direct mapping of prediction directions to complex adaptive schemes including intra-prediction modes, codewords, most probable modes, and similar techniques. However, in any case, there may be certain intra-prediction directions in video content that are statistically less likely than certain other directions. Since the goal of video compression is to reduce redundancy, these less likely directions can be represented with more bits than more likely directions in a well-designed video coding technique.

[0018]

[0018] Inter-picture prediction or inter-prediction can be based on motion compensation. In motion compensation, sample data from a previously reconstructed picture or a portion thereof (reference picture) can be used to predict the newly reconstructed picture or a portion of a picture (e.g., a block) after being spatially shifted in the direction indicated by a motion vector (hereinafter referred to as MV). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or three dimensions, the third dimension being an indication of the reference picture in use (the latter being analogous to the time dimension).

[0019]

[0019] In some video compression techniques, the current MV applicable to a particular area of ​​sample data can be predicted from other MVs, for example, those relating to other areas of sample data spatially adjacent to the area being reconstructed and preceding the current MV in the order of decoding. In this way, by relying on the reduction of redundancy in interrelated MVs, the overall amount of data required to code the MV can be greatly reduced, thereby increasing compression efficiency. MV prediction can work effectively because, for example, when coding an input video signal derived from a camera (known as natural video), there is a statistical likelihood that an area larger than the area to which a single MV is applicable will move in a similar direction in the video sequence, and therefore, in some cases, can be predicted using similar motion vectors derived from the MVs of adjacent areas. This results in an actual MV for a given area that is similar to or identical to the MV predicted from the surrounding MVs. Such an MV can be represented after entropy coding with fewer bits than would be used if the MV were coded directly rather than being predicted from nearby MVs. In some cases, MV prediction can be an example of lossless compression of the signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, MV prediction itself can be lossless due to rounding errors, for example, when calculating the predictor from several surrounding MVs.

[0020]

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec.H.265, “High Efficiency Video Coding”, December 2016). Of the many MV prediction mechanisms specified in H.265, the one described below is a technique hereafter referred to as “spatial merge”.

[0021]

[0021] Specifically, referring to Figure 2, the current block (201) contains samples discovered by the encoder during the motion search process so that it is predictable from previous blocks of the same size that have been spatially shifted. Instead of directly coding its MV, the MV can be derived from the most recent reference picture (in the order of decoding) using the MV associated with any of the five surrounding samples (202 to 206, respectively) designated as A0, A1 and B0, B1, B2, from the metadata associated with one or more reference pictures. In H.265, the MV prediction can use predictors from the same reference pictures used by the adjacent blocks. [Overview of the project]

[0022]

[0022] This disclosure describes various embodiments of methods, apparatus, and computer-readable storage media for encoding and / or decoding video.

[0023]

[0023] In one aspect, an embodiment of the present disclosure provides a method for decoding interpredicted video blocks. The method is: a device including memory for storing instructions and a processor for communicating with memory receives a coded video bitstream; the device retrieves from the coded video bitstream a flag indicating whether the interpredicted mode is JOINT_NEWMV mode with respect to the current block of the current frame, wherein JOINT_NEWMV mode is a first delta motion vector for a first reference frame from reference list 0 Step 1830, indicating that the joint delta motion vector (MV) and the second delta MV for the second reference frame from reference list 1 are signaled together; in response to the flag indicating that the interprediction mode is JOINT_NEWMV mode: the device retrieves the joint delta motion vector (MV) for the current block from the coded video / bitstream; and the device derives a first delta MV and a second delta MV based on the joint delta MV; and / or in step 1840, the device decodes the current block of the current frame based on the first delta MV and the second delta MV.

[0024]

[0024] In another aspect, embodiments of the present disclosure provide an apparatus for video coding and / or decoding. The apparatus includes a memory for storing instructions and a processor for communicating with the memory. When the processor executes an instruction, the processor is configured to cause the apparatus to perform the above method for video coding and / or decoding.

[0025]

[0025] Furthermore, aspects of the present disclosure provide a video coding or decoding device or apparatus that includes a circuit configured to perform any of the above-described implementations.

[0026]

[0026] In another aspect, embodiments of the present disclosure provide a non-transitory computer-readable storage medium storing instructions, which, when executed by a computer for video encoding and / or decoding, cause the computer to perform the above method for video encoding and / or decoding.

[0027]

[0027] The above and other aspects and their implementations are described in more detail in the drawings, the specification, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0028]

[0028] Further features, properties and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Figure 1A]

[0029] FIG. 1A is a schematic diagram of an exemplary subset of intra prediction direction modes. [Figure 1B]

[0030] FIG. 1B shows an example of an exemplary intra prediction direction. [Figure 2]

[0031] FIG. 2 shows a schematic diagram of a current block and neighboring spatial merge candidates for motion vector prediction in one example. [Figure 3]

[0032] FIG. 3 shows a schematic block diagram of a simplified communication system (300) according to an exemplary embodiment. [Figure 4]

[0033] FIG. 4 shows a schematic block diagram of a simplified communication system (400) according to an exemplary embodiment. [Figure 5]

[0034] FIG. 5 shows a schematic block diagram of a simplified video decoder according to an exemplary embodiment. [Figure 6]

[0035] FIG. 6 shows a schematic block diagram of a simplified video encoder according to an exemplary embodiment. [Figure 7]

[0036] Figure 7 shows a block diagram of a video encoder according to another exemplary embodiment. [Figure 8]

[0037] Figure 8 shows a block diagram of a video decoder according to another exemplary embodiment. [Figure 9]

[0038] Figure 9 shows a coding block partitioning scheme according to an exemplary embodiment of the present disclosure. [Figure 10]

[0039] Figure 10 shows another method of coding block partitioning according to an exemplary embodiment of the present disclosure. [Figure 11]

[0040] Figure 11 shows another method of coding block partitioning according to an exemplary embodiment of the present disclosure. [Figure 12]

[0041] Figure 12 shows an exemplary partitioning scheme for the base block into the coding block. [Figure 13]

[0042] Figure 13 shows an exemplary ternary partitioning scheme. [Figure 14]

[0043] Figure 14 shows an exemplary quadtree binary tree coding block partitioning scheme. [Figure 15]

[0044] Figure 15 shows a method for partitioning a coding block into a plurality of transformation blocks, and the coding order of the transformation blocks, according to an exemplary embodiment of the present disclosure. [Figure 16]

[0045] Figure 16 shows another method for partitioning a coding block into multiple transformation blocks, and the coding order of the transformation blocks, according to exemplary embodiments of the present disclosure. [Figure 17]

[0046] Figure 17 illustrates another method for partitioning a coding block into multiple transformation blocks, according to an exemplary embodiment of the present disclosure. [Figure 18]

[0047] Figure 18 shows a flowchart of the method according to an exemplary embodiment of the present disclosure. [Figure 19]

[0048] Figure 16 shows a schematic diagram of a computer system according to an embodiment of the present disclosure. [Modes for carrying out the invention]

[0029]

[0049] The present invention will be described in detail below with reference to the accompanying drawings, which constitute part of the invention and illustrate specific examples of embodiments. However, it should be noted that the invention can be carried out in a variety of different forms, and therefore the subject matter covered or claimed is intended to be construed as not being limited to any embodiment described below. It should also be noted that the invention can be carried out as a method, device, component, or system. Accordingly, embodiments of the invention can take the form of, for example, hardware, software, firmware, or any combination thereof.

[0030]

[0050] Throughout the specification and claims, terms may have nuances implied or suggested in context beyond their expressly stated meaning. The phrases “in one embodiment” or “in some embodiments” used herein do not necessarily refer to the same embodiment, and the phrases “in another embodiment” or “in other embodiments” used herein do not necessarily refer to different embodiments. Similarly, the phrases “in one implementation” or “in some implementations” used herein do not necessarily refer to the same implementation, and the phrases “in another implementation” or “in other implementations” used herein do not necessarily refer to different implementations. For example, the claimed subject matter is intended to include, in whole or in part, exemplary combinations of embodiments / implementations.

[0031]

[0051] In general, terms can be understood, at least partially, from their usage in context. For example, terms like "and," "or," or "and / or," when used in this context, can have various meanings that may depend at least partially on the context in which such terms are used. Typically, when used to relate a list such as A, B, or C, "or" is intended to mean "A, B, or C" in an exclusive sense, as well as "A, B, or C" in an inclusive sense. Furthermore, the terms "one or more" or "at least one" as used in this context may be used, at least partially context-dependent, to describe some feature, structure, or characteristic in a singular sense, or to describe a combination of features, structures, or characteristics in a plural sense. Similarly, terms like "a," "an," or "the" can also be understood, at least partially context-dependent, to convey either a singular or plural usage. Furthermore, the terms "based on" or "determined by" are not necessarily intended to convey an exclusive set of factors; rather, here too, at least partially context-dependent, they may allow for the presence of additional factors that are not necessarily explicitly stated.

[0032]

[0052] Figure 3 shows a simplified block diagram of a communication system (300) according to an embodiment of the present disclosure. The communication system (300) includes a plurality of terminal devices capable of communicating with each other, for example, via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of Figure 3, the first pair of terminal devices (310) and (320) are capable of one-way transmission of data. For example, terminal device (310) can encode video data (e.g., video data of a stream of video pictures captured by terminal device (310)) for transmission to other terminal devices (320) via the network (350). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. The terminal device (320) can receive coded video data from the network (350), decode the coded video data to restore the video picture, and display the video picture according to the restored video data. One-way data transmission may be implemented in media serving applications, etc.

[0033]

[0053] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) for bidirectional transmission of coded video data, for example, during a video conferencing application. With regard to bidirectional transmission of data, for example, each terminal device of terminal devices (330) and (340) can code video data (e.g., video data of a stream of video pictures captured by the terminal device) for transmission to the other terminal device of terminal devices (330) and (340) via the network (350). Each terminal device of terminal devices (330) and (340) is also capable of receiving coded video data transmitted by the other terminal device of terminal devices (330) and (340), decoding the coded video data to restore the video pictures, and displaying the video pictures on an accessible display device according to the restored video data.

[0034]

[0054] In the example in Figure 3, terminal devices (310), (320), (330), and (340) are implemented as a server, a personal computer, and a smartphone, respectively, but the applicability of the principles of this disclosure is not limited thereto. Embodiments of this disclosure may be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing equipment, and / or similar devices. Network (350) represents any number or type of network that carries coded video data between terminal devices (310), (320), (330), and (340), including, for example, wired and / or wireless communication networks. Communication network (350) can exchange data via circuit switching, packet switching, and / or other types of channels. Typical networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this disclosure, the architecture and topology of the network (350) may not be important to the operation of this disclosure unless expressly described below.

[0035]

[0055] Figure 4 shows an example of the application of the disclosed subject matter, illustrating the arrangement of a video encoder and video decoder in a video streaming environment. The disclosed subject matter may be equally applicable to other video applications, including, for example, video conferencing, digital television broadcasting, gaming, virtual reality, and the storage of compressed video on digital media (including CDs, DVDs, memory sticks, etc.).

[0036]

[0056] The video streaming system may include a video source (401), a video capture subsystem (413) which may include, for example, a digital camera, for generating a stream of uncompressed video pictures or images (402). In one example, the stream of video pictures (402) includes samples recorded by the digital camera of the video source 401. The stream of video pictures (402), which is drawn as a thicker line to emphasize the larger amount of data compared to encoded video data (404) (or encoded video bitstream), can be processed by an electronic device (420) which includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof, and can operate or realize aspects of the disclosed subject matter as described in detail below. The encoded video data (404) (or encoded video bitstream (404)), which is depicted as a thin line to emphasize the smaller amount of data compared to the uncompressed video picture (402) stream, can be stored in a streaming server (405) or directly in a downstream video device (not shown) for future use. One or more streaming client subsystems, such as client subsystems (406) and (408) in Figure 4, can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include a video decoder (410) within, for example, an electronic device (430). A video decoder (410) decodes an incoming copy (407) of the encoded video data to produce an output stream (411) of a decompressed video picture that can be rendered on a display (412) (e.g., a display screen) or other rendering device (not shown).The video decoder 410 may be configured to perform some or all of the various functions described herein. In some streaming systems, encoded video data (404), (407), and (409) (e.g., video bitstream) can be encoded according to a specific video coding / compression standard. Examples of these standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC). The disclosures may be used in the context of VVC and other video coding standards.

[0037]

[0057] It should be noted that electronic devices (420) and (430) may include other components (not shown). For example, electronic device (420) may include a video decoder (not shown), and electronic device (430) may include a video encoder (not shown).

[0038]

[0058] Figure 5 shows a block diagram of a video decoder (510) according to some embodiment of the present disclosure. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used in place of the video decoder (410) in the example of Figure 4.

[0039]

[0059] The receiver (531) is capable of receiving one or more coded video sequences that are to be decoded by the video decoder (510). In the same or different embodiments, it is possible to receive one coded video sequence at a time, provided that the decoding of each coded video sequence is independent of other coded video sequences. Each video sequence may relate to multiple video frames or images. The coded video sequences can be received from a channel (501), which may be a hardware / software link to a storage device that stores coded video data or to a streaming source that transmits coded video data. The receiver (531) is capable of receiving coded video data together with other data, such as coded audio data and / or auxiliary data streams, which can be transferred to their respective processing circuits (not shown). The receiver (531) can isolate the coded video sequences from other data. To address network jitter, the buffer memory (515) may be located between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In certain applications, the buffer memory (515) may be implemented as part of the video decoder (510). In other applications, it may be separate from and outside the video decoder (510) (not shown). In yet another application, for example to address network jitter, the buffer memory (not shown) may exist outside the video decoder (510), and another additional buffer memory (515) may exist inside the video decoder (510) to handle playback timing, for example.If the receiver (531) is receiving data from a store / forward device with sufficient bandwidth and controllability, or from an isochronous network, the buffer memory (515) may not be required or can be made smaller. For use in best-effort packet networks such as the Internet, a sufficiently sized buffer memory (515) may be required, and that size may be relatively large. Such a buffer memory may be implemented in an adaptive size and may also be implemented at least partially in an operating system or similar element (not shown) outside of the video decoder (510).

[0040]

[0060] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the coded video sequence. These categories of symbols may include information used to manage the operation of the video decoder (510), and potentially information for controlling a rendering device such as a display (512) (e.g., a display screen), which may or may not be an integrated part of an electronic device (530), as shown in Figure 5, but may be coupled to an electronic device (530). The rendering device control information may be in the form of Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not shown). The parser (520) may parse / entropy decode the coded video sequence received by the parser (520). The entropy coding of the video sequence to be coded can follow video coding techniques or standards, and can follow various principles including variable-length coding, Huffman coding, and arithmetic coding with or without context influence. The parser (520) can extract from the coded video sequence a set of subgroup parameters for at least one subgroup of pixels in the video decoder, based on at least one parameter corresponding to a subgroup. Subgroups can include groups of pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), predictive units (PU), etc. The parser (520) can also extract information from the coded video sequence, such as transform coefficients (e.g., Fourier transform coefficients), quantization parameter values, and motion vectors.

[0041]

[0061] The parser (520) can perform entropy decoding / analysis on the video sequence received from buffer memory (515) in order to generate symbols (521).

[0042]

[0062] The reconstruction of the symbol (521) may include multiple different processing or functional units, depending on the type of coded video picture or part thereof (inter and intra picture, inter and intra block) and other factors. How and which units are included can be controlled by subgroup control information parsed from the coded video sequence by the parser (520). The flow of such subgroup control information between the parser (520) and the following multiple processing or functional units is not depicted for clarity.

[0043]

[0063] Beyond the functional blocks already described, the video decoder (510) can be conceptually subdivided into several functional units, as described below. In a practical implementation operating under commercial constraints, many of these units can interact closely with each other and be at least partially integrated. However, for the purpose of illustrating the various functions of the disclosed subject matter, a conceptual subdivision into functional units is employed in the following disclosure.

[0044]

[0064] The first unit may be a scaler / inverse unit (551). The scaler / inverse unit (551) can receive not only quantized transformation coefficients but also control information (including which type of inverse transform to use, block size, quantization factors / parameters, quantization scaling matrix, etc.) from the parser (520) as symbols (521). The scaler / inverse unit (551) can output a block containing sample values ​​that can be input to the aggregator (555).

[0045]

[0065] In some cases, the output samples of the scaler / inverse transform (551) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from previously reconstructed pictures, but can use prediction information from previously reconstructed portions of the current picture. Such prediction information can be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) can generate blocks of the same size and shape as the block being reconstructed by using information from surrounding blocks that have already been reconstructed and stored in the current picture buffer (558). The current picture buffer (558) buffers, for example, partially reconstructed current pictures and / or fully reconstructed current pictures. The aggregator (555) may, in some cases, add the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) on a sample-by-sample basis.

[0046]

[0066] Otherwise, the output samples of the scaler / inverse unit (551) may be associated with intercoded, motion-compensated blocks. In such cases, the motion-compensated prediction unit (553) can access the reference picture memory (557) to retrieve samples to be used for interpicture prediction. After motion-compensating the retrieved samples according to the symbols (521) associated with the blocks, these samples are added by the aggregator (555) to the output of the scaler / inverse unit (551) (the output of unit 551 may be called residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (557) from which the motion-compensated prediction unit (553) fetches prediction samples can be controlled by motion vectors available to the motion-compensated prediction unit (553), in the form of symbols (521) which may have, for example, X,Y components (shift) and a reference picture component (time). Furthermore, motion compensation may involve interpolation of sample values, such as those retrieved from reference picture memory (557), when accurate motion vectors of sub-samples are used, and may also be related to a motion vector prediction mechanism, etc.

[0047]

[0067] The output samples of the aggregator (555) may be subject to various loop filtering techniques within the loop filter unit (556). The video compression technique may include an in-loop filtering technique controlled by parameters contained in the coded video sequence (also called the coded video bitstream), which are made available to the loop filter unit (556) as symbols (521) from the parser (520). The in-loop filtering technique may respond to metadata obtained during the decoding of earlier parts (in the order of decoding) of the coded picture or coded video sequence, as well as to previously reconstructed loop-filtered sample values. Several types of loop filters may be included as part of the loop filter unit 556 in various orders, as will be described in more detail below.

[0048]

[0068] The output of the loop filter unit (556) can be a sample stream that can be output to the rendering device (512) as well as stored in reference picture memory (557) for use in future inter-picture prediction.

[0049]

[0069] Once a given coded picture is fully reconfigured, it can be used as a reference picture for future inter-picture predictions. For example, once the coded picture corresponding to the current picture is fully reconfigured and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a fresh current picture buffer can be reallocated before the reconfiguration of subsequent coded pictures begins.

[0050]

[0070] The video decoder (510) is capable of performing decoding operations according to a predetermined video compression technique adopted in a standard such as ITU-T Rec.H.265. A coded video sequence can conform to the syntax specified by the video compression technique or standard being used, in the sense that the coded video sequence conforms to both the syntax of the video compression technique or standard and a profile documented in the video compression technique or standard. Specifically, a profile allows for the selection of a particular tool from among all tools available in the video compression technique or standard, with that tool being the only tool usable under that profile. Furthermore, to conform to the standard, the complexity of the coded video sequence may be within a range defined by the level of the video compression technique or standard. In some cases, this level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further restricted by the Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.

[0051]

[0071] In some exemplary embodiments, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, a time, space, or signal-to-noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, forward error correction code, etc.

[0052]

[0072] Figure 6 shows a block diagram of a video encoder (603) according to an exemplary embodiment of the present disclosure. The video encoder (603) may be included in an electronic device (620). The electronic device (620) may include a transmitter (640) (e.g., a transmitting circuit). The video encoder (603) can be used in place of the video encoder (403) in the example of Figure 4.

[0053]

[0073] The video encoder (603) can receive video samples from a video source (601) (not part of the electronic device (620) in the example of Figure 6) that is capable of capturing video images to be coded by the video encoder (603). In another example, the video source (601) may be implemented as part of the electronic device (620).

[0054]

[0074] The video source (601) can provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 YCrCB, RGB, ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (601) may be a storage device capable of storing pre-prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a series of individual pictures or images that convey motion when viewed in sequence. The picture itself can be organized as a spatial array of pixels, where each pixel may contain one or more samples depending on the sampling structure, color space, etc., in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following explanation focuses on samples.

[0055]

[0075] According to some exemplary embodiments, the video encoder (603) can encode and compress pictures of a source video sequence into a coded video sequence (643) in real time or under any other time constraints required by the application. Enforcing an appropriate coding speed is one function of the controller (650). In some embodiments, the controller (650) can be functionally coupled to and control other functional units as described below. The coupling is not depicted for clarity. Parameters set by the controller (650) can include rate control-related parameters (picture skip, quantizer, lambda value of rate distortion optimization technique, ...), picture size, picture group (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other appropriate functions related to the video encoder (603) optimized for a particular system design.

[0056]

[0076] In some exemplary embodiments, the video encoder (603) can be configured to operate in a coding loop. In an extremely simplified explanation, in one example, the coding loop may include a source coder (630) (for example, responsible for generating symbols such as a symbol stream based on the input picture and reference picture to be coded) and a (local) decoder (533) built into the video encoder (603). Even if the built-in decoder 633 processes the coded video stream by the source coder 630 without entropy coding, the decoder (633) reconstructs the symbols to generate sample data in a similar manner to that produced by a (remote) decoder (because any compression between the symbols and the coded video bitstream in entropy coding may be lossless in the video compression techniques considered in the disclosed subject matter). The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream yields bit-exact results that are independent of the decoder's location (local or remote), the contents of the reference picture memory (634) are also bit-exact between the local and remote encoders. In other words, the predictive part of the encoder "sees" as reference picture samples exactly the same sample values ​​that the decoder would "see" if it were using predictions during decoding. This fundamental principle of reference picture synchronization (which results in drift if synchronization cannot be maintained, for example, due to channel errors) is used to improve coding quality.

[0057]

[0077] It is possible to assume that the operation of the “local” decoder (633) is the same as that of a “remote” decoder, such as the video decoder (510) which has already been described in detail above in relation to Figure 5. However, as can be briefly seen in Figure 5, since symbols are available and the encoding / decoding of the symbols to the coded video sequence by the entropy coder (645) and parser (520) is lossless, the entropy decoding unit of the video decoder (510), which includes the buffer memory (515) and parser (520), may not be fully realized in the local decoder (633) in the encoder.

[0058]

[0078] An insight that can be gained at this point is that any decoder technique other than analysis / entropy decoding, which can only exist within the decoder, may also need to exist in substantially the same functional form within the corresponding encoder. For this reason, the disclosed subject matter can often focus on the decoder portion of the encoder and the associated decoder operation. The description of encoder techniques can be omitted as it is the inverse of the comprehensively described decoder techniques. A more detailed description of encoders is given below only in specific areas or aspects.

[0059]

[0079] During operation, in some exemplary implementations, the source coder (630) may perform motion-compensated predictive coding, predicting and coding the input picture by referencing one or more previously coded pictures from a video sequence designated as “reference pictures”. In this way, the coding engine (632) codes the difference (or residual) in the chroma channel between the pixel blocks of the input picture and the pixel blocks of the reference pictures that can be selected as predictive references for the input picture. The terms “residue” and its derivative “residual” may be used interchangeably.

[0060]

[0080] The local video decoder (633) can decode coded video data of a picture that can be designated as a reference picture based on symbols generated by the source coder (630). The operation of the coding engine (632) may, advantageously, be a non-lossless process. If the coded video data can be decoded by a video decoder (not shown in Figure 6), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (633) can repeat the decoding process that can be performed by the video decoder with respect to the reference picture, and this process may cause the reconstructed reference picture to be stored in the reference picture cache (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference picture having common content as the reconstructed reference picture obtained by the remote video decoder (assuming no transmission errors).

[0061]

[0081] The predictor (635) can perform predictive searches for the coding engine (632). That is, for a new picture to be coded, the predictor (635) can search the reference picture memory (534) for sample data (such as candidate reference pixel blocks) or given metadata (such as reference picture motion vectors, block shapes), which may serve as appropriate predictive references for the new picture. The predictor (535) can operate on a sample block-pixel block basis to find appropriate predictive references. In some cases, the input picture may have predictive references drawn from multiple reference pictures stored in the reference picture memory (634), as determined by the search results obtained by the predictor (635).

[0062]

[0082] The controller (650) can manage the coding operations of the source coder (630), including, for example, setting parameters and subgroup parameters used to encode video data.

[0063]

[0083] All outputs of the aforementioned functional units can be subjected to entropy coding in the entropy coder (645). The entropy coder (645) converts the symbols generated by the various functional units into coded video sequences by lossless compression of symbols, following techniques such as Huffman coding, variable-length coding, and arithmetic coding.

[0064]

[0084] The transmitter (640) can buffer coded video sequences, such as those created by the entropy coder (645), to prepare them for transmission over the communication channel (660), which may be a hardware / software link to a storage device that stores coded video data. The transmitter (640) can merge coded video data from the video encoder (603) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0065]

[0085] The controller (650) can manage the operation of the video encoder (603). During coding, the controller (650) can assign a specific coded picture type to each coded picture, which may affect the coding techniques that can be applied to each picture. For example, a picture may often be assigned as one of the following picture types:

[0066]

[0086] An intra-picture (I-picture) can be defined as one that can be encoded and decoded without using any other picture in the sequence as a source of prediction. Some video codecs allow different types of intra-pictures, including, for example, Independent Decoder Refresh ("IDR") pictures. Those skilled in the art are aware of these variations of I-pictures, as well as their respective uses and characteristics.

[0067]

[0087] A predictive picture (P-picture) can be encoded and decoded using intra-prediction or inter-prediction, which uses at most one motion vector and a reference index to predict the sample values ​​of each block.

[0068]

[0088] A bidirectional prediction picture (B-picture) can be encoded and decoded using intra-prediction or inter-prediction, which uses at most two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple prediction pictures can use more than two reference pictures and associated metadata to reconstruct a single block.

[0069]

[0089] A source picture can typically be spatially subdivided into multiple sample coding blocks (e.g., 4x4, 8x8, 4x8, or 16x16 sample blocks, respectively) and coded block by block. Blocks can be predictively coded by referencing other (already coded) blocks, as determined by the coding assignment applied to each picture in the block. For example, blocks of picture I may be coded non-predictively, or they may be coded predictively by referencing already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of picture P may be coded predictively by spatial or temporal prediction, referencing one previously coded reference picture. Blocks of picture B may be coded predictively by spatial or temporal prediction, referencing one or two previously coded reference pictures. A source picture or a picture undergoing intermediate processing may be subdivided into other types of blocks for other purposes. The subdivision of coding blocks and other types of blocks may or may not follow the same methods as those described in further detail below.

[0070]

[0090] The video encoder (603) can perform coding operations in accordance with a specified video coding technique or standard, such as ITU-T Rec.H.265. In this operation, the video encoder (603) can perform various compression operations, including predictive coding operations that leverage temporal and spatial redundancy in the input video sequence. The coded video data can conform to the syntax specified by the video coding technique or standard used.

[0071]

[0091] In some exemplary embodiments, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, and other forms of redundant data (such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.).

[0072]

[0092] Video can be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) utilizes spatial correlations within a given picture, while inter-picture prediction utilizes temporal or other correlations between pictures. For example, a particular picture under encoding / decoding, referred to as the current picture, may be partitioned into blocks. Blocks within the current picture can be coded by a vector called a motion vector if they are analogous to reference blocks in a reference picture that has been previously coded and is still buffering in the video. The motion vector points to a reference block in the reference picture and may have a third dimension to identify the reference picture if multiple reference pictures are used.

[0073]

[0093] In some exemplary embodiments, a bi-prediction technique can be used for inter-picture prediction. According to such a bi-prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture, both of which precede the current picture in the video in decoding order (although they may be past and future in display order, respectively). Blocks in the current picture can be coded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. Blocks can be predicted together by the combination of the first and second reference blocks.

[0074]

[0094] Furthermore, merge mode technology can be used for inter-picture prediction to improve coding efficiency.

[0075]

[0095] According to some embodiments of this disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and the CTUs within a picture can have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU can contain three parallel coding tree blocks (CTBs), which are one lumen CTB and two chroma CTBs. Each CTU can be recursively quadtree-partitioned into one or more coding units (CUs). For example, a 64x64 pixel CTU can be divided into one 64x64 pixel CU and four 32x32 pixel CUs. Each of the one or more 32x32 blocks can be further divided into four 16x16 pixel CUs. In some exemplary embodiments, each CU can be analyzed during coding to determine the prediction type of the CU among various prediction types, such as inter-prediction type or intra-prediction type. A CU can be divided into one or more prediction units (PUs) depending on its temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, the prediction process in coding (encoding / decoding) is performed in units of prediction blocks. The division of a CU into PUs (or PBs for different color channels) can be performed in various spatial patterns. For example, a luma or chroma PB can include a matrix of values ​​(e.g., luma values) for pixels, such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.

[0076]

[0096] Figure 7 shows a diagram of a video encoder (703) according to another embodiment of the present disclosure. The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values ​​in the current video picture in a sequence of video pictures, and to encode the processing block into a coded picture which is part of a coded video sequence. For example, the video encoder (703) can be used instead of the video encoder (403) in the example of Figure 4.

[0077]

[0097] For example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as an 8x8 sample prediction block. The video encoder (703) then determines whether the processing block should be best coded, for example using rate-distortion optimization (RDO), using intra-mode, inter-mode, or bi-predictive mode. If it is determined that the processing block should be coded in intra-mode, the video encoder (703) can use intra-predictive technique to encode the processing block into a coded picture; and if it is determined that the processing block should be coded in inter-mode or bi-predictive mode, the video encoder (703) can use inter-predictive technique or bi-predictive technique, respectively, to encode the processing block into a coded picture. In some exemplary embodiments, merge mode may be used as a submode of inter-picture prediction, in which case the motion vector is derived from one or more motion vector predictors without the benefit of coded motion vector components outside the predictors. In some other exemplary embodiments, there may be motion vector components applicable to the target block. Thus, the video encoder (703) may include components not explicitly shown in Figure 7, such as a mode determination module, to determine the predictive mode of the processing block.

[0078]

[0098] In the example shown in Figure 7, the video encoder (703) includes an interencoder (730), an intraencoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general-purpose controller (721), and an entropy encoder (725), all coupled together as shown in the exemplary arrangement in Figure 7.

[0079]

[0099] The inter-encoder (730) is configured to receive a sample of the current block (e.g., a processing block), compare that block to one or more reference blocks in the reference picture (e.g., blocks in earlier and later pictures in display order), generate inter-prediction information (e.g., descriptions of redundant information by inter-coding techniques, motion vectors, merge mode information), and compute inter-prediction results (e.g., predicted blocks) based on the inter-prediction information using any appropriate technique. In some examples, the reference picture is a decoded reference picture decoded based on encoded video information using a decoding unit 633 (shown as a residual decoder 728 in Figure 7, which will be described in more detail below) incorporated into the exemplary encoder 620 in Figure 6.

[0080]

[0100] The intra encoder (722) is configured to receive a sample of the current block (e.g., a processing block), compare that block to a block already coded within the same picture, generate quantized coefficients after transformation, and optionally also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques). Based on the intra prediction information and reference block within the same picture, the intra encoder (722) can compute an intra prediction result (e.g., a prediction block).

[0081]

[0101] The general-purpose controller (721) can be configured to determine general control data and control other components of the video encoder (703) based on that general control data. For example, the general-purpose controller (721) determines the prediction mode of a block and provides control signals to the switch (726) based on that prediction mode. For example, if the prediction mode is intra-mode, the general-purpose controller (721) controls the switch (726) to select the intra-mode result for use by the residual calculator (723) and the entropy encoder (725) to select the intra-prediction information and include the intra-prediction information in the bitstream; and if the prediction mode of a block is inter-mode, the general-purpose controller (721) controls the switch (726) to select the inter-prediction result for use by the residual calculator (723) and the entropy encoder (725) to select the inter-prediction information and include the inter-prediction information in the bitstream.

[0082]

[0102] The residual calculator (723) can be configured to calculate the difference (residual data) between the received block and the prediction result for a block selected from the intra-encoder (722) or inter-encoder (730). The residual encoder (724) can be configured to encode the residual data to generate conversion coefficients. For example, the residual encoder (724) may be configured to convert the residual data from the spatial domain to the frequency domain to generate conversion coefficients. The conversion coefficients are then quantized to obtain quantized conversion coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse conversion to generate decoded residual data. The decoded residual data can be appropriately used by the intra-encoder (722) and inter-encoder (730). For example, an inter-encoder (730) can generate a decoded block based on decoded residual data and inter-prediction information, and an intra-encoder (722) can generate a decoded block based on decoded residual data and intra-prediction information. The decoded block is appropriately processed to generate a decoded picture, which is buffered in a memory circuit (not shown) and can be used as a reference picture.

[0083]

[0103] The entropy encoder (725) can be configured to format a bitstream to include encoded blocks and to perform entropy coding. The entropy encoder (725) can be configured to include various types of information in the bitstream. For example, the entropy encoder (725) can be configured to include general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information in the bitstream. Residual information may not be present when coding blocks in either inter-mode or bi-prediction mode merge submodes.

[0084]

[0104] Figure 8 shows an exemplary video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive a coded picture which is part of a coded video sequence, decode the coded picture, and produce a reconstructed picture. In one example, the video decoder (810) may be used instead of the video decoder (410) in the example of Figure 4.

[0085]

[0105] In the example shown in Figure 8, the video decoder (810) includes an entropy decoder (871), an inter-decoder (880), a residual decoder (873), a reconfiguration module (874), and an intra-decoder (872), coupled together as shown in the exemplary arrangement in Figure 8.

[0086]

[0106] The entropy decoder (871) can be configured to reconstruct from the coded picture certain symbols representing the syntactic elements in which the coded picture is constructed. Such symbols may include, for example, the mode in which the block is coded (e.g., intra-mode, inter-mode, bi-predictive mode, merge sub-mode, or another sub-mode), prediction information (e.g., intra-predictive information or inter-predictive information) that can identify specific samples or metadata used for prediction by the intra-decoder (872) or inter-decoder (880), residual information (e.g., in the form of quantized transformation coefficients), etc. In one example, if the prediction mode is inter or bi-predictive mode, inter-predictive information is provided to the inter-decoder (880); if the prediction type is intra-predictive type, intra-predictive information is provided to the intra-decoder (872). The residual information can be inversely quantized and provided to the residual decoder (873).

[0087]

[0107] The inter-decoder (880) can be configured to receive inter-prediction information and generate inter-prediction results based on the inter-prediction information.

[0088]

[0108] The intra decoder (872) can be configured to receive intra prediction information and generate prediction results based on the intra prediction information.

[0089]

[0109] The residual decoder (873) can be configured to perform inverse quantization to extract de-quantization conversion coefficients and process these de-quantization conversion coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also use certain control information (including quantization parameters (QP)), which may be provided by the entropy decoder (871) (the data path is not drawn because this may only be a small amount of control information).

[0090]

[0110] The reconstruction module (874) can be configured to combine the residuals as output from the residual decoder (873) and the prediction results (which may be output by an inter or intra prediction module) in the spatial domain to form reconstructed blocks that form part of the reconstructed picture as part of the reconstructed video. It should be noted that other appropriate processes, such as deblocking, may be performed to improve visual quality.

[0091]

[0111] The video encoders (403), (603), and (703), as well as the video decoders (410), (510), and (810), can be implemented using any suitable technology. In some exemplary embodiments, the video encoders (403), (603), and (703), as well as the video decoders (410), (510), and (810), can be implemented using one or more integrated circuits. In other embodiments, the video encoders (403), (603), and (703), as well as the video decoders (410), (510), and (810), can be implemented using one or more processors that execute software instructions.

[0092]

[0112] Turning to block partitioning for encoding and decoding, general partitioning can start with a base block and follow a predefined set of rules, a specific pattern, a partition tree, or some partition structure or scheme. Partitioning can be hierarchical and recursive. After dividing or partitioning the base block according to one of the exemplary partitioning procedures or other procedures, or a combination thereof, as described below, a final set of partitions or coding blocks can be obtained. Each of these partitions can be one of various partitioning levels in the partitioning hierarchy and can be of various shapes. Each partition may be called a coding block (CB). With respect to the various exemplary partitioning implementations described further below, each resulting CB can be of some acceptable size and partitioning level. Such partitions are called coding blocks because they form a unit over which some basic coding / decoding decisions may be made, and coding / decoding parameters can be signaled in the optimized, determined, and encoded video bitstream. The highest or deepest level in the final partition represents the depth of the tree's coding block partition structure. A coding block may be a luma coding block or a chroma coding block. Each color's CB tree structure may be called a coding block tree (CBT).

[0093]

[0113] The coding blocks for all color channels may be collectively referred to as a coding unit (CU). The hierarchical structure of all color channels may be collectively referred to as a coding tree unit (CTU). The partitioning patterns or structures of the various color channels within a CTU may or may not be the same.

[0094]

[0114] In some implementations, the partition tree scheme or structure used for luma and chroma channels may not need to be the same. In other words, luma and chroma channels may have separate coding tree structures or patterns. Furthermore, whether luma and chroma channels use the same or different coding partition tree structures and the actual coding partition tree structures may depend on whether the slice being coded is a P, B, or I slice. For example, in the case of an I slice, chroma and luma channels may have separate coding partition tree structures or coding partition tree structure modes, but in the case of a P or B slice, luma and chroma channels may share the same coding partition tree scheme. When separate coding partition tree structures or modes are applied, luma channels may be partitioned into CBs by one coding partition tree structure, and chroma channels may be partitioned into chroma CBs by another coding partition tree structure.

[0095]

[0115] In some exemplary implementations, a given partitioning pattern may be applied to the base block. As shown in Figure 9, an exemplary four-way partition tree may begin at a first predefined level (e.g., a 64x64 block level or other size as the base block size), and the base block may be partitioned hierarchically downwards to a predefined lowest level (e.g., a 4x4 level). For example, the base block can follow four predefined partitioning options or patterns shown in 902, 904, 906, and 908, and the partition designated as R allows for recursive partitioning, in which the same partitioning option as shown in Figure 9 may be repeated at a lower scale down to the lowest level (e.g., a 4x4 level). In some implementations, additional restrictions may be applied to the partitioning scheme in Figure 9. In the implementation of Figure 9, rectangular partitions (e.g., 1:2 / 2:1 rectangular partitions) may be allowed, but they may not be allowed to be recursive, while square partitions are allowed to be recursive. Partitioning following Figure 9, with recursion as needed, generates the final set of coding blocks. The coding tree depth may be further defined to indicate the partitioning depth from the root node or root block. For example, the coding tree depth of the root node or root block, e.g., a 64x64 block, may be set to 0, and after the root block is further partitioned once according to Figure 9, the coding tree depth is incremented by only one. The maximum or deepest level from the 64x64 base block to the smallest 4x4 partition is 4 for the above scheme (starting from level 0). Such a partitioning scheme may be applied to one or more color channels.Each color channel may be partitioned independently according to the method shown in Figure 9 (for example, a partitioning pattern or option from a predefined pattern may be determined independently for each color channel at each hierarchy level). Alternatively, two or more color channels may share the same hierarchy pattern tree in Figure 9 (for example, the same partitioning pattern or option from a predefined pattern may be selected for two or more color channels at each hierarchy level).

[0096]

[0116] Figure 10 shows another exemplary predefined partitioning pattern that allows recursive partitioning to form a partitioning tree. An exemplary 10-way partitioning structure or pattern may be predefined, as shown in Figure 10. The root block may start from a predefined level (e.g., from a base block at the 128x128 level or 64s64 level). The exemplary partitioning structure in Figure 10 includes various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. The partition type with three subpartitions, indicated as 1002, 1004, 1006, and 1008 in the second row of Figure 10, may be called a "T-type" partition. The "T-type" partitions 1002, 1004, 1006, and 1008 may be called the left T-type, upper T-type, right T-type, and lower T-type, respectively. In some exemplary implementations, none of the rectangular partitions in Figure 10 are allowed to be further subdivided. The coding tree depth may be further defined to indicate the partitioning depth from the root node or root block. For example, the coding tree depth of the root node or root block (e.g., a 128x128 block) may be set to 0, and after the root block is further partitioned once according to Figure 10, the coding tree depth is incremented by only one. In some implementations, only the 1010 all-square partitions may be allowed to recursively partition to the next level of the partitioning tree following the pattern in Figure 10. In other words, recursive partitioning may not be allowed for the square partitions of the T-shaped patterns 1002, 1004, 1006, and 1008. The partitioning procedure following Figure 10, with recursion as needed, produces the final set of coding blocks. Such a scheme may be applied to one or more color channels. In some implementations, more flexibility may be added for the use of partitions at 8x8 levels or less. For example, 2x2 chroma interpretation may be used in certain cases.

[0097]

[0117] In some other exemplary implementations for coding block partitioning, a quadtree structure may be used to divide a base block or intermediate block into quadtree partitions. Such quadtree partitioning may be applied hierarchically and recursively to any square partition. Whether the base block or intermediate block or partition is further quadtree partitioned can be adapted to various local characteristics of the base block or intermediate block / partition. Quadtree partitioning at picture boundaries can be further adapted. For example, implicit quadtree partitioning may be performed at picture boundaries so that a block maintains the quadtree partition until its size fits the picture boundary.

[0098]

[0118] In some other exemplary implementations, hierarchical binary partitioning from a base block may be used. In such a scheme, the base block or intermediate-level block may be divided into two partitions. Binary partitioning can be horizontal or vertical. For example, horizontal binary partitioning can divide the base block or intermediate block into equal left and right partitions. Similarly, vertical binary partitioning can divide the base block or intermediate block into equal upper and lower partitions. Such binary partitioning may be hierarchical and recursive. At each base block or intermediate block, a decision may be made as to whether the binary partitioning scheme should continue, and if so, whether horizontal or vertical binary partitioning should be used. In some implementations, further partitioning may be stopped at a predefined minimum partition size (in one or both dimensions). Alternatively, further partitioning may be stopped when a predefined partitioning level or depth from the base block is reached. In some implementations, the aspect ratio of the partitions may be restricted. For example, the aspect ratio of a partition may not be less than 1:4 (or may not be greater than 4:1). Thus, a vertical strip partition with a vertical-to-horizontal aspect ratio of 4:1 may simply be further divided vertically into an upper partition and a lower partition, each having a vertical-to-horizontal aspect ratio of 2:1.

[0099]

[0119] In some other examples, a ternary partitioning scheme may be used to partition a base block or any intermediate block, as shown in Figure 13. The ternary pattern may be implemented vertically, as shown in 1302 of Figure 13, or horizontally, as shown in 1304 of Figure 13. The exemplary partition ratios in Figure 13 are shown as 1:2:1 for both the vertical and horizontal directions, but other ratios may be predefined. In some implementations, two or more different ratios may be predefined. Such a ternary partitioning scheme can be used to complement quadtree or binary partitioning structures, where such ternary tree partitioning can incorporate objects located at the center of the block into a single, unbroken partition, whereas quadtrees and binary trees are always partitioned along the center of the block, thus dividing such objects into separate partitions. In some implementations, to avoid additional transformations, the width and height of the partitions in the exemplary ternary tree are always powers of 2.

[0100]

[0120] The above partitioning schemes can be combined in any way at different partitioning levels. For example, the quadtree and binary tree schemes described above may be combined to partition the base block into a quadtree-binary-tree (QTBT) structure. In such a scheme, the base block or intermediate block / partition may be either quadtree-partitioned or binary-partitioned according to a set of predefined conditions, if specified. A specific example is shown in Figure 14. In the example in Figure 14, the base block is first quadtree-partitioned into four partitions, as shown in 1402, 1404, 1406, and 1408. Each resulting partition is then quadtree-partitioned into four further partitions (as in 1408), binary-partitioned into two further partitions (e.g., horizontally or vertically, as in 1402 or 1406, both symmetrically), or unpartitioned (as in 1404). Binary or quadtree partitions may be recursively permissible for square-shaped partitions, as shown by the overall exemplary partition pattern in 1410 and the corresponding tree structure / representation in 1420 (solid lines represent quadtree partitions, dashed lines represent binary partitions). A flag may be used for each binary partition node (non-leaf binary partition) to indicate whether the binary partition is horizontal or vertical. For example, as shown in 1420, consistent with the partition structure in 1410, a flag "0" can represent a horizontal bipartite and a flag "1" can represent a vertical bipartite. In the case of quadtree partitions, there is no need to indicate the partition type, as the quadtree partition always divides a block or partition both horizontally and vertically to produce four subblocks / partitions of equal size. In some implementations, a flag "1" can represent a horizontal bipartite and a flag "0" can represent a vertical bipartite.

[0101]

[0121] In some exemplary implementations of QTBT, the quadtree-to-binary tree splitting ruleset can be represented by the following pre-configured parameters and their associated corresponding functions: - CTU size: The size of the root node (base block) of a quadtree. - MinQTSize: Minimum allowable quadtree leaf node - MaxBTSize: Maximum allowable binary tree root node size - MaxBTDepth: Maximum allowable binary tree depth - MinBTSize: Minimum allowable binary tree leaf node size In some exemplary implementations of the QTBT partitioning structure, the CTU size may be set as a 128x128 chroma sample with two corresponding 64x64 blocks of chroma samples (when exemplary chroma subsampling is considered and used), MinQTSize may be set as 16x16, MaxBTSize may be set as 64x64, MinBTSize (for both width and height) may be set as 4x4, and MaxBTDepth may be set as 4. The quadtree partitioning may be applied to the CTU first to generate quadtree leaf nodes. A quadtree leaf node can have a size from its minimum allowable size of 16x16 (i.e., MinQTSize) up to 128x128 (i.e., CTU size). If a node is 128x128, its size exceeds MaxBTSize (i.e., 64x64) and therefore it is not initially partitioned by the binary tree. Otherwise, nodes that do not exceed MaxBTSize can be partitioned by a binary tree. In the example in Figure 14, the base block is 128x128. The base block can only be quadruped according to a predefined set of rules. The base block has a partitioning depth of 0. Each of the four resulting partitions is 64x64, does not exceed MaxBTSize, and can be further quadruped or binary at level 1. The process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), further partitioning is not considered. If a binary tree node has a width equal to MinBTSize (i.e., 4), further horizontal partitioning is not considered. Similarly, if a binary tree node has a height equal to MinBTSize, further vertical partitioning is not considered.

[0102]

[0122] In some exemplary implementations, the above QTBT scheme can be configured to support the flexibility for lumens and chromians to have the same QTBT structure or separate QTBT structures. For example, in the case of P and B slices, lumens and chromians within a single CTU can share the same QTBT structure. However, in the case of I slices, the lumens CTB may be partitioned into CBs by a QTBT structure, and the chromens CTB may be partitioned into chromians CBs by a separate QTBT structure. This means that CUs can be used to indicate different color channels in an I slice; for example, an I slice may consist of a coding block for the lumens component or a coding block for two chromians, and a CU in a P or B slice may consist of coding blocks for all three color components.

[0103]

[0123] In some other implementations, the QTBT scheme can be supplemented by the ternary scheme described above. Such implementations may be called multi-type-tree (MTT) structures. For example, in addition to node bipartition, one of the ternary partition patterns shown in Figure 13 may be selected. In some implementations, only square nodes may be subject to ternary partitioning. Additional flags may be used to indicate whether the ternary partitioning is horizontal or vertical.

[0104]

[0124] Designing two-level or multi-level trees, such as the QTBT implementation and the QTBT implementation supplemented by tripartitioning, can be primarily motivated by complexity reduction. Theoretically, the complexity of traversing a tree is T D Here, T represents the number of partition types and D is the depth of the tree. A trade-off may be made by using multiple types (T) while reducing the depth (D).

[0105]

[0125] In some implementations, the CB may be further partitioned. For example, the CB may be further partitioned into multiple prediction blocks (PBs) for the purpose of intra or inter-frame prediction during the coding and decoding processes. In other words, the CB may be further divided into different sub-partitions, in which case individual prediction decisions / settings may be made. In parallel, the CB may be further partitioned into multiple transformation blocks (TBs) for the purpose of delimiting the level at which transformation or inverse transformation of video data is performed. The partitioning scheme of the CB into PBs and TBs may be the same or different. For example, each partitioning scheme may be performed using its own procedure, for example, based on various characteristics of the video data. The partitioning schemes of PBs and TBs may be independent in some exemplary implementations. The partitioning schemes and boundaries of PBs and TBs may be related in other exemplary implementations. In some implementations, for example, TBs may be partitioned after PB partitioning, and in particular, each PB may be further partitioned into one or more TBs after being determined after the coding block partitioning. For example, in some implementations, a PB may be divided into one, two, four, or any other number of TBs.

[0106]

[0126] In some implementations, luma channels and chroma channels may be treated differently in order to divide the base block into coding blocks, and further partition them into prediction and / or transformation blocks. For example, in some implementations, partitioning coding blocks into prediction and / or transformation blocks may be permitted for luma channels, but not for chroma channels. Thus, in such implementations, transformation and / or prediction of luma blocks may only be performed at the coding block level. In another example, the minimum transformation block sizes for luma channels and chroma channels may differ; for example, coding blocks of luma channels may be allowed to be partitioned into smaller transformation and / or prediction blocks than chroma channels. In yet another example, the maximum partitioning depth of coding blocks into transformation blocks and / or prediction blocks may differ between lumar and chroma channels. For example, coding blocks for lumar channels may be allowed to be partitioned into deeper transformation and / or prediction blocks than chroma channels. As a specific example, lumar coding blocks may be partitioned into transformation blocks of multiple sizes that can be represented by recursive partitioning descending to at most two levels, with transformation block shapes such as square, 2:1 / 1:2, and 4:1 / 1:4, and transformation block sizes from 4x4 to 64x64 being allowed. However, for chroma blocks, only the maximum possible transformation blocks specified for lumar blocks may be allowed.

[0107]

[0127] In some exemplary implementations for partitioning coding blocks into PBs, the depth, shape, and / or other characteristics of the PB partitioning may depend on whether the PB is intra-coded or inter-coded.

[0108]

[0128] The partitioning of coding blocks (or prediction blocks) into transformation blocks can be carried out in a variety of exemplary ways, including, but not limited to, quadtree partitioning and predefined pattern partitioning, with additional considerations for the transformation blocks at the boundaries of the coding or prediction blocks. In general, the resulting transformation blocks may be at different partitioning levels, may not be the same size, and may not be required to be in a square shape (for example, they may be rectangles with some acceptable size and aspect ratio). Further examples are described in more detail below in relation to Figures 15, 16, and 17.

[0109]

[0129] However, in some other implementations, the CB obtained by any of the above partitioning schemes may be used as the basic or smallest coding block for prediction and / or transformation. In other words, no further partitioning is performed for the purpose of inter-prediction / intra-prediction and / or transformation. For example, the CB obtained from the above QTBT scheme may be used directly as the unit for performing prediction. Specifically, such a QTBT structure eliminates the concept of multiple partition types, i.e., eliminates the separation of CU, PU, ​​and TU, and supports more flexibility for the CU / CB partition shape as described above. In such a QTBT block structure, the CU / CB can have either a square or rectangular shape. The leaf nodes of such a QTBT are used as the unit for prediction and transformation processing without any further partitioning. This means that in such an exemplary QTBT coding block structure, the CU, PU, ​​and TU have the same block size.

[0110]

[0130] The various CB partitioning schemes described above, and further partitioning of CBs into PBs and / or TBs (excluding PB / TB partitioning), may be combined in any manner. The following specific implementations are provided as non-limiting examples.

[0111]

[0131] The following describes specific examples of coding block and transform block partitioning. In such exemplary implementations, the base block may be partitioned into coding blocks using recursive quadtree partitioning or the predefined partitioning patterns described above (as shown in Figures 9 and 10). At each level, whether further quadtree partitioning of a particular partition should continue can be determined by local video data characteristics. The resulting CBs may be at various quadtree partitioning levels and of various sizes. Decision on whether to code a picture area using inter-picture (temporal) or intra-picture (spatial) prediction may be made at the CB level (or at the CU level for all three color channels). Each CB may be further partitioned into one, two, four, or other number of PBs according to a predefined PB partitioning type. Within a single PB, the same prediction process may be applied, and relevant information may be transmitted to the decoder at the PB level. By applying a prediction process based on the PB partitioning type, after obtaining the residual block, the CB can be partitioned into TBs according to another quadtree structure similar to the coding tree for the CB. In this particular implementation, the CB or TB may be, but is not limited to, a square shape. Furthermore, in this particular example, the PB may be square or rectangular for interpretations, and square only for intrapredictions. The coding block may be partitioned, for example, into four square TBs. Each TB may be further partitioned recursively (using quadtree partitioning) into smaller TBs called residual quadtrees (RQTs).

[0112]

[0132] Further examples of implementations for partitioning a base block into CBs, PBs, and / or TBs are described below. For example, instead of using multiple partition unit types as shown in Figures 9 and 10, a quadtree with nested multi-type trees using 2-partition and 3-partition segmentation structures (e.g., a QTBT or QTBT performing 3-partition as described above) may be used. The separation of CBs, PBs, and TBs (i.e., partitioning CBs into PBs and / or TBs, and partitioning PBs into TBs) may be abandoned unless required for CBs that are too large for the maximum transformation length (such CBs may require further partitioning). This exemplary partitioning scheme can be designed to support more flexibility with respect to the CB partition shape, and both prediction and transformation can be performed at the CB level without further partitioning. In such a coding tree structure, CBs can be either square or rectangular in shape. Specifically, a coding tree block (CTB) may first be partitioned by a quadtree structure. Next, the quadtree leaf nodes can be further partitioned by nested multi-type tree structures. Examples of nested multi-type tree structures using bipartitioning or tripartitioning are shown in Figure 11. Specifically, the exemplary multi-type tree structure in Figure 11 is: Vertical bisection (SPLIT_BT_VER) (1102), Horizontal bisection (SPLIT_BT_HOR) (1104), Vertical trisection (SPLIT_TT_VER)(1106), and This includes four partitioning types, including horizontal trisection (SPLIT_TT_HOR)(1108). Therefore, CB corresponds to a leaf in a multi-type tree. In this implementation example, this segmentation is used for both prediction and transformation without any further partitioning, as long as CB is not too large relative to the maximum transformation length. This means that in most cases, in a quadtree with a nested multi-type tree coding block structure, CB, PB, and TB have the same block size. The exception occurs when the supported maximum transformation length is smaller than the width or height of the color component of CB. In some implementations, in addition to bipartite or tripartite, the nested pattern in Figure 11 may further include quadtree partitioning.

[0113]

[0133] A concrete example of a quadtree with a nested multi-type tree coding block structure for block partitioning with respect to a single base block is shown in Figure 12 (including the options of quadtree partitioning, binary partitioning, and ternary partitioning). More specifically, Figure 12 shows that the base block 1200 is quadtree partitioned into four square partitions 1202, 1204, 1206, and 1208. The decision to use the quadtree and the multi-type tree structure of Figure 11 for further partitioning is made for each partition that is quadtree partitioned. In the example in Figure 12, partition 1204 is not further partitioned. Partitions 1202 and 1208 each employ different quadtree partitioning. For partition 1202, the second-level quadtree-divided upper-left, upper-right, lower-left, and lower-right partitions employ the third-level quadtree, the horizontal binary tree division 1104 in Figure 11, undivided, and the horizontal ternary tree division 1108 in Figure 11, respectively. Partition 1208 employs a different quadtree division, with the second-level quadtree-divided upper-left, upper-right, lower-left, and lower-right partitions employing the third-level vertical ternary tree division 1106 in Figure 11, undivided, undivided, and the horizontal binary tree division 1104 in Figure 11, respectively. Two of the sub-partitions of the third-level upper-left partition of 1208 are further divided according to the horizontal binary division 1104 and the horizontal ternary division 1108 in Figure 11, respectively. Partition 1206 is divided into two partitions following a second-level division pattern, specifically a vertical bipartition 1102 in Figure 11. These two patterns are further divided at a third level according to the horizontal tripartition 1108 and vertical bipartition 1102 in Figure 11. A fourth-level division is then applied to one of these partitions, according to the horizontal bipartition 1104 in Figure 11.

[0114]

[0134] In the specific example above, the maximum luma conversion size may be 64x64, and the maximum supported chroma conversion size may differ from that of luma, for example, 32x32. The exemplary CB in Figure 12 above is generally not further divided into smaller PBs and / or TBs, but even so, if the width or height of a luma coding block or chroma coding block is greater than the maximum conversion width or height, the luma coding block or chroma coding block may be automatically divided horizontally and / or vertically to satisfy the conversion size limitations in that direction.

[0115]

[0135] In specific examples of partitioning the base block into the above-mentioned CBs, and as described above, the coding tree scheme can support the ability to have separate block tree structures for lumens and chromens. For example, in the case of P and B slices, lumens and chromens CTBs within a single CTU can share the same coding tree structure. For example, in the case of an I slice, lumens and chromens can have separate coding block tree structures. When separate block tree structures are applied, lumens CTBs are partitioned into lumens CBs by one coding tree structure, and chromens CTBs are partitioned into chromens CBs by another coding tree structure. This means that a CU in an I slice may consist of coding blocks for the lumens component, or coding blocks for two chromens component, and that a CU in a P or B slice will always consist of coding blocks for all three color components, unless the video is monochrome.

[0116]

[0136] When a coding block is further partitioned into multiple transformation blocks, these transformation blocks may be ordered within the bitstream according to various orders or scanning schemes. Exemplary implementations for partitioning coding or prediction blocks into transformation blocks, and the coding order of the transformation blocks, are described in further detail below. In some exemplary implementations, as described above, transformation partitioning can support transformation blocks of multiple shapes, e.g., 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1, with transformation block sizes ranging, for example, from 4x4 to 64x64. In some implementations, if the coding block is 64x64 or smaller, transformation block partitioning may be applied only to the chroma component, and for the chroma block, the transformation block size may match the coding block size. Otherwise, if the width or height of a coding block is greater than 64, both the luma and chroma coding blocks may be implicitly divided into transformation blocks that are multiples of min(W,64) × min(H,64) and min(W,32) × min(H,32), respectively.

[0117]

[0137] In some exemplary implementations of transform block partitioning, both intra and intercoded blocks can be further partitioned into multiple transform blocks with a predefined number of partitioning depths (e.g., 2 levels). The partitioning depth and size of the transform blocks may be associated. For some exemplary implementations, the mapping from the transform size at the current depth to the transform size at the next depth is shown in Table 1 below.

[0118] Table 1: Conversion Partition Size Settings

[0119] [Table 1]

[0138] Based on the mapping example in Table 1, for a 1:1 square block, the next level of transformation partitioning could potentially create four 1:1 square sub-transformation blocks. The transformation partitioning could stop at, for example, 4x4. Thus, the transformation size of 4x4 at the current depth corresponds to the same 4x4 size for the next depth. In the example in Table 1, for a 1:2 / 2:1 non-square block, the next level of transformation partitioning could create two 1:1 square sub-transformation blocks, while for a 1:4 / 4:1 non-square block, the next level of transformation partitioning could create two 1:2 / 2:1 sub-transformation blocks.

[0120]

[0139] In some exemplary implementations, additional restrictions may apply to transform block partitioning with respect to the luma components of the intra-coded blocks. For example, for each level of transform partitioning, all sub-transformation blocks may be restricted to having equal sizes. For instance, for a 32x16 coding block, a level 1 transform partition creates two 16x16 sub-transformation blocks, and a level 2 transform partition creates eight 8x8 sub-transformation blocks. In other words, the second level partition must be applied to all first-level sub-blocks to keep the transform units of equal size. An example of transform block partitioning for an intra-coded square block according to Table 1 is shown in Figure 15, along with the coding order indicated by arrows. Specifically, 1502 shows a square coding block. A first-level partition into four equally sized transform blocks according to Table 1 is shown in 1504, along with the coding order indicated by arrows. A second-level partition that transforms all first-level blocks of equal size into 16 equally sized transformation blocks according to Table 1 is shown in 1506, along with the coding order indicated by the arrows.

[0121]

[0140] In some exemplary implementations, the above restrictions on intra-coding may not apply to the luma components of an inter-coded block. For example, after a first level of transformation partitioning, any one of the sub-transformation blocks may be further partitioned independently at another further level. The resulting transformation blocks may or may not be of the same size. An example of partitioning an inter-coded block into transformation blocks in their coding order is shown in Figure 16. In the example in Figure 16, the inter-coded block 1602 is partitioned into transformation blocks at two levels according to Table 1. At the first level, the inter-coded block is partitioned into four transformation blocks of equal size. Then, only one of the four transformation blocks (but not all of them) is further partitioned into four sub-transformation blocks, resulting in a total of seven transformation blocks of two different sizes, as shown in 1604. The exemplary coding order of these seven transformation blocks is indicated by the arrow in 1604 of Figure 16.

[0122]

[0141] In some exemplary implementations, additional restrictions may be applied to the chroma component(s) on the transformation block. For example, the transformation block size for a chroma component(s) can be the same size as the coding block size, but cannot be smaller than a predefined size (e.g., 8x8).

[0123]

[0142] In some other exemplary implementations, for coding blocks where either the width (W) or height (H) is greater than 64, both the luma and chroma coding blocks may be implicitly divided into transformation units that are multiples of min(W,64) × min(H,64) and min(W,32) × min(H,32), respectively. Hereinafter, "min(a,b)" may return a value smaller than a and b.

[0124]

[0143] Figure 17 further illustrates another alternative example of a scheme for partitioning coding blocks or prediction blocks into transformation blocks. As shown in Figure 17, instead of using recursive transformation partitioning, a predefined set of partitioning types may be applied to coding blocks according to the transformation type of the coding block. In the particular example shown in Figure 17, one of six exemplary partitioning types may be applied to divide the coding block into a varying number of transformation blocks. Such a scheme resulting in transformation block partitioning may be applied to either coding blocks or prediction blocks.

[0125]

[0144] More specifically, the partitioning scheme in Figure 17 provides up to six exemplary partition types for any given transformation type (where transformation type refers to primary transformations such as ADST and other types). In this scheme, all coding blocks or prediction blocks can be assigned to a transformation partition type, for example, based on rate-distortion cost. In one example, the transformation partition type assigned to a coding block or prediction block may be determined based on the transformation type of the coding block or prediction block. A particular transformation partition type may correspond to a transformation block partition size and pattern, as shown by the six transformation partition types shown in Figure 17. The correspondence between various transformation types and various transformation partition types may be predefined. In the example shown below, uppercase labels indicate transformation partition types, which may be assigned to coding blocks or prediction blocks based on rate-distortion cost:

[0145] • PARTITION_NONE: Allocates a conversion size equal to the block size.

[0126]

[0146] • PARTITION_SPLIT: Assigns a conversion size that is half the width of the block size and half the height of the block size.

[0127]

[0147] • PARTITION_HORZ: Assigns a conversion size that is the same width as the block size and half the height of the block size.

[0128]

[0148] • PARTITION_VERT: Assigns a conversion size that is half the width of the block size and the same height as the block size.

[0129]

[0149] • PARTITION_HORZ4: Assigns a conversion size that is the same width as the block size and 1 / 4 of the block size's height.

[0130]

[0150] • PARTITION_VERT4: Assigns a conversion size that is 1 / 4 the width of the block size and the same height as the block size.

[0131]

[0151] In the example above, all conversion partition types, as shown in Figure 17, include a uniform conversion size with respect to the partitioned conversion block. This is merely an example, not an limitation. Some other implementations may allow mixed conversion block sizes to be used for partitioned conversion blocks of a particular partition type (or pattern).

[0132]

[0152] The PBs (or CBs, also called PBs if not further partitioned into prediction blocks) obtained from any of the above partitioning schemes can then become individual blocks for coding by either intra-prediction or inter-prediction. In the case of inter-prediction on the current PB, a residual is generated between the current block and the prediction block, which can be coded and included in the coded bitstream.

[0153] Interpretation can be implemented, for example, in single-reference mode or composite-reference mode. In some implementations, a skip flag may first be included in the bitstream of the current block (or at a higher level) to indicate whether the current block is intercoded and not skipped. If the current block is intercoded, another flag may be included in the bitstream as a signal indicating whether single-reference mode or composite-reference mode is used for the current block. In single-reference mode, it is possible to generate a predicted block for the current block using one reference block. In composite-reference mode, it is possible to generate a predicted block using two or more reference blocks, for example, by a weighted average. Composite-reference mode may also be called more-than-one-reference mode, two-reference mode, or multiple-reference mode. A reference block or multiple reference blocks may be identified using a reference frame index or index, and additionally, using a corresponding motion vector or motion vector(s) indicating the shift(s) between the reference block(s) and the current block in position (e.g., in horizontal and vertical pixels). For example, in single-reference mode, an interpretation block for the current block can be generated from a single reference block identified as a prediction block by a single motion vector within the reference frame, whereas in composite-reference mode, the prediction block can be generated by a weighted average of two reference blocks in two reference frames indicated by two motion vectors. Motion vectors can be coded in various ways and included in the bitstream.

[0133]

[0154] In some implementations, the coding or decoding system may maintain a decoded picture buffer (DPB). Some images / pictures may be held in the DPB waiting to be displayed (in the decoding system), and some images / pictures in the DPB may be used as reference frames to enable interpretation. In some implementations, reference frames in the DPB may be tagged as either short-term or long-term references to the current image being encoded or decoded. For example, short-term reference frames may include frames used for interpretation of blocks in the current frame or in a predetermined number of subsequent video frames (e.g., two) closest to the current frame in the decoding order. Long-term reference frames may include frames in the DPB that can be used to predict image blocks within a frame, and may include a predefined number of frames farther away from the current frame in the decoding order. Information regarding the tags of such short-term and long-term reference frames may be called a reference picture set (RPS) and may be appended to the header of each frame in the coded bitstream. Each frame in the encoded video stream can be identified by a Picture Order Counter (POC), which is numbered according to the playback sequence, either in an absolute way or relative to a group of pictures, for example, starting with frame I.

[0134]

[0155] In some exemplary implementations, one or more reference picture lists, including the identification of short-term and long-term reference frames for interpretation, can be formed based on information in the RPS. For example, for unidirectional interpretation, a single picture reference list can be formed so as to be indicated as L0 reference (or reference list 0), while for bidirectional interpretation, two picture reference lists can be formed so as to be indicated as L0 (or reference list 0) and L1 (or reference list 1) for each of the two prediction directions. The reference frames included in the L0 and L1 lists can be ordered in various predetermined ways. The lengths of the L0 and L1 lists may be signaled in the video bitstream. Unidirectional interpretation can be either a composite reference mode or a single reference mode, where multiple references for the generation of predicted blocks by weighted average in the composite prediction mode are on the same side of the predicted block. Bidirectional interpretation may be only a composite mode, in which case the bidirectional interpretation includes at least two reference blocks.

[0135]

[0156] In some implementations, a merge mode (MM) may be implemented for interpretation. Generally, in merge mode, the motion vector in a single reference prediction or one or more motion vectors in a composite reference prediction for the current PB can be derived from other motion vectors rather than being independently computed and signaled. For example, in an encoding system, the current motion vector(s) for the current PB may be converted into a difference(s) between the current motion vector(s) and one or more other already encoded motion vectors (called reference motion vectors). Such a difference(s) of motion vector(s), not the entirety of the current motion vector(s), can be encoded and included in the bitstream and linked to the reference motion vector(s). Correspondingly in a decoding system, the motion vector(s) corresponding to the current PB may be derived based on the decoded motion vector difference and the decoded reference motion vector(s) linked to it. As a specific form of general merge-mode (MM) interpretation, such interpretation based on motion vector differences is sometimes referred to as Merge Mode with Motion Vector Difference (MMVD). Generally, MM, or more specifically MMVD, may therefore be implemented to utilize correlations between motion vectors related to different PBs in order to improve coding efficiency. For example, adjacent PBs may have similar motion vectors. In another example, motion vectors may correlate temporally (between frames) for blocks that are positioned / located, as they may be in space.

[0136]

[0157] In some exemplary implementations, the MM flag may be included in the bitstream during the encoding process to indicate whether the current PB is in merge mode. Additionally or alternatively, the MMVD flag may be included during the encoding process and signaled in the bitstream to indicate whether the current PB is in MMVD mode. The MM and / or MMVD flags or indicators may be provided at the PB level, CB level, CU level, CTB level, CTU level, slice level, picture level, etc. In certain examples, both the MM and MMVD flags may be included for the current CU, and the MMVD flag may be signaled immediately after the skip flag and MM flag to specify whether MMVD mode is used for the current CU.

[0137]

[0158] In some exemplary implementations of MMVD, a list of merge candidates for motion vector prediction can be formed for a block to be predicted. The list of merge candidates may contain a predetermined number (e.g., two) of MV predictor candidate blocks whose motion vectors can be used to predict the current motion vector. MVD candidate blocks may include blocks selected from adjacent blocks within the same frame and / or temporal blocks (e.g., identically located blocks in the current frame's forward-movement frame or subsequent frames). These options represent blocks at spatial or temporal locations relative to the current block that are likely to have the same or similar motion vectors as the current block. The size of the list of MV predictor candidates may be predetermined. For example, the list may contain two candidates. To be in the list of merge candidates, candidate blocks may be required to have, for example, the same reference frame (or multiple frames) as the current block (e.g., boundary checking must be performed if the current block is close to a frame edge), and to have already been encoded during the encoding process and / or decoded during the decoding process. In some implementations, the list of merge candidates may first include spatially adjacent blocks (especially those scanned in a predefined order), provided they are available and meet the above conditions, and then, if space is still available in the list, temporal blocks may be included. Adjacent candidate blocks may be selected, for example, from the blocks to the left and above the current block. The list of merge MV predictor candidates may be signaled within the bitstream.

[0138]

[0159] In some implementations, the actual merge candidate used as a reference motion vector for predicting the motion vector of the current block may be signaled. If the merge candidate list contains two candidates, a one-bit flag called the merge candidate flag may be used to indicate the selection of the reference merge candidate. For the current block predicted in composite mode, each of the multiple motion vectors predicted using the MV predictor may be associated with a reference motion vector from the merge candidate list.

[0139]

[0160] In some exemplary implementations of MMVD, after a merge candidate is selected and used as a base motion vector predictor for the predicted motion vector, a motion vector difference (MVD or deltaMV, representing the difference between the predicted motion vector and the reference candidate motion vector) may be calculated in the encoding system. Such an MVD may include information representing the magnitude and direction of the MV difference, which may be signaled in the bitstream. The magnitude and direction of the motion difference can be signaled in various ways.

[0140]

[0161] In some exemplary implementations of MMVD, the distance index may be used to specify the magnitude of the motion vector difference and indicate one of a set of predefined offsets that represent a predefined motion vector difference from the starting point (reference motion vector). The MV offset by the signaled index may then be added to either the horizontal or vertical component of the starting (reference) motion vector. Whether the horizontal or vertical component of the reference motion vector should be offset is determined by the exemplary directional information of the MVD. An example of a predefined relationship between the distance index and the predefined offset is specified in Table 2.

[0141] Table 2 - Example of the relationship between distance index and predefined MV offset

[0142] [Table 2]

[0162] In some exemplary implementations of MMVD, the direction index may be further signaled and used to represent the direction of the MVD relative to the reference motion vector. In some implementations, the direction may be restricted to either the horizontal or vertical direction. Table 3 shows exemplary 2-bit direction indices. In the examples in Table 3, the interpretation of the MVD can vary depending on the information of the start / reference MV. For example, if the start / reference MV corresponds to a uni-prediction block, or if it corresponds to a bi-directional block and both reference frame lists point to the same side of the current picture (i.e., the POCs of both reference pictures are either greater than or less than the POC of the current picture), the sign in Table 3 may specify the sign (direction) of the MV offset added to the start / reference MV. If the start / reference MV corresponds to a bidirectional block and the two reference pictures are on different sides of the current picture (i.e., the POC of one reference picture is greater than the POC of the current picture and the POC of the other reference picture is less than the POC of the current picture), and if the difference between the reference POC in picture reference list 0 and the current frame is greater than the difference between the reference POC in picture reference list 1 and the current frame, then the sign in Table 3 can specify the sign of the MV offset added to the reference MV corresponding to the reference picture in picture reference list 0, and the sign of the offset of the MV corresponding to the reference picture in picture reference list 1 can have the opposite value (the opposite sign of the offset). Otherwise, if the difference between the reference POC in picture reference list 1 and the current frame is greater than the difference between the reference POC in picture reference list 0 and the current frame, the sign in Table 3 can specify the sign of the MV offset added to the reference MV associated with picture reference list 1, while the sign of the offset of the reference MV associated with picture reference list 0 may have the opposite value.

[0143] Table 3 - Implementation example regarding the sign of the MV offset specified by the direction index

[0144] [Table 3]

[0163] In some exemplary implementations, the MVD may be scaled according to the difference in the POCs in each direction. If the difference in the POCs in both lists is the same, scaling is not necessary. Otherwise, if the difference in the POCs in reference list 0 is greater than that of reference list 1, the MVD of reference list 1 is scaled. If the difference in the POCs in reference list 1 is greater than that of reference list 0, the MVD of reference list 0 may be scaled similarly. If the starting MV is uni-predicted, the MVD is added to the available or referenced MV.

[0145]

[0164] In some exemplary implementations of MVD coding and signaling for bidirectional composite prediction, symmetric MVD coding may be implemented, in addition to or instead of coding and signaling two MVDs separately, requiring only one MVD to be signaled and allowing the other MVD to be derived from the signaled MVD. In such implementations, motion information including reference picture indices in both List 0 and List 1 is signaled. However, only the MVD associated with reference list 0 is signaled, and the MVD associated with reference list 1 is derived without being signaled. Specifically, at the slice level, a flag may be included in the bitstream, called ("mvd_l1_zero_flag"), indicating whether reference list 1 is not signaled in the bitstream. If this flag is 1, indicating that reference list 1 is equal to 0 (and therefore not signaled), then a bidirectional prediction flag called "BiDirPredFlag" may be set to 0, meaning that there is no bidirectional prediction. If not, and mvd_l1_zero_flag is 0, BiDirPredFlag may be set to 1 if the nearest reference picture in list-0 and the nearest reference picture in list-1 form a forward-backward or backward-backward pair of reference pictures, and both reference pictures in list-0 and list-1 are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. A BiDirPredFlag of 1 indicates that a symmetric mode flag is additionally signaled in the bitstream. The decoder can extract the symmetric mode flag from the bitstream if BiDirPredFlag is 1. The symmetric mode flag may be signaled, for example, at the CU level (if necessary), indicating whether a symmetric MVD coding mode is being used for the corresponding CU.If the symmetric mode flag is 1, it indicates the use of the symmetric MVD coding mode, where only the reference picture indices of both list-0 and list-1 (referred to as "mvp_l0_flag" and "mvp_l1_flag") are signaled with the MVD associated with list-0 (referred to as "MVD0"), while the other motion vector difference "MVD1" is derived rather than signaled. For example, MVD1 may be derived as -MVD0. Thus, in the exemplary symmetric MVD mode, only one MVD is signaled. In other exemplary implementations of MV prediction, harmonized schemes may be used to implement the general merge mode, MMVD, and some other type of MV prediction for both single-reference mode and composite-reference mode MV prediction. Various syntactic elements can be used to signal the scheme in which the MV of the current block is predicted.

[0146]

[0165] For example, in single-reference mode, the following MV prediction modes may be signaled:

[0166] NEARMV - Without directly using any MVD, it uses one of the motion vector predictors (MVPs) in a list indicated by a Dynamic Reference List (DRL) index.

[0147]

[0167] NEWMV - Uses one of the motion vector predictors (MVPs) in a list signaled by a DRL index as a reference and applies a delta to the MVP (e.g., using an MVP).

[0148]

[0168] GLOBALMV - Uses motion vectors based on frame-level global motion parameters.

[0149]

[0169] Similarly, with respect to a composite reference interpretation mode that uses two reference frames corresponding to two predicted MVs, the following MV prediction modes may be signaled:

[0170] NEAR_NEARMV - For each of the two predicted MVs, use one of the motion vector predictors (MVPs) in a list signaled by a DRL index, without using MVD.

[0150]

[0171] NEAR_NEWMV - To predict the first of two motion vectors, use one of the motion vector predictors (MVPs) in a list signaled by a DRL index as the reference MV, without using MVD; to predict the second of two motion vectors, use one of the motion vector predictors (MVPs) in a list signaled by a DRL index as the reference MV, in combination with an additionally signaled delta MV (MVD).

[0151]

[0172] NEW_NEARMV - To predict the second of two motion vectors, use one of the motion vector predictors (MVPs) in a list signaled by a DRL index as the reference MV, without using MVD; to predict the first of two motion vectors, use one of the motion vector predictors (MVPs) in a list signaled by a DRL index as the reference MV, in combination with an additionally signaled delta MV (MVD).

[0152]

[0173] NEW_NEWMV - Uses one of the motion vector predictors (MVPs) in a list signaled by the DRL index as the reference MV, and uses it in combination with an additionally signaled delta MV to make predictions for each of the two MVs.

[0153]

[0174] GLOBAL_GLOBALMV - Uses MV from each reference based on frame-level global motion parameters.

[0154]

[0175] Therefore, the term "NEAR" above refers to MV prediction using a reference MV without using an MVD, as a general merge mode, while the term "NEW" refers to MV prediction that uses a reference MV and supplements it with a signaled MVD, as in MMVD mode. In the case of composite interpretation, both the reference-based motion vector and the motion vector difference described above can generally differ or be independent between the two references, even if they are correlated, and such correlation can be used to reduce the amount of information required to signal the two motion vector differences. In such situations, joint signaling of the two MVDs may be implemented and specified in the bitstream.

[0155]

[0176] The above Dynamic Reference List (DRL) can be used to hold a set of indexed motion vectors that are dynamically maintained and considered as candidate motion vector predictors.

[0156]

[0177] In some exemplary implementations, an optical flow-based approach can be used to refine motion vectors (MVs) at the sub-block level for composite prediction. In particular, the optical flow equation can be applied to formulate a least-squares problem from which precise motions can be derived from the gradients of composite interprediction samples. These fine motions allow for the refinement of the MVs per sub-block within the prediction block, which can improve interprediction quality. Some coding features may be extensions of the concept of bi-directional optical flow (BDOF) because it supports MV refinement when the two reference blocks are at any temporal distance from the current block.

[0157]

[0178] Some implementations may include four additional intercomplex modes, listed below: NEAR_NEARMV_OPTFLOW, NEAR_NEWMV_OPTFLOW, NEW_NEARMV_OPTFLOW, and / or NEW_NEWMV_OPTFLOW.

[0158]

[0179] These modes are sometimes called optical flow modes, and the reference MV type can be defined in the same way as in normal composite modes (for example, NEAR_NEWMV_OPTFLOW has the same reference MV type as NEAR_NEWMV). Composite prediction can be performed based on a refined MV for each sub-block instead of the original MV.

[0159]

[0180] The various embodiments and / or implementations described herein may be used separately or in any combination. Furthermore, some, all, or any partial or overall combination of these embodiments and / or implementations may be embodied as part of an encoder and / or decoder and may be implemented in hardware and / or software. For example, they may be hardcoded in dedicated processing circuits (e.g., one or more integrated circuits). In another example, they may be implemented by one or more processors that execute a program stored on a non-temporary computer-readable medium.

[0160]

[0181] There may be several challenges / issues related to the implementation of some motion vector difference signaling methods, such as how delta MV(s) are signaled in NEW_NEARMV mode, NEAR_NEWMV mode, or NEW_NEWMV mode. One of these challenges / issues may be that the correlation of motion vector differences in the two reference lists is not utilized, thus reducing the efficiency and performance of coding / decoding.

[0161]

[0182] This disclosure describes various embodiments for signaling motion vector differences (MVD or delta-MV) for interpredictive mode coding and / or decoding, addressing at least one of the challenges / problems described above and enabling improved and efficient software / hardware implementations for interpredictive mode coding / decoding.

[0162]

[0183] Referring to Figure 18 in various embodiments, a method 1800 for decoding interpredicted video blocks is shown. Method 1800 may include all or part of the following steps: step 1810, a device including memory for storing instructions and a processor for communicating with memory receives a coded video bitstream; step 1820, the device retrieves from the coded video bitstream a flag indicating whether the interpredicted mode is JOINT_NEWMV mode with respect to the current block of the current frame, wherein JOINT_NEWMV mode is a first delta motion vector (MV) for a first reference frame from reference list 0 and a second reference from reference list 1 Steps include: indicating that a second delta MV for a frame is signaled together; step 1830, in response to a flag indicating that the interprediction mode is JOINT_NEWMV mode: the device retrieves a joint delta motion vector (MV) for the current block from the coded video / bitstream; and the device derives a first delta MV and a second delta MV based on the joint delta MV; and / or 1840, the device decodes the current block of the current frame based on the first delta MV and the second delta MV.

[0163]

[0184] In some implementations, step 1820 may include the step of the device extracting an inter-prediction mode for the current block of the current frame from the coded video bitstream; and / or, in response to the inter-prediction mode being a mode indicating that a first delta MV for a first reference frame and a second delta MV for a second reference frame are signaled together, step 1830 may include the step of the device extracting a joint delta motion vector (MV) for the current block from the coded video bitstream, and the device deriving a first delta MV and a second delta MV based on the joint delta MV. The mode indicating that a first delta MV for a first reference frame and a second delta MV for a second reference frame are signaled together may be the JOINT_NEWMV mode.

[0164]

[0185] In various embodiments of this disclosure, the size of a block (e.g., a coding block, a prediction block, or a transformation block) may refer to the width or height of the block. The width or height of the block may be an integer in pixels. In various embodiments of this disclosure, the size of a block may refer to the area size of the block. The area size of the block may be an integer in pixels, calculated by multiplying the width of the block by the height of the block. In some various embodiments of this disclosure, the size of a block may refer to the maximum width or height of the block, the minimum width or height of the block, or the aspect ratio of the block. The aspect ratio of the block may be calculated by dividing the width of the block by the height, or by dividing the height of the block by the width.

[0165]

[0186] Herein, in some embodiments of the present disclosure, “first” reference frame may refer not only to “one” reference frame, but also to “first” reference frame among multiple reference frames (e.g., the one with the smallest index or the one that appears earliest in the sequence), and “second” reference frame may refer not only to “another” reference frame, but also to “second” reference frame among multiple reference frames (e.g., the one with the second smallest index or the one that appears second earliest in the sequence).

[0166]

[0187] Herein, in various embodiments of the present disclosure, “XYZ is signaled” may mean that XYZ is encoded into a coded bitstream during the coding process; and / or, after the coded bitstream has been transmitted from one device to another, “XYZ is signaled” may mean that XYZ is decoded / extracted from the coded bitstream during the decoding process.

[0167]

[0188] In various embodiments of this disclosure, the orientation of a reference frame may be determined by whether the reference frame is before or after the current frame in the display order. In some implementations of composite reference modes, the orientation of two reference frames is the same if the picture order counts (POCs) of both reference frames for a given motion vector pair are greater than or less than the POC of the current frame. Otherwise, the orientation of the two reference frames is different if the POC of one reference frame is greater than the POC of the current frame and the POC of the other reference frame is less than the POC of the current frame.

[0168]

[0189] Herein, in various embodiments of this disclosure, “block” may refer to a prediction block, a coding block, a transformation block, or a coding unit (CU).

[0169]

[0190] With respect to step 1810, the device may be the electronic device (530) in Figure 5 or the video decoder (810) in Figure 8. In some implementations, the device may be the decoder (633) within the encoder (620) in Figure 6. In other implementations, the device may be part of the electronic device (530) in Figure 5, part of the video decoder (810) in Figure 8, or part of the decoder (633) within the encoder (620) in Figure 6. The coded video bitstream may be the coded video sequence in Figure 8, or it may be the intermediate coded data in Figure 6 or Figure 7.

[0170]

[0191] With respect to step 1820, the device can extract the interprediction mode for the current block of the current frame from the coded video bitstream. In some implementations, the current block may be a composite reference mode. The interprediction mode may be a new intercoded mode, e.g., JOINT_NEWMV, which indicates that the first delta MV for the first reference frame and the second delta MV for the second reference frame are signaled together.

[0171]

[0192] With respect to step 1830, in response to the interpretation mode being a mode indicating that the first delta MV for the first reference frame and the second delta MV for the second reference frame are signaled together: the device can extract the joint delta motion vector (MV) for the current block from the coded video bitstream and derive the first delta MV and the second delta MV based on the joint delta MV.

[0172]

[0193] In some implementations, the mode may be a new intercoding mode named JOINT_NEWMV mode, which indicates that delta MVs for multiple reference lists (e.g., reference list 0 and reference list 1) are jointly signaled and / or transmitted. When the inter-prediction mode is equal to JOINT_NEWMV mode, in a situation where the current frame is inter-predicted based on two reference frames, the delta MVs for the first reference frame indicated by the first reference list (reference list 0) and the second reference frame indicated by the second reference list (reference list 1) are jointly signaled. Thus, for example, a single delta MV named joint_delta_mv may be signaled and transmitted to decode the current block, in which case the delta MVs for reference list 0 and reference list 1 may be derived from joint delta MV (joint_delta_mv). Joint delta MV may also be called MV difference (MVD).

[0173]

[0194] In some other implementations, the mode is one of the composite reference modes; the syntax terms (sometimes abbreviated as syntax) extracted from the coded video bitstream indicate the interpredictive mode.

[0174]

[0195] In some other implementations, the mode is JOINT_NEWMV; the prediction mode is: JOINT_NEWMV mode, NEAR_NEARMV mode, NEAR_NEWMV mode, NEW_NEARMV mode, NEW_NEWMV mode, and GLOBAL_GLOBALMV mode It is one of them.

[0175]

[0196] In some other implementations, whether the interpretation mode is in that mode (JOINT_NEWMV mode) is based on at least one of the following: the first picture order count (POC) distance between the first reference frame and the current frame, the second POC distance between the second reference frame and the current frame, or the directional relationship between the first and second reference frames with respect to the current frame.

[0176]

[0197] In some other implementations, the interprediction mode is such that the directional relationship indicates opposite directions, or the first POC distance is equal to the second POC distance.

[0177]

[0198] In some other implementations, method 1800 may further include the step of the device deriving a context for deriving an interprediction mode based on at least one of: a first picture order count (POC) distance between a first reference frame and the current frame, a second POC distance between a second reference frame and the current frame, or the directional relationship between the first and second reference frames with respect to the current frame.

[0178]

[0199] In some other implementations, if the current block is coded as a composite reference mode, one syntax can be used to signal whether or not JOINT_NEWMV mode is used for the current block.

[0179]

[0200] In some other implementations, the JOINT_NEWMV mode may be a new mode added to the list of composite reference modes, and as a result, the list of composite reference modes is JOINT_NEWMV mode, NEAR_NEARMV mode, NEAR_NEWMV mode, NEW_NEARMV mode, NEW_NEWMV mode, and GLOBAL_GLOBALMV mode, This may include the following. Therefore, JOINT_NEWMV mode can be signaled and / or transmitted in a manner similar to any of the NEAR_NEARMV mode, NEAR_NEWMV mode, NEW_NEARMV mode, NEW_NEWMV mode, and / or GLOBAL_GLOBALMV mode (for example, via a common syntax).

[0180]

[0201] In some other implementations, the JOINT_NEWMV mode may be conditionally signaled based on the orientation of the two reference frames and / or the distance of the two reference frames relative to the current frame. For example, the JOINT_NEWMV mode may be signaled only if the orientations of the two reference frames are different and / or the absolute values ​​of the distance between the two reference frames relative to the current frame are the same.

[0181]

[0202] In some other implementations, the context for signaling the JOINT_NEWMV mode may depend on the directional relationship of the two reference frames and / or the distance of the two reference frames to the current frame. The directional relationship of the two reference frames to the current frame may refer to whether the two reference frames are oriented in the same direction or opposite directions. The distance of the two reference frames to the current frame may refer to the POC distance of the two reference frames to the current frame, which may include at least one of the following: whether the absolute values ​​of the POC distances of the two reference frames to the current frame are the same or different; the ratio of the POC distances of the two reference frames to the current frame; and / or whether the absolute value of the first POC distance of the first reference frame in the two reference frames is greater than the absolute value of the second POC distance of the second reference frame in the two reference frames.

[0182]

[0203] In various embodiments / implementations of this disclosure, step 1803, which involves deriving a first delta MV and a second delta MV based on the joint delta MV, includes the steps of determining one of the two delta MVs as the joint delta MV and scaling the joint delta MV to the following: The first POC distance between the first reference frame and the current frame, The second POC distance between the second reference frame and the current frame, or The orientation of the first and second reference frames relative to the current frame. The process may include the step of obtaining the other of the delta-MV according to at least one of the following. In some implementations, the scaling may be a linear scaling method, i.e., the absolute value of the scaled delta-MV may be proportional to the ratio obtained by dividing the second POC distance by the first POC distance, and the sign of the scaled delta-MV may be determined according to the directional relationship. For example, if the first POC distance is 4 and the second POC distance is 8, and the directional relationship is the same for the first and second reference frames with respect to the current frame, then the second delta MV is obtained by scaling / multiplying the joint delta MV by a coefficient of 2 (=8 / 4); the second delta MV has the same sign as the joint delta MV because the directional relationship is the same.

[0183] As another example, if the first POC distance is 3 and the second POC distance is -9, and the directional relationship is opposite with respect to the current frame for the first and second reference frames, the second delta MV is obtained by scaling / multiplying the joint delta MV by a coefficient of -3 (=-9 / 3); the second delta MV has the opposite sign to the joint delta MV because the directional relationship is opposite.

[0184]

[0204] In some other implementations, in response to the interpredictive mode being that mode, the device performs optical flow motion refinement on the current block.

[0185]

[0205] In some other implementations, method 1800 may further include the steps of: the device extracting a flag from a coded video bitstream indicating whether optical flow motion refinement is to be performed for that mode; and the device performing optical flow motion refinement for the current block in response to the flag indicating that optical flow motion refinement is to be performed for that mode. In some other implementations, the flag may be signaled and / or transmitted in a high-level syntax including at least one of a sequence parameter set (SPS), video parameter set (VPS), picture parameter set (PPS), picture header, tile header, slice header, frame header, coding tree unit (CTU) header, or superblock header.

[0186]

[0206] In some other implementations, method 1800 may further include a step in which the device extracts a flag from the coded video bitstream indicating whether the mode is applicable to the current block.

[0187]

[0207] In some other implementations, optical flow motion refinement is always applied to JOINT_NEWMV mode.

[0188]

[0208] In some other implementations, a flag / syntax is signaled to indicate whether optical flow motion refinement is applied to JOINT_NEWMV mode.

[0189]

[0209] In some other implementations, a flag is signaled with high-level syntax to indicate whether JOINT_NEWMV mode can be used, and the flag includes, but is not limited to, SPS, VPS, PPS, picture header, tile header, slice header, frame header, and CTU (or superblock) header.

[0190]

[0210] In some implementations of composite reference modes, a first flag / syntax, which may be named optflow_flag, is transmitted to the device to indicate whether optical flow motion refinement applies to JOINT_NEWMV mode. In some implementations, a value of 0 for the first flag may indicate that optical flow motion refinement applies to JOINT_NEWMV mode (e.g., JOINT_NEWMV_OPTFLOW mode); a value of 1 for the first flag may indicate that optical flow motion refinement does not apply to JOINT_NEWMV mode. Conversely, in some other implementations, a value of 1 for the first flag may indicate that optical flow motion refinement applies to JOINT_NEWMV mode (e.g., JOINT_NEWMV_OPTFLOW mode); a value of 0 for the first flag may indicate that optical flow motion refinement does not apply to JOINT_NEWMV mode.

[0191]

[0211] In some implementations of composite reference mode, a second flag / syntax, which may be named joint_mvd_flag, can be transmitted to the device to indicate whether or not JOINT_NEWMV mode is available.

[0192]

[0212] In some other implementations, in response to a value of a second flag (e.g., joint_mvd_flag) indicating that a delta MV for the JOINT_NEWMV mode may be used, only one joint delta MV, which may be named joint_delta_mv, is signaled and transmitted to the decoder, and the delta MVs for reference list 0 and reference list 1 can be derived from the joint delta MV (e.g., joint_delta_mv). In response to a value of a second flag (e.g., joint_mvd_flag) indicating that the JOINT_NEWMV mode may not be used, zero, one, or two delta MVs may be signaled separately for reference list 0 and / or reference list 1 based on the inter-prediction mode. For example, if a single inter-mode is "NEAR" or two are "NEAR" for a combined mode (in which case, under certain circumstances, delta MV may not be required), zero delta MVs may be signaled / transmitted; if a single inter-mode is "NEW" or one is "NEW" and one is "NEAR" for other inter-predictive modes, one delta MV may be signaled / transmitted; if two are "NEW" for inter-predictive modes, two delta MVs may be signaled / transmitted.

[0213] In some implementations, a value of 0 for the second flag may indicate that JOINT_NEWMV mode is available, for example, the delta MVs of reference list 0 and reference list 1 are signaled together, and only one joint delta MV is signaled and transmitted; a value of 1 for the second flag may indicate that JOINT_NEWMV mode is not available, for example, the delta MVs of reference list 0 and reference list 1 are not signaled together, and zero (or one or two) joint delta MVs are signaled and transmitted. Conversely, in some other implementations, a value of 1 for the second flag may indicate that JOINT_NEWMV mode is available, and a value of 0 for the flag may indicate that JOINT_NEWMV mode is not available.

[0193]

[0214] In some implementations, if the interpretation mode of the current block is JOINT_NEWMV mode and / or if a flag (e.g., joint_mvd_flag) indicates that the delta MVs of reference list 0 and reference list 1 are signaled together, then the delta MV for one reference list (e.g., reference list 0 or reference list 1) may be determined as joint_delta_mv; and the delta MV for another reference list may be derived from joint_delta_mv based on the POC distances of the first and second reference frames to the current frame and the directions of the two reference frames.

[0194]

[0215] In some other implementations, the interpretation mode of the current block is JOINT_NEWMV mode; step 1830 is the step of determining the first delta MV as joint delta MV; and, The first picture order count (POC) distance between the first reference frame and the current frame, The second POC distance between the second reference frame and the current frame, or The orientation of the first and second reference frames relative to the current frame. This may include a step of determining a second delta MV by scaling the joint delta MV according to at least one of the following:

[0195]

[0216] In some other implementations, the interpretation mode of the current block is JOINT_NEWMV mode; step 1830 is the step of determining the second delta MV as joint delta MV; and, The first picture order count (POC) distance between the first reference frame and the current frame, The second POC distance between the second reference frame and the current frame, or The orientation of the first and second reference frames relative to the current frame. This may include the step of determining a first delta MV by scaling the joint delta MV according to at least one of the following:

[0196]

[0217] In some embodiments, the delta MV of reference list 0 (or list 1) may always be set to be equal to joint_delta_mv, and the delta MV of reference list 1 (or list 0) may be scaled from joint_delta_mv according to the POC distance of the reference frame to the current frame and / or the direction of the two reference frames.

[0197]

[0218] In some other implementations, the interpretation mode of the current block is JOINT_NEWMV mode; step 1830 is performed in response to the first absolute POC distance between the first reference frame and the current frame being greater than the second absolute POC distance between the second reference frame and the current frame: The first Delta MV was designated as the Joint Delta MV; To determine the second delta MV by scaling the joint delta MV: The first POC distance between the first reference frame and the current frame, The second POC distance between the second reference frame and the current frame, or The orientation of the first and second reference frames relative to the current frame. It is possible to include a step that is carried out according to at least one of the following.

[0198]

[0219] In some other implementations, the interpretation mode of the current block is JOINT_NEWMV mode; step 1830 is performed in response to the first absolute POC distance between the first reference frame and the current frame being smaller than the second absolute POC distance between the second reference frame and the current frame: The second Delta MV was designated as the Joint Delta MV; To determine the first delta MV by scaling the joint delta MV: The first POC distance between the first reference frame and the current frame, The second POC distance between the second reference frame and the current frame, or The orientation of the first and second reference frames relative to the current frame. It is possible to include a step that is carried out according to at least one of the following.

[0199]

[0220] In some embodiments, if the absolute POC distance between reference list 0 (or list 1) and the current frame is greater than the absolute POC distance between reference list 1 (or list 0) and the current frame, the delta MV of reference list 0 (or list 1) with the greater absolute POC distance can be set equal to joint_delta_mv. The other delta MV of reference list 1 (or list 0) with the smaller absolute POC distance can be scaled from joint_delta_mv according to the POC distance of the reference frame to the current frame and / or the direction of the two reference frames.

[0200]

[0221] In some other implementations, the interpretation mode of the current block is JOINT_NEWMV mode; step 1830 is, In response to the first absolute POC distance between the first reference frame and the current frame being equal to the second absolute POC distance between the second reference frame and the current frame: the first delta MV is determined as the joint delta MV; In response to the fact that the directional relationship of the first and second reference frames with respect to the current frame is the same: the second delta MV is determined as the joint delta MV; In response to the directional relationship between the first and second reference frames with respect to the current frame being opposite, the second delta MV may include a step of determining the joint delta MV multiplied by -1.

[0201]

[0222] In some other implementations, if the absolute POC distance between reference list 1 and the current frame is the same as the absolute POC distance between reference list 0 and the current frame, the delta MV of reference list 0 can be set to equal to joint_delta_mv. If the orientations of the two reference frames are the same, the delta MV of reference list 1 may similarly be set to equal to joint_delta_mv. Otherwise, if the orientations of the two reference frames are different, the delta MV of reference list 1 is set to joint_delta_mv multiplied by -1.

[0202]

[0223] In some other implementations, an optical flow-based approach is used to refine motion vectors (MVs) at the sub-block level for composite prediction. In particular, the optical flow equation may be applied to formulate a least-squares problem from which fine motions can be derived from the gradients of composite interprediction samples. These fine motions allow for the refinement of the MVs per sub-block within the prediction block, which can improve interprediction quality. One coding feature may be an extension of the concept of bi-directional optical flow (BDOF) because it supports MV refinement when the two reference blocks have an arbitrary temporal distance from the current block.

[0203]

[0224] Using BDOF, the bi-prediction of the current block is enhanced through a more accurate motion vector derived from its two reference blocks. Using the more accurate motion vector (e.g., via optical flow refinement), it is possible to reduce the motion vector prediction error, resulting in better coding performance. Some other implementations may include one or more additional composite intermodes, which are: NEAR_NEARMV_OPTFLOW mode (e.g., NEAR_NEARMV mode with optical flow motion vector refinement), NEAR_NEWMV_OPTFLOW mode (e.g., NEAR_NEWMV mode with optical flow motion vector refinement), NEW_NEARMV_OPTFLOW mode (e.g., NEW_NEARMV mode with optical flow motion vector refinement), and / or This may include the NEW_NEWMV_OPTFLOW mode (for example, the NEW_NEWMV mode with optical flow motion vector refinement). These composite intermodes with optical flow motion vector refinements may also be called optical flow modes, and the reference MV type may be defined in the same way as in normal composite modes (for example, NEAR_NEWMV_OPTFLOW has the same reference MV type as NEAR_NEWMV). Composite prediction may be performed based on the refined MV for each subblock instead of the original MV.

[0204]

[0225] The embodiments in this disclosure may be used separately or in any combination of any order. In this disclosure, any step or process in any embodiment may be combined in any quantity or in any order as desired. In this disclosure, two or more steps or processes in any embodiment may be performed in parallel. Furthermore, each method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors, or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-temporary computer-readable medium. Embodiments of this disclosure may be applied to luma blocks or chroma blocks.

[0205]

[0226] The technologies described above can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, Figure 19 shows a computer system (2000) suitable for realizing a particular embodiment of the disclosed subject matter.

[0206]

[0227] Computer software can be coded using any suitable machine code or computer language that may be subject to assembly, compilation, linking, or similar mechanisms, to create code that includes instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or instructions that are executed via interpretation, microcode execution, etc.

[0207]

[0228] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, and Internet of Things devices.

[0208]

[0229] The components shown in Figure 19 with respect to the computer system (2000) are essentially illustrative and are not intended to imply any limitations on the scope or functionality of the computer software that implements embodiments of this disclosure. Furthermore, the configuration of the components should not be construed as having any dependency or requirement on any one or combination of components shown in the exemplary embodiment of the computer system (2000).

[0209]

[0230] The computer system (2000) may include certain human interface input devices. Such human interface input devices may respond to input from one or more human users, for example, through tactile input (e.g., keystrokes, swipes, data glove movements), auditory input (e.g., voice, applause), visual input (e.g., gestures), or olfactory input (not shown). Human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., conversations, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still image cameras), or video (e.g., 2D video, 3D video including stereoscopic pictures).

[0210]

[0231] Input human interface devices may include one or more of the following (though only one of each is depicted): keyboard (2001), mouse (2002), trackpad (2003), touchscreen (2010), data glove (not shown), joystick (2005), microphone (2006), scanner (2007), and camera (2008).

[0211]

[0232] The computer system (2000) may also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via touch screens (2010), data gloves (not shown), joysticks (2005), although there may also be haptic feedback devices that do not function as inputs), auditory output devices (e.g., speakers (2009), headphones (not shown)), visual output devices (e.g., screens (2010), including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touch screen input functionality, each with or without haptic feedback functionality, some of which may be capable of outputting two-dimensional visual output, three-dimensional or more output by means such as stereoscopic output; virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0212]

[0233] Computer systems (2000) may also include human-accessible storage devices and associated media, such as optical media including CD / DVD ROM / RW (2020) using media such as CD / DVD (2021), thumb drives (2022), removable hard drives or solid-state drives (2023), legacy magnetic media such as tapes and floppy disks (not shown), and specialized ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0213]

[0234] Those skilled in the art will also understand that the term “computer-readable medium” as used in relation to the subject matter disclosed herein does not include a transmission medium, carrier wave, or other transient signal.

[0214]

[0235] A computer system (2000) may also include a network interface (2054) to one or more communication networks (2655). The network may be, for example, wireless, wired, or optical. The network may further be local, wide-area, metropolitan, vehicle and industrial, real-time, latency-tolerant, etc. Examples of networks include Ethernet, wireless LANs, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), wired or wireless wide-area digital networks for television (including cable TV, satellite TV, and terrestrial TV), vehicle and industrial networks including CANBus, etc. Certain networks generally require an external network interface adapter attached to a specific general-purpose data port or peripheral bus (2049) (e.g., a USB port on a computer system (2000)); others are generally integrated into the core of the computer system (2000) by being attached to a system bus, as described below (e.g., an Ethernet interface is integrated into a PC computer system, and a cellular network interface is integrated into a smartphone computer system). Using any of these networks, the computer system (2000) can communicate with other entities. Such communication can be one-way receive-only (e.g., broadcast television), one-way transmit-only (e.g., CANbus to a specific CANbus device), or bi-way, to other computer systems using, for example, local or wide-area digital networks. Specific protocols and protocol stacks can be used for each of these networks and network interfaces, as described above.

[0215]

[0236] The aforementioned human interface devices, human-accessible storage devices, and network interfaces can be mounted on the core (2040) of the computer system (2000).

[0216]

[0237] The core (2040) may include one or more central processing units (CPUs) (2041), graphics processing units (GPUs) (2042), specialized programmable processing devices in the form of field-programmable gate arrays (FPGAs) (2043), hardware accelerators for specific tasks (2044), graphics adapters (2050), etc. These devices, along with read-only memory (ROM) (2045), random-access memory (2046), and internal mass storage devices (e.g., internal non-user-accessible hard drives, SSDs, etc.) (2047), may be connected via a system bus (2048). In some computer systems, the system bus (2048) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be connected directly to the core's system bus (2048) or via a peripheral bus (2049). For example, a screen (2010) can be connected to a graphics adapter (2050). The peripheral bus architecture includes PCI, USB, etc.

[0217]

[0238] The CPU (2041), GPU (2042), FPGA (2043), and accelerator (2044) can be combined to execute specific instructions that constitute the aforementioned computer code. The computer code can be stored in ROM (2045) or RAM (2046). Temporary data can be stored in RAM (2046), while persistent data can be stored, for example, in internal mass storage (2047). High-speed storage and retrieval of any memory device may be possible by utilizing cache memory, which can be closely associated with one or more CPUs (2041), GPUs (2042), mass storage (2047), ROMs (2045), RAM (2046), etc.

[0218]

[0239] Computer-readable media can have computer code therein for performing various computer implementation operations. The media and computer code may be specifically designed and constructed for the purposes of this disclosure, or they may be of a type that is well known and available to those skilled in the art in the field of computer software.

[0219]

[0240] As a non-limiting example, a computer system having architecture (2000), specifically a core (2040), can provide the ability to run software embodied in one or more tangible computer-readable media as a result of a processor (including CPU, GPU, FPGA, accelerator, etc.). Such computer-readable media can be media related to user-accessible mass storage as described above, as well as specific storage of the core (2040) of a non-transient nature, such as mass storage (2047) or ROM (2045) within the core. Software that implements various embodiments of the present disclosure can be stored in such devices and run by the core (2040). The computer-readable media can include one or more memory devices or chips, depending on the specific needs. The software can cause the core (2040), specifically the processor (including CPU, GPU, FPGA, etc.) therein, to run a specific process or a specific part of a specific process described herein, including defining data structures stored in RAM (2046) and modifying such data structures according to a process defined by the software. Furthermore, or alternatively, a computer system may provide functionality as a result of logic wired or otherwise embodied within a circuit (e.g., an accelerator (2044)), which may perform a particular process or a particular part of a particular process described herein, either in place of or in conjunction with software. References to software may include logic, and vice versa, as appropriate. References to computer-readable media may include circuits (such as integrated circuits (ICs)) that store software for execution, circuits that embody logic for execution, or both, where appropriate. This disclosure encompasses any appropriate combination of hardware and software.

[0220]

[0241] While certain inventions are described in relation to exemplary embodiments, the descriptions are not intended to be restrictive. Various modifications of the exemplary embodiments and additional embodiments of the invention will be apparent to those skilled in the art from this description. Those skilled in the art will readily recognize that these and various other modifications may be made to the exemplary embodiments illustrated and described herein without departing from the spirit and scope of the invention. Accordingly, it is assumed that the appended claims will cover any such modifications and alternative embodiments. Certain proportions in the drawings may be exaggerated, while other proportions may be minimized. Accordingly, this disclosure and drawings should be construed as illustrative, not restrictive.

[0221] The following are illustrative examples of the solutions provided by this matter. (Note 1) A method for decoding an interpreted video block: A device including a memory for storing instructions and a processor for communicating with the memory receives a coded video bitstream; Steps include: the device extracting a flag from the coded video bitstream indicating whether the interprediction mode is JOINT_NEWMV mode with respect to the current block of the current frame, wherein JOINT_NEWMV mode indicates that a first delta motion vector (MV) for a first reference frame from reference list 0 and a second delta MV for a second reference frame from reference list 1 are signaled together; In response to the flag indicating that the interpretation mode is JOINT_NEWMV mode: The device extracts the joint delta MV for the current block from the coded video / bitstream; and The device derives the first delta MV and the second delta MV based on the joint delta MV; and The device decodes the current block of the current frame based on the first delta MV and the second delta MV; A method that includes this. (Note 2) In the method described in Appendix 1: The mode is one of the composite reference modes; and The syntax elements extracted from the coded video bitstream represent the interprediction mode, in a method. (Note 3) In the method described in Appendix 1: The JOINT_NEWMV mode is signaled together with one of the following modes: NEAR_NEARMV mode, NEAR_NEWMV mode, NEW_NEARMV mode, NEW_NEWMV mode, and GLOBAL_GLOBALMV mode. (Note 4) In the method described in Appendix 1: Whether the aforementioned interpretation mode is JOINT_NEWMV mode is: The first picture order count (POC) distance between the first reference frame and the current frame, The second POC distance between the second reference frame and the current frame, or A method based on at least one of the directional relationships between the first reference frame and the second reference frame with respect to the current frame. (Note 5) In the method described in Appendix 4: A method in which the interpretation mode is the JOINT_NEWMV mode, depending on whether the directional relationship indicates opposite directions or whether the first POC distance is equal to the second POC distance. (Note 6) In the method described in Appendix 4: A method wherein, depending on the directional relationship indicating opposite directions and the first POC distance being equal to the second POC distance, the interpretation mode is the JOINT_NEWMV mode. (Note 7) In the method described in Appendix 1, further: The device provides a context for deriving the interprediction mode: The first picture order count (POC) distance between the first reference frame and the current frame, The second POC distance between the second reference frame and the current frame, or A method comprising the step of deriving based on at least one of the directional relationships of the first reference frame and the second reference frame with respect to the current frame. (Note 8) In the method described in Appendix 1, further: A method comprising the step of the device performing optical flow motion refinement for the current block, depending on whether the interpretation mode is the JOINT_NEWMV mode. (Note 9) In the method described in Appendix 1, further: The steps include: the device extracting a flag from the coded video bitstream indicating whether optical flow motion refinement is performed for the JOINT_NEWMV mode; and Depending on whether the flag indicates that the optical flow motion refinement is performed for the JOINT_NEWMV mode, the device performs the optical flow motion refinement for the current block; A method that includes this. (Note 10) In the method described in Appendix 1, further: The device retrieves a flag from the coded video bitstream indicating whether the JOINT_NEWMV mode is applicable to the current block; A method that includes this. (Note 11) In the method described in Appendix 9: A method in which the aforementioned flag is signaled in a high-level syntax element that includes at least one of the following: a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a tile header, a slice header, a frame header, a coding tree unit (CTU) header, or a superblock header. (Note 12) A device for decoding interpreted video blocks: Memory for storing instructions; and A processor that communicates with the aforementioned memory; A device comprising, wherein when the processor executes the instruction, the processor is configured to cause the device to perform the method described in any one of the appendices 1 to 11. (Note 13) A non-temporary, computer-readable storage medium for storing instructions, wherein, when an instruction is executed by a processor, the instruction is configured to cause the processor to execute the method described in any one of the appendices 1 to 11. (Note 14) A computer program that causes a processor to perform any one of the methods described in Appendix 1 through 11.

[0222]

[0242] The following is a list of abbreviations, some of which may appear in this disclosure: JEM: Joint Exploration Model; Collaborative Exploration Model VVC: versatile video coding; general-purpose video coding BMS: benchmark set; benchmark set MV: Motion Vector; motion vector HEVC: High Efficiency Video Coding SEI:Supplementary Enhancement Information;Supplementary Enhancement Information VUI:Video Usability Information;Video Usability Information GOPs:Groups of Pictures;Groups of Pictures TUs:Transform Units;Transform Units PUs:Prediction Units;Prediction Units CTUs:Coding Tree Units;Coding Tree Units CTBs:Coding Tree Blocks;Coding Tree Blocks PBs:Prediction Blocks;Prediction Blocks HRD:Hypothetical Reference Decoder;Hypothetical Reference Decoder SNR:Signal Noise Ratio;Signal-to-Noise Ratio CPUs:Central Processing Units;Central Processing Units GPUs:Graphics Processing Units;Graphics Processing Units CRT:Cathode Ray Tube;Cathode Ray Tube LCD:Liquid-Crystal Display;Liquid-Crystal Display OLED:Organic Light-Emitting Diode;Organic Light-Emitting Diode CD:Compact Disc;Compact Disc DVD:Digital Video Disc;Digital Video Disc ROM:Read-Only Memory;Read-Only Memory RAM:Random Access Memory;Random Access Memory ASIC:Application-Specific Integrated Circuit;Application-Specific Integrated Circuit PLD: Programmable Logic Device; Programmable Logic Device LAN: Local Area Network; Local Area Network GSM: Global System for Mobile communications; Global System for Mobile Communications LTE: Long-Term Evolution; Long-Term Evolution CANBus: Controller Area Network Bus; Controller Area Network Bus USB: Universal Serial Bus; Universal Serial Bus PCI: Peripheral Component Interconnect; Peripheral Component Interconnect FPGA: Field Programmable Gate Array; Field Programmable Gate Array SSD: solid-state drive; Solid State Device IC: Integrated Circuit; Integrated Circuit HDR: high dynamic range; High Dynamic Range SDR: standard dynamic range; Standard Dynamic Range JVET: Joint Video Exploration Team; Joint Video Exploration Team MPM: most probable mode; Most Probable Mode WAIP: Wide-Angle Intra Prediction; Wide-Angle Intra Prediction CU: Coding Unit; Coding Unit PU: Prediction Unit; Prediction Unit TU: Transform Unit; Transform Unit CTU: Coding Tree Unit; Coding Tree Unit PDPC: Position Dependent Prediction Combination ISP: Intra Sub-Partitions SPS: Sequence Parameter Set PPS: Picture Parameter Set APS: Adaptation Parameter Set VPS: Video Parameter Set; Video Parameter Set DPS: Decoding Parameter Set ALF: Adaptive Loop Filter SAO: Sample Adaptive Offset CC-ALF: Cross-Component Adaptive Loop Filter CDEF: Constrained Directional Enhancement Filter CCSO: Cross-Component Sample Offset LSO: Local Sample Offset LR: Loop Restoration Filter AV1: AOMedia Video 1; AOMedia Video 1 AV2: AOMedia Video 2; AOMedia Video 2 MVD: Motion Vector Difference CfL: Chroma from Luma; Chroma from Luma SDT: Semi-Decoupled Tree SDP: Semi-Decoupled Partitioning SST: Semi-Separate Tree SB: Super Block; Super Block IBC (or IntraBC): Intra Block Copy CDF: Cumulative Density Function SCC: Screen Content Coding GBI: Generalized Bi-prediction BCW: Bi-prediction with CU-level weights CIIP: Combined intra-inter prediction POC: Picture Order Count RPS: Reference Picture Set DPB: Decoded Picture Buffer MMVD: Merge mode using motion vector difference

Claims

[Claim 1] A method for decoding an interpreted video block: A device including a memory for storing instructions and a processor for communicating with the memory receives a coded video bitstream; Steps include: the device extracting a flag from the coded video bitstream indicating whether the interprediction mode is JOINT_NEWMV mode with respect to the current block of the current frame, wherein JOINT_NEWMV mode indicates that a first delta motion vector (MV) for a first reference frame from reference list 0 and a second delta MV for a second reference frame from reference list 1 are signaled together; In response to the flag indicating that the interpretation mode is JOINT_NEWMV mode: The device extracts the joint delta MV for the current block from the coded video / bitstream; and The device derives the first delta MV and the second delta MV based on the joint delta MV; and The device decodes the current block of the current frame based on the first delta MV and the second delta MV; A method that includes this.