Joint motion vector difference coding
Patent Information
- Application Number
- KR1020237029405
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-03-22
- Filing Date
- 2022-04-15
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-04-15
Smart Images

Figure 112023095251517-PCT00022_ABST
Abstract
Description
Technology Field
[0001] Integration by citation
[0002] This application is based on U.S. Provisional Application No. 63 / 245,655 filed September 17, 2021, and claims the benefit of priority thereto, the entirety of which is incorporated herein by reference. This application is also based on U.S. Regular Application No. 17 / 700,745 filed March 22, 2022, and claims the benefit of priority thereto, the entirety of which is incorporated herein by reference.
[0003] Technology field
[0004] The present disclosure relates to video coding and / or decoding techniques, and in particular, to an improved design and signaling of joint motion vector difference for coding and / or decoding. Background Technology
[0005] The background description provided in this specification is intended to provide a general context for the present disclosure. The research of the currently registered inventors described in this background section and the modes of description that may not otherwise be considered prior art at the time of filing this application are not recognized as prior art to the present disclosure, either explicitly or implicitly.
[0006] Video coding and decoding can be performed using inter-picture prediction accompanied by motion compensation. Uncompressed digital video may contain a series of pictures, each having spatial dimensions of, for example, 1920x1080 luminance samples and associated whole or subsampled chrominance samples. The series of pictures may have a fixed or variable picture rate (alternatively referred to as the frame rate), for example, 60 pictures per second or 60 frames per second. Uncompressed video has specific bitrate requirements for streaming or data processing. For example, video with a pixel resolution of 1920 x 1080, a frame rate of 60 frames / second, and 4:2:0 chroma subsampling at 8 bits per pixel per color channel requires bandwidth approaching 1.5 Gbit / s. One hour of such video requires more than 600 GBytes of storage space.
[0007] One objective of video coding and decoding may be to reduce redundancy in the uncompressed input video signal through compression. Compression can help reduce the previously described bandwidth and / or storage requirements by more than two orders of magnitude in some cases. Both lossless and lossy compression, as well as combinations thereof, can be utilized. Lossless compression refers to techniques where an exact copy of the original signal can be reconstructed from the compressed original signal through the decoding process. Lossy compression refers to coding / decoding processes where the original video information is not fully preserved during coding and is not fully recoverable during decoding. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signals is made small enough to render the reconstructed signal useful for the intended application, despite some loss of information. For video, lossy compression is widely used in many applications. The amount of acceptable distortion depends on the application. For example, users of certain consumer video streaming applications may tolerate higher distortion than users of cinematic or television broadcast applications. The compression ratios achievable by specific coding algorithms can be selected or adjusted to reflect varying distortion tolerances: higher acceptable distortion generally allows for coding algorithms that yield higher losses and higher compression ratios.
[0008] Video encoders and decoders can utilize techniques from several broad categories and steps, including, for example, motion compensation, Fourier transform, quantization, and entropy coding.
[0009] Video codec technologies may include techniques known as intra coding. In intra coding, sample values are represented without referencing samples from previously reconstructed reference pictures or other data. In some video codecs, a picture is spatially subdivided into blocks of samples. When all blocks of samples are coded in intra mode, the picture may be referred to as an intra picture. Intra pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and thus can be used as the first picture or a still image in a coded video bitstream and video session. Samples in a block after intra prediction can then be transformed into the frequency domain, and the resulting transform coefficients can be quantized before entropy coding. Intra prediction refers to a technique that minimizes sample values in the pre-transform domain. In some cases, the smaller the DC value after conversion and the smaller the AC coefficients, the fewer bits are required at a given quantization step size to represent the block after entropy coding.
[0010] For example, traditional intra-coding, such as that known from MPEG-2 generation coding techniques, does not use intra-prediction. However, some newer video compression techniques include methods that attempt to code / decode blocks based on surrounding sample data and / or metadata that are acquired, for example, during the encoding and / or decoding of spatial neighbors and that precede the data blocks being intra-coded or decoded in the decoding order. These methods are hereinafter referred to as "intra-prediction" techniques. Note that, in at least some cases, intra-prediction uses only reference data from the current picture being reconstructed, rather than from other reference pictures.
[0011] There may be many different forms of intra-prediction. When more than one of these techniques is available in a given video coding technique, the technique in use may be referred to as an intra-prediction mode. One or more intra-prediction modes may be provided in a specific codec. In certain cases, modes may have submodes and / or may be associated with various parameters, and mode / submode information and intra-coding parameters for blocks of video may be coded individually or collectively included in mode codewords. The codewords used for a given combination of mode, submode, and / or parameters can influence the coding efficiency gain through intra-prediction, and the entropy coding technique used to convert the codewords into a bitstream can do so as well.
[0012] Specific modes of intra prediction were introduced with H.264, improved in H.265, and further enhanced in newer coding techniques such as JEM (joint exploration model), VVC (versatile video coding), and BMS (benchmark set). Generally, for intra prediction, a predictor block can be formed using available neighbor sample values. For example, available values from a specific set of neighbor samples following a particular direction and / or lines can be copied into the predictor block. References to the direction in use can be encoded in the bitstream or predicted themselves.
[0013] Referring to FIG. 1a, a subset of 9 predictor directions specified from the 33 possible intra predictor directions of H.265 (corresponding to 33 angle modes out of 35 intra modes specified in H.265) is depicted in the lower right. The point (101) where the arrows converge indicates the sample being predicted. The arrows indicate the direction in which neighboring samples are used to predict the sample from 101. For example, arrow (102) indicates that sample (101) is predicted from neighboring samples or samples to the upper right at a 45-degree angle from the horizontal direction. Similarly, arrow (103) indicates that sample (101) is predicted from neighboring samples or samples to the lower left of sample (101) at a 22.5-degree angle from the horizontal direction.
[0014] Referring again to FIG. 1a, a square block (104) of 4x4 samples (indicated by a bold dashed line) is depicted in the upper left. The square block (104) contains 16 samples, each labeled "S", a position in the Y dimension (e.g., row index), and a position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from the top) and the first sample in the X dimension (from the left). Similarly, sample S44 is the fourth sample in both the Y and X dimensions in the block (104). Since the block is 4x4 samples in size, S44 is located in the lower right. Exemplary reference samples following a similar numbering scheme are additionally illustrated. For the block (104), the reference sample is labeled R, its Y position (e.g., row index) and X position (column index). In both H.264 and H.265, neighboring prediction samples adjacent to the block being reconstructed are used.
[0015] Intra-picture prediction of block (104) can be initiated by copying reference sample values from neighboring samples according to the signaled prediction direction. For example, assume that the coded video bitstream includes signaling indicating the prediction direction of the arrow (102) for this block (104)—that is, samples are predicted from the prediction sample or samples at an angle of 45 degrees from the horizontal direction to the upper right. In this case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Subsequently, sample S44 is predicted from reference sample R08.
[0016] In certain cases, the values of multiple reference samples can be combined, for example, through interpolation to calculate a reference sample, particularly when the directions are not evenly divided by 45 degrees.
[0017] As video coding technology continues to advance, the number of possible directions has increased. In H.264 (2003), for example, nine different directions are available for intra-prediction. This increased to 33 in H.265 (2013), and JEM / VVC / BMS can support up to 65 directions at the time of this publication. Experimental studies have been conducted to help identify the most suitable intra-prediction directions, and specific bit penalties for directions can be tolerated by encoding these most suitable directions with a small number of bits using specific techniques in entropy coding. Additionally, the directions themselves can sometimes be predicted from the neighbor directions used in the intra-prediction of decoded neighbor blocks.
[0018] FIG. 1b illustrates a schematic diagram (180) depicting 65 intra-predicted directions according to JEM to illustrate an increasing number of predicted directions in various encoding technologies developed over time.
[0019] The method of mapping bits representing intra-prediction directions to prediction directions in a coded video bitstream can vary depending on the video coding technique; for example, it can range from simple direct mappings of prediction directions to intra-prediction modes to complex adaptive methods involving codewords, most probable modes, and similar techniques. However, in all cases, there may be specific directions for intro prediction that are statistically less likely to occur in the video content than certain other directions. Since the goal of video compression is to reduce redundancy, in a well-designed video coding technique, these less likely directions can be represented by a larger number of bits than the more likely directions.
[0020] Inter-picture prediction, or inter-prediction, may be based on motion compensation. In motion compensation, sample data from a previously reconstructed picture or part thereof (reference picture) may be used to predict a newly reconstructed picture or part of a picture (e.g., a block) after being spatially shifted in the direction indicated by a motion vector (hereinafter MV). In some cases, the reference picture may be the same as the picture currently being reconstructed. MVs may have two dimensions X and Y, or three dimensions, and the third dimension is an indication of the reference picture in use (similar to a time dimension).
[0021] In some video compression techniques, the current MV applicable to a specific region of sample data can be predicted from other MVs, for example, from other regions of sample data spatially adjacent to the region being reconstructed, and from those other MVs that precede the current MV in the decoding order. By doing so, the total amount of data required to code the MVs can be substantially reduced by relying on the elimination of redundancy in correlated MVs, thereby increasing compression efficiency. MV prediction can work effectively, for example when coding an input video signal derived from a camera (known as natural video), because there is a statistical probability that regions larger than the area where a single MV is applicable move in similar directions within the video sequence; thus, in some cases, prediction can be made using similar motion vectors derived from the MVs of neighboring regions. As a result, the actual MV for a given region becomes similar or identical to the MV predicted from the surrounding MVs. These MVs can eventually be represented with fewer bits than when the MV is coded directly rather than predicted from neighboring MV(s) after entropy coding. In some cases, MV prediction can be an example of lossless compression of signals (i.e., MVs) derived from the original signal (i.e., sample stream). In other cases, MV prediction itself can be lossy, for example, due to rounding errors when calculating the predictor from several surrounding MVs.
[0022] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms specified by H.265, a technique referred to as "spatial merge" is described below.
[0023] Specifically, referring to FIG. 2, the current block (201) includes samples discovered by the encoder during the motion search process that are predictable from a previous block of the same size that is spatially shifted. Instead of directly coding the MV, the MV may be derived from metadata associated with one or more reference pictures, for example, from the most recent (in decoding order) reference picture, using an MV associated with any one of five surrounding samples represented as A0, A1, and B0, B1, B2 (202 to 206, respectively). In H.265, the MV prediction may use predictors from the same reference picture used by neighboring blocks.
[0024] The present disclosure describes various embodiments of methods, apparatus, and computer-readable storage media for video encoding and / or decoding.
[0025] According to one aspect, an embodiment of the present disclosure provides a method for decoding an inter-predicted video block. The method comprises receiving a coded video bitstream by a device. The device comprises a memory storing instructions and a processor communicating with the memory. The method also comprises extracting an inter-predicted mode and a joint delta motion vector (MV) for a current block in a current frame from the coded video bitstream by the device; extracting a flag from the coded video bitstream by the device indicating whether a first delta MV for a first reference frame and a second delta MV for a second reference frame are jointly signaled; deriving a first delta MV and a second delta MV based on the joint delta MV by the device in response to the flag indicating that the first delta MV and the second delta MV are jointly signaled by the device; and decoding a current block in a current frame based on the first delta MV and the second delta MV by the device.
[0026] According to another aspect, an embodiment of the present disclosure provides an apparatus for video encoding and / or decoding. The apparatus includes a memory storing instructions; and a processor communicating with the memory. When the processor executes the instructions, the processor is configured to cause the apparatus to perform the above methods for video decoding and / or encoding.
[0027] Aspects of the present disclosure also provide a video encoding or decoding device or apparatus comprising a circuit configured to perform any of the above method implementations.
[0028] In another aspect, embodiments of the present disclosure provide non-transient computer-readable media storing instructions that, when executed by a computer for video decoding and / or encoding, cause the computer to perform the above methods for video decoding and / or encoding.
[0029] The above and other embodiments and implementations are described in more detail in the drawings, descriptions, and claims. Brief explanation of the drawing
[0030] Additional features, nature, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. Figure 1a illustrates a schematic example of an exemplary subset of intra-predicted directional modes. FIG. 1b illustrates examples of exemplary intra-predicted directions. Figure 2 illustrates a schematic example of the current block and spatial merge candidates around it for motion vector prediction in one example. FIG. 3 illustrates a schematic example of a simplified block diagram of a communication system (300) according to an exemplary embodiment. FIG. 4 illustrates a schematic example of a simplified block diagram of a communication system (400) according to an exemplary embodiment. FIG. 5 illustrates a schematic example of a simplified block diagram of a video decoder according to an exemplary embodiment. FIG. 6 illustrates a schematic example of a simplified block diagram of a video encoder according to an exemplary embodiment. FIG. 7 illustrates a block diagram of a video encoder according to another exemplary embodiment. FIG. 8 illustrates a block diagram of a video decoder according to another exemplary embodiment. FIG. 9 illustrates a method of coding block partitioning according to exemplary embodiments of the present disclosure. FIG. 10 illustrates another method of coding block partitioning according to exemplary embodiments of the present disclosure. FIG. 11 illustrates another method of coding block partitioning according to exemplary embodiments of the present disclosure. FIG. 12 illustrates an example of partitioning a base block into coding blocks according to an exemplary partitioning method. Figure 13 illustrates an exemplary ternary partitioning method. FIG. 14 illustrates an exemplary quadtree binary tree coding block partitioning method. FIG. 15 illustrates a method for partitioning a coding block into a plurality of conversion blocks and a coding order of conversion blocks according to exemplary embodiments of the present disclosure. FIG. 16 illustrates another method for partitioning a coding block into a plurality of conversion blocks and the coding order of the conversion blocks according to exemplary embodiments of the present disclosure. FIG. 17 illustrates another method for partitioning a coding block into a plurality of conversion blocks according to exemplary embodiments of the present disclosure. FIG. 18 illustrates a flowchart of a method according to an exemplary embodiment of the present disclosure. FIG. 19 illustrates a schematic diagram of a computer system according to exemplary embodiments of the present disclosure. Specific details for implementing the invention
[0031] Now, the invention will be described in detail below with reference to the accompanying drawings, which form part of the invention and illustrate specific examples of embodiments. However, it should be noted that the invention may be embodied in various different forms, and thus, the subject matter covered or claimed is intended not to be limited to any of the embodiments presented below. It should also be noted that the invention may be embodied in methods, devices, components, or systems. Accordingly, embodiments of the invention may take the form of, for example, hardware, software, firmware, or any combination thereof.
[0032] Throughout the specification and claims, terms may have nuanced meanings implied or suggested by the context beyond their explicitly stated meanings. Phrases such as "in one embodiment" or "in some embodiments" as used herein do not necessarily refer to the same embodiment, nor do phrases such as "in another embodiment" or "in other embodiments" as used herein refer to different embodiments. Likewise, phrases such as "in one implementation" or "in some implementations" as used herein do not necessarily refer to the same implementation, nor do phrases such as "in another implementation" or "in other implementations" as used herein refer to different implementations. For example, the claimed subject matter is intended to include combinations of exemplary embodiments / implements, wholly or in part.
[0033] Generally, terms may be understood at least partially from their use in context. For example, terms such as “and,” “or,” or “and / or” as used herein may include various meanings that may depend at least partially on the context in which these terms are used. Typically, when used to associate a list such as A, B, or C, “or” is intended to mean A, B, and C as used herein in an inclusive sense, as well as A, B, or C as used herein in an exclusive sense. Additionally, terms such as “one or more” or “at least one” as used herein may, depending at least partially on the context, be used to describe any feature, structure, or characteristic in a singular sense, or to describe combinations of features, structures, or characteristics in a plural sense. Similarly, terms such as “one,” “one,” or “that” may again, depending at least partially on the context, be understood to convey a singular usage or a plural usage. Additionally, the terms "based on" or "determined by" may be understood as not necessarily intended to convey an exclusive set of arguments, but instead, again, depending at least partially on the context, may allow for the existence of additional arguments that are not necessarily explicitly described.
[0034] FIG. 3 illustrates a simplified block diagram of a communication system (300) according to an embodiment of the present disclosure. The communication system (300) includes a plurality of terminal devices capable of communicating with each other, for example, through a network (350). For example, the communication system (300) includes a first pair of terminal devices (310 and 320) interconnected through the network (350). In the example of FIG. 3, the first pair of terminal devices (310 and 320) can perform unidirectional transmission of data. For example, the terminal device (310) can encode video data (for example, a stream of video pictures captured by the terminal device (310)) for transmission to another terminal device (320) through the network (350). The encoded video data may be transmitted in the form of one or more coded video bitstreams. The terminal device (320) receives coded video data from the network (350), decodes the coded video data to recover video pictures, and can display the video pictures according to the recovered video data. Unidirectional data transmission can be implemented in media serving applications, etc.
[0035] In another example, the communication system (300) includes a second pair of terminal devices (330 and 340) that perform bidirectional transmission of coded video data that can be implemented, for example, during a video conferencing application. For bidirectional transmission of data, in one example, each terminal device among the terminal devices (330 and 340) can code video data (e.g., a stream of video pictures captured by the terminal device) for transmission to another terminal device among the terminal devices (330 and 340) via a network (350). Each terminal device among the terminal devices (330 and 340) can also receive coded video data transmitted by the other terminal device among the terminal devices (330 and 340), can decode the coded video data to recover video pictures, and can display the video pictures on an accessible display device according to the recovered video data.
[0036] In the example of FIG. 3, the terminal devices (310, 320, 330, and 340) may be implemented as servers, personal computers, and smartphones, but the applicability of the basic principles of the present disclosure may not be so limited. Embodiments of the present disclosure may be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing equipment, and / or similar devices. Network (350) represents any number or type of networks that transmit coded video data between terminal devices (310, 320, 330, and 340), including, for example, wired and / or wireless communication networks. The communication network (350)9 may exchange data on circuit-switched, packet-switched, and / or other types of channels. Representative networks include communication networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of the network (350) may not be important to the operation of the present disclosure unless explicitly described in this specification.
[0037] FIG. 4 illustrates the arrangement of a video encoder and a video decoder in a video streaming environment as an example of an application for the disclosed subject. The disclosed subject may be equally applicable to other video applications, such as, for example, video conferencing, digital TV broadcasting, gaming, virtual reality, and storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0038] A video streaming system may include a video capture subsystem (413) which may include a video source (401), e.g., a digital camera, for generating a stream (402) of uncompressed video pictures or images. In one example, the stream (402) of video pictures includes samples recorded by the digital camera of the video source (401). The stream (402) of video pictures, depicted in bold lines to emphasize the high data capacity compared to encoded video data (404) (or coded video bitstream), may be processed by an electronic device (420) comprising a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof to enable or implement the embodiments of the disclosed subject matter as described in more detail below. Encoded video data (404) (or encoded video bitstream (404)), depicted as a thin line to emphasize the lower data capacity compared to a stream (402) of uncompressed video pictures, may be stored on a streaming server (405) or directly on downstream video devices (not shown) for future use. One or more streaming client subsystems, such as the client subsystems (406 and 408) in FIG. 4, may access the streaming server (405) to retrieve copies (407 and 409) of the encoded video data (404). The client subsystem (406) may include a video decoder (410) within an electronic device (430), for example. A video decoder (410) decodes an incoming copy (407) of encoded video data and generates an outgoing stream (411) of video pictures that are uncompressed and can be rendered on a display (412) (e.g., a display screen) or other rendering devices (not described).The video decoder (410) may be configured to perform some or all of the various functions described in this disclosure. In some streaming systems, encoded video data (404, 407, and 409) (e.g., video bitstreams) may be encoded according to specific video coding / compression standards. Examples of such standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally known as VVC (Versatile Video Coding). The disclosed subject matter may be used in the context of VVC and other video coding standards.
[0039] Note that the electronic devices (420 and 430) may include other components (not shown). For example, the electronic device (420) may also include a video decoder (not shown) and the electronic device (430) may also include a video encoder (not shown).
[0040] FIG. 5 illustrates a block diagram of a video decoder (510) according to any embodiment of the present disclosure below. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used instead of the video decoder (410) in the example of FIG. 4.
[0041] The receiver (531) may receive one or more coded video sequences to be decoded by the video decoder (510). In the same or different embodiments, one coded video sequence may be decoded at a time, wherein the decoding of each coded video sequence is independent of other coded video sequences. Each video sequence may be associated with multiple video frames or images. The coded video sequence may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data or to a streaming source transmitting the encoded video data. The receiver (531) may receive the encoded video data along with other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective processing circuits (not described). The receiver (531) may separate the coded video sequence from the other data. To prevent network jitter, a buffer memory (515) may be placed between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In certain applications, the buffer memory (515) may be implemented as part of the video decoder (510). In other applications, it may be outside the video decoder (510) (not described) and separated from it. In other applications, for example, for the purpose of preventing network jitter, a buffer memory (not described) outside the video decoder (510) may exist, and for example, to handle playback timing, another additional buffer memory (515) inside the video decoder (510) may exist. When the receiver (531) is receiving data from a storage / forward device of sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory (515) may not be needed or may be small.For use on best-effort packet networks such as the Internet, a buffer memory (515) of sufficient size may be required, and the size may be relatively large. This buffer memory may be implemented with an adaptive size and may be implemented at least partially in an operating system or similar elements (not described) outside the video decoder (510).
[0042] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from a coded video sequence. The categories of the symbols include information used to manage the operation of the video decoder (510), and information for controlling a rendering device, such as a display (512) (e.g., a display screen), which may or may not be an integral part of the electronic device (530) as illustrated in FIG. 5, but may be coupled to the electronic device (530). The control information for the rendering device(s) may be in the form of Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not depicted). The parser (520) may parse / entropy decode the coded video sequence received by the parser (520). Entropy coding of a coded video sequence may follow video coding techniques or standards and may follow various principles including variable-length coding, Huffman coding, and arithmetic coding with or without context sensitivity. The parser (520) may extract a set of subgroup parameters for at least one of a subgroup of pixels in a video decoder based on at least one parameter corresponding to the subgroups from the coded video sequence. The subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CU), blocks, Transform Units (TU), Prediction Units (PU), etc.The parser (520) can also extract information such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, etc. from the coded video sequence.
[0043] The parser (520) can generate symbols (521) by performing an entropy decoding / parsing operation on a video sequence received from a buffer memory (515).
[0044] The reconstruction of the symbols (521) may involve a number of different processing or functional units depending on the type of the coded video picture or its parts (e.g., inter- and intra-picture, inter- and intra-block), and other factors. The units involved and how they are involved may be controlled by subgroup control information parsed from the video sequence coded by the parser (520). The flow of this subgroup control information between the parser (520) and the number of processing or functional units below is not depicted for the sake of simplification.
[0045] In addition to the functional blocks already mentioned, the video decoder (510) may be conceptually subdivided into a number of functional units as described below. In actual implementations operating under commercial constraints, many of these functional units may interact closely with one another and be at least partially integrated with one another. However, to clearly explain the various functions of the disclosed subject matter, the conceptual subdivision into functional units is adopted in the following disclosure.
[0046] The first unit may include a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) may receive quantized transform coefficients as well as control information, including information indicating which type of inverse transform to use, block size, quantization factor / parameters, quantization scaling matrices, etc., as symbol(s) (521) from the parser (520). The scaler / inverse transform unit (551) may output blocks containing sample values that can be input to an aggregator (555).
[0047] In some cases, the output samples of the scaler / inverse transform (551) may relate to an intra-coded block, that is, a block that does not use prediction information from previously reconstructed pictures but can use prediction information from previously reconstructed parts of the current picture. Such prediction information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) may generate a block of the same size and shape as the block being reconstructed by using surrounding block information that has already been reconstructed and stored in the current picture buffer (558). The current picture buffer (558) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. In some implementations, the aggregator (555) may add the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) on a per-sample basis.
[0048] In other cases, the output samples of the scaler / inverse conversion unit (551) may be intercoded and potentially associated with a motion-compensated block. In such cases, the motion compensation prediction unit (553) may access the reference picture memory (557) to fetch samples used for inter-picture prediction. After motion-compensating the fetched samples according to the symbols (521) associated with the block, these samples may be added by the aggregator (555) to the output of the scaler / inverse conversion unit (551) (the output of the unit (551) may be referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (557) where the motion compensation prediction unit (553) fetches prediction samples may be controlled by motion vectors available to the motion compensation prediction unit (553) in the form of symbols (521) that may have, for example, X, Y components (shift), and reference picture components (time). Motion compensation may also include interpolation of sample values fetched from reference picture memory (557) when subsample accurate motion vectors are in use, and may also be associated with motion vector prediction mechanisms, etc.
[0049] Various loop filtering techniques within the loop filter unit (556) may be performed on the output samples of the aggregator (555). Video compression techniques may include in-loop filter techniques that are made available to the loop filter unit (556) as symbols (521) from the parser (520) and are controlled by parameters included in the coded video sequence (also referred to as the coded video bitstream), but which respond not only to meta-information acquired during the decoding of the previous (in decoding order) parts of the coded picture or coded video sequence, but also to previously reconstructed and loop-filtered sample values. Several types of loop filters may be included in various order as part of the loop filter unit (556), which will be described in more detail below.
[0050] The output of the loop filter unit (556) may be a sample stream that is not only output to the rendering device (512) but may also be stored in the reference picture memory (557) for use in future inter-picture predictions.
[0051] Certain coded pictures, when fully reconstructed, can be used as reference pictures for future inter-picture prediction. For example, when a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a new current picture buffer can be reallocated before starting the reconstruction of the next coded picture.
[0052] A video decoder (510) can perform decoding operations according to a predetermined video compression technique adopted in a standard such as ITU-T Rec. H.265. In that the coded video sequence adheres to both the syntax of the video compression technique or standard and the profiles documented in the video compression technique or standard, the coded video sequence may comply with the syntax specified by the video compression technique or standard in use. Specifically, the profile may select specific tools from all tools available in the video compression technique or standard as the only tools available to use under that profile. To comply with the standard, the complexity of the coded video sequence may be within the boundaries defined by the levels of the video compression technique or standard. In some cases, the levels limit the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the levels may, in some cases, be further restricted through HRD (Hypothetical Reference Decoder) specifications and metadata for managing HRD buffers signaled in the coded video sequence.
[0053] In some exemplary embodiments, the receiver (531) may receive additional (redundant) data along with the encoded video. This additional data may be included as part of the encoded video sequence(s). This additional data may be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0054] FIG. 6 illustrates a block diagram of a video encoder (603) according to an exemplary embodiment of the present disclosure. The video encoder (603) may be included in an electronic device (620). The electronic device (620) may further include a transmitter (640) (e.g., a transmitting circuit). The video encoder (603) may be used instead of the video encoder (403) in the example of FIG. 4.
[0055] The video encoder (603) can receive video samples from a video source (601) (which is not part of the electronic device (620) in the example of FIG. 6) capable of capturing video image(s) to be encoded by the video encoder (603). In another example, the video source (601) may be implemented as part of the electronic device (620).
[0056] A video source (601) may provide a source video sequence to be coded by a video encoder (603) in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 YCrCb, RGB, XYZ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (601) may be a storage device capable of storing previously prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures or images that impart motion when viewed sequentially. The pictures themselves may be organized as a spatial array of pixels, wherein each pixel may contain one or more samples depending on the sampling structure, color space, etc. being used. A person skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.
[0057] According to some exemplary embodiments, the video encoder (603) can code and compress pictures of a source video sequence into a coded video sequence (643) in real time or under any other time constraints required by the application. Implementing an appropriate coding rate constitutes one function of the controller (650). In some embodiments, the controller (650) may be functionally coupled to other functional units to control them as described below. The coupling is not depicted for the sake of simplification. Parameters set by the controller (650) may include rate control related parameters (picture skip, quantizer, lambda values of rate-distortion optimization techniques, ...), picture size, picture group (GOP) layout, maximum motion vector search range, etc. The controller (650) may be configured to have other suitable functions related to the video encoder (603) optimized for a specific system design.
[0058] In some exemplary embodiments, the video encoder (603) may be configured to operate in a coding loop. For the sake of oversimplification, in one example, the coding loop may include a source coder (630) (responsible for generating symbols, such as a symbol stream, based on, for example, the input picture to be coded and reference picture(s)), and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to generate sample data in a manner similar to that generated by the (remote) decoder, even though the embedded decoder (633) processes the video stream coded by the source coder (630) without entropy coding (since any compression between the video bitstream and symbols coded in entropy coding may be lossless in the video compression techniques considered in the subject matter disclosed). The reconstructed sample stream (sample data) is input into the reference picture memory (634). Because the decoding of the symbol stream produces bit-exact results independently of the decoder location (local or remote), the content within the reference picture memory (634) is also bit exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the exact same sample values as the reference picture samples that the decoder "sees" when using prediction during decoding. This fundamental principle of reference picture synchronicity (and the resulting drift when synchronicity cannot be maintained, for example, due to channel errors) is used to improve coding quality.
[0059] The operation of the “local” decoder (633) may be the same as the operation of the “remote” decoder, such as the video decoder (510) already described in detail above in relation to FIG. 5. However, with brief reference to FIG. 5, since symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder (645) and the parser (520) may be lossless, the entropy decoding parts of the video decoder (510), including the buffer memory (515) and the parser (520), may not be fully implemented in the local decoder (633) in the encoder.
[0060] An observation that can be made at this point is that any decoder technique, excluding parsing / entropy decoding which may exist only in the decoder, may also inevitably need to exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter may sometimes focus on decoder operations associated with the decoding part of the encoder. Since encoder techniques are the inverse of the comprehensively described decoder techniques, their description may therefore be condensed. A more detailed description of the encoder is provided below only in specific domains or modalities.
[0061] During operation, in some exemplary implementations, the source coder (630) may perform motion-compensated predictive coding, which predictively codes the input picture by referencing one or more previously coded pictures from a video sequence designated as a “reference picture.” In this way, the coding engine (632) codes the differences (or residuals) in color channels between the pixel blocks of the input picture and the pixel blocks of the reference picture(s) that can be selected as predictive reference(s) for the input picture. The term “residue” and its adjective form “residual” may be used interchangeably.
[0062] The local video decoder (633) can decode the coded video data of pictures that may be designated as reference pictures based on the symbols generated by the source coder (630). The operations of the coding engine (632) may advantageously be lossy processes. If the coded video data can be decoded in a video decoder (not shown in FIG. 6), the reconstructed video sequence may be a replica of the source video sequence, typically having some errors. The local video decoder (633) can replicate the decoding processes that can be performed by the video decoder on the reference pictures and allow the reconstructed reference pictures to be stored in the reference picture cache (634). In this way, the video encoder (603) can locally store copies of reconstructed reference pictures having common content as reconstructed reference pictures to be acquired by the original (remote) video decoder (without transmission errors).
[0063] The predictor (635) can perform prediction searches for the coding engine (632). That is, for a new picture to be coded, the predictor (635) can search the reference picture memory (634) for specific metadata or sample data (as candidate reference pixel blocks), such as reference picture motion vectors, block shapes, etc., which can serve as appropriate prediction references for the new pictures. The predictor (635) can operate on a sample block-by-pixel block basis to find appropriate prediction references. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have prediction references drawn from a number of reference pictures stored in the reference picture memory (634).
[0064] The controller (650) can manage the coding operation of the source coder (630), including, for example, the settings of parameters and subgroup parameters used to encode video data.
[0065] The outputs of all the aforementioned function units may undergo entropy coding in an entropy coder (645). The entropy coder (645) converts the symbols generated by the various function units into a coded video sequence by lossless compression of the symbols according to techniques such as Huffman coding, variable-length coding, and arithmetic coding.
[0066] The transmitter (640) may buffer the coded video sequence(s) generated by the entropy coder (645) to prepare for transmission over a communication channel (660), which may be a hardware / software link to a storage device for storing the encoded video data. The transmitter (640) may merge the coded video data from the video coder (603) with other data to be transmitted, e.g., coded audio data and / or auxiliary data streams (sources not shown).
[0067] The controller (650) can manage the operation of the video encoder (603). During coding, the controller (650) can assign a specific coded picture type to each coded picture, which can influence the coding techniques that can be applied to each picture. For example, pictures can often be assigned as one of the following picture types:
[0068] An Intra Picture (I Picture) may be one that can be coded and decoded without using any other picture in the sequence as a source of prediction. Some video codecs allow different types of Intra Pictures, including, for example, Independent Decoder Refresh ("IDR") pictures. A person skilled in the art recognizes such variations of I Pictures and their respective applications and features.
[0069] The predictive picture (P picture) may be coded and decoded using intra prediction or inter prediction, using at most one motion vector and reference index to predict the sample values of each block.
[0070] A bi-directionally predictive picture (Picture B) may be coded and decoded using intra-prediction or inter-prediction, using at most two motion vectors and reference indices to predict sample values of each block. Similarly, multi-prediction pictures may use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0071] Source pictures are generally spatially subdivided into multiple sample coding blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples, respectively) and can be coded on a block-by-block basis. Blocks can be predictively coded by referencing other (already coded) blocks determined by the coding assignment applied to each of the blocks' pictures. For example, blocks of pictures I can be non-predictively coded, or they can be predictively coded by referencing already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of picture P can be predictively coded through spatial prediction or temporal prediction by referencing one previously coded reference picture. Blocks of pictures B can be predictively coded through spatial prediction or temporal prediction by referencing one or two previously coded reference pictures. Source pictures or intermediately processed pictures may be subdivided into different types of blocks for different purposes. The division of coding blocks and other types of blocks may or may not follow the same method, as described in more detail below.
[0072] The video encoder (603) can perform coding operations according to a predetermined video coding technology or standard, such as ITU-T Rec. H.265. In the operation, the video encoder (603) can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancies in the input video sequence. Thus, the coded video data can comply with the syntax specified by the video coding technology or standard in use.
[0073] In some exemplary embodiments, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0074] Video can be captured as multiple source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often abbreviated as intra prediction) utilizes spatial correlations within a given picture, while inter-picture prediction utilizes temporal or other correlations between pictures. For example, a specific picture currently being encoded / decoded, referred to as the current picture, can be partitioned into blocks. A block within the current picture can be coded by a vector referred to as a motion vector when it is similar to a reference block within a previously coded and still buffered reference picture in the video. The motion vector points to the reference block within the reference picture and, if multiple reference pictures are in use, may have a third dimension identifying the reference picture.
[0075] In some exemplary embodiments, a bi-prediction technique may be used for inter-picture prediction. According to this bi-prediction technique, two reference pictures, such as a first reference picture and a second reference picture, are used, both of which advance the current picture in the decoding order in the video (but may be in the past or future, respectively, in the display order). A block in the current picture may be coded by a first motion vector pointing to a first reference block in the first reference picture, and a second motion vector pointing to a second reference block in the second reference picture. A block may be jointly predicted by a combination of the first reference block and the second reference block.
[0076] In addition, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.
[0077] According to some exemplary embodiments of the present disclosure, predictions such as inter-picture predictions and intra-picture predictions are performed on a block basis. For example, a picture within a sequence of video pictures is partitioned into coding tree units (CTUs) for compression, and the CTUs within the picture may have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU may include three parallel coding tree blocks (CTBs): one luma CTB and two chroma CTBs. Each CTU may be recursively quadtree split into one or more coding units (CUs). For example, a CTU of 64x64 pixels may be split into one CU of 64x64 pixels or four CUs of 32x32 pixels. One or more of the 32x32 blocks may be further subdivided into four CUs of 16x16 pixels each. In some exemplary embodiments, each CU may be analyzed during encoding to determine a prediction type for the CU among various prediction types, such as inter-prediction type or intra-prediction type. The CU may be subdivided into one or more prediction units (PUs) based on temporal and / or spatial predictability. Generally, each PU includes a luminance prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation during coding (encoding / decoding) is performed on a unit of prediction blocks. The subdivision of the CU into PUs (or PBs of different color channels) may be performed in various spatial patterns. The luminance or chroma PB may include a matrix of values (e.g., luminance values) for samples, such as, for example, 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 samples, etc.
[0078] FIG. 7 illustrates a video encoder (703) according to another exemplary embodiment of the present disclosure. The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values within a current video picture in a sequence of video pictures and to encode the processing block into a coded picture that is part of a coded video sequence. The exemplary video encoder (703) may be used instead of the video encoder (403) in the example of FIG. 4.
[0079] For example, the video encoder (703) receives a matrix of sample values for a processing block, such as a prediction block of 8x8 samples. Then, the video encoder (703) determines which of the intra mode, inter mode, or two prediction mode the processing block is best coded, for example, using rate-distortion optimization (RDO). When it is determined that the processing block is coded in intra mode, the video encoder (703) may encode the processing block into the coded picture using an intra prediction technique; when it is determined that the processing block is coded in inter mode or two prediction mode, the video encoder (703) may encode the processing block into the coded picture using an inter prediction technique or a two prediction technique, respectively. In some exemplary embodiments, a merge mode may be used as a submode of inter-picture prediction in which motion vectors are derived from one or more motion vector predictors without the benefit of coded motion vector components outside the predictors. In some other exemplary embodiments, there may be motion vector components applicable to the target block. Accordingly, the video encoder (703) may include components not explicitly shown in FIG. 7, such as a mode determination module, to determine the prediction mode of the processing blocks.
[0080] In the example of FIG. 7, the video encoder (703) includes an inter encoder (730), an intra encoder (722), a residue calculator (723), a switch (726), a residue encoder (724), a general controller (721), and an entropy encoder (725) combined together as shown in the exemplary arrangement of FIG. 7.
[0081] The inter-encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks within reference pictures (e.g., blocks within previous and subsequent pictures in the display order), generate inter-prediction information (e.g., description of redundant information according to the inter-encoding technique, motion vectors, merge mode information), and calculate inter-prediction results (e.g., predicted blocks) based on the inter-prediction information using any suitable technique. In some examples, the reference pictures are decoded reference pictures that are decoded based on encoded video information using a decoding unit (633) built into the exemplary encoder (620) of FIG. 6 (shown as the residual decoder (728) of FIG. 7, as described in more detail below).
[0082] The intra-encoder (722) is configured to receive samples of the current block (e.g., processing block), compare the block with already coded blocks within the same picture, generate quantized coefficients after transformation, and in some cases also generate intra-prediction information (e.g., intra-prediction direction information according to one or more intra-encoding techniques). The intra-encoder (722) can calculate intra-prediction results (e.g., prediction blocks) based on reference blocks within the same picture and intra-prediction information.
[0083] A general controller (721) may be configured to determine general control data and to control other components of the video encoder (703) based on the general control data. In one example, the general controller (721) determines the prediction mode of a block and provides a control signal to a switch (726) based on the prediction mode. For example, when the prediction mode is intra mode, the general controller (721) controls the switch (726) to select an intra mode result for use by the residual calculator (723), and controls the entropy encoder (725) to select intra prediction information and include the intra prediction information in the bitstream; when the prediction mode for a block is inter mode, the general controller (721) controls the switch (726) to select an inter prediction result for use by the residual calculator (723), and controls the entropy encoder (725) to select inter prediction information and include the inter prediction information in the bitstream.
[0084] The residual calculator (723) may be configured to calculate the difference (residual data) between the received block and the prediction results for the block selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) may be configured to encode the residual data to generate transformation coefficients. For example, the residual encoder (724) may be configured to generate transformation coefficients by converting the residual data from the spatial domain to the frequency domain. Then, quantization processing is performed on the transformation coefficients to obtain quantized transformation coefficients. In various exemplary embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transformation and generate decoded residual data. The decoded residual data may be suitably used by the intra-encoder (722) and the inter-encoder (730). For example, the inter encoder (730) can generate decoded blocks based on decoded residual data and inter prediction information, and the intra encoder (722) can generate decoded blocks based on decoded residual data and intra prediction information. The decoded blocks are appropriately processed to generate decoded pictures, and the decoded pictures are buffered in a memory circuit (not shown) and can be used as reference pictures.
[0085] The entropy encoder (725) may be configured to format the bitstream to include the encoded block and to perform entropy coding. The entropy encoder (725) is configured to include various information in the bitstream. For example, the entropy encoder (725) may be configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other suitable information within the bitstream. When coding the block in the inter mode or the merged submode of the two prediction modes, residual information may not be present.
[0086] FIG. 8 illustrates an exemplary video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive coded pictures that are part of a coded video sequence and to decode the coded pictures to generate reconstructed pictures. In one example, the video decoder (810) may be used instead of the video decoder (410) in the example of FIG. 4.
[0087] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) combined together as shown in the exemplary arrangement of FIG. 8.
[0088] The entropy decoder (871) may be configured to reconstruct specific symbols representing the syntax elements that constitute the coded picture from the coded picture. Such symbols may include, for example, the mode in which the block is coded (e.g., intra mode, inter mode, bi-predicted mode, merged submode, or other submode), prediction information (e.g., intra prediction information or inter prediction information) capable of identifying specific samples or metadata used for prediction by the intra decoder (872) or inter decoder (880), for example, residual information in the form of quantized transformation coefficients. In one example, when the prediction mode is an inter or bi-predicted mode, the inter prediction information is provided to the inter decoder (880); and when the prediction type is an intra prediction type, the intra prediction information is provided to the intra decoder (872). Inverse quantization may be performed on the residual information and is provided to the residual decoder (873).
[0089] The inter decoder (880) can be configured to receive inter prediction information and generate inter prediction results based on the inter prediction information.
[0090] The intra decoder (872) can be configured to receive intra prediction information and generate prediction results based on the intra prediction information.
[0091] The residual decoder (873) may be configured to perform inverse quantization to extract dequantized transform coefficients and process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) may also utilize specific control information (to include quantizer parameters (QP)) that may be provided by the entropy decoder (871) (this may be merely low data capacity control information so that the data path is not described).
[0092] The reconstruction module (874) may be configured to form a reconstructed block that forms part of a reconstructed picture as part of a reconstructed video by combining the residuals output by the residual decoder (873) and the prediction results (output by inter or intra prediction modules, depending on the case) in the spatial domain. Note that other suitable operations, such as deblocking operations, may also be performed to improve visual quality.
[0093] It should be noted that the video encoders (403, 603, and 703), and video decoders (410, 510, and 810) can be implemented using any suitable technique. In some exemplary embodiments, the video encoders (403, 603, and 703), and video decoders (410, 510, and 810) can be implemented using one or more integrated circuits. In other embodiments, the video encoders (403, 603, and 603), and video decoders (410, 510, and 810) can be implemented using one or more processors that execute software instructions.
[0094] Referring to block partitioning for coding and decoding, general partitioning may start from a base block and may follow a predefined set of rules, specific patterns, partition trees, or any partition structure or method. Partitioning may be hierarchical and recursive. After partitioning the base block according to exemplary partitioning procedures or any of the other procedures described below, or a combination thereof, a final set of partitions or coding blocks may be obtained. Each of these partitions may be at one of various partitioning levels within the partitioning hierarchy and may have various shapes. Each of the partitions may be referred to as a coding block (CB). For various exemplary partitioning implementations further described below, each resulting CB may be of any of the allowed sizes and partitioning levels. These partitions are referred to as coding blocks because they can form units where some basic coding / decoding decisions can be made, coding / decoding parameters can be optimized, determined, and signaled in the encoded video bitstream. The highest or deepest level in the final partitions indicates the depth of the tree's coding block partitioning structure. A coding block can be a luminance coding block or a chroma coding block. Each color's CB tree structure can be referred to as a coding block tree (CBT).
[0095] Coding blocks of all color channels can be collectively referred to as Coding Units (CUs). The hierarchical structure of all color channels can be collectively referred to as Coding Tree Units (CTUs). Partitioning patterns or structures for various color channels within a CTU may or may not be identical.
[0096] In some implementations, the partition tree schemes or structures used for the Luma and Chroma channels may not need to be identical. In other words, the Luma and Chroma channels may have distinct coding tree structures or patterns. Additionally, whether the Luma and Chroma channels use the same or different coding partition tree structures, and the actual coding partition tree structures to be used, may depend on whether the slice being coded is a P, B, or I slice. For example, in the case of an I slice, the Chroma channels and Luma channels may have distinct coding partition tree structures or modes of coding partition tree structures, whereas in the case of a P or B slice, the Luma and Chroma channels may share the same coding partition tree scheme. When distinct coding partition tree structures or modes are applied, the Luma channel may be partitioned into CBs by one coding partition tree structure, and the Chroma channel may be partitioned into Chroma CBs by another coding partition tree structure.
[0097] In some exemplary implementations, a predetermined partitioning pattern may be applied to the base block. As illustrated in FIG. 9, an exemplary 4-way partition tree may start from a first predefined level (e.g., a 64x64 block level as the base block size or other sizes), and the base block may be hierarchically partitioned down to a predefined lowest level (e.g., a 4x4 level). For example, the base block may be subject to four predefined partitioning options or patterns indicated by 902, 904, 906, and 908, and partitions designated as R are permitted for recursive partitioning in that the same partition options shown in FIG. 9 may be repeated at a lower scale down to the lowest level (e.g., a 4x4 level). In some implementations, additional restrictions may be applied to the partitioning method of FIG. 9. In the implementation of FIG. 9, rectangular partitions (e.g., 1:2 / 2:1 rectangular partitions) may be allowed, but they may not be recursive, whereas square partitions are allowed to be recursive. Partitioning according to FIG. 9 using recursion generates a final set of coding blocks if necessary. The coding tree depth may be further defined to indicate the splitting depth from the root node or root block. For example, the coding tree depth for the root node or root block, e.g., a 64x64 block, may be set to 0, and after the root block is split once more according to FIG. 9, the coding tree depth is increased by 1. The maximum or deepest level from the 64x64 base block to the minimum 4x4 partition will be 4 (starting from level 0) in the case of the above method. This partitioning method may be applied to one or more of the color channels.Each color channel can be partitioned independently according to the method of FIG. 9 (for example, a partitioning pattern or option between predefined patterns can be determined independently for each color channel at each hierarchical level). Alternatively, two or more color channels may share the same hierarchical pattern tree of FIG. 9 (for example, the same partitioning pattern or option can be selected from predefined patterns for two or more color channels at each hierarchical level).
[0098] FIG. 10 illustrates another exemplary predefined partitioning pattern that allows recursive partitioning to form a partitioning tree. As illustrated in FIG. 10, an exemplary 10-way partitioning structure or pattern may be predefined. The root block may start from a predefined level (e.g., from a base block at a 128x128 level, or at a 64x64 level). The exemplary partitioning structure of FIG. 10 includes various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. Partition types having three sub-partitions labeled 1002, 1004, 1006, and 1008 in the second row of FIG. 10 may be referred to as "T-type" partitions. The "T-type" partitions (1002, 1004, 1006, and 1008) may be referred to as the left T-type, top T-type, right T-type, and bottom T-type. In some exemplary implementations, none of the rectangular partitions of FIG. 10 are allowed to be further subdivided. The coding tree depth may be further defined to indicate the splitting depth from the root node or root block. For example, the coding tree depth for the root node or root block, e.g., a 128x128 block, may be set to 0, and after the root block is split once more according to FIG. 10, the coding tree depth is increased by 1. In some implementations, only 1010 all-square partitions may be allowed for recursive partitioning to the next level of the partitioning tree following the pattern of FIG. 10. In other words, recursive partitioning may not be allowed for square partitions within the T-type patterns (1002, 1004, 1006, and 1008). The partitioning procedure according to FIG. 10 using recursion generates a final set of coding blocks if necessary. This method can be applied to one or more of the color channels.In some implementations, more flexibility can be added to the use of partitions below the 8x8 level. For example, 2x2 chroma inter-prediction can be used in certain cases.
[0099] In some other exemplary implementations for coding block partitioning, a quadtree structure may be used to partition a base block or intermediate block into quadtree partitions. Such quadtree partitioning can be applied hierarchically and recursively to partitions of any square shape. Whether a base block, intermediate block, or partition is an additional quadtree partition may be adapted to various local characteristics of the base block or intermediate block / partition. Quadtree partitioning at picture boundaries may be further adapted. For example, an implicit quadtree split may be performed at picture boundaries so that the block maintains the quadtree split until its size fits within the picture boundary.
[0100] In some other exemplary implementations, hierarchical binary partitioning from a base block may be used. In this manner, a base block or an intermediate level block may be partitioned into two partitions. Binary partitioning can be horizontal or vertical. For example, horizontal binary partitioning may divide a base block or an intermediate block into equal right and left partitions. Similarly, vertical binary partitioning may divide a base block or an intermediate block into equal top and bottom partitions. Such binary partitioning can be hierarchical and recursive. A decision regarding whether the binary partitioning method should continue, and if so, whether horizontal binary partitioning or vertical binary partitioning should be used, can be made at the base block or the intermediate block, respectively. In some implementations, additional partitioning may stop at a predefined minimum partition size (in one or both dimensions). Alternatively, additional partitioning may stop once a predefined partitioning level or depth is reached from the base block. In some implementations, the aspect ratio of the partitions may be limited. For example, the aspect ratio of the partitions may not be less than 1:4 (or greater than 4:1). As such, a vertical strip partition with a vertical-to-horizontal aspect ratio of 4:1 can be further binary partitioned vertically into only upper and lower partitions, each with a vertical-to-horizontal aspect ratio of 2:1.
[0101] In some other examples, as illustrated in FIG. 13, a ternary partitioning scheme may be used to partition a base block or any intermediate block. The ternary pattern may be implemented vertically, as illustrated in 1302 of FIG. 13, or horizontally, as illustrated in 1304 of FIG. 13. The exemplary split ratio in FIG. 13 is shown as 1:2:1 vertically or horizontally, but other ratios may be predefined. In some implementations, two or more different ratios may be predefined. Since this triple-tree partitioning can capture objects located at the center of a block into a single contiguous partition, whereas quadtrees and binary trees always partition along the center of a block and thus divide objects into distinct partitions, this ternary partitioning scheme may be used to complement quadtree or binary partitioning structures. In some implementations, the width and height of the partitions of the exemplary tripletrees are always powers of 2 to avoid additional transformations.
[0102] The above partitioning methods can be combined in any way at different partitioning levels. As an example, the quadtree and binary partitioning methods described above can be combined to partition a base block into a quadtree-binary-tree (QTBT) structure. In this manner, the base block or intermediate block / partition can be quadtree partitioned or binary partitioned according to a set of predefined conditions, where specified. A specific example is illustrated in FIG. 14. In the example of FIG. 14, as shown in 1402, 1404, 1406, and 1408, the base block is first quadtree partitioned into four partitions. Subsequently, each of the resulting partitions is quadtree partitioned into four additional partitions (e.g., 1408), binary partitioned into two additional partitions at the next level (e.g., 1402 or 1406, horizontally or vertically, both symmetrically), or not partitioned (e.g., 1404). As illustrated by the full exemplary partition pattern in 1410 and the corresponding tree structure / representation in 1420, binary or quadtree partitioning may be recursively allowed for square-shaped partitions, where solid lines indicate quadtree partitioning and dashed lines indicate binary partitioning. Flags may be used for each binary partition node (non-leaf binary partitions) to indicate whether the binary partitioning is horizontal or vertical. For example, as illustrated in 1420, consistent with the partitioning structure of 1410, flag "0" can represent a horizontal binary partition, and flag "1" can represent a vertical binary partition. For quadtree-partition partitions, there is no need to indicate the partition type, because quadtree partitions always partition a block or partition both horizontally and vertically to create four sub-blocks / partitions of the same size.In some implementations, flag "1" may indicate horizontal binary partitioning, and flag "0" may indicate vertical binary partitioning.
[0103] In some exemplary implementations of QTBT, the quadtree and binary partitioning rule set can be represented by the following predefined parameters and their associated corresponding functions:
[0104] - CTU size: Size of the quadtree's root node (size of the base block)
[0105] - MinQTSize: Minimum allowed quadtree leaf node size
[0106] - MaxBTSize: Maximum allowed binary tree root node size
[0107] - MaxBTDepth: Maximum allowed binary tree depth
[0108] - MinBTSize: Minimum allowed binary tree leaf node size
[0109] In some exemplary implementations of the QTBT partitioning structure, the CTU size can be set to 128x128 luminance samples with two corresponding 64x64 blocks of chroma samples (when exemplary chroma subsampling is considered and used), MinQTSize can be set to 16x16, MaxBTSize can be set to 64x64, MinBTSize (for both width and height) can be set to 4x4, and MaxBTDepth can be set to 4. Quadtree partitioning can first be applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes can have a size of their minimum allowable size of 16x16 (i.e., MinQTSize) to 128x128 (i.e., CTU size). If a node is 128x128, it will not be partitioned first by a binary tree because its size exceeds MaxBTSize (i.e., 64x64). Otherwise, nodes that do not exceed MaxBTSize can be partitioned by a binary tree. In the example of Fig. 14, the base block is 128x128. The base block can only be partitioned by a quadtree according to a predefined set of rules. The base block has a partitioning depth of 0. Each of the resulting four partitions is 64x64 and does not exceed MaxBTSize, and can be further partitioned by a quadtree or binary tree at Level 1. The process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), no further partitioning can be considered. If a binary tree node has a width equal to MinBTSize (i.e., 4), no further horizontal partitioning can be considered. Similarly, if a binary tree node has a height equal to MinBTSize, no further vertical partitioning is considered.
[0110] In some exemplary implementations, the above QTBT scheme may be configured to support flexibility, allowing the Luma and Chroma to have the same QTBT structure or distinct QTBT structures. For example, in the case of P and B slices, the Luma and Chroma CTBs within a single CTU may share the same QTBT structure. However, in the case of I slices, the Luma CTBs may be partitioned into CBs by the QTBT structure, and the Chroma CTBs may be partitioned into Chroma CBs by a different QTBT structure. This means that a CU may be used to reference different color channels in the I slice; for example, the I slice may consist of coding blocks of the Luma component or coding blocks of two Chroma components, and the CU in the P or B slice may consist of coding blocks of all three color components.
[0111] In some other implementations, the QTBT method can be supplemented with the ternary method described above. These implementations may be referred to as multi-type-tree (MTT) structures. For example, in addition to binary partitioning of nodes, one of the ternary partition patterns of Fig. 13 may be selected. In some implementations, only square nodes may undergo ternary partitioning. An additional flag may be used to indicate whether the ternary partitioning is horizontal or vertical.
[0112] The design of 2-level or multi-level trees, such as QTBT implementations and QTBT implementations complemented by ternary partitioning, can be primarily motivated by complexity reduction. Theoretically, the complexity of traversing a tree is T D And, where T represents the number of partition types and D is the depth of the tree. A trade-off can be made by using multiple types (T) while reducing the depth (D).
[0113] In some implementations, the CB may be further partitioned. For example, the CB may be further partitioned into multiple prediction blocks (PBs) for the purpose of intra- or inter-frame prediction during coding and decoding processes. In other words, the CB may be further divided into different subpartitions where individual prediction decisions / configurations can be performed. At the same time, the CB may be further partitioned into multiple transformation blocks (TBs) for the purpose of describing the levels at which the transformation or inverse transformation of the video data is performed. The partitioning method of the CB into PBs and TBs may be identical or different. For example, each partitioning method may be performed using its own procedure based, for example, on various characteristics of the video data. The PB and TB partitioning methods may be independent in some exemplary implementations. The PB and TB partitioning methods and boundaries may be correlated in some other exemplary implementations. In some implementations, for example, TBs may be partitioned after PB partitions, and in particular, after the partitioning of the coding block is determined, each PB may subsequently be further partitioned into one or more TBs. For example, in some implementations, a PB may be divided into 1, 2, 4, or other number of TBs.
[0114] In some implementations, the Luma Channel and the Chroma Channel may be treated differently in order to partition the base block into coding blocks and additionally into prediction blocks and / or transformation blocks. For example, in some implementations, partitioning the coding block into prediction blocks and / or transformation blocks may be permitted for the Luma Channel, whereas partitioning these coding blocks into prediction blocks and / or transformation blocks may not be permitted for the Chroma Channel(s). Therefore, in these implementations, transformation and / or prediction of the Luma Blocks may be performed only at the coding block level. As another example, the minimum transformation block size for the Luma Channel and the Chroma Channel(s) may differ; for example, the coding blocks for the Luma Channel may be allowed to be partitioned into smaller transformation and / or prediction blocks than those for the Chroma Channels. As another example, the maximum depth for partitioning coding blocks into transformation blocks and / or prediction blocks may differ between the lumina channel and the chroma channel; for example, coding blocks for the lumina channel may be allowed to be partitioned into transformation and / or prediction blocks deeper than the chroma channel(s). In a specific example, lumina coding blocks may be partitioned into transformation blocks of multiple sizes that can be represented by recursive partitioning down to a maximum of two levels, and transformation block shapes such as square, 2:1 / 1:2, and 4:1 / 1:4, and transformation block sizes from 4x4 to 64x64 may be allowed. However, for chroma blocks, only the largest possible transformation blocks specified for the lumina blocks may be allowed.
[0115] In some exemplary implementations for partitioning coding blocks into PBs, the depth, shape, and / or other characteristics of the PB partitioning may depend on whether the PBs are intra-coded or inter-coded.
[0116] Partitioning a coding block (or prediction block) into transformation blocks can be implemented in various exemplary ways, including but not limited to quadtree partitioning and predefined pattern partitioning, recursively or non-recursively, with additional consideration for transformation blocks at the boundaries of the coding block or prediction block. Generally, the resulting transformation blocks may be at different partition levels, may not have the same size, and may not need to be square in shape (e.g., they may be rectangular with some allowed sizes and aspect ratios). Additional examples are described in more detail below in relation to FIGS. 15, 16, and 17.
[0117] However, in some other implementations, CBs obtained through any of the above partitioning schemes can be used as basic or minimal coding blocks for prediction and / or transformation. In other words, no additional partitioning is performed for inter-prediction / intra-prediction purposes and / or transformation purposes. For example, CBs obtained from the above QTBT scheme can be used directly as units for performing predictions. Specifically, this QTBT structure eliminates the concept of multiple partition types, namely the separation of CU, PU, and TU, and supports greater flexibility regarding CU / CB partition shapes as previously mentioned. In this QTBT block structure, CU / CB can have a square or rectangular shape. The leaf nodes of this QTBT are used as units for prediction and transformation processing without any additional partitioning. This means that in this exemplary QTBT coding block structure, CU, PU, and TU have the same block size.
[0118] The various CB partitioning schemes above and additional partitioning of CBs into PBs and / or TBs (excluding PB / TB partitioning) may be combined in any way. The following specific implementations are provided as non-limiting examples.
[0119] Specific exemplary implementations of coding block and transform block partitioning are described below. In these exemplary implementations, the base block may be partitioned into coding blocks using recursive quadtree partitioning or the aforementioned predefined partitioning patterns (such as those in FIGS. 9 and 10). At each level, whether further quadtree partitioning of a specific partition is required may be determined by local video data characteristics. The resulting CBs may be of various quadtree partitioning levels and various sizes. The decision on whether to code a picture region using inter-picture (temporal) or intra-picture (spatial) prediction may be made at the CB level (or CU level for all 3-color channels). Each CB may be further partitioned into 1, 2, 4, or other numbers of PBs according to a predefined PB partitioning type. Within a single PB, the same prediction process may be applied, and relevant information may be transmitted to the decoder in PB units. After obtaining residual blocks by applying a prediction process based on the PB partitioning type, the CBs can be partitioned into TBs according to another quadtree structure similar to the coding tree for the CBs. In this particular implementation, the CBs or TBs may be restricted to square shapes, but are not required to be. Also, in this particular example, the PBs can be square or rectangular shapes for inter-predictions, and only square for intra-predictions. The coding blocks can be partitioned, for example, into four square-shaped TBs. Each TB can be recursively further partitioned (using quadtree partitioning) into smaller TBs referred to as Residual Quadtrees (RQTs).
[0120] Other exemplary implementations for partitioning a base block into CBs, PBs, and / or TBs are further described below. For example, rather than using multi-partition unit types such as those shown in FIG. 9 or FIG. 10, a quadtree having a nested multi-type tree using binary and ternary splits segmentation structures (e.g., a QTBT as described above or a QTBT with ternary splits) may be used. The separation of CBs, PBs, and TBs (i.e., partitioning of CBs into PBs and / or TBs, and partitioning of PBs into TBs) may be abandoned except when CBs that are too large for the maximum transformation length are required, where these CBs may require additional partitioning. This exemplary partitioning method can be designed to support greater flexibility regarding CB partition shapes so that both prediction and transformation can be performed on the CB level without additional partitioning. In this coding tree structure, the CB can have a square or rectangular shape. Specifically, the coding tree block (CTB) can first be partitioned by a quadtree structure. Then, the quadtree leaf nodes can be further partitioned by a nested multi-type tree structure. An example of a nested multi-type tree structure using binary or ternary partitioning is shown in FIG. 11. Specifically, the exemplary multi-type tree structure of FIG. 11 includes four split types referred to as vertical binary split (SPLIT_BT_VER) (1102), horizontal binary split (SPLIT_BT_HOR) (1104), vertical ternary split (SPLIT_TT_VER) (1106), and horizontal ternary split (SPLIT_TT_HOR) (1108). The CBs then correspond to the leaves of the multi-type tree.In this exemplary implementation, as long as the CB is not too large for the maximum transformation length, this segmentation is used for both prediction and transformation processing without any additional partitioning. This means that, in most cases, the CB, PB, and TB have the same block size in a quadtree having a nested multi-type tree coding block structure. An exception occurs when the maximum supported transformation length is smaller than the width or height of the color component of the CB. In some implementations, in addition to binary or ternary partitioning, the nested patterns of FIG. 11 may additionally include quadtree partitioning.
[0121] A specific example of a quadtree having a nested multi-type tree coding block structure of block partitions (including quadtree, binary, and ternary partition options) for a single base block is illustrated in FIG. 12. More specifically, FIG. 12 illustrates a base block (1200) being quadtree-partitioned into four square partitions (1202, 1204, 1206, and 1208). A decision to further use the multi-type tree structure and quadtree of FIG. 11 for additional partitioning is made for each of the quadtree-partitioned partitions. In the example of FIG. 12, partition (1204) is not further partitioned. Each of the partitions (1202 and 1208) adopts a different quadtree partitioning. In the case of partition (1202), the second-level quadtree-partition top-left, top-right, bottom-left, and bottom-right partitions each adopt a third-level partition of quadtree, horizontal binary partition (1104) of FIG. 11, non-partition, and horizontal ternary partition (1108) of FIG. 11. Partition (1208) adopts a different quadtree partition, and the second-level quadtree-partition top-left, top-right, bottom-left, and bottom-right partitions each adopt a third-level partition of vertical ternary partition (1106) of FIG. 11, non-partition, non-partition, and horizontal binary partition (1104) of FIG. 11. Two of the sub-partitions of the third-level top-left partition of 1208 are each further partitioned according to the horizontal binary partition (1104) and horizontal ternary partition (1108) of FIG. 11. Partition (1206) adopts a second level partitioning pattern according to the vertical binary partitioning (1102) of FIG. 11 into two partitions, and the two partitions are further divided into a third level according to the horizontal binary partitioning (1108) and vertical binary partitioning (1102) of FIG. 11. A fourth level partitioning is further applied to one of these according to the horizontal binary partitioning (1104) of FIG. 11.
[0122] For the specific example above, the maximum luma transform size may be 64x64 and the maximum supported chroma transform size may differ from the luma at, for example, 32x32. Even if the exemplary CBs above in FIG. 12 are not generally further subdivided into smaller PBs and / or TBs, when the width or height of a luma coding block or chroma coding block is greater than the maximum transform width or height, the luma coding block or chroma coding block may be automatically subdivided in the horizontal and / or vertical directions to meet the transform size limit in that direction.
[0123] In a specific example for partitioning the above base block into CBs, and as described above, the coding tree method can support the ability for Luma and Chroma to have separate block tree structures. For example, in the case of P and B slices, Luma and Chroma CTBs within a single CTU may share the same coding tree structure. For example, in the case of I slices, Luma and Chroma may have separate coding block tree structures. When separate block tree structures are applied, Luma CTBs may be partitioned into Luma CBs by a single coding tree structure, and Chroma CTBs are partitioned into Chroma CBs by a different coding tree structure. This means that the CU in Slice I may consist of coding blocks of the Luma component or coding blocks of two Chroma components, and that the CU in Slice P or B is always composed of coding blocks of all three color components, unless the video is monochrome.
[0124] When a coding block is further partitioned into multiple transformation blocks, the transformation blocks within it may be aligned into bitstreams according to various orders or scanning methods. Exemplary implementations for partitioning a coding block or a prediction block into transformation blocks, and the coding order of the transformation blocks, are described in more detail below. In some exemplary implementations, as described above, transformation partitioning may support transformation blocks of multiple shapes, e.g., 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1, and the transformation block sizes range, e.g., from 4x4 to 64x64. In some implementations, if the coding block is less than or equal to 64x64, transformation block partitioning may be applied only to the luminance component, so that for the chroma blocks, the transformation block size is the same as the coding block size. Otherwise, if the coding block width or height is greater than 64, both the luma and chroma coding blocks can be implicitly divided into multiples of min(W, 64) x min(H, 64) and min(W, 32) x min(H, 32) transformation blocks, respectively.
[0125] In some exemplary implementations of transformation block partitioning, for both intra and intercoded blocks, the coding block may be further partitioned into multiple transformation blocks having a partitioning depth of up to a predefined number of levels (e.g., 2 levels). Transformation block partitioning depths and sizes may be related. For some exemplary implementations, an exemplary mapping from the transformation size of the current depth to the transformation size of the next depth is shown in Table 1 as follows.
[0126] Table 1: Convert Partition Size Settings
[0127]
[0128] Based on the exemplary mapping in Table 1, for a 1:1 square block, the next level transformation partition can generate four 1:1 square sub-transform blocks. The transformation partition can stop at, for example, 4x4. In this way, the transformation size for the current depth of 4x4 corresponds to the same size of 4x4 for the next depth. In the example in Table 1, for a 1:2 / 2:1 non-square block, the next level transformation partition can generate two 1:1 square sub-transform blocks, and for a 1:4 / 4:1 non-square block, the next level transformation partition can generate two 1:2 / 2:1 sub-transform blocks.
[0129] In some exemplary implementations, additional restrictions may be applied to the luma component of an intra-coded block regarding transformation block partitioning. For example, for each level of transformation partitioning, all sub-transformation blocks may be restricted to having the same size. For example, in the case of a 32x16 coding block, Level 1 transformation partitioning creates two 16x16 sub-transformation blocks, and Level 2 transformation partitioning creates eight 8x8 sub-transformation blocks. In other words, Level 2 partitioning must be applied to all Level 1 sub-blocks to maintain the transformation units of the same size. An example of transformation block partitioning for an intra-coded square block according to Table 1 is illustrated in FIG. 15 with the coding order indicated by arrows. Specifically, 1502 illustrates a square coding block. Level 1 partitioning into four transformation blocks of the same size according to Table 1 is illustrated in 1504 with the coding order indicated by arrows. The second level division of all first-level blocks of the same size into 16 conversion blocks of the same size according to Table 1 is shown in 1506 with the coding order indicated by the arrow.
[0130] In some exemplary implementations, the above restrictions on intra-coding may not apply to the luma components of the intercoded block. For example, after the first level of transformation division, any of the sub-transformation blocks may be independently divided one more level. Thus, the resulting transformation blocks may or may not be of the same size. An exemplary division of the intercoded block into transformation blocks, along with their coding order, is illustrated in FIG. 16. In the example of FIG. 16, the intercoded block (1602) is divided into transformation blocks at two levels according to Table 1. At the first level, the intercoded block is divided into four transformation blocks of the same size. Then, only one of the four transformation blocks (but not all of them) is further divided into four sub-transformation blocks, so that, as illustrated in 1604, a total of seven transformation blocks have two different sizes. An exemplary coding sequence of these seven transformation blocks is illustrated by an arrow at 1604 in Fig. 16.
[0131] In some exemplary implementations, some additional restrictions on transformation blocks may apply to the chroma component(s). For example, for the chroma component(s), the transformation block size may be as large as the coding block size, but may not be smaller than a predefined size, e.g., 8x8.
[0132] In some other exemplary implementations, for a coding block where the width (W) or height (H) is greater than 64, both the luma and chroma coding blocks may be implicitly divided into multiples of the min(W, 64) x min(H, 64) and min(W, 32) x min(H, 32) transformation units, respectively. Here, in the present disclosure, "min(a, b)" may return the smaller value between a and b.
[0133] FIG. 17 further illustrates another alternative exemplary method for partitioning a coding block or a prediction block into transformation blocks. As illustrated in FIG. 17, instead of using recursive transformation partitioning, a predefined set of partitioning types may be applied to the coding block according to the transformation type of the coding block. In the specific example illustrated in FIG. 17, one of six exemplary partitioning types may be applied to divide the coding block into a varying number of transformation blocks. This method of generating transformation block partitioning may be applied to a coding block or a prediction block.
[0134] More specifically, the partitioning method of FIG. 17 provides up to six exemplary partition types for any given transformation type (transformation type refers to a type of first-order transformation, such as ADST, etc.). In this method, a transformation partition type may be assigned to each coding block or prediction block, for example, based on a rate-distortion cost. In one example, the transformation partition type assigned to a coding block or prediction block may be determined based on the transformation type of the coding block or prediction block. A specific transformation partition type may correspond to the transformation block partition size and pattern, as shown by the six transformation partition types exemplified in FIG. 17. Correspondence relationships between various transformation types and various transformation partition types may be predefined. An example with uppercase labels indicating transformation partition types that may be assigned to a coding block or prediction block based on a rate-distortion cost is shown below:
[0135] ● PARTITION_NONE: Allocates a transformation size equal to the block size.
[0136] ● PARTITION_SPLIT: Allocates a transformation size that is half the width of the block and half the height of the block.
[0137] ● PARTITION_HORZ: Assigns a transformation size equal to the block size width and half the block size height.
[0138] ● PARTITION_VERT: Assigns a transformation size that is half the width of the block size and the same height as the block size.
[0139] ● PARTITION_HORZ4: Assigns a transformation size equal to the block size width and 1 / 4 of the block size height.
[0140] ● PARTITION_VERT4: Assigns a transformation size that is 1 / 4 of the block size's width and has a height equal to the block size.
[0141] In the above example, the transformation partition types as illustrated in FIG. 17 all include uniform transformation sizes for the partitioned transformation blocks. This is merely an example rather than a limitation. In some other implementations, mixed transformation block sizes may be used for partitioned transformation blocks of a specific partition type (or pattern).
[0142] PBs obtained from any of the above partitioning methods (or CBs, also called PBs when not further partitioned into prediction blocks) can then become individual blocks for coding through intra- or inter-predictions. For inter-prediction of the current PB, the residual between the current block and the prediction block is generated, coded, and can be included in the coded bitstream.
[0143] Inter-prediction can be implemented, for example, in single-reference mode or compound-reference mode. In some implementations, a skip flag may be included first in the bitstream for the current block (or at a higher level) to indicate whether the current block is inter-coded and will not be skipped. If the current block is inter-coded, another flag may be additionally included in the bitstream as a signal to indicate whether single-reference mode or compound-reference mode is used for the current block. In the case of single-reference mode, one reference block may be used to generate the prediction block for the current block. In the case of compound-reference mode, two or more reference blocks may be used to generate the prediction block, for example, by weighted average. Compound-reference mode may be referred to as more-than-one-reference mode, two-reference mode, or multiple-reference mode. Reference blocks or reference blocks may be identified using reference frame indices or indices and additionally using corresponding motion vectors or motion vectors that indicate the position, e.g., the shift(s) between the reference block(s) and the current blocks in horizontal and vertical pixels. For example, an inter-prediction block for the current block may be generated from a single-reference block identified by a single motion vector in the reference frame as a prediction block in single-reference mode, whereas for composite-reference mode, the prediction block may be generated by the weighted average of two reference blocks in two reference frames indicated by two motion vectors. Motion vector(s) may be coded in various ways and included in the bitstream.
[0144] In some implementations, the encoding or decoding system may maintain a decoded picture buffer (DPB). Some images / pictures may be maintained in the DPB awaiting display (in the decoding system), and some images / pictures within the DPB may be used as reference frames to enable inter-prediction. In some implementations, reference frames within the DPB may be tagged as short-term or long-term references for the current image being encoded or decoded. For example, short-term reference frames may include frames used for inter-prediction of blocks within the current frame or within a predefined number (e.g., 2) of video frames closest to the current frame in the decoding order. Long-term reference frames may include frames within the DPB that may be used to predict image blocks within frames located further away from the current frame than a predefined number of frames in the decoding order. Information regarding these tags for short and long reference frames may be referred to as a Reference Picture Set (RPS) and may be added to the header of each frame within the encoded bitstream. Each frame within the encoded video stream may be numbered in an absolute manner according to the playback sequence, or identified by a Picture Order Counter (POC) associated with a picture group, for example, starting from the I-frame.
[0145] In some exemplary implementations, one or more reference picture lists, including the identification of short-term and long-term reference frames for inter-prediction, may be formed based on information within the RPS. For example, a single picture reference list may be formed for unidirectional inter-prediction, denoted as the L0 reference (or reference list 0), whereas two picture reference lists may be formed for bidirectional inter-prediction, denoted as L0 (or reference list 0) and L1 (or reference list 1) for each of the two prediction directions. The reference frames included in the L0 and L1 lists may be ordered in various predetermined ways. The lengths of the L0 and L1 lists may be signaled in the video bitstream. Unidirectional inter-prediction may be in a single-reference mode, or in a composite-reference mode when multiple references for generating a prediction block by weighted average in a composite prediction mode are on the same side of the block to be predicted. Bidirectional inter-prediction can only be a composite mode in that bidirectional inter-prediction involves at least two reference blocks.
[0146] In some implementations, a merge mode (MM) for inter-prediction may be implemented. Generally, in the case of a merge mode, one or more of the motion vectors in a single-reference prediction or a complex-reference prediction for the current PB may be derived from other motion vector(s) rather than being calculated and signaled independently. For example, in an encoding system, the current motion vector(s) for the current PB may be reduced to the difference(s) between the current motion vector(s) and one or more other already encoded motion vectors (referred to as reference motion vectors). These difference(s) in the motion vector(s), rather than the entire current motion vector(s), may be encoded and included in the bitstream and linked to the reference motion vector(s). Correspondingly, in a decoding system, the motion vector(s) corresponding to the current PB may be derived based on the decoded motion vector difference and the decoded reference motion vector(s) linked thereto. As a specific form of general Merge Mode (MM) inter-prediction, such inter-prediction based on motion vector differences can be referred to as MMVD (Merge Mode with Motion Vector Difference). Therefore, MM in general, or specifically MMVD, can be implemented to leverage correlations between motion vectors associated with different PBs to improve coding efficiency. For example, neighboring PBs may have similar motion vectors. As another example, motion vectors may be temporally correlated (between frames) for blocks that are similarly located / positioned in space.
[0147] In some exemplary implementations, an MM flag may be included in the bitstream during the encoding process to indicate whether the current PB is in merge mode. Additionally, or alternatively, an MMVD flag may be included in the bitstream during the encoding process to signal whether the current PB is in MMVD mode. The MM and / or MMVD flags or indicators may be provided at the PB level, CB level, CU level, CTB level, CTU level, slice level, picture level, etc. As a specific example, both the MM flag and the MMVD flag may be included for the current CU, and the MMVD flag may be signaled immediately after the skip flag and the MM flag to specify whether MMVD mode is being used for the current CU.
[0148] In some exemplary implementations of MMVD, a list of merge candidates for motion vector prediction for a predicted block may be formed. The list of merge candidates may include a predetermined number (e.g., 2) of MV predictor candidate blocks from which motion vectors can be used to predict the current motion vector. The MVD candidate blocks may include blocks selected from neighboring blocks and / or temporal blocks within the same frame (e.g., blocks located identically within the preceding or subsequent frames of the current frame). These options represent blocks at spatial or temporal locations relative to the current block that are likely to have motion vectors similar or identical to the current block. The size of the list of MV predictor candidates may be predetermined. For example, the list may include 2 candidates. To be in the list of merge candidates, for example, a candidate block may be required to have the same reference frame (or frames) as the current block, must exist (e.g., when the current block is near the edge of a frame, a boundary check needs to be performed), and must have already been encoded during the encoding process and / or already decoded during the decoding process. In some implementations, the list of merge candidates may first be filled with spatially adjacent blocks (scanned in a specific predefined order) if available and satisfying the above conditions, and then with temporal blocks if space is still available in the list. For example, neighbor candidate blocks may be selected from the blocks to the left and above the current block. The list of merge MV predictor candidates may be signaled in the bitstream.
[0149] In some implementations, an actual merge candidate used as a reference motion vector to predict the motion vector of the current block may be signaled. If the merge candidate list contains two candidates, a 1-bit flag, referred to as the merge candidate flag, may be used to indicate the selection of the reference merge candidate. For the current block predicted in composite mode, each of the multiple motion vectors predicted using the MV predictor may be associated with a reference motion vector from the merge candidate list.
[0150] In some exemplary implementations of MMVD, after a merge candidate is selected and used as a base motion vector predictor for the motion vector to be predicted, a motion vector difference (MVD or delta MV, representing the difference between the motion vector to be predicted and the reference candidate motion vector) can be calculated in the encoding system. This MVD may include information indicating the magnitude of the MV difference and the direction of the MV difference, which can be signaled in the bitstream. The magnitude of the motion difference and the direction of the motion difference can be signaled in various ways.
[0151] In some exemplary implementations of MMVD, a distance index may be used to specify the magnitude information of the motion vector difference and to indicate one of a set of predefined offsets representing a predefined motion vector difference from a starting point (reference motion vector). Then, an MV offset according to the signaled index may be added to the horizontal or vertical component of the starting (reference) motion vector. The horizontal or vertical components of the reference motion vector to be offset are determined by the exemplary orientation information of the MVD. An exemplary predefined relationship between the distance index and the predefined offsets is specified in Table 2.
[0152] Table 2 - Exemplary Relationship between Distance Index and Predefined MV Offset
[0153]
[0154] In some exemplary implementations of MMVD, a direction index may be additionally signaled and used to indicate the direction of the MVD relative to the reference motion vector. In some implementations, the direction may be restricted to either horizontal or vertical directions. An exemplary 2-bit direction index is shown in Table 3. In the examples in Table 3, the interpretation of the MVD may vary depending on the information of the start / reference MVs. For example, when the start / reference MV corresponds to a single-predicted block or when both reference frame lists correspond to a two-predicted block pointing to the same side of the current picture (i.e., when both POCs of the two reference pictures are greater than or both are smaller than the POC of the current picture), the sign in Table 3 may specify the sign (direction) of the MV offset added to the start / reference MV. When the start / reference MV corresponds to a two-predicted block having two reference pictures on different sides of the current picture (i.e., the POC of one reference picture is greater than the POC of the current picture, and the POC of the other reference picture is smaller than the POC of the current picture), and the difference between the reference POC in picture reference list 0 and the current frame is greater than the difference between the reference POC in picture reference list 1 and the current frame, the sign in Table 3 may specify the sign of the MV offset added to the reference MV corresponding to the reference picture in picture reference list 0, and the sign for the offset of the MV corresponding to the reference picture in picture reference list 1 may have the opposite value (opposite sign for the offset). Otherwise, if the difference between the reference POC in picture reference list 1 and the current frame is greater than the difference between the reference POC in picture reference list 0 and the current frame, the sign in Table 3 may specify the sign of the MV offset added to the reference MV associated with picture reference list 1, and the sign for the offset of the reference MV associated with picture reference list 0 has the opposite value.
[0155] Table 3 - Exemplary implementations for the sign of the MV offset specified by the direction index
[0156]
[0157] In some exemplary implementations, the MVD can be scaled according to the difference in POCs in each direction. If the differences in POCs within the two lists are the same, no scaling is required. Otherwise, if the difference in POCs within reference list 0 is greater than that in reference list 1, the MVD for reference list 1 is scaled. If the difference in POCs in reference list 1 is greater than that in list 0, the MVD for list 0 can be scaled in the same way. If the starting MV is predicted to be single, the MVD is added to the available or reference MV.
[0158] In some exemplary implementations of MVD coding and signaling for bidirectional composite prediction, symmetric MVD coding may be implemented in addition to or as an alternative to coding and signaling two MVDs individually, such that only one MVD requires signaling and the other MVD can be derived from the signaled MVD. In these implementations, motion information containing reference picture indices of both List-0 and List-1 is signaled. However, for example, only the MVD associated with reference List-0 is signaled, and the MVD associated with reference List-1 is not signaled but is derived. Specifically, at the slice level, a flag referred to as "mvd_l1_zero_flag" may be included in the bitstream to indicate whether reference List-1 is not signaled in the bitstream. If this flag is 1, indicating that reference list-1 is equal to 0 (and therefore not signaled), the bidirectional-prediction flag, referred to as "BiDirPredFlag," may be set to 0, meaning that bidirectional prediction does not exist. Otherwise, if mvd_l1_zero_flag is 0, BiDirPredFlag may be set to 1 if the nearest reference picture in list-0 and the nearest reference picture in list-1 form a forward-reverse pair of reference pictures or a reverse-forward pair of reference pictures, and both the list-0 and list-1 reference pictures are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. A BiDirPredFlag of 1 indicates that the symmetric mode flag is additionally signaled in the bitstream. The decoder can extract the symmetric mode flag from the bitstream when BiDirPredFlag is 1. The symmetric mode flag can be signaled, for example, at the CU level (if necessary) and indicates whether the symmetric MVD coding mode is being used for the corresponding CU.When the symmetric mode flag is 1, this indicates the use of the symmetric MVD coding mode and indicates that only the reference picture indices of both List-0 and List-1 (referred to as "mvp_l0_flag" and "mvp_l1_flag") are signaled along with the MVD associated with List-0 (referred to as "MVD0"), and that the other motion vector difference, "MVD1," should be derived rather than signaled. For example, MVD1 can be derived as -MVD0. Thus, only one MVD is signaled in the exemplary symmetric MVD mode. In some other exemplary implementations for MV prediction, a harmonized approach may be used to implement general merge mode, MMVD, and some other types of MV prediction for both single-reference and multi-reference mode MV prediction. Various syntax elements may be used to signal how the MV for the current block is predicted.
[0159] For example, in the case of single-reference mode, the following MV prediction modes can be signaled:
[0160] NEARMV - Uses one of the motion vector predictors (MVP) in the list indicated by the DRL (Dynamic Reference List) index directly without any MVD.
[0161] NEWMV - Use one of the motion vector predictors (MVP) in the list signaled by the DRL index as a reference and apply a delta to the MVP (e.g., use MVD).
[0162] GLOBALMV - Uses motion vectors based on frame-level global motion parameters.
[0163] Likewise, in the case of a composite-reference inter-prediction mode using two reference frames corresponding to two predicted MVs, the following MV prediction modes may be signaled:
[0164] NEAR_NEARMV - For each of the two predicted MVs, use one of the motion vector predictors (MVP) from the list signaled by the DRL index without MVD.
[0165] NEAR_NEWMV - To predict the first of two motion vectors, use one of the motion vector predictors (MVP) in the list signaled by the DRL index as a reference MV without an MVD; to predict the second of two motion vectors, use one of the motion vector predictors (MVP) in the list signaled by the DRL index as a reference MV with an additionally signaled delta MV (MVD).
[0166] NEW_NEARMV - To predict the second motion vector of two motion vectors, use one of the motion vector predictors (MVP) in the list signaled by the DRL index as a reference MV without an MVD; to predict the first motion vector of two motion vectors, use one of the motion vector predictors (MVP) in the list signaled by the DRL index as a reference MV with an additionally signaled delta MV (MVD).
[0167] NEW_NEWMV - Use one of the motion vector predictors (MVP) in the list signaled by the DRL index as a reference MV, and use it along with an additionally signaled delta MV to predict for each of the two MVs.
[0168] GLOBAL_GLOBALMV - Uses MVs from each reference based on their frame-level global motion parameters.
[0169] Therefore, the term "NEAR" above refers to MV prediction using a reference MV without an MVD as a general merge mode, whereas the term "NEW" refers to MV prediction involving the use of a reference MV and offsetting it to a signaled MVD, as in MMVD mode. In the case of composite inter-prediction, both the motion vector deltas above and the reference base motion vectors can generally be different or independent between the two references, but they can be correlated, and such correlation can be utilized to reduce the amount of information required to signal the two motion vector deltas. In these situations, joint signaling of the two MVDs can be implemented and displayed in the bitstream.
[0170] The above dynamic reference list (DRL) can be used to hold a set of indexed motion vectors that are dynamically maintained and considered as candidate motion vector predictors.
[0171] In some exemplary implementations, an optical flow-based approach can be used to refine motion vectors (MVs) at a subblock-wise manner for complex prediction. Specifically, optical flow equations can be applied to formulate a least-squares problem from which fine motions can be derived from the gradients of complex inter-prediction samples. Using such fine motions, MVs per subblock can be refined within the prediction block, which can improve inter-prediction quality. Some coding features may be an extension of the concept of bi-directional optical flow (BDOF) because it supports MV refinement when two reference blocks have arbitrary temporal distances from the current block.
[0172] In some implementations, the four additional inter compound modes listed below may be added: NEAR_NEARMV_OPTFLOW, NEAR_NEWMV_OPTFLOW, NEW_NEARMV_OPTFLOW, and / or NEW_NEWMV_OPTFLOW.
[0173] These modes may be referred to as optical flow modes, and reference MV types may be defined as in conventional composite modes (e.g., NEAR_NEWMV_OPTFLOW has the same reference MV types as in NEAR_NEWMV). Composite prediction may be performed based on subblock-wise refined MVs instead of the original MVs.
[0174] The various embodiments and / or implementations described in this disclosure may be used individually or combined in any order. Additionally, some, all, or any partial or whole combination of these embodiments and / or implementations may be embodied as part of an encoder and / or decoder and may be implemented in hardware and / or software. For example, they may be hard-coded in a dedicated processing circuit (e.g., one or more integrated circuits). In one other example, they may be implemented by one or more processors that execute a program stored on a non-transient computer-readable medium.
[0175] There may be some issues / problems related to some implementations of signaling methods for motion vector differences, for example, how delta MV(s) are signaled in NEW_NEARMV mode, NEAR_NEWMV mode, or NEW_NEWMV mode. One of the issues / problems is that the correlation of motion vector differences within two reference lists is not utilized, which may reduce the efficiency and performance of coding / decoding.
[0176] The present disclosure describes various embodiments for signaling motion vector differences (MVD or delta MV) for inter-predicted mode coding and / or decoding, which solve at least one of the issues / problems discussed above and achieve an efficient software / hardware implementation for improved inter-predicted mode coding / decoding.
[0177] In various embodiments, with reference to FIG. 18, a method (1800) for decoding an inter-predicted video block is illustrated. The method (1800) may include some or all of the following steps: receiving a coded video bitstream by a device comprising a memory storing instructions and a processor communicating with the memory (1810); extracting an inter-predicted mode and a joint delta motion vector (MV) for a current block within a current frame from the coded video bitstream by the device (1820); extracting a flag from the coded video bitstream indicating whether a first delta MV for a first reference frame and a second delta MV for a second reference frame are jointly signaled (1830); and in response to the flag indicating that the first delta MV and the second delta MV are jointly signaled, deriving the first delta MV and the second delta MV based on the joint delta MV by the device (1840). and / or by the device, a step (1850) of decoding the current block within the current frame based on the first delta MV and the second delta MV.
[0178] In some implementations, the joint delta MV may be denoted as joint_delta_mv, which may be an element indicating the joint delta MV. In some implementations, a flag indicating whether the first delta MV for the first reference frame and the second delta MV for the second reference frame are jointly signaled may be denoted as joint_mvd_flag. In some implementations, the first reference frame may be a frame in reference list (reference list 0) and / or; and the second reference frame may be a frame in another reference list (reference list 1).
[0179] In some implementations, step 1830 may include the step of extracting a flag (joint_mvd_flag) from a video bitstream coded by the device, indicating whether a first delta MV for a first reference frame in reference list 0 and a second delta MV for a second reference frame in reference list 1 are jointly signaled.
[0180] In various embodiments of the present disclosure, the size of a block (e.g., a coding block, a prediction block, or a transformation block (but not limited thereto)) may refer to the width or height of the block. The width or height of the block may be an integer in pixels. In various embodiments of the present disclosure, the size of a block may refer to the area size of the block. The area size of the block may be an integer calculated by multiplying the width of the block by the height of the block in pixels. In some various embodiments of the present disclosure, the size of a block may refer to a maximum value of the width or height of the block, a minimum value of the width or height of the block, or an aspect ratio of the block. The aspect ratio of the block may be calculated by dividing the width of the block by the height, or by dividing the height of the block by the width.
[0181] Herein, in some embodiments of the present disclosure, the “first” reference frame refers not only to a “one” reference frame but also to the “first” reference frame among a plurality of reference frames (e.g., having the smallest index or appearing earliest in the sequence), and the “second” reference frame refers not only to another” reference frame but also to the “second” reference frame among a plurality of reference frames (e.g., having the second smallest index or appearing second earliest in the sequence).
[0182] Herein, in various embodiments of the present disclosure, “XYZ is signaled” may refer to XYZ being encoded into a bitstream coded during an encoding process and / or; after the coded bitstream is transmitted from one device to another, “XYZ is signaled” may refer to XYZ being decoded / extracted from the coded bitstream during a decoding process.
[0183] Here, in various embodiments of the present disclosure, the orientation of a reference frame may be determined by whether the reference frame is before the current frame in the display order or after the current frame in the display order. In some implementations of a composite reference mode, if the picture order counts (POCs) of two reference frames for a pair of motion vectors are greater than or less than the POC of the current frame, the orientations of the two reference frames are the same. Otherwise, if the POC of one reference frame is greater than the POC of the current frame while the POC of the other reference frame is less than the POC of the current frame, the orientations of the two reference frames are different.
[0184] Herein, in various embodiments of the present disclosure, "block" may refer to a prediction block, a coding block, a transformation block, or a coding unit (CU).
[0185] Referring to step 1810, the device may be the electronic device (530) of FIG. 5 or the video decoder (810) of FIG. 8. In some implementations, the device may be the decoder (633) within the encoder (620) of FIG. 6. In other implementations, the device may be part of the electronic device (530) of FIG. 5, part of the video decoder (810) of FIG. 8, or part of the decoder (633) within the encoder (620) of FIG. 6. The coded video bitstream may be the coded video sequence of FIG. 8, or the intermediate coded data of FIG. 6 or FIG. 7.
[0186] Referring to step 1820, the device can extract an inter-predict mode and a joint delta motion vector (MV) for the current block within the current frame from the coded video bitstream. The current block may be in a composite reference mode. The inter-predict mode may include one of NEAR_NEAR mode, NEW_NEARMV mode, NEAR_NEWMV mode, or NEW_NEWMV mode. The joint delta MV may be referred to as the MV difference (MVD).
[0187] Referring to step 1830, the device can extract a flag from the coded video bitstream indicating whether the first delta MV for the first reference frame and the second delta MV for the second reference frame are jointly signaled.
[0188] In some implementations for some inter prediction mode(s), the flag is encoded in the coded video bitstream, and the device can extract the flag from the coded video bitstream.
[0189] In some implementations for a specific inter-prediction mode(s), the flag is not encoded in the coded video bitstream, and the device may derive the flag according to a specific inter-prediction model(s) based on a default value. For example, the inter-prediction mode of the current block is NEAR_NEARMV; step 1830 may include the step of determining the flag as a default value. In some implementations, the default value is 0, indicating that the first delta MV for the first reference frame and the second delta MV for the second reference frame are not jointly signaled.
[0190] In some implementations of the composite reference mode, a flag named joint_mvd_flag may be transmitted to the device to indicate whether delta MVs for the first reference list (reference list 0) and the second reference list (reference list 1) are jointly signaled.
[0191] In some other implementations, in response to the value of a flag (joint_mvd_flag) indicating that the delta MVs for reference list 0 and reference list 1 are jointly signaled, only one joint delta MV, which may be named joint_delta_mv, is signaled and transmitted to the decoder, and the delta MVs for reference list 0 and reference list 1 can be derived from the joint delta MV (joint_delta_mv). In response to the value of a flag (joint_mvd_flag) indicating that the delta MVs for reference list 0 and reference list 1 are not jointly signaled, zero, one, or two delta MVs may be signaled individually for reference list 0 and / or reference list 1 based on inter-prediction modes.
[0192] In some implementations, a flag value of 0 may indicate that delta MVs for reference list 0 and reference list 1 are jointly signaled, and only one joint delta MV is signaled and transmitted; a flag value of 1 may indicate that delta MVs for reference list 0 and reference list 1 are not jointly signaled, and zero (or one or two) joint delta MVs may be signaled and transmitted. Conversely, in some other implementations, a flag value of 1 may indicate that delta MVs for reference list 0 and reference list 1 are jointly signaled, and only one joint delta MV is signaled and transmitted; a flag value of 0 may indicate that delta MVs for reference list 0 and reference list 1 are not jointly signaled, and zero (or one or two) joint delta MVs may be signaled and transmitted.
[0193] Referring to step 1840, in response to a flag indicating that the first delta MV and the second delta MV are jointly signaled, the device can derive the first delta MV and the second delta MV based on the joint delta MV.
[0194] In some implementations, when the inter-prediction mode of the current block is NEW_NEWMV and a flag (e.g., joint_mvd_flag) indicates that delta MVs for reference list 0 and reference list 1 are jointly signaled, delta MVs for reference list 0 and / or reference list 1 can be derived from joint_delta_mv based on the POC distance of the first and second reference frames up to the current frame and the orientation of the two reference frames.
[0195] In some other implementations, the inter-prediction mode of the current block is NEW_NEWMV; and step 1840 may include the step of determining a first delta MV as a joint delta MV, and the step of determining a second delta MV by scaling the joint delta MV according to at least one of a first picture order count (POC) distance between a first reference frame and a current frame, a second POC distance between a second reference frame and a current frame, or the directional relationship between the first reference frame and the second reference frame with respect to the current frame.
[0196] In some other implementations, the inter-prediction mode of the current block is NEW_NEWMV; and step 1840 may include determining a second delta MV as a joint delta MV, and determining a first delta MV by scaling the joint delta MV according to at least one of a first picture order count (POC) distance between a first reference frame and a current frame, a second POC distance between a second reference frame and a current frame, or the orientation relationship between the first reference frame and the second reference frame with respect to the current frame.
[0197] In some embodiments, the delta MV in reference list 0 (or list 1) may always be set to be equal to joint_delta_mv, and the delta MV in reference list 1 (or list 0) may be scaled from joint_delta_mv according to the POC distances of the reference frames up to the current frame and / or the orientation of the two reference frames.
[0198] In some other implementations, the inter-prediction mode of the current block is NEW_NEWMV; and step 1840 may include: determining a first delta MV as a joint delta MV in response to a first absolute POC distance between a first reference frame and a current frame being greater than a second absolute POC distance between a second reference frame and a current frame; and determining a second delta MV by scaling the joint delta MV according to at least one of the first POC distance between a first reference frame and a current frame, the second POC distance between a second reference frame and a current frame, or the orientation relationship between the first reference frame and the second reference frame with respect to the current frame.
[0199] In some other implementations, the inter-prediction mode of the current block is NEW_NEWMV; and step 1840 may include: determining a second delta MV as a joint delta MV in response to the first absolute POC distance between the first reference frame and the current frame being smaller than the second absolute POC distance between the second reference frame and the current frame; and determining a first delta MV by scaling the joint delta MV according to at least one of the first POC distance between the first reference frame and the current frame, the second POC distance between the second reference frame and the current frame, or the orientation relationship between the first reference frame and the second reference frame with respect to the current frame.
[0200] In some embodiments, when the absolute POC distance between reference list 0 (or list 1) and the current frame is greater than the absolute POC distance between reference list 1 (or list 0) and the current frame, the delta MV in reference list 0 (or list 1) having the greater absolute POC distance may be set to be equal to joint_delta_mv. The delta MV in reference list 1 (or list 0) having the smaller absolute POC distance may be scaled from joint_delta_mv according to the POC distances of the reference frames up to the current frame and / or the orientation of the two reference frames.
[0201] In some other implementations, the inter-prediction mode of the current block is NEW_NEWMV; and step 1840 may include: determining a first delta MV as a joint delta MV in response to the first absolute POC distance between the first reference frame and the current frame being equal to the second absolute POC distance between the second reference frame and the current frame; determining a second delta MV as a joint delta MV in response to the first reference frame and the second reference frame being equal to the current frame; and determining a second delta MV as a joint delta MV multiplied by -1 in response to the first reference frame and the second reference frame being opposite to the current frame.
[0202] In some embodiments, when the absolute POC distance between reference list 1 and the current frame is the same as the absolute POC distance between reference list 0 and the current frame, the delta MV in reference list 0 may be set to be equal to joint_delta_mv. When the orientations of the two reference frames are the same, the delta MV in reference list 1 may also be set to be equal to joint_delta_mv. Otherwise, when the orientations of the two reference frames are different, the delta MV for reference list 1 is set to joint_delta_mv multiplied by -1.
[0203] In various embodiments / implements of the present disclosure, scaling the joint delta MV to obtain the second delta MV according to at least one of a first POC distance between the first reference frame and the current frame, a second POC distance between the second reference frame and the current frame, or an orientation relationship between the first reference frame and the second reference frame with respect to the current frame may include a linear scaling method, that is, the absolute value of the scaled delta MV may be proportional to the ratio of the second POC distance divided by the first POC distance, and the sign of the scaled delta MV may be determined according to the orientation relationship. As an example, when the first POC distance is 4 and the second POC distance is 8 and the orientation relationship between the first reference frame and the second reference frame with respect to the current frame is the same direction, the second delta MV is obtained by scaling / multiplying the joint delta MV by a factor of 2 (=8 / 4); and because the orientation relationship is the same direction, the second delta MV has the same sign as the joint delta MV. As another example, when the first POC distance is 3 and the second POC distance is -9 and the orientation relationship between the first reference frame and the second reference frame with respect to the current frame is opposite, the second delta MV is obtained by scaling / multiplying the joint delta MV by a factor of -3 (= -9 / 3); and because the orientation relationship is opposite, the second delta MV has the opposite sign to the joint delta MV.
[0204] In some other implementations, the inter-prediction mode of the current block is NEW_NEARMV; and the second delta MV may be slightly adjusted by a predefined weight. Step 1840 may include: determining the first delta MV as a joint delta MV; and determining the second delta MV by scaling the joint delta MV according to at least one of a first POC distance between a first reference frame and a current frame, a second POC distance between a second reference frame and a current frame, the directional relationship between the first reference frame and the second reference frame with respect to the current frame, or a predefined weighting factor.
[0205] In some other implementations, the inter-prediction mode of the current block is NEAR_NEWMV; and the first delta MV may be slightly adjusted by a predefined weight. Step 1840 may include: determining the second delta MV as a joint delta MV; and determining the first delta MV by scaling the joint delta MV according to at least one of a first POC distance between the first reference frame and the current frame, a second POC distance between the second reference frame and the current frame, the directional relationship between the first reference frame and the second reference frame with respect to the current frame, or a predefined weighting factor.
[0206] In some other implementations, the predefined weighting factor is a fraction between -1 and 1.
[0207] In some other implementations, a predefined weighting factor is signaled to a high-level syntax including at least one of: a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a tile header, a slice header, a frame header, a coding tree unit (CTU) header, or a superblock header.
[0208] In some embodiments, when the current block is in NEW_NEAR mode (or NEAR_NEW mode) and joint_mvd_flag indicates that the delta MVs for reference list 0 and reference list 1 are jointly signaled, the delta MV for reference list 0 (or list 1) may be set to be equal to joint_delta_mv, and the delta MV for list 1 (or list 0) is scaled from joint_delta_mv based on one or more coded information including, but not limited to, the POC distance of the reference frames to the current frame, the orientation of the two reference frames, the difference between the MV predictors of the two MVs, and / or a predefined weighting factor w. In some implementations, this predefined weighting factor may be a number between -1 and 1, e.g., 1 / 2. In some other implementations, these predefined weighting factors may be signaled to high-level syntax including, but not limited to, SPS, VPS, PPS, picture header, tile header, slice header, frame header, and CTU (or superblock) header.
[0209] In various embodiments / implements of the present disclosure, scaling the joint delta MV to obtain a weighted-adjusted delta MV according to at least one of a first POC distance between a first reference frame and a current frame, a second POC distance between a second reference frame and a current frame, an directional relationship between the first reference frame and the second reference frame with respect to the current frame, or a predefined weighting factor may include a linear scaling method, that is, the absolute value of the scaled delta MV may be proportional to the ratio of the second POC distance divided by the first POC distance, and then multiplied by a predefined weighting factor. The sign of the scaled delta MV may be determined according to the directional relationship. As an example, when the first POC distance is 4 and the second POC distance is 8, and the orientation relationship between the first reference frame and the second reference frame with respect to the current frame is the same direction, and the predefined weighting factor is 1 / 2, the weighted-adjusted delta MV is obtained by scaling / multiplying the joint delta MV by a total factor of 1, calculated according to the value of multiplying the factor of 2 (= 8 / 4) by the weighting factor of 1 / 2; and because the orientation relationship is the same direction, the weighted-adjusted delta MV has the same sign as the joint delta MV. As another example, when the first POC distance is 3 and the second POC distance is -9, and the orientation relationship between the first reference frame and the second reference frame with respect to the current frame is opposite, and the predefined weighting factor is 1 / 2, the weighted-adjusted delta MV is obtained by scaling / multiplying the joint delta MV by a total factor of -3 / 2, calculated according to the value of multiplying the factor of -3 (= -9 / 3) by the weighting factor of 1 / 2. Because the directional relationship is opposite, the weighted-adjusted delta MV has the opposite sign to the joint delta MV.
[0210] The embodiments of the present disclosure may be used individually or combined in any order. Any steps and / or operations in any of the embodiments of the present disclosure may be combined or arranged in any amount or order as desired. Two or more of the steps and / or operations in any of the embodiments of the present disclosure may be performed in parallel. Additionally, each of the methods (or embodiments), the encoder, and the decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transient computer-readable medium. The embodiments of the present disclosure may be applied to a lumina block or a chroma block.
[0211] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, FIG. 19 illustrates a computer system (2000) suitable for implementing specific embodiments of the disclosed subject matter.
[0212] Computer software may be coded using any suitable machine code or computer language in which assembly, compilation, linking, or similar mechanisms can be performed to generate code containing instructions that can be executed directly or through interpretation, micro-code execution, etc., by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.
[0213] The commands can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0214] The components illustrated in FIG. 19 for the computer system (2000) are exemplary in nature and are not intended to imply any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should not be interpreted as having any dependency or requirement with respect to any one or a combination thereof of the components illustrated in the exemplary embodiments of the computer system (2000).
[0215] The computer system (2000) may include specific human interface input devices. Such human interface input devices may respond to input by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input (not described). Human interface devices may also be used to capture specific media that are not necessarily directly related to conscious input by humans, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., 2D video, 3D video including stereoscopic video).
[0216] Input human interface devices may include one or more of: a keyboard (2001), a mouse (2002), a trackpad (2003), a touch screen (2010), a data glove (not shown), a joystick (2005), a microphone (2006), a scanner (2007), and a camera (2008) (only one of each is depicted).
[0217] The computer system (2000) may also include specific human interface output devices. Such human interface output devices may stimulate one or more human user's senses through, for example, tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., touch-screen (2010), data-glove (not shown), or haptic feedback via joystick (2005), but there may also be haptic feedback devices that do not serve as input devices), audio output devices (e.g., speakers (2009), headphones (not depicted)), visual output devices (e.g., screens (2010) including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touch-screen input capability and with or without haptic feedback capability—some of which may output two-dimensional visual output or output beyond three dimensions through means such as stereographic output—; virtual reality glasses (not depicted), holographic displays and smoke tanks (not depicted)), and printers (not depicted).
[0218] The computer system (2000) may also include human-accessible storage devices and their associated media, such as optical media including a CD / DVD ROM / RW (2020) having a CD / DVD ROM / RW (2021), a thumb drive (2022), a removable hard drive or solid-state drive (2023), legacy magnetic media such as tape and floppy disk (not described), specialized ROM / ASIC / PLD-based devices such as a security dongle (not described).
[0219] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not include transmission media, carriers, or other transient signals.
[0220] The computer system (2000) may also include an interface (2054) to one or more communication networks (2055). The networks may be, for example, wireless, wired, optical. The networks may additionally be local, wide-area, metropolitan, automotive and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks, such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, automotive and industrial including CAN buses, etc. Certain networks generally require external network interface adapters attached to specific general-purpose data ports or peripheral buses (2049) (e.g., USB ports of the computer system (2000)); Others are generally integrated into the core of the computer system (2000) by attachment to a system bus (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system) as described below. Using any of these networks, the computer system (2000) can communicate with other entities. Such communication may be unidirectional receive-only (e.g., broadcast TV), unidirectional transmit-only (e.g., CANbus to specific CANbus devices), or bidirectional with other computer systems using, for example, local area or wide area digital networks. Specific protocols and protocol stacks may be used for each of the networks and network interfaces described above.
[0221] The aforementioned human interface devices, human-accessible storage devices, and network interfaces can be attached to the core (2040) of the computer system (2000).
[0222] The core (2040) may include one or more central processing units (CPUs) (2041), specialized programmable processing units in the form of graphics processing units (GPUs) (2042), field programmable gate areas (FPGAs) (2043), hardware accelerators (2044) for specific tasks, graphics adapters (2050), etc. These devices may be connected via a system bus (2048), along with internal mass storage (2047), such as read-only memory (ROM) (2045), random access memory (2046), internal non-user accessible hard drives, SSDs, etc. In some computer systems, the system bus (2048) may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (2048) or via a peripheral bus (2049). In one example, the screen (2010) can be connected to a graphics adapter (2050). Architectures for peripheral buses include PCI, USB, etc.
[0223] CPUs (2041), GPUs (2042), FPGAs (2043), and accelerators (2044) can be combined to execute specific instructions that can constitute the aforementioned computer code. The computer code may be stored in ROM (2045) or RAM (2046). Transient data may also be stored in RAM (2046), while persistent data may be stored, for example, in internal mass storage (2047). High-speed storage and retrieval of any of the memory devices may be made possible through the use of cache memory, which may be closely associated with one or more CPUs (2041), GPUs (2042), mass storage (2047), ROM (2045), RAM (2046), etc.
[0224] A computer-readable medium may have computer code for performing various computer-implemented operations. The medium and the computer code may be those specifically designed and configured for the purposes of this disclosure, or they may be of a kind well known and available to those skilled in the field of computer software technology.
[0225] As a non-limiting example, a computer system (2000) having an architecture, and specifically a core (2040), may provide functionality as a result of processor(s) (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software implemented on one or more types of tangible computer-readable media. Such computer-readable media may be media associated with specific storage of the core (2040) that is of a non-transient nature, such as core-internal mass storage (2047) or ROM (2045), as well as user-accessible mass storage as described above. Software implementing various embodiments of the present disclosure may be stored on such devices and executed by the core (2040). The computer-readable media may include one or more memory devices or chips as needed. Software may enable the core (2040) and, specifically, the processors within it (including a CPU, GPU, FPGA, etc.) to execute specific processes or specific parts of specific processes described herein, including defining data structures stored in RAM (2046) and modifying such data structures according to processes defined by the software. Additionally, or alternatively, the computer system may provide functionality as a result of logic hardwired or otherwise implemented in a circuit (e.g., accelerator (2044)) that can operate instead of or with the software to execute specific processes or specific parts of specific processes described herein. Reference to software may include logic, where appropriate, and vice versa. Reference to a computer-readable medium may include, where appropriate, a circuit storing software for execution (e.g., an integrated circuit (IC)), or a circuit implementing logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.
[0226] Although specific inventions have been described with reference to exemplary embodiments, this description is not intended to be limiting. Various modifications of exemplary and additional embodiments of the invention will be apparent to those skilled in the art from this description. Those skilled in the art will readily recognize that such and various other modifications may be made to the exemplary embodiments illustrated and described herein without departing from the spirit and scope of the invention. Accordingly, the appended claims are considered to cover any such modifications and alternative embodiments. Certain proportions in the drawings may be exaggerated, and other proportions may be minimized. Accordingly, the disclosure and drawings should be considered as examples rather than limitations.
[0227] The following is a list of acronyms, some of which may appear in the present disclosure:
[0228] JEM: joint exploration model
[0229] VVC: versatile video coding
[0230] BMS: benchmark set
[0231] MV: Motion Vector
[0232] HEVC: High Efficiency Video Coding
[0233] SEI: Supplementary Enhancement Information
[0234] VUI: Video Usability Information
[0235] GOPs: Groups of Pictures
[0236] TUs: Transform Units,
[0237] PUs: Prediction Units
[0238] CTUs: Coding Tree Units
[0239] CTBs: Coding Tree Blocks
[0240] PBs: Prediction Blocks
[0241] HRD: Hypothetical Reference Decoder
[0242] SNR: Signal Noise Ratio
[0243] CPUs: Central Processing Units
[0244] GPUs: Graphics Processing Units
[0245] CRT: Cathode Ray Tube
[0246] LCD: Liquid-Crystal Display
[0247] OLED: Organic Light-Emitting Diode
[0248] CD: Compact Disc
[0249] DVD: Digital Video Disc
[0250] ROM: Read-Only Memory
[0251] RAM: Random Access Memory
[0252] ASIC: Application-Specific Integrated Circuit
[0253] PLD: Programmable Logic Device
[0254] LAN: Local Area Network
[0255] GSM: Global System for Mobile communications
[0256] LTE: Long-Term Evolution
[0257] CANBus: Controller Area Network Bus
[0258] USB: Universal Serial Bus
[0259] PCI: Peripheral Component Interconnect
[0260] FPGA: Field Programmable Gate Areas
[0261] SSD: solid-state drive
[0262] IC: Integrated Circuit
[0263] HDR: high dynamic range
[0264] SDR: standard dynamic range
[0265] JVET: Joint Video Exploration Team
[0266] MPM: most probable mode
[0267] WAIP: Wide-Angle Intra Prediction
[0268] CU: Coding Unit
[0269] PU: Prediction Unit
[0270] TU: Transform Unit
[0271] CTU: Coding Tree Unit
[0272] PDPC: Position Dependent Prediction Combination
[0273] ISP: Intra Sub-Partitions
[0274] SPS: Sequence Parameter Setting
[0275] PPS: Picture Parameter Set
[0276] APS: Adaptation Parameter Set
[0277] VPS: Video Parameter Set
[0278] DPS: Decoding Parameter Set
[0279] ALF: Adaptive Loop Filter
[0280] SAO: Sample Adaptive Offset
[0281] CC-ALF: Cross-Component Adaptive Loop Filter
[0282] CDEF: Constrained Directional Enhancement Filter
[0283] CCSO: Cross-Component Sample Offset
[0284] LSO: Local Sample Offset
[0285] LR: Loop Restoration Filter
[0286] AV1: AOMedia Video 1
[0287] AV2: AOMedia Video 2
[0288] MVD: Motion Vector difference
[0289] CfL: Chroma from Luma
[0290] SDT: Semi Decoupled Tree
[0291] SDP: Semi Decoupled Partitioning
[0292] SST: Semi Separate Tree
[0293] SB: Super Block
[0294] IBC (or IntraBC): Intra Block Copy
[0295] CDF: Cumulative Density Function
[0296] SCC: Screen Content Coding
[0297] GBI: Generalized Bi-prediction
[0298] BCW: Bi-prediction with CU-level Weights
[0299] CIIP: Combined intra-inter prediction
[0300] POC: Picture Order Count
[0301] RPS: Reference Picture Set
[0302] DPB: Decoded Picture Buffer
[0303] MMVD: Merge Mode with Motion Vector Difference.
Claims
Claim 1 A method for decoding an inter-predicted video block comprises: receiving a coded video bitstream by a device comprising a memory storing instructions and a processor communicating with said memory; extracting an inter-predicted mode and a joint delta motion vector (MV) (joint_delta_mv) for a current block within a current frame by said device from the coded video bitstream; extracting a flag (joint_mvd_flag) from said device from the coded video bitstream indicating whether a first delta MV for a first reference frame in reference list 0 and a second delta MV for a second reference frame in reference list 1 are jointly signaled; and in response to said flag indicating that the first delta MV and the second delta MV are jointly signaled by said device, deriving the first delta MV and the second delta MV based on said joint delta MV. and the device includes the step of decoding the current block within the current frame based on the first delta MV and the second delta MV, wherein the inter-prediction mode of the current block is NEW_NEWMV;A method comprising the step of deriving the first delta MV and the second delta MV based on the joint delta MV, wherein in response to the first absolute POC distance between the first reference frame and the current frame being equal to the second absolute POC distance between the second reference frame and the current frame: determining the first delta MV as the joint delta MV; in response to the directional relationship between the first reference frame and the second reference frame with respect to the current frame being the same: determining the second delta MV as the joint delta MV; and in response to the directional relationship between the first reference frame and the second reference frame with respect to the current frame being opposite: determining the second delta MV as the joint delta MV multiplied by -1. Claim 2 The method of claim 1, wherein the step of deriving the first delta MV and the second delta MV based on the joint delta MV comprises: determining the first delta MV as the joint delta MV; and determining the second delta MV by scaling the joint delta MV according to at least one of a first picture order count (POC) distance between the first reference frame and the current frame, a second POC distance between the second reference frame and the current frame, or the directional relationship between the first reference frame and the second reference frame with respect to the current frame. Claim 3 The method of claim 1, wherein the step of deriving the first delta MV and the second delta MV based on the joint delta MV comprises: determining the second delta MV as the joint delta MV; and determining the first delta MV by scaling the joint delta MV according to at least one of a first picture order count (POC) distance between the first reference frame and the current frame, a second POC distance between the second reference frame and the current frame, or the directional relationship between the first reference frame and the second reference frame with respect to the current frame. Claim 4 The method of claim 1, wherein the step of deriving the first delta MV and the second delta MV based on the joint delta MV comprises: determining the first delta MV as the joint delta MV in response to the first absolute POC distance between the first reference frame and the current frame being greater than the second absolute POC distance between the second reference frame and the current frame; and determining the second delta MV by scaling the joint delta MV according to at least one of the first POC distance between the first reference frame and the current frame, the second POC distance between the second reference frame and the current frame, or the directional relationship between the first reference frame and the second reference frame with respect to the current frame. Claim 5 The method of claim 1, wherein the step of deriving the first delta MV and the second delta MV based on the joint delta MV comprises: determining the second delta MV as the joint delta MV in response to the first absolute POC distance between the first reference frame and the current frame being smaller than the second absolute POC distance between the second reference frame and the current frame; and determining the first delta MV by scaling the joint delta MV according to at least one of the first POC distance between the first reference frame and the current frame, the second POC distance between the second reference frame and the current frame, or the directional relationship between the first reference frame and the second reference frame with respect to the current frame. Claim 6 An apparatus for decoding an inter-predicted video block, comprising: a memory storing instructions; and a processor communicating with said memory, wherein when said processor executes said instructions, said processor is configured to cause said apparatus to perform the method of any one of claims 1 to 5. Claim 7 A non-transient computer-readable storage medium storing instructions, wherein when the instructions are executed by a processor, the instructions are configured to cause the processor to perform the method of any one of claims 1 to 5. Claim 8 A method for encoding video data comprises: determining whether an inter-predication mode and a delta motion vector (MV) are applied to a current block within a current frame of the video data; generating a flag (joint_mvd_flag) indicating whether a first delta MV for a first reference frame in reference list 0 and a second delta MV for a second reference frame in reference list 1 are jointly signaled for composite inter-predication of the current block; and encoding the flag in a bitstream for the video data, wherein when it is determined that the inter-predication mode of the current block is NEW_NEWMV and the first delta MV and the second delta MV are jointly signaled: when the first absolute POC distance between the first reference frame and the current frame is equal to the second absolute POC distance between the second reference frame and the current frame, the first delta MV is determined as a joint delta MV; A method for determining the second delta MV as the joint delta MV when the directional relationship between the first reference frame and the second reference frame with respect to the current frame is the same; and determining the second delta MV as the joint delta MV multiplied by -1 when the directional relationship between the first reference frame and the second reference frame with respect to the current frame is opposite. Claim 9 A non-transient computer-readable medium storing a video bitstream generated according to the method of paragraph 8. Claim 10 delete Claim 11 delete Claim 12 delete Claim 13 delete Claim 14 delete
Citation Information
Patent Citations
Video coding using adaptive motion vector resolution
KR1020140043807A
Dynamic reference motion vector coding mode
US20170223350A1